CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Can removing a zero-width space change word segmentation?
Yes. Our English Intl.Segmenter fixture splits alpha and beta at U+200B; deleting that character yields one word-like segment, alphabeta. Inspect intended boundaries before cleanup.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Three destination checks use separately recorded browser observations; those pages explain the distinction. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Selected boundary direct string / keep / invisible true | "alphabeta"U+0061 U+006C U+0070 U+0068 U+0061 U+200B U+0062 U+0065 U+0074 U+0061 | "alphabeta"Removed 1 invisible character | {"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":"","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}{"segments":[{"segment":"alphabeta","index":0,"isWordLike":true}]} |
| Ordinary space direct string / keep / invisible true | "alpha beta"U+0061 U+006C U+0070 U+0068 U+0061 U+0020 U+0062 U+0065 U+0074 U+0061 | "alpha beta"No selected invisible characters found | {"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":" ","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":" ","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]} |
| Hyphen boundary direct string / keep / invisible true | "alpha-beta"U+0061 U+006C U+0070 U+0068 U+0061 U+002D U+0062 U+0065 U+0074 U+0061 | "alpha-beta"No selected invisible characters found | {"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":"-","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":"-","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]} |
| Plain combined word direct string / keep / invisible true | "alphabeta"U+0061 U+006C U+0070 U+0068 U+0061 U+0062 U+0065 U+0074 U+0061 | "alphabeta"No selected invisible characters found | {"segments":[{"segment":"alphabeta","index":0,"isWordLike":true}]}{"segments":[{"segment":"alphabeta","index":0,"isWordLike":true}]} |
The observer reports segments, not guessed words
The selected-boundary fixture initially yields alpha, the U+200B boundary segment and beta, with the word-like flags recorded for each segment. After cleanup, the observer returns the combined word-like segment alphabeta. The ordinary-space reference retains its separation, and the hyphen reference also remains unchanged under the keep-dashes setting. The plain combined reference shows the same word-like output as the cleaned first fixture. Saving indexes and segment text makes the changed boundary visible without assigning a search volume, linguistic quality score or editorial word count to the result.
Segmentation depends on an explicit policy
Our observer requests English word granularity through Intl.Segmenter in the saved Node runtime. Locale data and implementation versions can affect segmentation, so the result record includes the runtime and reproducer. This is a destination segmentation experiment, while our earlier joining-words page concerned raw string deletion and our regex page concerned pattern interpretation. The homepage does not invoke a segmenter or replace a removed boundary with an ordinary space. No tokenizer used by Claude, search engine or language model was tested. A word-like flag describes this API output and does not establish a universal definition of a word.
Choose a replacement only when meaning supports it
Keep the original sequence and determine whether the boundary separates intended words or interrupts a single word. If it separates two words, deleting it alone may create an unwanted combined token; an editorial repair might require a visible space, reviewed in context. If it interrupts one identifier, a space could instead be wrong. Rerun the destination segmenter or tokenizer with its documented locale and settings after the deliberate change. The cleaner cannot infer this decision from an invisible code point alone. The four fixtures provide exact evidence of one segmentation change, not a guarantee about indexing, model tokens or all languages.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
ECMA-402 Intl.Segmenter provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.