Zerowidth CleanerAll measured guides

CLAUDE WATERMARK REMOVER · PRACTICAL TEST

Can removing a zero-width space change word segmentation?

Yes. Our English Intl.Segmenter fixture splits alpha and beta at U+200B; deleting that character yields one word-like segment, alphabeta. Inspect intended boundaries before cleanup.

Tested October 8, 2026 · v24.19.0 · 4 measured runs

Open the Claude text cleaner · Full measured data

Code-point comparison for Selected boundary

Measured inputs and outputs

These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Three destination checks use separately recorded browser observations; those pages explain the distinction. We make no detector-score or statistical-watermark removal claim.

Before and after observations; full code points and strings are in the JSON download.
Fixture and modeInput string and code pointsOutput and cleaner statusObserved before / after
Selected boundary
direct string / keep / invisible true
"alpha​beta"
U+0061 U+006C U+0070 U+0068 U+0061 U+200B U+0062 U+0065 U+0074 U+0061
"alphabeta"
Removed 1 invisible character
{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":"​","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}
{"segments":[{"segment":"alphabeta","index":0,"isWordLike":true}]}
Ordinary space
direct string / keep / invisible true
"alpha beta"
U+0061 U+006C U+0070 U+0068 U+0061 U+0020 U+0062 U+0065 U+0074 U+0061
"alpha beta"
No selected invisible characters found
{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":" ","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}
{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":" ","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}
Hyphen boundary
direct string / keep / invisible true
"alpha-beta"
U+0061 U+006C U+0070 U+0068 U+0061 U+002D U+0062 U+0065 U+0074 U+0061
"alpha-beta"
No selected invisible characters found
{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":"-","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}
{"segments":[{"segment":"alpha","index":0,"isWordLike":true},{"segment":"-","index":5,"isWordLike":false},{"segment":"beta","index":6,"isWordLike":true}]}
Plain combined word
direct string / keep / invisible true
"alphabeta"
U+0061 U+006C U+0070 U+0068 U+0061 U+0062 U+0065 U+0074 U+0061
"alphabeta"
No selected invisible characters found
{"segments":[{"segment":"alphabeta","index":0,"isWordLike":true}]}
{"segments":[{"segment":"alphabeta","index":0,"isWordLike":true}]}

The observer reports segments, not guessed words

The selected-boundary fixture initially yields alpha, the U+200B boundary segment and beta, with the word-like flags recorded for each segment. After cleanup, the observer returns the combined word-like segment alphabeta. The ordinary-space reference retains its separation, and the hyphen reference also remains unchanged under the keep-dashes setting. The plain combined reference shows the same word-like output as the cleaned first fixture. Saving indexes and segment text makes the changed boundary visible without assigning a search volume, linguistic quality score or editorial word count to the result.

Segmentation depends on an explicit policy

Our observer requests English word granularity through Intl.Segmenter in the saved Node runtime. Locale data and implementation versions can affect segmentation, so the result record includes the runtime and reproducer. This is a destination segmentation experiment, while our earlier joining-words page concerned raw string deletion and our regex page concerned pattern interpretation. The homepage does not invoke a segmenter or replace a removed boundary with an ordinary space. No tokenizer used by Claude, search engine or language model was tested. A word-like flag describes this API output and does not establish a universal definition of a word.

Choose a replacement only when meaning supports it

Keep the original sequence and determine whether the boundary separates intended words or interrupts a single word. If it separates two words, deleting it alone may create an unwanted combined token; an editorial repair might require a visible space, reviewed in context. If it interrupts one identifier, a space could instead be wrong. Rerun the destination segmenter or tokenizer with its documented locale and settings after the deliberate change. The cleaner cannot infer this decision from an invisible code point alone. The four fixtures provide exact evidence of one segmentation change, not a guarantee about indexing, model tokens or all languages.

Reproduce this test

Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.

node reproduce.cjs

Reference and next check

ECMA-402 Intl.Segmenter provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.