Zerowidth CleanerAll measured guides

CLAUDE WATERMARK REMOVER · PRACTICAL TEST

Why can a saved sentence index be wrong after cleanup?

Our second sentence starts at UTF-16 index 5 before marker deletion and 4 afterward. Recompute offsets against the edited string; a retained emoji also shows why code-point counts are different.

Tested October 11, 2026 · v24.19.0 · 4 measured runs

Open the Claude text cleaner · Full measured data

Code-point comparison for Marker in first sentence

Measured inputs and outputs

These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Destination observations run independently on the exact input and cleaned strings in the recorded Node runtime. We make no detector-score or statistical-watermark removal claim.

Before and after observations; full code points and strings are in the JSON download.
Fixture and modeInput string and code pointsOutput and cleaner statusObserved before / after
Marker in first sentence
direct string / keep / invisible true
"Hi​. Next."
U+0048 U+0069 U+200B U+002E U+0020 U+004E U+0065 U+0078 U+0074 U+002E
"Hi. Next."
Removed 1 invisible character
{"segments":[{"segment":"Hi​. ","index":0},{"segment":"Next.","index":5}],"utf16Units":10,"codePoints":10}
{"segments":[{"segment":"Hi. ","index":0},{"segment":"Next.","index":4}],"utf16Units":9,"codePoints":9}
Plain two sentences
direct string / keep / invisible true
"Hi. Next."
U+0048 U+0069 U+002E U+0020 U+004E U+0065 U+0078 U+0074 U+002E
"Hi. Next."
No selected invisible characters found
{"segments":[{"segment":"Hi. ","index":0},{"segment":"Next.","index":4}],"utf16Units":9,"codePoints":9}
{"segments":[{"segment":"Hi. ","index":0},{"segment":"Next.","index":4}],"utf16Units":9,"codePoints":9}
Retained joiner
direct string / keep / invisible true
"Hi‍. Next."
U+0048 U+0069 U+200D U+002E U+0020 U+004E U+0065 U+0078 U+0074 U+002E
"Hi‍. Next."
No selected invisible characters found
{"segments":[{"segment":"Hi‍. ","index":0},{"segment":"Next.","index":5}],"utf16Units":10,"codePoints":10}
{"segments":[{"segment":"Hi‍. ","index":0},{"segment":"Next.","index":5}],"utf16Units":10,"codePoints":10}
Emoji sentence
direct string / keep / invisible true
"😀. Next."
U+1F600 U+002E U+0020 U+004E U+0065 U+0078 U+0074 U+002E
"😀. Next."
No selected invisible characters found
{"segments":[{"segment":"😀. ","index":0},{"segment":"Next.","index":4}],"utf16Units":9,"codePoints":8}
{"segments":[{"segment":"😀. ","index":0},{"segment":"Next.","index":4}],"utf16Units":9,"codePoints":8}

The sentence count stays stable while its position moves

The marked Hi sentence and Next sentence form two segments in the saved English sentence segmenter. The second segment begins at index five before cleanup and four afterward. Its wording is unchanged; the removed U+200B shifted its position in the whole string. The unmarked reference starts at four throughout. The retained-joiner case still starts at five. The smile reference places the second sentence at index four even though the complete text has eight code points and nine UTF-16 units. We record every segment string and start index rather than assuming that an unchanged sentence count makes old offsets reusable.

Sentence indices are measured in one concrete runtime

Intl.Segmenter uses en and sentence granularity in this observer. The older segmentation guide uses word granularity and records wordlike information. Our test instead concerns a position that could be stored for highlights or annotations. We do not attach a real annotation, inspect a browser selection or claim universal linguistic sentence boundaries. The Node and ICU versions can influence segmentation rules, so preserve runtime metadata with a rerun. The cleaner does not know about saved indices and cannot remap them for an editor. It changes the underlying string before the observer computes positions; both exact stages are retained in the JSON record.

Keep annotations associated with a text version

When a destination stores offsets, record the precise string version and indexing unit that produced them. After any edit, recompute or deliberately remap positions against the new text according to that application’s rules. Do not subtract a global removal total from every position: a removal after an annotation would affect it differently from a removal before it. Our small matrix only deletes one marker before the second sentence. It establishes a concrete shift and unchanged controls, without implementing a general offset mapping algorithm. Existing word and grapheme measurements can help choose a unit, but neither substitutes for validating the stored sentence range on the edited representation.

Reproduce this test

Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.

node reproduce.cjs

Reference and next check

ECMA-402 Intl.Segmenter provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.