CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Why do separately decoded chunks create replacement characters?
Our two-byte split cuts inside the UTF-8 marker, accent and emoji. Separate decoders produce U+FFFD; one TextDecoder with stream:true preserves each complete string. Cleanup does not fix chunk handling.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Destination observations run independently on the exact input and cleaned strings in the recorded Node runtime. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Marker straddles chunks direct string / keep / invisible true | "AB"U+0041 U+200B U+0042 | "AB"Removed 1 invisible character | {"leftHex":"41e2","rightHex":"808b42","separate":"A���B","streamFirst":"A","streamFinal":"B","streamJoined":"AB"}{"leftHex":"4142","rightHex":"","separate":"AB","streamFirst":"AB","streamFinal":"","streamJoined":"AB"} |
| Accent straddles chunks direct string / keep / invisible true | "AéB"U+0041 U+00E9 U+0042 | "AéB"No selected invisible characters found | {"leftHex":"41c3","rightHex":"a942","separate":"A��B","streamFirst":"A","streamFinal":"éB","streamJoined":"AéB"}{"leftHex":"41c3","rightHex":"a942","separate":"A��B","streamFirst":"A","streamFinal":"éB","streamJoined":"AéB"} |
| Emoji straddles chunks direct string / keep / invisible true | "A😀B"U+0041 U+1F600 U+0042 | "A😀B"No selected invisible characters found | {"leftHex":"41f0","rightHex":"9f988042","separate":"A����B","streamFirst":"A","streamFinal":"😀B","streamJoined":"A😀B"}{"leftHex":"41f0","rightHex":"9f988042","separate":"A����B","streamFirst":"A","streamFinal":"😀B","streamJoined":"A😀B"} |
| Plain reference direct string / keep / invisible true | "AB"U+0041 U+0042 | "AB"No selected invisible characters found | {"leftHex":"4142","rightHex":"","separate":"AB","streamFirst":"AB","streamFinal":"","streamJoined":"AB"}{"leftHex":"4142","rightHex":"","separate":"AB","streamFirst":"AB","streamFinal":"","streamJoined":"AB"} |
The same bytes have different chunk outcomes
We encode each exact string as UTF-8 and split its bytes after the second byte. For A-marker-B, that split leaves A and the first marker byte on the left, with the remaining marker bytes and B on the right. Independent TextDecoder instances yield A followed by three replacement characters and B. A single decoder retains state after its stream:true call, then completes the marker during the final call. Cleaning first removes that particular marker and leaves ASCII AB, so the split no longer cuts a multibyte sequence. The accent and emoji controls still cross the boundary after cleanup and continue to require stateful decoding.
A complete-buffer test cannot establish a streaming result
This observer explicitly creates two chunks and compares two decoder lifecycles. The final call on the shared decoder omits stream:true, so the stream is completed. The earlier BOM and wrong-encoding pages supply complete buffers; neither tests state carried between chunks. We do not read a real fetch body, rely on actual packet sizes or exercise a network reader loop. The two-byte split is an authored adversarial boundary, not a prediction of browser chunk lengths. The untouched ASCII reference verifies that independent decoding can appear correct when there is no split multibyte sequence. That success does not generalize to the retained accent or emoji.
Finish decoding before inspecting the resulting text
When consuming a real byte stream, preserve one decoder for its chunks and finish that decoder when the stream ends, following the relevant API contract. Inspect the completed Unicode text before making targeted character edits. The order in this experiment is deliberately string cleanup followed by re-encoding, allowing the two destination observers to receive the exact before and after strings. It is not a prescribed fetch pipeline. Once independent decoding has inserted U+FFFD, deleting selected invisible characters cannot reconstruct the original chunk bytes. These four records establish lifecycle differences on known valid encodings, without testing truncated final input, fatal mode or arbitrary network failures.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
WHATWG Encoding Standard provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.