CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Can Latin1 encoding turn an invisible marker into a control byte?
Yes in our Node Buffer fixture. U+200B becomes byte 0B under latin1, while UTF-8 uses E2 80 8B. Cleanup removes this marker, but retained U+0100 still becomes byte 00 under latin1.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Destination observations run independently on the exact input and cleaned strings in the recorded Node runtime. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Marked ASCII direct string / keep / invisible true | "AB"U+0041 U+200B U+0042 | "AB"Removed 1 invisible character | {"utf8Hex":"41e2808b42","latin1Hex":"410b42","latin1Roundtrip":"A\u000bB"}{"utf8Hex":"4142","latin1Hex":"4142","latin1Roundtrip":"AB"} |
| Latin1 accented direct string / keep / invisible true | "café"U+0063 U+0061 U+0066 U+00E9 | "café"No selected invisible characters found | {"utf8Hex":"636166c3a9","latin1Hex":"636166e9","latin1Roundtrip":"café"}{"utf8Hex":"636166c3a9","latin1Hex":"636166e9","latin1Roundtrip":"café"} |
| Outside Latin1 direct string / keep / invisible true | "Ā"U+0100 | "Ā"No selected invisible characters found | {"utf8Hex":"c480","latin1Hex":"00","latin1Roundtrip":"\u0000"}{"utf8Hex":"c480","latin1Hex":"00","latin1Roundtrip":"\u0000"} |
| Plain reference direct string / keep / invisible true | "AB"U+0041 U+0042 | "AB"No selected invisible characters found | {"utf8Hex":"4142","latin1Hex":"4142","latin1Roundtrip":"AB"}{"utf8Hex":"4142","latin1Hex":"4142","latin1Roundtrip":"AB"} |
The encoder changes more than byte length
Our A-marker-B fixture produces UTF-8 hex 41e2808b42. Node Latin1 encoding instead produces 410b42 and decoding those bytes as Latin1 yields A, a vertical-tab character and B. After deletion, both encoders produce 4142. The accented café reference remains unchanged by cleanup but has distinct UTF-8 and Latin1 byte sequences; its Latin1 roundtrip preserves the accented letter. U+0100 is also retained, yet Latin1 emits 00 and returns a NUL character. The plain AB reference matches in both encodings. These paired hex records distinguish selected deletion from lossy conversion of characters the cleaner leaves intact.
This is Node output encoding, not browser decoding
We invoke Buffer.from on JavaScript strings with explicit utf8 and latin1 names, then show a Latin1 roundtrip. The test does not load a file, use TextDecoder, set an HTTP charset or inspect a spreadsheet export. Our UTF-16 and BOM articles examine incoming bytes; this question concerns bytes generated after a cleaned string is sent to an encoder. Node defines its Latin1 behavior, and another environment using a similarly named encoding needs its own test. The controls were deliberately selected for readable hex. They do not characterize arbitrary supplementary characters or establish recovery of information after a lossy conversion.
Keep the encoding contract with the original string
Choose the byte encoding required by the destination rather than expecting invisible-character deletion to make every retained character representable. Preserve the original Unicode string and compare generated bytes or a roundtrip when investigating an export discrepancy. In this matrix, deleting one selected marker repairs only that authored marker case; the U+0100 loss remains. A changed control byte can be caused by encoding rules even when the cleaner reports no selected characters. These measurements provide reproducible byte examples and a clear boundary, not a guarantee of lossless export, a file-format validator or evidence that Claude produced a particular byte sequence.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
Node.js Buffer encodings provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.