CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Can cleaning repair UTF-16 bytes decoded as UTF-8?
Not in our byte fixtures. Decoding UTF-16LE bytes as UTF-8 leaves NULs or replacement characters; the cleaner does not reconstruct the intended text. Use the correct decoder before cleaning.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Wrong UTF-8 ASCII body direct string / keep / invisible true | "A\u0000B\u0000"U+0041 U+0000 U+0042 U+0000 | "A\u0000B\u0000"No selected invisible characters found | {"hasNul":true,"hasReplacement":false,"equalsAB":false,"equalsAccent":false}{"hasNul":true,"hasReplacement":false,"equalsAB":false,"equalsAccent":false} |
| Correct UTF-16LE ASCII body direct string / keep / invisible true | "AB"U+0041 U+0042 | "AB"No selected invisible characters found | {"hasNul":false,"hasReplacement":false,"equalsAB":true,"equalsAccent":false}{"hasNul":false,"hasReplacement":false,"equalsAB":true,"equalsAccent":false} |
| Wrong UTF-8 non-ASCII body direct string / keep / invisible true | "�\u0000"U+FFFD U+0000 | "�\u0000"No selected invisible characters found | {"hasNul":true,"hasReplacement":true,"equalsAB":false,"equalsAccent":false}{"hasNul":true,"hasReplacement":true,"equalsAB":false,"equalsAccent":false} |
| Correct UTF-16LE non-ASCII body direct string / keep / invisible true | "é"U+00E9 | "é"No selected invisible characters found | {"hasNul":false,"hasReplacement":false,"equalsAB":false,"equalsAccent":true}{"hasNul":false,"hasReplacement":false,"equalsAB":false,"equalsAccent":true} |
The wrong decoder loses the intended representation
The ASCII body AB has UTF-16LE bytes 41 00 42 00. Reading those bytes as UTF-8 produces A, NUL, B, NUL, which cleanup retains. Reading the same bytes as UTF-16LE produces AB. The non-ASCII body é has bytes E9 00; the UTF-8 decoder with its default error policy emits a replacement character and NUL, while the UTF-16LE decoder produces é. We save these exact decoded inputs in results.json. No selected character in the malformed representation is available for the cleaner to delete.
This tests bodies without a BOM
Both byte fixtures omit a byte-order mark intentionally. That isolates wrong character decoding from the leading UTF-8 BOM behavior covered in our earlier guide. The harness constructs the byte buffers locally and uses named TextDecoder encodings. It does not measure the browser File.text path, detect a file encoding automatically or infer the original bytes from pasted text. Once a replacement character has replaced an undecodable byte, cleaning that string alone cannot recover which byte was lost. The original byte buffer is the evidence needed for recovery.
Return to the original file bytes
Preserve the file and confirm its encoding from the producer, format declaration or a reviewed byte inspection. Decode those bytes with the intended decoder, compare the decoded content, and only then consider selective Unicode cleanup. Deleting every NUL from a wrong decoding may appear to repair ASCII while leaving other text corrupted. Our homepage receives text and has no general encoding-selection interface. These two short bodies demonstrate a decoding boundary, not support for every UTF-16 file, endian order or error policy.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
WHATWG Encoding Standard provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.