CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Did the cleaner remove a BOM, or did the decoder already consume it?
Either stage can change a leading BOM. Our Node.js tests show that the default UTF-8 TextDecoder consumes a leading BOM before cleanup, while a preserving decoder exposes it. Record bytes and decoded strings separately.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. Invisible removal is enabled; the dash mode appears in each row. The observations below test this page's specific question. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Leading UTF-8 BOM default decoder / keep | "name"U+006E U+0061 U+006D U+0065 | "name"No selected invisible characters found | {"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4} |
| Leading UTF-8 BOM ignoreBOM true / keep | "name"U+FEFF U+006E U+0061 U+006D U+0065 | "name"Removed 1 invisible character | {"hasLeadingBOM":true,"hasInternalBOM":false,"codePointCount":5}{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4} |
| Internal BOM default decoder / keep | "name"U+006E U+0061 U+FEFF U+006D U+0065 | "name"Removed 1 invisible character | {"hasLeadingBOM":false,"hasInternalBOM":true,"codePointCount":5}{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4} |
| Internal BOM ignoreBOM true / keep | "name"U+006E U+0061 U+FEFF U+006D U+0065 | "name"Removed 1 invisible character | {"hasLeadingBOM":false,"hasInternalBOM":true,"codePointCount":5}{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4} |
| No BOM default decoder / keep | "name"U+006E U+0061 U+006D U+0065 | "name"No selected invisible characters found | {"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4} |
| No BOM ignoreBOM true / keep | "name"U+006E U+0061 U+006D U+0065 | "name"No selected invisible characters found | {"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4} |
Two decoding paths for the same bytes
We encode each fixture as UTF-8 and save the byte sequence as hexadecimal. The leading-BOM fixture begins EF BB BF. Default TextDecoder returns name with no leading U+FEFF; decoding with ignoreBOM: true retains U+FEFF in the string despite that option name. The cleaner therefore reports zero removals for the default path and one for the preserving path. Both end at name. The saved record separates input bytes, decoded input, cleaner output and status, making the stage of change explicit.
Position changes the decoding result
The internal fixture places EF BB BF between na and me. Both decoding paths retain an internal U+FEFF; the cleaner deletes it in both. The no-BOM reference stays name throughout. We tested Node.js TextDecoder, not the browser file picker, Excel or every encoding library. The homepage file loader calls File.text(), which is a distinct browser API; these results should not be advertised as a measured file-picker experiment. A future browser test should save the same bytes and capture the loaded field before pressing Clean.
Keep a record at each boundary
When investigating an import problem, save the original byte length and leading bytes, name the decoder and options, then inspect the resulting string. If U+FEFF was already consumed, a no-removal message from the cleaner is expected for that character. If it remains inside a header or field, assess whether it is accidental before editing it. The WHATWG Encoding Standard specifies BOM handling and the ignore-BOM flag; our downloaded script reproduces these particular measurements. Character cleanup cannot restore an original encoding marker or promise byte-identical re-export. This page addresses decoding attribution rather than comparing trim or normalization methods.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
WHATWG Encoding Standard provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.