Zerowidth CleanerAll measured guides

CLAUDE WATERMARK REMOVER · PRACTICAL TEST

Did the cleaner remove a BOM, or did the decoder already consume it?

Either stage can change a leading BOM. Our Node.js tests show that the default UTF-8 TextDecoder consumes a leading BOM before cleanup, while a preserving decoder exposes it. Record bytes and decoded strings separately.

Tested October 4, 2026 · v24.19.0 · 6 measured runs

Open the Claude text cleaner · Full measured data

Code-point comparison for Leading UTF-8 BOM

Measured inputs and outputs

These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. Invisible removal is enabled; the dash mode appears in each row. The observations below test this page's specific question. We make no detector-score or statistical-watermark removal claim.

Before and after observations; full code points and strings are in the JSON download.
Fixture and modeInput string and code pointsOutput and cleaner statusObserved before / after
Leading UTF-8 BOM
default decoder / keep
"name"
U+006E U+0061 U+006D U+0065
"name"
No selected invisible characters found
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
Leading UTF-8 BOM
ignoreBOM true / keep
"name"
U+FEFF U+006E U+0061 U+006D U+0065
"name"
Removed 1 invisible character
{"hasLeadingBOM":true,"hasInternalBOM":false,"codePointCount":5}
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
Internal BOM
default decoder / keep
"name"
U+006E U+0061 U+FEFF U+006D U+0065
"name"
Removed 1 invisible character
{"hasLeadingBOM":false,"hasInternalBOM":true,"codePointCount":5}
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
Internal BOM
ignoreBOM true / keep
"name"
U+006E U+0061 U+FEFF U+006D U+0065
"name"
Removed 1 invisible character
{"hasLeadingBOM":false,"hasInternalBOM":true,"codePointCount":5}
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
No BOM
default decoder / keep
"name"
U+006E U+0061 U+006D U+0065
"name"
No selected invisible characters found
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
No BOM
ignoreBOM true / keep
"name"
U+006E U+0061 U+006D U+0065
"name"
No selected invisible characters found
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}
{"hasLeadingBOM":false,"hasInternalBOM":false,"codePointCount":4}

Two decoding paths for the same bytes

We encode each fixture as UTF-8 and save the byte sequence as hexadecimal. The leading-BOM fixture begins EF BB BF. Default TextDecoder returns name with no leading U+FEFF; decoding with ignoreBOM: true retains U+FEFF in the string despite that option name. The cleaner therefore reports zero removals for the default path and one for the preserving path. Both end at name. The saved record separates input bytes, decoded input, cleaner output and status, making the stage of change explicit.

Position changes the decoding result

The internal fixture places EF BB BF between na and me. Both decoding paths retain an internal U+FEFF; the cleaner deletes it in both. The no-BOM reference stays name throughout. We tested Node.js TextDecoder, not the browser file picker, Excel or every encoding library. The homepage file loader calls File.text(), which is a distinct browser API; these results should not be advertised as a measured file-picker experiment. A future browser test should save the same bytes and capture the loaded field before pressing Clean.

Keep a record at each boundary

When investigating an import problem, save the original byte length and leading bytes, name the decoder and options, then inspect the resulting string. If U+FEFF was already consumed, a no-removal message from the cleaner is expected for that character. If it remains inside a header or field, assess whether it is accidental before editing it. The WHATWG Encoding Standard specifies BOM handling and the ignore-BOM flag; our downloaded script reproduces these particular measurements. Character cleanup cannot restore an original encoding marker or promise byte-identical re-export. This page addresses decoding attribution rather than comparing trim or normalization methods.

Reproduce this test

Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.

node reproduce.cjs

Reference and next check

WHATWG Encoding Standard provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.