CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Why can four counted characters exceed a four-byte field?
Our marked ABC fixture has four code points but six UTF-8 bytes. Cleaning reduces it to three bytes. A retained emoji still uses four bytes despite counting as one code point.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Destination observations run independently on the exact input and cleaned strings in the recorded Node runtime. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Marked ASCII direct string / keep / invisible true | "ABC"U+0041 U+0042 U+200B U+0043 | "ABC"Removed 1 invisible character | {"codePoints":4,"utf16Units":4,"utf8Bytes":6,"fitsFourByteBudget":false}{"codePoints":3,"utf16Units":3,"utf8Bytes":3,"fitsFourByteBudget":true} |
| Plain ASCII direct string / keep / invisible true | "ABC"U+0041 U+0042 U+0043 | "ABC"No selected invisible characters found | {"codePoints":3,"utf16Units":3,"utf8Bytes":3,"fitsFourByteBudget":true}{"codePoints":3,"utf16Units":3,"utf8Bytes":3,"fitsFourByteBudget":true} |
| Accented letter direct string / keep / invisible true | "é"U+00E9 | "é"No selected invisible characters found | {"codePoints":1,"utf16Units":1,"utf8Bytes":2,"fitsFourByteBudget":true}{"codePoints":1,"utf16Units":1,"utf8Bytes":2,"fitsFourByteBudget":true} |
| Supplementary symbol direct string / keep / invisible true | "😀"U+1F600 | "😀"No selected invisible characters found | {"codePoints":1,"utf16Units":2,"utf8Bytes":4,"fitsFourByteBudget":true}{"codePoints":1,"utf16Units":2,"utf8Bytes":4,"fitsFourByteBudget":true} |
A small budget exposes the counting mismatch
We set an invented four-byte budget and measure the exact UTF-8 encoding length at both stages. The authored AB-marker-C string initially has four code points and four UTF-16 units, yet its UTF-8 representation needs six bytes. Deleting U+200B leaves three ordinary ASCII letters, requiring three bytes, and changes the budget decision to true. Plain ABC is already within the limit. The accented letter has one code point and two bytes; the supplementary smile has one code point, two UTF-16 units and four bytes. Both reference characters remain unchanged by the cleaner. These controls show why visible length cannot stand in for an encoded size.
The budget belongs to this experiment
Buffer.byteLength receives an explicit utf8 encoding. The four-byte threshold is a fixture parameter, not a reported limit for Claude, a database, Cloudflare or any other service. The earlier character-counter article explains display units; this page adds a concrete downstream acceptance decision using generated byte length. We do not count an entire JSON request, HTTP headers, compression or a file BOM. Those layers can add bytes beyond the value measured here. The saved callback removes selected characters and reports code points; it never checks this byte budget. All observations are local and require no network request or destination account.
Measure the exact representation that is limited
Keep the original value and identify whether the destination limits characters, UTF-16 units or encoded bytes. If its contract specifies UTF-8 bytes for the field alone, measure that field with the same encoding after any justified edit. If the limit applies to a whole payload, serialize and measure that payload instead. Do not remove a meaningful accented letter or emoji merely to force a smaller count. These four inputs establish exact byte sizes and one changed acceptance result; they do not prove that an unnamed importer will accept the text. Retain the encoding label and threshold with a reproduction so its pass decision remains interpretable.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
Node.js Buffer.byteLength provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.