Zerowidth CleanerAll measured guides

CLAUDE WATERMARK REMOVER · PRACTICAL TEST

Can removing U+034F affect subsequent Unicode normalization?

Yes. In our combining-mark fixture, U+034F prevents canonical reordering across its position. The cleaner removes it; a later NFD pass then yields a different mark order.

Tested October 5, 2026 · v24.19.0 · 3 measured runs

Open the Claude text cleaner · Full measured data

Code-point comparison for CGJ between marks

Measured inputs and outputs

These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. Invisible removal is enabled; the dash mode appears in each row. The observations below test this page's specific question. We make no detector-score or statistical-watermark removal claim.

Before and after observations; full code points and strings are in the JSON download.
Fixture and modeInput string and code pointsOutput and cleaner statusObserved before / after
CGJ between marks
direct string / keep
"á͏̣"
U+0061 U+0301 U+034F U+0323
"ạ́"
Removed 1 invisible character
{"nfd":["U+0061","U+0301","U+034F","U+0323"],"nfc":["U+00E1","U+034F","U+0323"]}
{"nfd":["U+0061","U+0323","U+0301"],"nfc":["U+1EA1","U+0301"]}
Same marks without CGJ
direct string / keep
"ạ́"
U+0061 U+0301 U+0323
"ạ́"
No selected invisible characters found
{"nfd":["U+0061","U+0323","U+0301"],"nfc":["U+1EA1","U+0301"]}
{"nfd":["U+0061","U+0323","U+0301"],"nfc":["U+1EA1","U+0301"]}
Baseline base and mark
direct string / keep
"á"
U+0061 U+0301
"á"
No selected invisible characters found
{"nfd":["U+0061","U+0301"],"nfc":["U+00E1"]}
{"nfd":["U+0061","U+0301"],"nfc":["U+00E1"]}

The boundary changes before normalization

The first fixture places acute accent, CGJ and dot below after a. The cleaner removes only U+034F; it does not call normalize. We then separately apply NFD and NFC to the before and after strings. The output record shows the exact mark ordering and composed code points at both stages. The second fixture is a reference with the same marks and no CGJ, allowing direct comparison. The third contains only the base and acute mark, isolating ordinary composition behavior.

CGJ has a Unicode purpose

Unicode normalization describes canonical ordering and how CGJ can block reordering. Deleting it may matter for a text whose mark ordering is intentional, even if the character has no visible glyph. The measured observation concerns code-point order in these synthetic strings, not font rendering, linguistic correctness or Claude provenance. We did not test a full language corpus. The homepage currently includes CGJ in its deletion map, which creates a capability boundary worth checking before processing combining-mark-heavy material.

Keep the normalization stage explicit

Retain the original sequence and identify why CGJ is present before removing it. Record whether the destination later applies NFC or NFD, then compare that destination result with the intended order. If CGJ is intentional, turn off invisible-character deletion or edit only the accidental characters in a working copy. A visible-text review alone can miss a later normalization change. Our fixture downloads let readers reproduce each stage and confirm which operation caused the difference.

Reproduce this test

Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.

node reproduce.cjs

Reference and next check

Unicode normalization forms provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.