Zerowidth CleanerAll measured guides

CLAUDE WATERMARK REMOVER · PRACTICAL TEST

Why does a cleaned capital I still fail a Turkish i comparison?

Turkish lowercasing maps our plain I to dotless ı, while English lowercasing maps it to i. Deleting U+200B leaves that language rule intact. Dotted capital İ maps to i in the Turkish test.

Tested October 10, 2026 · v24.19.0 · 4 measured runs

Open the Claude text cleaner · Full measured data

Code-point comparison for Marked capital I

Measured inputs and outputs

These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. Destination observations run independently on the exact input and cleaned strings in the recorded Node runtime. We make no detector-score or statistical-watermark removal claim.

Before and after observations; full code points and strings are in the JSON download.
Fixture and modeInput string and code pointsOutput and cleaner statusObserved before / after
Marked capital I
direct string / keep / invisible true
"I​"
U+0049 U+200B
"I"
Removed 1 invisible character
{"defaultLower":"i​","englishLower":"i​","turkishLower":"ı​","turkishMatchesI":false}
{"defaultLower":"i","englishLower":"i","turkishLower":"ı","turkishMatchesI":false}
Plain capital I
direct string / keep / invisible true
"I"
U+0049
"I"
No selected invisible characters found
{"defaultLower":"i","englishLower":"i","turkishLower":"ı","turkishMatchesI":false}
{"defaultLower":"i","englishLower":"i","turkishLower":"ı","turkishMatchesI":false}
Dotted capital I
direct string / keep / invisible true
"İ"
U+0130
"İ"
No selected invisible characters found
{"defaultLower":"i̇","englishLower":"i̇","turkishLower":"i","turkishMatchesI":true}
{"defaultLower":"i̇","englishLower":"i̇","turkishLower":"i","turkishMatchesI":true}
Ordinary capital A
direct string / keep / invisible true
"A"
U+0041
"A"
No selected invisible characters found
{"defaultLower":"a","englishLower":"a","turkishLower":"a","turkishMatchesI":false}
{"defaultLower":"a","englishLower":"a","turkishLower":"a","turkishMatchesI":false}

The remaining difference is visible linguistic data

Our marked capital-I fixture loses its U+200B during cleanup. English and default lowercase then return i, but the explicit Turkish operation returns dotless ı, so comparison with ordinary i remains false. The plain-I control already behaves that way. Dotted capital İ produces i plus a combining dot with the default and English operations, while Turkish produces a single i and the comparison succeeds. Ordinary A becomes a in every operation. The full outputs are stored as strings in the data download, and the input/output code-point columns describe the cleanup stage rather than incorrectly labeling a case conversion as another deletion.

Case mapping differs from collation and normalization

We invoke toLowerCase and two explicitly tagged toLocaleLowerCase calls. The observer does not use Intl.Collator, normalize the resulting strings, fold all case distinctions or test a database search. Our previous collation page asks whether comparison can consider distinct strings equivalent; here we measure the strings produced by a language-dependent transformation and a strict equality against i. ECMA-402 defines the locale-sensitive operation, with the saved runtime delimiting the actual outputs. Neither the cleaner nor the test infers the intended language from the input. Dotted and dotless letters are deliberate reference characters, not accidental markers selected for removal.

Keep language intent separate from marker inspection

When investigating a case-insensitive lookup, retain the original spelling and record the locale or folding policy used by the destination. Compare the transformed strings and their code points after any justified cleanup instead of assuming that every capital I should become the same lowercase letter. Avoid deleting combining marks or substituting letters merely to force equality without confirming the intended text. These four fixtures explain a specific remaining mismatch and a contrasting dotted-letter result. They do not choose a universal username policy, implement Unicode case folding or promise that a cleanup operation preserves every locale-dependent identity rule in a larger application.

Reproduce this test

Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.

node reproduce.cjs

Reference and next check

ECMA-402 locale-sensitive case conversion provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.