CLAUDE WATERMARK REMOVER · PRACTICAL TEST
Can cleaning recover text already replaced by U+FFFD?
No. Our U+FFFD fixtures stay unchanged except for a separate U+200B. A replacement character does not contain the original missing byte sequence.
Open the Claude text cleaner · Full measured data
Measured inputs and outputs
These are locally constructed test strings, not evidence that Claude inserts these characters. We executed the saved homepage script snapshot with a minimal DOM harness and compared exact strings. The invisible-checkbox state and dash mode appear in each row. The observations below test this page's specific question. File API and event-lifecycle tests include an additional stage described in the analysis. We make no detector-score or statistical-watermark removal claim.
| Fixture and mode | Input string and code points | Output and cleaner status | Observed before / after |
|---|---|---|---|
| Invalid FF decoded direct string / keep / invisible true | "�"U+FFFD | "�"No selected invisible characters found | {"replacementCount":1,"utf8Hex":"efbfbd"}{"replacementCount":1,"utf8Hex":"efbfbd"} |
| Invalid FE decoded direct string / keep / invisible true | "�"U+FFFD | "�"No selected invisible characters found | {"replacementCount":1,"utf8Hex":"efbfbd"}{"replacementCount":1,"utf8Hex":"efbfbd"} |
| Replacement plus marker direct string / keep / invisible true | "A�B"U+0041 U+FFFD U+200B U+0042 | "A�B"Removed 1 invisible character | {"replacementCount":1,"utf8Hex":"41efbfbde2808b42"}{"replacementCount":1,"utf8Hex":"41efbfbd42"} |
| Intentional literal replacement direct string / keep / invisible true | "A�B"U+0041 U+FFFD U+0042 | "A�B"No selected invisible characters found | {"replacementCount":1,"utf8Hex":"41efbfbd42"}{"replacementCount":1,"utf8Hex":"41efbfbd42"} |
Different original bytes can produce one string
We decode a one-byte buffer containing FF and another containing FE with the default UTF-8 TextDecoder error policy. Both produce the same U+FFFD string. The cleaner retains that string in both cases. A mixed input loses its separate U+200B but keeps U+FFFD; the intentional-literal reference also remains. The observer counts replacement characters and records the output UTF-8 encoding. Those output bytes encode the replacement character itself; they are not a reconstruction of either original invalid byte. This demonstrates an information boundary rather than the wrong-UTF-16-decoder scenario measured on another page.
A retained symbol cannot identify its history
Once two distinct source byte sequences have become the same string, a cleaner receiving only that string cannot distinguish their origins. U+FFFD can also have been entered deliberately, as our final fixture illustrates. The mere presence of the symbol therefore does not prove a particular encoding fault. We save the explicit byte preparation in the downloadable reproducer so the two decoding cases can be checked independently. No real file was corrupted for this test, and we do not infer that Claude, an editor or a server caused a replacement in a user document.
Return to original bytes when they exist
Preserve the original file or response buffer before trying another decoder. Determine its declared and actual encoding through the originating system, then decode a working copy with the intended rules. If only the already-replaced string remains, state that recovery is unavailable from that representation alone. Replacing every U+FFFD with a guessed letter creates new content and cannot be verified as restoration. Rerun the destination check after any evidence-based repair. The homepage can remove its selected formatting characters but cannot reverse a lossy decoding decision or infer missing source bytes from a glyph.
Reproduce this test
Save reproduce.cjs and tested-app.js in the same folder. Run the command below with Node.js. The harness prints its runtime, script SHA-256 and every measured row. Compare those rows with the original record. Using a newer script or runtime creates a new experiment; retain the version information with your rerun.
node reproduce.cjsReference and next check
WHATWG Encoding Standard provides the relevant primary definition. The table and fixture analysis are original measurements. For broader inspection, use our Unicode inspector. Read the scope distinction before interpreting cleanup as a watermark result.