REPRODUCIBLE SOURCE TEST
Invisible Character Source Test
Five saved UTF-8 text files were read as raw bytes, decoded without discarding a byte-order mark, and counted by Unicode code point. Positions below are one-based.
Method and sample path
- Read
C:\tests\chatgpt.txt,claude.txt,gemini.txt,docs.txt, andnotepad.txtas bytes. - Decode each file as strict UTF-8 while retaining a leading
U+FEFF. - Count Unicode code points rather than UTF-16 code units. The first character is position 1.
- Check
U+200B,U+200C,U+200D,U+2060,U+FEFF,U+00AD,U+034F, andU+00A0.
| Saved sample | Bytes | Code points | U+200B | U+200C | U+200D | U+2060 | U+FEFF | U+00AD | U+034F | U+00A0 |
|---|---|---|---|---|---|---|---|---|---|---|
| chatgpt.txt | 1490 | 1490 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| claude.txt | 1363 | 1363 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| gemini.txt | 830 | 830 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| docs.txt (derived from gemini.txt) | 835 | 833 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| notepad.txt (derived from gemini.txt) | 832 | 832 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
Character-by-character transfer differences
The comparison uses Unicode code-point Levenshtein alignment with insertion, deletion, and substitution costs of 1. Before alignment, only the leading U+FEFF was removed from docs.txt. Both transferred bodies were then aligned against the 830-code-point gemini.txt source with the same method.
| Transferred file | Difference type | Code point | Character | Position in transferred text | Count |
|---|---|---|---|---|---|
| docs.txt body | Added | U+000D | CARRIAGE RETURN | 76 | 1 |
| docs.txt body | Added | U+000A | LINE FEED | 77 | 1 |
| notepad.txt | Added | U+000D | CARRIAGE RETURN | 76 | 1 |
| notepad.txt | Added | U+000A | LINE FEED | 77 | 1 |
No code points were removed or changed in either aligned body. The alignment shows that the two line-ending characters appeared along each transfer path, but it cannot identify which individual copy or save step introduced them.
The raw docs.txt file also starts with a byte-order mark added during the transfer and file-writing path, not a character in the body text: U+FEFF, count 1, raw-file position 1, with a byte-count minus code-point-count difference of 2.
Sample boundary: 5 files, 5,350 bytes, 5,348 Unicode code points, and 8 specified invisible-character classes.