ZZerowidth CleanerMeasured from saved files

REPRODUCIBLE SOURCE TEST

Invisible Character Source Test

Five saved UTF-8 text files were read as raw bytes, decoded without discarding a byte-order mark, and counted by Unicode code point. Positions below are one-based.

Method and sample path

  1. Read C:\tests\chatgpt.txt, claude.txt, gemini.txt, docs.txt, and notepad.txt as bytes.
  2. Decode each file as strict UTF-8 while retaining a leading U+FEFF.
  3. Count Unicode code points rather than UTF-16 code units. The first character is position 1.
  4. Check U+200B, U+200C, U+200D, U+2060, U+FEFF, U+00AD, U+034F, and U+00A0.
Count and first position for the eight inspected code points. Each cell is count | first position.
Saved sampleBytesCode pointsU+200BU+200CU+200DU+2060U+FEFFU+00ADU+034FU+00A0
chatgpt.txt149014900 | 00 | 00 | 00 | 00 | 00 | 00 | 00 | 0
claude.txt136313630 | 00 | 00 | 00 | 00 | 00 | 00 | 00 | 0
gemini.txt8308300 | 00 | 00 | 00 | 00 | 00 | 00 | 00 | 0
docs.txt (derived from gemini.txt)8358330 | 00 | 00 | 00 | 01 | 10 | 00 | 00 | 0
notepad.txt (derived from gemini.txt)8328320 | 00 | 00 | 00 | 00 | 00 | 00 | 00 | 0

Character-by-character transfer differences

The comparison uses Unicode code-point Levenshtein alignment with insertion, deletion, and substitution costs of 1. Before alignment, only the leading U+FEFF was removed from docs.txt. Both transferred bodies were then aligned against the 830-code-point gemini.txt source with the same method.

Transferred fileDifference typeCode pointCharacterPosition in transferred textCount
docs.txt bodyAddedU+000DCARRIAGE RETURN761
docs.txt bodyAddedU+000ALINE FEED771
notepad.txtAddedU+000DCARRIAGE RETURN761
notepad.txtAddedU+000ALINE FEED771

No code points were removed or changed in either aligned body. The alignment shows that the two line-ending characters appeared along each transfer path, but it cannot identify which individual copy or save step introduced them.

The raw docs.txt file also starts with a byte-order mark added during the transfer and file-writing path, not a character in the body text: U+FEFF, count 1, raw-file position 1, with a byte-count minus code-point-count difference of 2.

Sample boundary: 5 files, 5,350 bytes, 5,348 Unicode code points, and 8 specified invisible-character classes.

Every non-ASCII code point found

FileCode pointCharacterCountFirst positionBytesCode pointsDifference
chatgpt.txt0 non-ASCII code points149014900
claude.txt0 non-ASCII code points136313630
gemini.txt0 non-ASCII code points8308300
docs.txtU+FEFF118358332
notepad.txt0 non-ASCII code points8328320

Sample provenance that can be verified

The test contains three independently requested model passages and two transfers of the gemini.txt passage. The transferred files are docs.txt and notepad.txt; they are not additional source passages.

Run the same measurement

Download the exact Node.js script used for this table, place the five named files in C:\tests, then run node measure-source-samples.js. The script writes source-character-report.json beside the samples.

Download the reproduction script

This sample cannot support the claim that these sources add the listed invisible characters.