The Strange Case of Unicode's 'Ghost Characters'
Long before Unicode existed, Japan built character encoding standards like JIS X 0208 by cataloging kanji from government records and place names. But somewhere in that process, mistakes crept in: mistranscribed characters, printing errors, or characters copied from illegible handwriting ended up encoded as if they were real, meaningful symbols.
These so-called 'ghost characters' (yūrei moji) have no known pronunciation, meaning, or historical usage — they're essentially digital fossils of human error. Yet because encoding standards prioritize backward compatibility, these ghosts were carried forward into JIS standards and eventually into Unicode itself, where they persist as valid, encodable characters today.
The piece traces specific examples, showing how one bad photocopy or misread stroke decades ago can become permanently enshrined in a global text standard used by billions of devices.