GKRootWire
AI Google Adds 'Preferred Source' Button to Help Publishers Fight AI Traffic LossesGadgets Linkdaze Launches a Smart Calendar Aimed at Running Your Whole HouseholdSecurity Popular Rust Crate arrayref Hijacked to Spread Infostealer MalwareCloud & Sysadmin GitHub Details Cause of August 17 Outage, Outlines Reliability FixesDev Tools Show HN: 'Huzzah' Proposes a Fresh Take on AI-Assisted CodingCloud & Sysadmin The Weird Science of Cooling Data Centers With UrineAI Google Adds 'Preferred Source' Button to Help Publishers Fight AI Traffic LossesGadgets Linkdaze Launches a Smart Calendar Aimed at Running Your Whole HouseholdSecurity Popular Rust Crate arrayref Hijacked to Spread Infostealer MalwareCloud & Sysadmin GitHub Details Cause of August 17 Outage, Outlines Reliability FixesDev Tools Show HN: 'Huzzah' Proposes a Fresh Take on AI-Assisted CodingCloud & Sysadmin The Weird Science of Cooling Data Centers With Urine
Dev Tools

The Strange Case of Unicode's 'Ghost Characters'

A deep dive into how clerical errors and typos decades ago created phantom kanji that still haunt text encoding standards today.

Long before Unicode existed, Japan built character encoding standards like JIS X 0208 by cataloging kanji from government records and place names. But somewhere in that process, mistakes crept in: mistranscribed characters, printing errors, or characters copied from illegible handwriting ended up encoded as if they were real, meaningful symbols.

These so-called 'ghost characters' (yūrei moji) have no known pronunciation, meaning, or historical usage — they're essentially digital fossils of human error. Yet because encoding standards prioritize backward compatibility, these ghosts were carried forward into JIS standards and eventually into Unicode itself, where they persist as valid, encodable characters today.

The piece traces specific examples, showing how one bad photocopy or misread stroke decades ago can become permanently enshrined in a global text standard used by billions of devices.

Why it matters: This is a great reminder that software standards inherit human mistakes and then ossify them forever in the name of compatibility. Anyone building text-processing, search, or font-rendering systems should know that not every character in Unicode represents something real — some are just historical accidents nobody can ever safely delete.

Sources: Hacker News