News

The Ghost Characters Haunting Unicode

In 1978, a series of cataloging mistakes created characters that never existed. They made it into JIS X 0208 and then Unicode, and now they're part of every computer on the planet.

August 16, 2026· 2 min read
The Ghost Characters Haunting Unicode

In 1978, Japan's Ministry of Economy, Trade and Industry established the encoding standard that would become JIS X 0208, a foundational reference for all Japanese character encodings. But shortly after its release, researchers noticed something odd: several characters in the standard had no clear origin. Nobody could say what they meant, how to pronounce them, or where they came from. These became known as yūrei moji — ghost characters.

For nearly two decades, the ghost characters were a forgotten curiosity. Then in 1997, an investigation was launched to trace their origins. Every character in the JIS standard was supposed to have a documented source, but the records were often vague — typically just a reference to a document, with no page number. One of the most common sources was the Overview of National Administrative Districts, a seven-volume set with roughly 900 pages per volume. Tracking down a single character without a page reference was a nightmare.

The investigation did uncover the truth for most of the ghosts — and it wasn't what anyone expected. Some characters were simply invented by mistake during the cataloging process. For example, the character 妛 was an error introduced while trying to record a character that was "山 over 女." Because the typesetting couldn't combine the two components into a single glyph, they were printed separately, cut out, and pasted onto a sheet of paper. When the copy was read, the seam between the two pieces looked like an extra stroke — and that stroke was added to the character. The correct character, 𡚴, wasn't added to JIS or Unicode until much later.

Of the core ghost characters — 妛挧暃椦槞蟐袮閠駲墸壥彁 — only one remains truly unexplained: 彁. The most plausible theory is that it was a misreading of 彊, but no specific incident has ever been confirmed.

These characters made their way into Unicode through CJK unification, and now they're part of every computer on the planet — lurking in the dark corners of character tables, waiting to be used by someone who doesn't know they're ghosts.

The story is a reminder that standards are made by humans, and human error can become immortalized in infrastructure. The ghosts of 1978 are now permanent residents of the digital world.

In 1978 a series of small mistakes created some characters out of nothing. The errors went undiscovered just long enough to be set in stone, and now these ghosts are, at least in potential, a part of every computer on the planet.
Manul X Editorial