Unicode 1.0: one table for every script on earth
A piece of Japanese text that left a workstation in Tokyo in 1988 could arrive at a desktop in Frankfurt as an unbroken stream of question marks, boxes, and characters that belonged to no language on earth. The Japanese had a word for this: mojibake, roughly “character transformation disease,” and it was the predictable outcome of thirty years of computing in which every nation, company, and standards body had invented its own scheme for turning characters into numbers — without asking anyone else.
In 1987, three engineers decided the proliferation had gone on long enough. Joe Becker at Xerox’s facility in El Segundo, California had been working the problem for years; he pulled in Lee Collins from the same building and Mark Davis from Apple in Cupertino. On August 29, 1988, Becker circulated a fifteen-page draft he called “Unicode 88” — the name, he wrote, was short for unique, universal, and uniform (Unicode Consortium). Its premise was simple enough to fit in a sentence: every character needed by modern computing would get one permanent number, and every piece of software would agree on what that number meant.
The Unicode Consortium was incorporated in California on January 3, 1991, with Apple, IBM, Microsoft, Sun Microsystems, and Novell among its founders. Nine months later, in October 1991, the first volume of the Unicode Standard 1.0 was published from Mountain View (Unicode.org). It encoded 7,161 characters across 24 scripts — Latin, Cyrillic, Greek, Hebrew, Arabic, Devanagari, Hiragana, Katakana, and others — in a fixed 16-bit format that reserved room for 65,536 distinct code points in total. Becker had argued this was sufficient for every language in active use. He was not wrong, exactly; but the ceiling turned out to be lower than anyone liked.
The decision that still generates arguments is Han Unification. Rather than give Chinese, Japanese, and Korean characters their own separate code points — which would have consumed most of those 65,536 slots — the designers merged 20,902 historically related ideographs into a single block, from U+4E00 to U+9FAF, relying on fonts to handle regional rendering differences (Wikipedia — CJK Unified Ideographs). Japan’s standards body found this intolerable. The same code point displayed through a Japanese font and a Chinese font can look perceptibly different, and the objection was not pedantic: typography is not merely decoration. The consortium held. The block at U+4E00 remains both Unicode’s largest and its most argued-over real estate.
What happened next moved fast. In 1992, Ken Thompson and Rob Pike sketched out UTF-8 during a single dinner — on a restaurant placemat in New Jersey, as later accounts have it — producing a variable-width encoding that kept ASCII’s original single-byte values intact while reaching the full Unicode range. By the mid-1990s, UTF-8 was the default encoding of the web. The mojibake that had seemed like an intractable feature of international computing was, for most purposes, gone.
Unicode 1.0 encoded 7,161 characters. Unicode 16.0, published in 2024, encodes 154,998 across 168 scripts — among them Egyptian hieroglyphs, cuneiform, and Linear B, writing systems that predate every other entry in this series by at least two millennia. They all share the same table now, addressed by the same scheme, indifferent to which language first put marks on clay.
Sources
- Unicode Version 1.0 — Unicode Consortium — founding timeline, January 1991 incorporation, October 1991 publication, scope.
- First Press Release, February 1991 — Unicode.org — founding members, 16-bit design rationale, industry backing.
- CJK Unified Ideographs — Wikipedia — Han Unification block (U+4E00–U+9FAF), 20,902 ideographs, controversy.
- The History of Unicode — SymbolFYI — “Unicode 88” draft, mojibake context, UTF-8 placemat origin story.