A Czech engineer and 100 billion words
King minus man plus woman equals queen. Not a riddle — a result. A computer in 2013 could solve that equation without being given any definition of “king,” any grammar rules, any understanding of gender or monarchy. The answer came from a geometry of words that a Czech engineer had trained on 100 billion words of raw text, and it was not supposed to work this well.
The paper was “Efficient Estimation of Word Representations in Vector Space,” posted to arXiv in January 2013 by Tomas Mikolov and three colleagues at Google Brain: Kai Chen, Greg Corrado, and Jeff Dean. Its stated goal was modest — train word embeddings faster than anyone had managed before. Not a new theory, not a breakthrough architecture. Just speed. The reviewers who first saw it accepted it only for a workshop session, not the main program of the ICLR 2013 conference. The paper went on to accumulate more than 37,000 citations.
Mikolov had finished his PhD at Brno University of Technology in the Czech Republic in 2012. His doctoral research — neural network language models — drew visits to Johns Hopkins and to Yoshua Bengio’s lab in Montreal. He joined Google Brain that year and set about a question that had dogged NLP: how do you teach a machine what a word means in a way that transfers across tasks (Wikipedia)?
Word2Vec’s key move was radical simplicity. Previous neural approaches to word embeddings used deep networks with multiple hidden layers. Mikolov stripped them out, keeping only a linear projection between input and output. The model had one task: predict which words appear near a given word. That constraint, applied at scale — up to 100 billion words of text, processed at several billion words per hour — produced a 300-number vector for every word in a vocabulary. Words used in similar contexts ended up in similar positions in that 300-dimensional space (arXiv:1301.3781).
The paper described two architectures. CBOW, Continuous Bag of Words, predicted a target word from its surrounding context. Skip-gram ran the problem in reverse: given a word, predict its neighbors. Both learned the same thing — the geometry of co-occurrence, encoded as coordinates.
What no one fully anticipated was how structured that geometry would turn out to be. Subtract the vector for “man” from the vector for “brother,” add the vector for “woman,” and the nearest point in the whole vocabulary is “sister.” The gender relationship had encoded itself as a consistent geometric direction — not programmed in, not annotated, extracted entirely from the patterns of ordinary text. King minus man plus woman lands near queen. Paris minus France plus Italy lands near Rome.
This was the discovery inside the discovery. It meant that the embedding space had semantic structure — that the coordinates reflected something real about the world. And it meant that the vectors were portable. Train Word2Vec once on a massive corpus and carry the learned representations to any downstream task: translation, sentiment analysis, named-entity recognition. Pretrain on large data, fine-tune on small data. That idea now underlies every large language model in existence; in 2013, Word2Vec was its proof of concept.
By 2014, Mikolov had moved to Facebook AI Research. By 2018, BERT had taken the same principle — reusable pretrained representations of language — and extended it through the full depth of a transformer. The embedding space grew wider and richer. But the king-minus-man equation that first showed it was possible came from one linear layer trained on text, on a cluster of CPUs, by a Czech engineer who had finished his PhD the year before.
Sources
- arXiv:1301.3781 — Mikolov et al., “Efficient Estimation of Word Representations in Vector Space” — the original Word2Vec paper; corpus size, training speed, CBOW and Skip-gram architectures, vector arithmetic demonstrations.
- Tomáš Mikolov — Wikipedia — biographical details, PhD at Brno University of Technology, research visits to Johns Hopkins and Bengio’s lab, career at Google Brain and Facebook AI Research.
- Ten Years of Word2Vec — David Strohmaier — field-historical perspective on Word2Vec’s impact, citation count, and what it unlocked for subsequent NLP research.