Hidden Markov models, 1989: probability beats the rule books
Fred Jelinek had a line he delivered at conferences that his audience did not always take as a compliment. “Every time I fire a linguist,” he reportedly said at a workshop in the late 1980s, “the performance of our speech recognition system goes up.” The linguists in the room were not, as far as the record shows, delighted.
Jelinek had been head of IBM’s Continuous Speech Recognition Group at the Thomas J. Watson Research Center in Yorktown Heights, New York, since 1972. The lab had a settled conviction that the rule-builders — the researchers writing phoneme dictionaries and grammatical constraints to teach machines the structure of English — were solving the wrong problem. Speech recognition, in Jelinek’s view, was not a linguistics problem. It was a statistical decoding problem: given an acoustic signal, find the word sequence most likely to have produced it.
The mathematical machinery for this was not new. Leonard Baum and colleagues at the Institute for Defense Analyses in Princeton had published the foundational theory in a sequence of papers between 1966 and 1972 — a training method, now called the Baum-Welch algorithm, for fitting probabilistic models to sequential data. James Baker brought this math to speech at Carnegie Mellon in the mid-1970s. But it was Jelinek’s IBM group that drove it through paper after paper, culminating in the 1983 landmark by Lalit Bahl, Jelinek, and Robert Mercer, “A Maximum Likelihood Approach to Continuous Speech Recognition”, published in IEEE Transactions on Pattern Analysis and Machine Intelligence. The framework they laid out would endure for thirty years.
A Hidden Markov Model treats speech as a sequence of hidden states — the underlying phonemes being produced — that generate observable acoustic measurements. You cannot hear the phonemes directly; you observe the waveform and infer what must have caused it. The Baum-Welch algorithm tunes the model’s parameters against a corpus of recorded speech until predicted features match actual recordings as closely as possible. No phoneme dictionary required. No hand-coded grammar rules. Just data, probability, and iteration.
The man who handed this to the rest of the world was Lawrence Rabiner at AT&T Bell Labs. In February 1989, Rabiner published “A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition” in the Proceedings of the IEEE — thirty pages, walking through every piece of the method: the forward-backward algorithm, the Viterbi decoder, the Baum-Welch training procedure. The paper was not reporting new theory. It was explaining existing theory in a form any researcher could implement over a weekend. It has since been cited more than fifteen thousand times.
Within a few years, every major speech recognition lab had switched architectures. Dragon Systems shipped a commercial dictation product in 1990 built on HMM infrastructure. The voice interfaces that preceded Siri ran on the same statistical skeleton. The linguists had not been fired from the field, exactly — their knowledge had been reframed as training data, as text corpora, as structure the models would recover on their own.
The larger argument this settled was not specific to speech. Jelinek’s team had demonstrated something stubborn and hard to argue with: given enough observed data and a model flexible enough to be wrong in useful ways, machines could learn structure they were never explicitly told was there. That argument would be had again, and again, and won the same way each time.
The linguists came back. They arrived labeled as training data.
Sources
- Frederick Jelinek — Wikipedia — career timeline at IBM Yorktown Heights, statistical approach to speech recognition, key collaborators Bahl and Mercer
- Frederick Jelinek 1932–2010 — National Academies Memorial Tributes — biographical tribute; Jelinek’s information-theoretic framing of speech recognition and IBM lab history
- Bahl, Jelinek, Mercer (1983) — IEEE TPAMI — the maximum-likelihood framework paper that established the IBM statistical approach to continuous speech
- Rabiner (1989) — Proceedings of the IEEE — the tutorial itself; source for the citation count and publication details