Things Have History
Sequence to sequence: how Sutskever put a sentence in a bottle

ai

Sequence to sequence: how Sutskever put a sentence in a bottle

Listen · 4:25

Ilya Sutskever had a strange idea in the summer of 2014: take an English sentence, feed it word by word into a neural network until the whole thing collapsed into a single vector of a thousand numbers, then hand that vector to a second network and ask it to write the sentence again — in French. Geoffrey Hinton, watching nearby at Google, called it a “thought vector.” If it worked, compressed meaning would cross the Atlantic without a phrase table.

It worked. The paper — arXiv:1409.3215, by Sutskever, Oriol Vinyals, and Quoc V. Le, all then at Google Brain — was submitted September 10, 2014, and presented at NeurIPS in Montreal that December. The architecture was deceptively plain: an encoder LSTM read the source sentence and compressed it into a fixed-length context vector; a decoder LSTM took that vector as its starting state and generated the target sentence one token at a time. Two independent networks, no linguistic features, no hand-crafted rules. Just gradients and time.

What they replaced was considerable. State-of-the-art machine translation in 2014 ran on phrase-based statistical pipelines — alignment models descended from IBM’s 1990s work, language models, phrase tables, and reordering algorithms tuned across two decades. These systems shuffled fragments according to probability; they could not reason about meaning. The seq2seq model replaced the entire pipeline with a single end-to-end network: four stacked LSTM layers, a thousand hidden units per layer, roughly 384 million parameters, trained on twelve million English-French sentence pairs across eight GPUs for about ten days.

On the WMT’14 English-to-French benchmark, an ensemble of five such networks scored a BLEU of 34.81. The phrase-based SMT baseline scored 33.30. A pure neural system had, for the first time, beaten a carefully tuned statistical machine at large-scale translation.

The strangest part of the paper is a confession in Section 3.3. To get those numbers, the authors had fed the source sentence to the encoder backwards — last word first. Performance jumped substantially. Their explanation was tentative: reversing the input placed the early source words close in time to the early target words, shortening the gradient’s path during backpropagation. But the authors admitted they had “no complete explanation” for why the effect was as large as it was. One of the most-cited papers in deep learning contains, at its center, an acknowledged mystery. The approach was retained because the experiments were unambiguous, even when the theory was not.

The architecture’s limitation was also clear from the start. Compressing an entire sentence into a single fixed-length vector is fine for twenty words; it becomes a bottleneck at forty. By the following year, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio had the remedy: instead of one context vector, the decoder could attend to all encoder hidden states at each step, learning soft alignment between source and target words. Attention arrived as a direct fix for the flaw seq2seq had invented.

In September 2016, Google replaced its phrase-based translation system with a neural model directly descended from this work, reporting a 55–85 percent reduction in translation errors on several language pairs. A prototype trained across eight GPUs in ten days had, within two years, ended a twenty-year engineering paradigm.

The thought vector did cross the Atlantic, after all. It just needed help knowing where to look on the other side.

Sources

Spot a mistake?

Wrong date, broken citation, a fact that doesn't hold? Tell us. It lands in an inbox a human reads and the post can be pulled or corrected.