Fourteen million photographs and a bedroom in Toronto
The score came in at Florence in October 2012, when the results of the annual ImageNet Large Scale Visual Recognition Challenge were read aloud at a computer vision conference. The second-place submission had achieved a top-5 error rate of 26.2 percent — roughly the expected ceiling, given years of grinding incremental progress. Then the first-place result: 15.3 percent. A gap of nearly eleven points. The audience sat very still for a moment.
The dataset those teams were competing on had been assembled by Fei-Fei Li, a computer scientist then at Stanford who had come to the United States at fifteen from Chengdu, China. In 2006, while a faculty member at the University of Illinois Urbana-Champaign, she had reached a conclusion most of her colleagues considered mildly eccentric: the bottleneck in machine vision was not the algorithms. It was the data. Every benchmark the field used was tiny, curated, and unrepresentative of the visual chaos of the real world. Fix the data problem, she argued, and the rest would follow.
The idea she settled on was deranged in its scale. She wanted an image dataset covering every noun in the English language — every category that WordNet, the Princeton lexical database, had ever classified. That meant not thousands of images but millions. She partnered with WordNet researcher Christiane Fellbaum, recruited human labelers worldwide through Amazon Mechanical Turk, and watched the numbers climb. In July 2008, ImageNet had zero photographs. By December it held three million, spread across six thousand categories. By April 2010 — when the first competition launched — it had passed eleven million images organized into fifteen thousand synsets.
For two years the ILSVRC competition yielded modest improvements, won by teams combining hand-crafted visual features with support vector machines — the methods that had dominated computer vision for a decade. Then Alex Krizhevsky, a graduate student working under Geoffrey Hinton at the University of Toronto, submitted a convolutional neural network trained on two NVIDIA GTX 580 GPUs installed in his bedroom at his parents’ house. The machine ran for weeks. The paper that Krizhevsky, Ilya Sutskever, and Hinton produced described a network eight layers deep with sixty-one million parameters, trained end-to-end using GPU parallelism — a technique so unfashionable in academic computer vision that it had barely been attempted since the 1980s.
The result at Florence was not a normal scientific finding. Before October 2012, almost no leading computer vision papers employed neural networks. After that presentation, almost all of them would. Hinton later summarized the collaboration with characteristic economy: “Ilya thought we should do it, Alex made it work, and I got the Nobel Prize.” He was joking about the Nobel Prize, but only barely — he received one in 2024.
ImageNet mattered because it changed what you could measure. Without a large, standardized benchmark, deep learning’s promise had been unverifiable — a compelling idea that nobody could prove at scale. With one, Krizhevsky demonstrated in a single competition that scale beat craft: raw compute plus raw data could outrun decades of hand-engineered feature detectors. The AlexNet paper has since been cited more than 172,000 times.
Every large language model you can speak to today traces a direct line back to that bedroom in Toronto and that tray of labeled photographs.
Sources
- ImageNet — Wikipedia — dataset scale, ILSVRC history, creation timeline, growth figures.
- Fei-Fei Li — Wikipedia — biographical details, UIUC and Stanford posts, motivations behind ImageNet.
- CHM Releases AlexNet Source Code — Computer History Museum — AlexNet technical setup, Hinton quote, citation count, field impact before and after.
- How AlexNet Transformed AI and Computer Vision — IEEE Spectrum — Krizhevsky’s GPU bedroom setup, the Florence 2012 presentation, shift in the field.
- ImageNet: A Pioneering Vision for Computers — History of Data Science — WordNet partnership with Christiane Fellbaum, Mechanical Turk labeling, growth from zero to three million images by December 2008.