The machine that learned to break itself
One evening in August 2008, firmware pushed to a disk array inside a Netflix data center corrupted the company’s production Oracle database. By the time engineers understood what had happened, three days had passed, DVD shipments had stopped, and nine million red-envelope customers were waiting. The embarrassment was manageable. The exposure was not.
Netflix Chief Product Officer Neil Hunt convened a meeting in a conference room the engineers had named “The Towering Inferno”. The decision reached there was not to patch the database or buy a redundant array. It was to leave data-center operations entirely to Amazon Web Services and rebuild the architecture from scratch. “Let’s rethink this completely,” Hunt recalled, “go back to first principles, and think about doing it in the cloud.”
Cloud Architect Adrian Cockcroft turned that directive into a technical program: decompose the monolithic Java application — NCCP internally, a single Oracle-backed API serving every client request — into independent microservices, each owning its own data, each deployable on its own schedule. The migration would be incremental, running the old system and the new in parallel as long as necessary. Netflix’s first move to AWS was, Hunt noted, “nicely symbolic”: a single webpage.
By 2009 the work was under way in earnest, beginning with non-critical workloads: video encoding, Hadoop analytics, batch processing. Core streaming services followed through 2010 and 2012. The Oracle database gave way to Cassandra. Each service that decoupled from the monolith had to define a clean API boundary, and those boundaries, accumulated over years, were the architecture the industry would later learn to call microservices.
In July 2011, engineers Yury Izrailevsky and Ariel Tseitlin published a blog post on the Netflix Tech Blog announcing Chaos Monkey: a tool that randomly killed production instances during business hours, while actual customers were actually watching movies. The reasoning was terse: “The best way to avoid failure is to fail constantly.” An AWS outage three months earlier had already proved the thesis — Netflix had survived it with minimal disruption, because the parts already migrated were already built to lose a node and keep running.
The Christmas Eve 2012 AWS elastic load balancer failure was a harder test. A maintenance process deleted ELB state in US-East; Netflix went dark on the busiest TV-watching night of the year for roughly seven hours. The response was Chaos Kong — a tool that simulates the loss of an entire AWS region — and a push toward multi-region active-active deployment, where multiple regions serve production traffic simultaneously and losing one changes nothing the customer can see.
The billing system, the last holdout, migrated off Netflix’s own data centers in January 2016. The journey had taken seven and a half years. Nine million subscribers had become 89 million. Twenty million API requests per day had become two billion. And the open-source tools built along the way — Eureka for service discovery, Hystrix for circuit-breaking, Zuul for routing — had become the default scaffolding for a generation of distributed-systems teams.
The name “microservices” would arrive in 2014. The practice had been running at production scale for five years by then — and carrying a third of North American internet traffic.
Sources
- A Brief History of Scaling Netflix — ByteByteGo — timeline of Netflix’s scaling decisions from founding through cloud migration, subscriber and API request growth figures.
- Case Studies in Cloud Migration: Netflix, Pinterest, and Symantec — Increment — Neil Hunt, the Towering Inferno meeting, Ruslan Meshenberg, and the incremental migration timeline.
- The Netflix Simian Army — Netflix TechBlog — original July 2011 announcement of Chaos Monkey by Yury Izrailevsky and Ariel Tseitlin.
- The Origin of Chaos Monkey — Gremlin — history and philosophy of chaos engineering at Netflix, including the Christmas Eve 2012 outage and Chaos Kong.