We spend a lot of time thinking about server infrastructure. It’s literally what we do. And one of the things that keeps hosting people up at night isn’t the day-to-day stuff, it’s the catastrophic, completely avoidable failures that happen when smart teams with real budgets make assumptions about scale that turn out to be spectacularly wrong. The three stories below are our favourites. Not because we enjoy watching things burn, but because every single one of them contains a lesson that is still being ignored by teams somewhere right now.
Healthcare.gov: $2 billion spent, six people enrolled on day one
Let’s set the scene. It’s October 1, 2013. The United States government has just launched Healthcare.gov, the centrepiece of the Affordable Care Act and the most politically significant website the federal government has ever tried to build. Millions of Americans without health insurance are waiting to use it. President Obama has staked a meaningful chunk of his second-term legacy on it. The budget for the project was somewhere north of $2 billion. The number of people who successfully enrolled in a health insurance plan on launch day was six. Not six thousand. Six.
By the end of day one, four million unique visitors had tried to use the site. The site had been designed to handle 50,000 to 60,000 concurrent users. It received 250,000 simultaneously, almost immediately after midnight, and collapsed within two hours. In the first week, with over eight million visitors hammering the system, approximately one percent of them managed to complete an enrollment. The site was taken offline that first weekend for emergency repairs. It was practically unusable for over a month.
The technical architecture was, to put it gently, a mess. The Department of Health and Human Services had awarded 60 separate contracts to 33 different vendors. No single contractor owned the integration of all those components. The largest contract went to CGI Federal, a subsidiary of a Canadian IT company with 70,000 employees and $10 billion in annual revenue, who had never built a consumer web application at anything close to this scale. When the components these vendors had built separately needed to work together, nobody was responsible for making that happen. They found out how well it worked when four million people showed up at midnight.
The single worst technical decision in the whole project was requiring users to create an account before they could browse insurance plans. It sounds like a minor UX call. Architecturally, it was catastrophic. Every visitor to the site had to pass through a single identity database before they could do anything at all. That database became the universal bottleneck. When 250,000 people hit it simultaneously, it saturated. The queue backed up across every downstream system. Pages timed out. Sessions dropped. Users who managed to fight through and complete an enrollment discovered their applications had been submitted with errors, as duplicates, or in many cases hadn’t been transmitted to insurers at all. The back-end transaction system failed to record purchases for tens of thousands of Americans who believed they had successfully signed up for health insurance.
The part that makes this genuinely painful to read about is that there had been warnings. An independent review conducted before the launch had found significant performance problems. The testing environment didn’t match production. Integration testing between the 33 vendors’ components had been minimal. A September stress test with a few hundred simulated users had revealed serious issues that were never fully resolved. Federal officials were briefed on all of this and launched on schedule anyway, because the political consequences of delaying the Affordable Care Act’s rollout were judged to be worse than the risk of a troubled launch. That calculation turned out to be wrong in both directions: the launch was disastrous, and the political fallout was severe.
The recovery is actually the good part of the story. A crisis team including engineers from Google, Red Hat, and Oracle was assembled and given authority to fix things without the bureaucratic constraints that had created them. They worked in sprint cycles, deployed continuously, and had the site functioning adequately within six weeks. The failure directly led to the creation of the US Digital Service and 18F, two government technology units that exist specifically to stop this kind of thing happening again. Healthcare.gov today handles millions of enrollments per year without drama. The 2013 launch remains the most expensive argument in history for doing load testing before you go live.
The Fastly outage: one config change, 49 minutes, 85% of the internet gone
June 8, 2021 started as a completely ordinary Tuesday morning. At 9:47 AM UTC, a Fastly customer made a routine change to their CDN configuration. The kind of change CDN customers make dozens of times a day. Adjusting a routing rule, tweaking cache behaviour, nothing unusual. Within 60 seconds of that change being processed, the UK Government’s website was showing a white error page. So were the New York Times, the Financial Times, the BBC, the Guardian, Reddit, GitHub, Stripe, Shopify, PayPal, Twitch, Spotify, and Amazon. Not degraded. Down. Returning 503 errors to every user on earth who tried to reach them. Fastly had taken 85% of its global infrastructure offline in under a minute, with a single customer’s configuration change as the trigger.
It took 49 minutes to fix. Those 49 minutes cost the affected companies an estimated $67 million in lost revenue collectively, with Amazon alone accounting for a significant portion of that. The UK government’s emergency communications infrastructure, which runs on gov.uk, was unreachable. News organisations covering the outage couldn’t publish their own coverage because their sites were down. The financial cost was significant. The demonstration of how concentrated internet infrastructure had become was more significant.
Here’s what actually happened technically. Fastly had deployed a software update on May 12, three and a half weeks before the outage. That update contained a bug associated with a feature flag, buried in a code path that was not active. The bug sat dormant for 26 days. Then the customer made their routine configuration change, which triggered the feature flag, which activated the dormant code path, which hit the bug, which caused a single edge node to crash. Under normal circumstances, one node crashing is a recoverable event. What happened next is the part that turned a local failure into a global catastrophe.
Fastly’s configuration management system was built for consistency and speed. When a customer changes their CDN configuration, it propagates to every edge node in the global network simultaneously and instantly. This is actually the right design for a CDN: you want your cache rules and routing logic to be consistent everywhere, and you want changes to take effect immediately. The problem is that this same system, when it processed the triggering configuration change, propagated that change to every node simultaneously before anyone knew the first node had crashed. There was no staged rollout. No canary deployment. No regional rings. No circuit breaker that could detect the first node failing and pause the propagation. The mechanism built to guarantee consistency became the mechanism that guaranteed the crash would be universal. Every node received the same configuration. Every node crashed.
The fix was fast once they found it because it was a configuration revert rather than a code rollback. Reverting a configuration change takes minutes. The 49-minute window was almost entirely investigation time: confirming what the problem was, tracing it to the specific change, and verifying that reverting it would actually fix things. Within minutes of the revert, 95% of Fastly’s network was back online. The remaining time was spent waiting for caches to refill across the edge network.
What the Fastly outage really exposed was how few major websites had built any redundancy against losing their CDN provider. CDNs work so reliably that most organisations treat them like the power grid: something you depend on completely without building a fallback for. The organisations that came through the outage best were the ones load-balancing across multiple CDN providers simultaneously, who could remove Fastly from their DNS responses and route around the failure within minutes. Most were not doing that. They were entirely dependent on a single provider, and when that provider went down, they went with it.
Pokemon Go: fifty times the expected traffic, because nobody modelled “the entire world goes outside”
Here’s a number that still seems impossible even knowing the outcome: Pokemon Go generated 50 times more traffic than Niantic expected on launch day. Not 50% more. Not twice as much. Fifty times. Google’s Customer Reliability Engineering team later disclosed that the game was pinging their Cloud Platform 10 times more often than the most extreme worst-case scenario Niantic and Google had jointly modelled before launch. Whatever methodology they used to forecast demand, reality beat it by an order of magnitude before lunch on day one.
The game launched on July 6, 2016 and reached 10 million installs in its first week, faster than any mobile app in history to that point. Within days it had more daily active users than Twitter. It was being played by people who had never touched a video game. It was causing traffic accidents. People were walking into fountains. And the servers, built for a fraction of this load, were completely unable to cope. The game was unplayable for most of the first week. Login failures were constant. The in-game Pokemon tracking system broke almost immediately and stayed broken for months. Users who managed to get into the game found it freezing, dropping their sessions, and refusing to register catches. The Pokemon Go launch was the rare tech disaster where the failure was caused by the product being too successful too fast, which somehow made it more embarrassing rather than less.
The specific technical problems compounded each other. Network analysis during the outage showed that Pokemon Go’s server infrastructure was centralised in a single US location, without the geographic sharding that would distribute load across regional data centres. Players in Australia, Japan, and Europe were routing requests to servers in the United States, adding significant latency on top of the raw capacity problem. At peak failure periods, independent monitoring showed packet loss to Pokemon Go servers reaching 100%, with connection failures occurring inside Google’s infrastructure rather than at the edge. The few players who managed to get in found an application that often froze entirely, because the servers that were intermittently accepting connections were too saturated to actually process requests reliably.
Google and Niantic worked around the clock after the US launch to fix it. They spun up tens of thousands of additional cloud cores. They rebuilt the architecture using the latest version of Google Container Engine to allow dynamic scaling as load patterns shifted. They moved from a centralised model to geographic sharding, splitting the game world so that players in each region were served by infrastructure in their own geography. They implemented a faster HTTPS protocol to reduce connection overhead at scale. By the time the Japan launch arrived two weeks later, the number of new users signing up tripled the US launch peak and the infrastructure held. The two weeks of chaos had produced a genuinely scalable architecture. They just hadn’t built it before the launch.
The detail that makes this story sting a little is the prior experience. Niantic’s co-founder John Hanke had been at Google when Google Earth launched in 2005, an event that crashed Google’s own servers when 100 million users arrived quickly and consumed more than half the company’s total bandwidth. That experience was supposed to inform how Niantic approached Pokemon Go. It clearly did, in that they thought carefully about scale and provisioned what they believed was generous capacity. It just didn’t inform them enough to model an outcome where the entire world decided to go outside and catch fictional animals on the same week.
The pattern behind all three
Three different failures, three different proximate causes. A procurement disaster with no systems integrator and a single database bottleneck. A dormant bug triggered by a configuration change that propagated globally with no blast radius limit. A capacity model that was wrong by fifty times because nobody could model a cultural phenomenon before it happened. They look very different on the surface.
Underneath, they’re the same problem. Every team involved made assumptions about what demand would look like, what failure modes were possible, and how resilient their architecture was under real conditions. Those assumptions turned out to be wrong in ways that were either discoverable before launch or built into the architecture as unavoidable risks. Healthcare.gov had a pre-launch review that found the problems. Pokemon Go had prior experience with exactly this kind of failure. Fastly had a configuration system with no staged rollout and knew it. In each case, the information needed to avoid the failure existed. The priority needed to act on it didn’t.
We think about this stuff constantly. Server infrastructure is only invisible when it’s working. When it isn’t, it’s the only thing anyone can think about, and the conversations that follow are always some version of “we knew this was a risk.” Load test past your worst case. Build circuit breakers into anything that propagates globally. Never make one component the universal bottleneck. And when a pre-launch review finds problems, treat it as the thing it actually is: the last warning you’re going to get before reality does the testing for you.