
How Datadog Lost 40,000 Servers in 13 Hours
Source: YouTube · Akhil Sharma · published Jul 6, 2026 · 26:25
Datadog’s March 2023 global outage—costing ~$5 million and taking 13+ hours to resolve—was caused by a software supply chain failure, proving that multi-region architectures don't protect against fleet-wide shared dependency bugs 0:00-0:32.
Key Takeaways:
• Multi-region redundancy only protects against geographic failures. Datadog's five regions across three clouds went down simultaneously because they all shared the same OS auto-updater running on the same UTC schedule 2:01-2:20.
• The root cause was a 27-month-old systemd change. When a routine security patch restarted the network service, systemd deleted all "foreign" routing rules it didn't create, wiping out Cilium's critical Kubernetes pod networking rules 8:07-8:51.
• Standard testing completely missed this bug. The failure only triggers when the network service restarts on a long-running machine with active Cilium rules, not during a clean reboot—meaning production was the only place it could manifest 10:30-11:21.
• The outage was self-compounding: Datadog's own monitoring tool was trapped inside the broken network, forcing engineers to rely on raw SSH and cloud provider consoles for recovery 13:30-14:04.
• Never build your recovery plan on top of the system that can fail. Always maintain an independent, out-of-band observability path that shares no dependencies with your primary infrastructure 22:26-23:25.
Your next major outage likely isn't hiding in a geographic region, but in the unmodeled shared supply chain feeding every machine in your fleet 25:00-25:48.
Sources:
- 0:00-0:32 Overview of the Datadog outage scope and financial impact
- 2:01-2:20 The shared auto-update mechanism across all regions
- 8:07-8:51 Systemd deleting Cilium's routing rules
- 10:30-11:21 Why clean boot tests missed the bug
- 13:30-14:04 Loss of monitoring during the crisis
- 22:26-23:25 The need for out-of-band recovery paths
Generate CPE Credits
Generate a professional CPE document from this video's transcript.
Estimated credit: 0.5 CPE hours
Estimate uses the video runtime (1 hour ≈ 1 CPE, rounded to the nearest 0.5, minimum 0.5, maximum 2.0). The final amount can be lower after review, never higher.
Topic: Cloud Security. Commonly maps to: Security Architecture and Engineering, Communication and Network Security. Exact CISSP domains are assigned during generation.
Free account. One generation at a time, with a daily limit.
CPEBuddy is independent and not affiliated with or endorsed by ISC2, ISACA, or any certification body. Exports are formatted for common CPE submissions; acceptance is at your certification body's discretion.
Transcript Preview
First 800 characters of the transcript
On March 8th, 2023, at 6:00 in the morning UTC, Datadog started going dark. In not just one region, but every region across all three clouds that they run on in the same minute. Within just 2 hours, the entire product was gone. Dashboard was dead, APIs were down, and the alerting that thousands of companies trust had stopped firing. That's 40,000 servers knocked out. It took 13 hours to bring most of them back, and almost a full day to bring back the rest of the servers online. And the Pragmatic Engineer, a very famous blogging newsletter, pegged the lost revenue at nearly $5 million. Now, Datadog had five regions across AWS, Azure, and Google Cloud with multiple availability zones in each. And this is the exact diagram you would draw if someone told you to build a system that no single fa…