Anatomy of a TLS Outage: 5 Famous Incidents and What We Learned
When a plane crashes, the industry doesn't quietly file the wreckage. It convenes investigators, builds a report, and circulates the findings so the same crash doesn't happen again. The crashes still happen. But the catalogue of how they happen, and why, keeps getting more useful.
TLS outages get less attention than plane crashes - for obvious reasons - but they follow the same pattern. Every few months, a major service goes down because of a forgotten certificate. The post-mortem gets written. The blog post goes up. The lessons get learned. And then, six months later, somewhere else, the exact same thing happens to someone else.
Here are five TLS-related incidents worth knowing, what actually went wrong, and what teams that don't want to be next can take from them.
Microsoft Teams, February 2020 - The Big One That Made the News
What happened: Microsoft Teams went offline for several hours on February 3, 2020. The outage was global. Coffee was spilled.
What caused it: an expired authentication certificate. Not the public-facing TLS cert on teams.microsoft.com - an internal certificate involved in service authentication. The cert expired and the service had no fallback.
What we learned: - "We have thousands of engineers" is not a defence. Microsoft has thousands of engineers. They still missed it. - Internal certs matter as much as public ones. Possibly more. Public certs have CT logs and external monitors watching. Internal certs often have neither. - The blast radius of an internal cert can be huge if the cert is in the auth path. Every dependent service goes down together.
The fix isn't "fewer certs." The fix is treating internal certs with the same monitoring and ownership rigour as public ones.
O2 / Ericsson, December 2018 - The Vendor Cert You Don't Own
What happened: O2 customers across the UK lost mobile data for most of a day on December 6, 2018. The incident also affected several other carriers using the same Ericsson software.
What caused it: an expired certificate inside Ericsson's network management software. Customers had no visibility into the cert. The operator had no visibility into the cert. The cert lived in the vendor's product.
What we learned: - You can't monitor what you can't see. Vendor-embedded certs are operational time bombs you inherit when you buy enterprise software. - The vendor's monitoring is not your monitoring. Vendors sometimes have monitoring that covers their certs. Sometimes they don't. Either way, the outage is yours. - Procurement contracts should ask the question: "what certs are embedded in this product, and what happens when they expire?"
This one is hard to fix as a customer. But you can ask the question before you buy, and you can demand transparency in contracts.
Equifax, 2017 - The Silent Failure That Hid a Breach
What happened: the 2017 Equifax breach exposed the data of roughly 148 million people. The story everyone remembers is the unpatched Apache Struts vulnerability. There's a less-remembered second part.
What caused the extended exposure: a digital certificate on a device used for SSL inspection - looking at outbound network traffic for signs of compromise - had been expired for nineteen months. The device was, in effect, blind. When the cert was finally renewed, the network monitoring tool came back online and almost immediately surfaced evidence of the breach in progress.
What we learned: - Cert outages don't always look like outages. Sometimes they look like "everything is working but the alarms are off." - Security tools that rely on certs need their own monitoring. The thing that watches your network for breaches needs to be watched too. - A cert expired for nineteen months is not a "we missed a renewal" failure. It's an "ownership and monitoring were broken for nineteen months" failure.
Documented in the US House Oversight Committee's 2018 report on the breach. Worth reading if you're responsible for any cert in a security-monitoring pathway.
Let's Encrypt, May 2026 - When Your CA Has a Bad Day
What happened: Let's Encrypt halted all certificate issuance on the evening of May 8, 2026, for several hours after detecting an issue with one of their cross-signed roots.
What caused it: a precautionary halt during incident response. No certs were misissued. But during the pause, any site with a healthy ACME automation that happened to need a new cert during the window couldn't get one.
What we learned: - "We have automation" and "we are protected" are not the same sentence. When the upstream has a problem, your automation can be working perfectly and still not get you a cert. - Single-CA dependence is a fragility. Sites that can issue from multiple CAs ride out these incidents better. - Monitoring needs to surface "we tried to renew and the CA said no" as a distinct event, not lump it with general renewal failures.
For most teams, the practical takeaway isn't "switch CAs." It's "know which of your renewals would have been stuck during a multi-hour CA pause."
Let's Encrypt, 2024 - The Cross-Sign That Wouldn't Cross
What happened: in 2024, a cross-signed Let's Encrypt root certificate caused chain validation issues on older clients - older Android devices, older OpenSSL versions, certain enterprise security appliances. Most modern browsers handled it fine. Specific subsets of clients did not.
What caused it: the legitimate complexity of cross-signing arrangements between CAs. Cross-signs are how new root CAs get bootstrapped into older trust stores. They have to expire eventually. When they do, the older clients lose the ability to validate the new chain.
What we learned: - "The cert is valid" depends on who you ask. Modern Chrome accepts what older Android refuses. Both are right. - Chain validation has to be tested across multiple client environments, not just the one your browser uses. - Mobile apps and B2B integrations are where these issues bite hardest. Consumer browsers tend to be the most forgiving; mobile and machine-to-machine clients are not.
This is one of the failure modes that pure expiry monitoring misses entirely. The cert was valid. The chain was broken for a subset of clients. The monitoring dashboard stayed green.
The pattern
Across all five incidents, the same themes appear:
- The thing that broke wasn't the cert most people were watching. The monitoring in place caught the loud failure modes (expiry) and missed the quiet ones (chain, dependency, vendor-embedded, silent failure). The fix wasn't "more engineers." It was better coverage of the failure modes that don't show up on a dashboard built around expiry.
A useful exercise: take your current cert monitoring and run it against the five incidents above. Would it have caught each one? If the answer is "no" for more than one, your monitoring is mostly a fuel-light dashboard.
One small ask
The lesson from every plane crash investigation is the same - the failure modes change, but the categories of failure repeat. The lesson from every TLS outage is similar. The interesting question isn't "is this one of these stories?" It's "do you have the systems to catch the next one before it becomes a story?"
Related reading
Get the next post in your inbox
TLS monitoring tips and product updates. No spam, unsubscribe anytime.