` reads). %> The Hidden Outage: When Your Cert is Valid But Your Site Still Breaks | TLS Radar Skip to main content
outage-prevention 6 min read By TLS Radar Team

The Hidden Outage: When Your Cert is Valid But Your Site Still Breaks

Imagine a small restaurant. The door is unlocked. The lights are on. The kitchen is running. The food is fine. But a "Closed" sign is still hanging in the window from last Tuesday's day off, and nobody has flipped it back. Every prospective customer walking past reads the sign and keeps walking. The restaurant is open. The restaurant is also empty. The closed sign and the actual state of the restaurant don't agree, and the sign wins because the sign is what people see.

A lot of cert outages look exactly like this.

The cert is valid. The expiry date is in the future. The chain validates. The cipher is strong. Your monitoring dashboard is green. But for some meaningful percentage of users, the site doesn't work - they get warnings, errors, or silent connection failures. The technical state of the cert and the user-facing state of the site don't agree, and the user-facing state is what matters.

These are the hidden outages. They're harder to detect than expired-cert outages because nothing looks wrong from inside. The cron job ran. The renewal succeeded. The dashboard says green. The user-facing experience says broken.

Here are the patterns to know, why they hide so well, and what monitoring actually catches.

The "valid but broken" failure modes

The deployment that didn't deploy. A new cert was issued. The automation reported success. But the cert never made it onto the load balancer, the web server, or the service it was supposed to protect. The old cert is still being served. It might still be valid for a few days - but it's the wrong cert, and the new one is sitting unused in a secrets store somewhere.

The reload that didn't happen. The cert was deployed to the right place. The web server was supposed to reload to pick it up. It didn't. nginx is still serving the old cert from memory; the new file is on disk; both are valid by their dates, but the served one is the older one with the imminent expiry.

The SNI mismatch. Multiple sites on one IP via Server Name Indication. The server is supposed to look at the requested hostname and serve the matching cert. A misconfiguration causes it to serve cert A for hostname B. Cert A is valid. Cert A doesn't cover hostname B. Users see the warning. The server sees nothing wrong.

The CDN cache that hasn't refreshed. Your origin has the new cert. Your CDN has the old one cached. The CDN serves what it has. Users see whichever the CDN serves. Could last minutes; could last hours depending on cache rules.

The mobile app that pinned the old cert. Mobile apps that pinned a specific certificate (cert pinning) as part of their security model. The cert rotates. The app still expects the old fingerprint. The app stops working. The server is fine. The cert is fine. The app sees a different cert than it was told to trust.

The chain that's broken for some clients. Covered in detail in our chain validation piece (an earlier piece). Worth a callback here: the cert is valid; the chain is broken for old Android, old Java, certain enterprise security appliances. Most users see no problem. A meaningful subset sees errors.

The wrong cert restored from backup. A server is restored from a backup. The backup includes a cert that was retired six months ago. The server is now serving a cert that, while technically still valid, has long been replaced everywhere else. Until someone notices, it's serving the wrong identity.

Each of these patterns shares one feature: from any internal check, the cert looks fine. The expiry is fine. The signature is fine. The dashboard is green. But the user experiences a broken site.

Why these are hard to detect

Internal monitoring shares the server's blind spots.

If your monitor lives on the same server as the cert, it checks the cert that's on disk. If the served cert is different from the on-disk cert (because the reload failed), the monitor sees disk; the user sees served. They disagree. The monitor is wrong.

If your monitor queries your renewal pipeline, it sees that the pipeline reported success. The pipeline doesn't know whether the deployment that was supposed to follow succeeded. The monitor reports green for the renewal; the deployment is broken; the monitor never tracks the gap.

If your monitor only checks one hostname (say, the apex), it doesn't catch the SNI mismatch on a subdomain.

If your monitor doesn't simulate clients across geographies, it doesn't catch the CDN-cached old cert that's still being served in one region.

The pattern: internal monitoring inherits the server's view of itself. The server's view of itself is exactly what you can't trust during a hidden outage.

How the user sees it vs how monitoring sees it

A useful exercise: imagine the same incident from two perspectives.

From inside: - Renewal automation log: success. - Cert file on disk: valid, new, future expiry. - Internal monitoring: green. - Dashboard: green. - Engineer on duty: nothing to do.

From outside: - User opens site: browser warning. - User goes to competitor. - Customer support: increasing tickets. - Twitter: people complaining. - Status page: still reporting all systems operational.

These two stories can persist for hours before they collide. The collision usually happens via a customer report - never via the internal monitoring that the team trusted to catch this.

What real monitoring catches

External monitoring - checks that talk to your endpoints from outside, the way real users do - catches what internal monitoring misses.

Specifically: - It connects to the endpoint and inspects the cert actually being served, not the cert on disk. - It tests from multiple geographic regions (catching CDN cache discrepancies). - It tests with multiple client TLS stacks (catching chain validation issues across audiences). - It checks every hostname individually (catching SNI mismatches and SAN coverage gaps). - It verifies that the cert being served is the cert you expect - by serial number, by fingerprint - not just that a valid cert is there.

When external monitoring disagrees with internal monitoring, the external view is usually correct. The user side is the side that pays the bills.

One small ask

The hidden outage is the failure mode that pure renewal-success monitoring can't catch. The cert is valid; the site is broken; the dashboard is green. The fix is monitoring that doesn't trust the server's view of itself. TLS Radar checks from outside, from multiple regions, against the cert actually being served - and flags when the served cert doesn't match the cert the renewal pipeline thinks is live. Free tier covers three domains, which is enough to see whether your current monitoring would catch a closed-sign-in-the-window incident.

Valid isn't the same as working

The most dangerous outages keep the dashboard green: the cert is valid, but a slice of your users still can't connect. TLS Radar checks what the visitor sees, not what the server thinks it's serving, and flags when the served cert doesn't match what your pipeline believes is live.

Related reading

Get the next post in your inbox

TLS monitoring tips and product updates. No spam, unsubscribe anytime.

Keep reading

Comparing tools? See how TLS Radar stacks up against DigiCert and SSL.com.