` reads). %> The 3am Cert Expiry: Why Outages Always Happen at the Worst Time | TLS Radar Skip to main content
outage-prevention 6 min read By TLS Radar Team

The 3am Cert Expiry: Why Outages Always Happen at the Worst Time

The car never breaks down in the driveway.

It breaks down on the motorway, in the rain, in the middle of nowhere, on the night you needed to be somewhere. The flat tyre announces itself fifteen minutes after the petrol station closed. The dead battery wakes up just in time for the school run. The engine fault appears as the satnav loses signal.

The fancy term is selection bias - you remember the bad-timing breakdowns because they hurt more - but most engineers will swear there's something else going on. The car knows.

TLS certificates have learned this trick.

They don't expire on Wednesday afternoons when the team is in the office. They expire at 3am on a Sunday during a bank holiday weekend. They expire while the on-call engineer is on the train home from their grandmother's birthday. They expire fifteen minutes after the cloud provider's status page goes amber for unrelated reasons, so nobody realises the cert thing is the cert thing.

It is not actually the certs. It is the structure of cert renewal, on-call rotations, and human attention. But the pattern is real, and worth understanding, because the fix isn't "be more careful." It's "have systems that don't depend on careful humans at 3am."

The math of expiry timing

Certs expire at the moment they were issued, but a year (or 200 days, or 90 days) later. The clock starts when you issue. It runs continuously. It does not pause for holidays.

So if you issued a cert at 2:47am during an incident in February 2024, that cert expires at 2:47am in February 2025. Or 2:47am three months later, with shorter validity periods.

This sounds harmless until you consider when most certs actually get issued. Two patterns:

One: during incidents. Someone provisions a cert in the middle of a 3am crisis because the old one was about to expire. The new cert inherits the 3am-issued timestamp. Now it expires at 3am next year too. The crisis perpetuates itself.

Two: during deployments. Many cert provisioning systems issue certs at deploy time. Deploys often happen outside business hours - overnight, on weekends - to minimise customer impact. Those certs now have expiry timestamps outside business hours.

The result: across a portfolio of certs, the expiry distribution is biased toward exactly the times when humans aren't watching. This is not paranoia. It's the consequence of how certs are issued.

Why monitoring fails worst at 3am

Three reasons monitoring effectiveness drops sharply outside business hours.

Alert fatigue is highest at 3am. A pager going off at 3am for a non-critical issue gets snoozed faster than the same alert at 11am. By the time the engineer is conscious enough to read it carefully, they've already triaged it as "deal with in the morning." Cert alerts disguised as one of many noisy alerts will lose this filter.

On-call rotations are sparser. Most companies have one on-call engineer at any time. During business hours, that engineer has the whole team nearby. At 3am, they have themselves, a laptop, and Slack. The diagnostic loop slows by an order of magnitude.

Communication channels degrade. During business hours, escalation is fast - a question gets answered in three minutes. At 3am, the question gets answered in fifteen, if at all. Decisions that need approval (issue a new cert, change DNS, contact the CA) take longer.

The cert problem doesn't get worse at 3am. The response to it does.

The on-call paradox

Most cert outages share a structural pattern. The cert was set up by an engineer. That engineer is the one who knows the renewal process. Two years later, the cert is expiring. The engineer is on holiday - or has changed teams, or has left the company.

The on-call engineer who picks up the alert has never touched this cert. They don't know which CA. They don't know which automation. They don't know who else might know. They look in the wiki - the wiki is out of date.

Compare this to a typical service outage. Most services have ten engineers who could debug them in a pinch. Cert provisioning, in many organisations, has one or two people who really understand the setup. When the alert fires, the probability that the on-call person is one of those one or two is low.

This is the cert ownership problem we keep talking about. It bites hardest at 3am.

What teams that don't get woken up do differently

Four practices.

Renewals happen well before expiry. A cert that's renewing thirty days early gives the renewal-related failures sixty days to manifest. A cert that's renewing two days early gives them two days. Push renewals as early as practical. There's almost no downside to a new cert that overlaps the old one for a few weeks.

Alerts escalate during business hours, not at 3am. A cert expiring in 30 days alerts on Monday morning, not Sunday at 3am. The alarm time should reflect when the team can actually act, not when the cert thinks it's interesting. Most monitoring systems support time-based filtering. Use it.

Multiple alarm tiers. Soft alerts (30 days out, 14, 7) during business hours. Hard alerts (24 hours, 1 hour) page someone regardless of time. The early warnings handle the routine case during the workday; the late warnings are the emergency backstop.

Runbooks that don't require Bob. The on-call engineer who's never touched this cert should be able to renew it from a documented procedure. Not "ask Bob." Not "check the wiki" (when the wiki is out of date). An actual runbook, tested, that an engineer who's only half awake can follow.

A team that does all four rarely gets woken up by cert expiries. Not because their certs don't expire - because the certs expire on Monday morning during business hours, by design.

The "is it me or is it always Friday?" pattern

A semi-serious observation: a disproportionate number of high-profile cert outages happen on Fridays, weekends, and holidays.

Some of this is selection bias - you remember those because they made the news. Some of it is real:

  • Releases push on Fridays (to "stabilise over the weekend"), introducing new services with fresh certs that all expire on Friday a year later. Year-end cert issuance spikes during the December slowdown, expiring during the next December slowdown when half the team is gone. Cron jobs run on their schedules without regard to the human calendar.

There is no actual conspiracy. The certs are not malicious. But the pattern is real enough that "Friday afternoon cert expiry" is its own genre of internal joke at most ops teams.

One small ask

Cert outages at 3am aren't a personality flaw of certs. They're a consequence of how certs get issued, how teams are staffed, and how monitoring is configured. The fix is structural - earlier renewals, smarter alarm scheduling, runbooks that don't depend on one person being available. TLS Radar handles the monitoring side: time-based alert escalation, team-level alerting, and full coverage of failure modes that often only surface when the on-call engineer is least equipped to handle them. Free tier covers three domains, which is enough to see whether your current setup would wake you up unnecessarily.

Related reading

Get the next post in your inbox

TLS monitoring tips and product updates. No spam, unsubscribe anytime.

Keep reading

Comparing tools? See how TLS Radar stacks up against DigiCert and SSL.com.