Inheriting a PKI Mess: The First 90 Days
When you buy an old house, you don't just inherit the walls. You inherit the previous owner's accumulated decisions. The wallpaper they chose in 1997. The boiler they replaced in 2014 with the cheapest available unit. The garden shed that's leaning slightly because they built it on uneven ground. The mysterious wire running through the attic that goes to nothing visible.
You do not have to fix all of it. You couldn't if you wanted to. What you have to do - in the first ninety days - is figure out what you've actually got, what's about to fall over, and what can wait until you have time and budget to do it properly.
Inheriting a certificate estate works the same way.
Maybe you joined a new company and just learned that "the cert person left six months ago." Maybe your team absorbed another team in a reorg and you now own their certs. Maybe your company acquired another company, and their PKI came with them. Whatever the path, you are now responsible for an inventory you didn't build, decisions you didn't make, and renewals you can't yet predict.
Here's a ninety-day plan that has worked for teams that have done this. Not a perfect plan. A survivable one.
The first week: inventory
You cannot fix what you cannot see. The first week is reconnaissance, not action.
Day 1–2: Find the working systems. - Get access to whatever cert management tool was in use. Spreadsheet, CMDB, dashboard, monitoring tool, scattered shell scripts. All of them. - Ask people: "what tools do you use for certs?" Write down every answer, even contradictory ones. Especially contradictory ones. - Find the CA account credentials. There may be several.
Day 3–5: Discover what's actually out there. - Query Certificate Transparency logs for your organisation's domains. You will find certs you didn't know existed. - Pull cert lists from every cloud account (AWS ACM, Azure Key Vault, GCP Certificate Manager). - Pull cert lists from every CDN (Cloudflare, Fastly). - Pull cert lists from every load balancer. - If there are internal CAs, find their issuance logs.
Day 5–7: Reconcile. - Compare what the existing inventory claims against what you found. - Make a list of discrepancies. Don't try to resolve them yet - just record them. - Identify certs expiring in the next thirty days. These are tomorrow's emergencies if you don't get to them.
By the end of week one, you should know roughly how many certs you have, where they live, which CA each came from, and which ones are about to expire. You will be wrong about some of this - but you'll be less wrong than you were on day one.
Weeks 2–4: ownership and triage
The second phase: figure out who owns what, and fix the most urgent fires.
an earlier piece - ownership. - For every cert in the inventory, identify the current owner. Not the historical owner - the current one. "The team that would handle it if it broke right now." - For certs with no clear owner, default to whichever team uses the service the cert protects. Document the assignment. - Surface orphan certs to leadership. Some of them protect services nobody remembers, which is a different conversation.
an earlier piece - urgent renewals. - Renew any certs expiring in the next 60 days. Don't wait. Push them through, even if the process is rough. - For certs you can't renew (lost credentials, expired CA accounts, missing access), document the blocker and escalate.
an earlier piece - first audit. - Run an external scan against every public endpoint. Note the grades, the weak ciphers, the chain issues. - This is your baseline. Don't try to fix everything yet. Just know what's there.
By the end of week four, you have an inventory with owners, you have triaged the immediate fires, and you have a baseline of TLS health to measure progress against.
Weeks 4–8: stabilisation
The first month was reconnaissance and triage. Weeks 4–8 are stabilisation - getting the system to a place where it doesn't depend on you personally to function.
an earlier piece - automation audit. - For every cert renewing automatically, find the automation. Verify it works. Check whether the credentials it uses are tied to a person who might leave. - For automation that's fragile, document how it works. Future-you will thank present-you.
an earlier piece - monitoring baseline. - If there's no external monitoring, install some. TLS Radar's free tier covers three domains; if you have more than three, the paid tiers cover bigger estates. Other options exist; the point is to have something running. - Configure alerts to go to a team channel, not your personal email.
an earlier piece - runbook draft. - Write a basic incident runbook for cert outages (we have a template - see Building a TLS Incident Runbook from earlier pieces). Even a rough one is better than none.
an earlier piece - communication. - Brief the team. Brief leadership. Surface what you found, what you fixed, and what's still on the list. - Establish a regular cadence - maybe monthly - for cert reviews going forward.
By the end of week eight, your cert estate has external monitoring, team-level ownership, basic runbooks, and a known set of outstanding issues. You are no longer the single point of failure for cert ops.
Weeks 8–12: setting up the future
The last month is about putting the system on a sustainable footing.
an earlier piece - process review. - How are new services added to the cert inventory? If the answer is "ad hoc," fix it. New service creation should include a cert registration step.
an earlier piece - backlog prioritisation. - Take the list of issues from week four. Prioritise. Most won't be fixed in 90 days. Pick the three or four that matter most.
an earlier piece - handoff plan. - Document everything you've done. Diagrams, runbooks, ownership maps, automation locations. Treat it as if you might be hit by a bus.
an earlier piece - review with leadership. - Show what's better, what's still outstanding, and what budget or process changes you need for the next ninety days.
What you can't fix in 90 days
Three things take longer.
Cultural change. The reason the previous owner left certs in a mess is usually not technical. It's that nobody made cert ownership someone's clear responsibility. Fixing this requires sustained effort over many months and probably the cooperation of leadership.
Inherited automation rot. Old scripts, old credentials, old assumptions baked into infrastructure. Untangling these often requires rebuilding from scratch rather than patching, which is a longer project.
External vendor dependencies. Vendors who embed certs in their products, BAAs that include cert obligations, third-party integrations that depend on specific cert configurations. These move at vendor speed, not yours.
Accept that these will be live for a year or more. Document them. Plan for them in next year's roadmap. Don't try to crash through them in 90 days.
One small ask
The first 90 days inheriting a PKI mess is mostly about visibility. You can't fix what you can't see, and the previous owner has carefully ensured that you can't see most of it. External monitoring is the cheapest way to build visibility fast. TLS Radar's free tier covers three domains and gives you continuous discovery from day one. Worth installing in the first week - you will surface certs you didn't know existed before you finish your second cup of coffee.
Related reading
Get the next post in your inbox
TLS monitoring tips and product updates. No spam, unsubscribe anytime.