` reads). %> ACME at Enterprise Scale: Pitfalls and Patterns | TLS Radar Skip to main content
lifecycle 6 min read By TLS Radar Team

ACME at Enterprise Scale: Pitfalls and Patterns

A revolving door is a small, elegant piece of engineering. It controls the flow of people. It keeps heat in. It looks nice. At a small office building, it works perfectly - twenty or thirty people pass through it on a typical day, never more than a few at once.

Now picture the same revolving door installed at Grand Central Terminal during rush hour. Same mechanism. Same elegance. Completely different problem. The constraints that were invisible at small scale - throughput, queue behaviour, what happens when one person stumbles - become the entire story.

ACME (Automatic Certificate Management Environment, RFC 8555) is the revolving door of certificate issuance.

It is, by any reasonable measure, the most successful protocol the cert ecosystem has ever produced. It powers Let's Encrypt. It runs on millions of small servers worldwide. It works beautifully at the scale it was designed for - one site, one cert, one renewal at a time, with a human or a simple cron in the loop.

At enterprise scale, ACME starts running into constraints that nobody saw in the small-site version. Rate limits, validation timing, multi-account management, post-issuance distribution, audit logging. Each one is solvable. None of them are obvious until you hit them.

Here's what changes when ACME goes from "the cert manager on my VPS" to "the cert manager for ten thousand internal services," what tends to break, and what patterns actually work.

What ACME does, briefly

For readers who haven't worked with it directly: ACME automates the conversation between a server that wants a cert and a CA that issues certs.

The shape: 1. Server tells the CA: "I want a cert for example.com." 2. CA tells the server: "Prove you control example.com. Either put this token at a specific URL on the site (HTTP-01) or publish this DNS record (DNS-01)." 3. Server does it. 4. CA verifies. Issues the cert. 5. Renewal is the same flow, automated.

The whole thing fits in a small script. certbot, acme.sh, lego, and win-acme all implement it in slightly different ways. Most modern web servers (Caddy, recent nginx with modules, Traefik) speak ACME natively.

For a single site, this is genuinely magical. The cert renews itself, you don't think about it, life goes on.

Where ACME breaks at scale

Four places.

Rate limits. Let's Encrypt has them. So does Google Trust Services, ZeroSSL, and any sane public CA - they're necessary to prevent abuse. Most rate limits are generous for small operators. At enterprise scale, you can hit them. Once you do, you wait. Wait time during a deploy or incident is bad.

Validation timing. HTTP-01 validation requires the CA to reach your server on port 80. If your firewall blocks it, or if your CDN intercepts it, or if your load balancer routes the request to a different backend than the one that wrote the token - validation fails. DNS-01 requires updating DNS, waiting for propagation, and not racing other DNS updates. At small scale, these timing issues happen rarely. At large scale, they happen daily.

Post-issuance distribution. ACME gets you a cert. It does not put the cert in your load balancer, your CDN, your service mesh, your secrets manager, and your container image registry. At small scale, the cert lives on the server that requested it. At large scale, the same cert needs to land in five or six places, in the right order, without breaking running services.

Audit logging. For SOC 2 and PCI compliance, you need a record of every cert issuance, who triggered it, and what it covers. ACME's standard logs are not designed for this. You're going to add a layer.

These aren't bugs in ACME. They're consequences of scaling a protocol designed for one-site simplicity to enterprise estate management.

Patterns that work

Four patterns, from teams that have actually done this.

Multiple accounts. Don't put all your eggs in one ACME account. Rate limits apply per registered domain, per account. Multiple accounts (one per business unit, or one per environment) give you more headroom and isolate blast radius when one account hits limits.

An internal ACME server. Tools like Smallstep, step-ca, and HashiCorp Vault PKI all implement ACME against an internal CA. For internal services, this removes the external CA from your hot path entirely. Your internal infrastructure issues its own certs, on your own rate limits, with your own validation rules. External CAs are reserved for external-facing endpoints where public trust matters.

A distribution layer. Don't let your ACME client touch production directly. Issue the cert to a secrets store (Vault, AWS Secrets Manager, Kubernetes secrets). Have the services pick up new certs from there. This decouples issuance from deployment, makes rollback easier, and centralises the place where the cert is "the real one."

Observability separate from the renewal pipeline. This is the part everyone misses. Your ACME automation is going to fail silently sooner or later - a credential expires, a DNS account gets disabled, a webhook stops firing. The only way to catch this is external monitoring that doesn't share infrastructure with the renewal pipeline. "Even with automation, having external validation helps avoid blind spots," as one operator put it on Hacker News.

The CA dependency

A separate scaling issue: depending on a single CA.

Let's Encrypt is wonderful. It is also a single point of failure if every cert in your estate comes from it. The May 2026 issuance halt - a few hours when no Let's Encrypt certs could be issued at all - was a useful reminder that the CA is part of your infrastructure, not a free utility you can take for granted.

Patterns for managing this:

  • Have a backup CA configured but not primary. Google Trust Services, ZeroSSL, and Buypass all speak ACME and can serve as failovers. For services where downtime during a CA outage would be catastrophic, issue from two CAs in parallel and pick whichever cert is fresher. Long-lived "emergency" certs from a separate CA, manually pre-staged, ready to swap in if the primary CA goes down for an extended period.

None of these are free. All of them are cheaper than the alternative - telling stakeholders "we couldn't issue a cert because Let's Encrypt was down."

One small ask

ACME at small scale works on its own. ACME at enterprise scale is a system: multiple accounts, distribution layer, observability, CA redundancy. The constraints aren't ACME's fault - they're the cost of using a small-scale-friendly protocol at large scale. Worth doing right. TLS Radar provides the observability piece for ACME-managed certs, including silent-failure detection that the renewal pipeline itself can't catch. Free tier covers three domains, which is enough to test whether your current ACME setup is actually doing what you think.

Related reading

Get the next post in your inbox

TLS monitoring tips and product updates. No spam, unsubscribe anytime.

Keep reading

Comparing tools? See how TLS Radar stacks up against DigiCert and SSL.com.