Workload identity & secret distribution

P58.service-to-service.02 · Audience: guest, it-ml, language-pro · Prerequisites: Service identity & mTLS

Real LLM grading for this pageLLM grading (this page):

mTLS answers how services prove themselves. This module answers the question underneath it: where do those credentials come from, and what happens when they expire?

The best answer is that services shouldn't hold long-lived credentials at all — the platform attests what a workload is and issues it something short-lived. That turns rotation from an incident into a routine, and it closes the P57 failure mode where a key ends up baked into an image or a repository.

One workload, from cold start to its second certificate

That paragraph is the claim. Here is the claim happening, to one service, over one day.

A container called checkout-7f9c starts up on one machine of a cluster, running the checkout service. A container is one service packaged with everything it needs to run and kept separate from whatever else shares the machine; a cluster is the pool of machines an organisation runs its services on and administers as a single system. The container was placed there by the orchestrator — the platform component that decides which machine runs what and actually starts each container — inside the payments namespace, which is simply a named partition of that cluster, grouping the services that belong to one team or one system. The container ships with no credential of any kind: no password in an environment variable, no certificate on disk, and no API key baked into the image it was built from — the read-only bundle of files that every copy of this container starts from. Unpack that image file by file and there is nothing in there worth stealing. That is the point of the design, and it is also, at this moment, a problem — the container has nothing with which to prove who it is.

The platform vouches for it, because the platform is the one that started it. The container asks the identity agent running on its own machine for a credential, over a local socket — a channel between two processes on the same machine, not a network connection. The agent does not take the container's word for anything, and it never looks at an address. The operating system kernel tells the agent which process is on the other end of that socket, and the orchestrator, which started that very process a second ago, says what it started it as: the container it launched to run checkout in the payments namespace. That is all attestation means — a party who already knows what a workload is says so, rather than the workload asserting its own name and being believed. An attacker elsewhere on the network cannot fake this, because there is nothing here to fake: they would have to be that process, on that machine.

The internal certificate authority signs a certificate that says so, and gives it 24 hours. The internal CA is a private certificate authority (P53) trusted by this cluster's services and by nobody outside it. The process inside the container generates its own keypair and sends out only a signing request — the public half plus the name it is asking for. The agent forwards that request with its attestation, the CA signs, and the signed certificate comes back through the agent over the same local socket. The private key is generated inside the container and never leaves it: the CA signs a request, it does not hand out keys, and the agent never sees the key either. What comes back is a certificate whose name is checkout.payments, valid for 24 hours, and it is precisely what the mTLS handshake you built in module 01 presents.

Now the identity does work. When checkout calls the ledger service, both sides present their certificates and both check that the other was signed by the internal CA. ledger then decides what the caller may do based on the name in the certificate, checkout.payments, and not on the address the connection arrived from. An attacker who takes over a neighbouring machine inherits an address, which now buys them nothing, because they cannot produce a certificate the internal CA signed.

At the 16-hour mark — not at 23:59 — the credential is replaced while it is still valid. Two thirds of the way through the lifetime the same exchange runs again: a fresh keypair is generated inside the container, a fresh signing request goes out through the agent, and a fresh certificate comes back over that same local socket. For the overlap window the old one and the new one are both good. The new material is loaded without restarting the process; connections already open finish on the old certificate, connections opened from now on use the new one, and ledger accepts either because the same CA signed both. Nothing is drained and nothing is restarted. If issuance fails, there are still eight hours of valid certificate left in which to raise an alarm and fix it, which is the entire reason the rotation is not attempted at the last minute.

Compare the day that container just had with the alternative it replaces: one API key, generated once, pasted into the configuration of six services, and rotated the next time somebody notices it in a screenshot. In the walkthrough above no two services ever hold the same credential, and none of the services being issued an identity — checkout, ledger, and every workload alongside them — holds a secret that outlives a day.

Be precise about what that buys, because it is easy to overclaim in both directions. The design does not abolish long-lived secrets: the internal CA's own signing key is long-lived, and it is the thing everything else rests on. What the design does is concentrate it — one key, in one place where it can be kept inside a hardware security module, a dedicated device that computes each signature internally and never hands the key back out, under narrow access and real auditing, instead of copies of credentials scattered across every machine that needs to talk to something.

And the 24-hour lifetime bounds what a leaked credential is worth. Note what the leaked thing has to be: the certificate alone is a public document, handed to every peer in every handshake, and it proves nothing without the matching private key — a certificate that gets out on its own is worthless immediately, not tomorrow. What is worth stealing is the pair: the private key sitting inside checkout-7f9c together with its certificate, which can escape together in a disk image, a backup, a crash dump, or a volume mounted somewhere it should not have been. That pair dies with the certificate — within the day, with nobody doing anything — and the thief cannot renew it, because renewal runs over that local socket, on that machine, on the orchestrator's word about the process at the other end. What the lifetime does not bound is an attacker who is running code inside checkout, because the platform goes on attesting that workload and cheerfully renews it — the intruder is standing where each new keypair is generated and receives every fresh certificate exactly as the legitimate process does. Evicting them is a separate job: detect the compromise, stop the workload, and withdraw its identity so the CA stops signing for it.

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send
Interview rung — workload identity & secret distribution (free-form)

Loading exercise…

My notes on this module

Loading your notes...

Workload identity & secret distribution — TransformerLab