P37 · Observability — Logging, Metrics & Traces
structured logging, metrics & distributed tracing
Make a service observable - the ask that started this discipline. Build the three pillars from scratch: structured logs you can follow by request, metrics you can alert on, and traces that show where a slow request spent its time.
A user reports that the service was slow at about eleven. You have logs. They are sentences, written by six different people over two years, with no shared field for which request they belong to — so answering a question as simple as "what else happened during that request" means reading by eye and guessing. This pillar is the answer to that situation, and it is the ask that started this whole discipline: make a service you can actually investigate.
Observability is conventionally three things, and you build all three by hand rather than importing them. Structured logging comes first: records made of fields rather than prose, so a log system can filter and aggregate them, with context binding that attaches a service name and a correlation id once and returns a logger whose every subsequent line carries them — which is what turns one request into a single followable story across every component it touched. Redaction sits at the logger itself, so a careless call site cannot leak a secret or a piece of PII.
Metrics and traces are the other two, and the pillar is as interested in when each is wrong as in how each works. A metrics registry — counters, gauges, histograms — gives you the aggregate you can alert on, which a log cannot; tracing instruments a request pipeline with spans so the waterfall shows where the time actually went, which neither logs nor metrics can. The judgement rung is the point of all of it: levels used with intent, awareness that log volume is a real cost, and knowing when a log is the wrong tool because the question is about a rate or a duration rather than an event.