Serving & scaling

P13.operations.01 · Audience: guest, it-ml, language-pro · Prerequisites: Packaging & deployment

Real LLM grading for this pageLLM grading (this page):

A trained, packaged model is worthless until it answers requests fast enough. This module is the serving half of Operations: how to put a latency budget on an endpoint (the URL a model answers on), size the hardware to the traffic with Little's law, and tune batching so throughput rises without the tail latency — the slowest few percent of requests — blowing past your SLO, the service-level objective that states your response-time promise. Real HTTP serving (uvicorn, a GPU box, an autoscaler) lives outside this sandbox — so we teach the behaviour with an in-process request-loop simulator you build yourself.

Step 1 / 4 — Latency budgets & the SLO

A serving SLO is a promise about latency, and the promise is almost never about the average. It is about a percentile: "p99 < 300 ms" means 99% of requests finish inside 300 ms. The p50 (median) tells you the typical experience; the p99 (the tail) tells you the worst experience your busiest 1% of calls actually get.

ⓘ Concept: Why the tail matters more than the mean
One user request often fans out into many backend calls (retrieval, re-rank, generation, a guardrail check). If a single page makes 10 sub-requests, the chance that all of them beat p99 is 0.9910 ≈ 0.90 — so ~10% of pages hit the tail of something. The tail is the experience, not the outlier. A healthy mean with a fat p99 still means angry users.

An end-to-end latency budget is the SLO split across the stages a request passes through. You decide the total first, then hand each stage a slice; the sum of the slices (plus headroom) must fit the budget. If a stage overspends, another must give it back — the budget is a zero-sum contract.

Worked example — a 300 ms p99 budget for a RAG endpoint

stagep99 budget (ms)note
gateway + auth15TLS, routing, quota check
embed query25one forward pass
vector search40ANN recall@k
re-rank50cross-encoder top-k
model inference140the packaged model itself
headroom30queueing + network jitter
total300= the p99 SLO

The model inference slice (140 ms) is where serving optimisations pay off. The KV cache in autoregressive decoding (cross-ref LS-1 — KV cache / serving) keeps per-token cost flat instead of quadratic, which is what makes that slice affordable at all. And the whole exercise of dividing a total into stage slices is the latency-budget discipline (cross-ref SY-1 — latency budget). Notice the explicit headroom line: it exists to absorb queueing delay, which the next steps quantify.

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send

🎓 Practice ladder

3 graded rungs · ~30 min

Build the serving math and simulator yourself: Little's law for capacity sizing, a largest-batch-under-budget selector, and a single-server FIFO queue that turns an arrival stream into a p99. The kernel has no network or GPU, so real HTTP serving stays prose — you implement the behaviour that governs the tail.

Rung 1 — capacity from Little's law

Loading exercise…

Rung 2 — the largest batch under budget

Loading exercise…

Rung 3 — a single-server FIFO queue p99 (senior)

Loading exercise…

⚡ Interview Ref — the quick-scan Reference face
  • SLOs are percentiles — p50 = typical, p99 = tail = the real experience. Fan-out (one page → many calls) makes the tail common: 0.99^10 ≈ 0.90, so ~10% of pages hit some tail.
  • Latency budget = split the SLO across stages (gateway, embed, search, re-rank, inference, headroom); the sum + headroom must fit the total. Overspend in one stage is repaid by another (SY-1). KV cache keeps the inference slice affordable (LS-1).
  • Little's law — L = λ·W; concurrency in flight = arrival rate × avg latency. Sizes the slot/replica count. Latency regression → capacity (and cost) regression at the same traffic.
  • Batching — bigger batch → higher throughput and higher p99. Ship the largest batch whose p99 still fits the budget; never max(throughput) past the SLO. Continuous/in-flight batching automates it.
  • The tail is queueing delay — completion[i] = max(arrival[i], completion[i-1]) + service_time; latency = completion − arrival. Colliding arrivals, not a slow model, build p99. Keep utilisation below saturation.
  • Real serving is out of sandbox — uvicorn + load balancer + autoscaler + GPU replicas is prose; the in-process FIFO sim models the queueing behaviour that governs p99.
📚 Go Further

Where the pieces of this module lead.

TypeResource
In-appP13.ml-lifecycle.02 — the packaged, promoted model this module serves
In-appP13.operations.02 — Monitoring & drift (reserved): watching live p99 and queue depth after ship
In-appP13.llmops.01 — Applied LLMOps: cost/latency routing on top of the serving tier
ConceptLS-1 (KV cache / serving) & SY-1 (latency budget) — the systems background this module applies
DocsvLLM / TGI continuous batching; Little's law for capacity planning; queueing theory (M/M/1) intuition

Try it yourself

A scratch console for this page's ideas — ungraded, nothing you run here is recorded.

Scratch console

A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.

Output appears here.

My notes on this module

Loading your notes...

Where next?

Continue

Monitoring & drift →

Operations · P13

Later in Operations

Serving & scaling — TransformerLab