P13 · Production ML & MLOps
taking models to production: deployment, monitoring, LLMOps
Taking models to production and keeping them alive: packaging, serving, monitoring, drift, and the LLM-era operational playbook — with real MLflow and real cases.
The model is finished. It scored well, the notebook is tidy, and the hard part is over — which is the belief this pillar exists to correct. A model in production is not an artefact, it is a service with a long life: it has to be reproducible when someone asks which version produced last quarter's numbers, it has to answer within a latency budget under real traffic, and it has to be noticed when the world it was trained on quietly stops resembling the world it now sees.
Four tracks follow that life. ML Lifecycle builds the reproducibility spine — experiment tracking through to a model registry, exercised on real MLflow rather than a diagram — so that "which model is this?" always has an answer. Operations treats the model as a service: serving, latency, queues, and the mathematics of p99 done honestly, since an average latency hides exactly the requests your users complain about. Drift belongs here too, because a model that was correct in March and is quietly wrong in September has not failed in any way your tests will report. LLMOps covers what generative systems added to the job — routing, cost control, evaluation in production, and failure modes that did not exist for a classifier. Production Cases then runs the whole discipline end to end on real models, including rescuing ones that are already in trouble.