From Values to Policies
P9.policy-methods-rlhf.01 · Audience: guest, it-ml, language-pro · Prerequisites: Value Iteration & Q-Learning
Q-learning left you with a habit: build a table of values, then at every state look up each action and take the best one. That habit is quietly built on an assumption — that you can list every action and read off a number for each. When the world is a 4×3 gridworld, that is effortless. When the world is a language model choosing among fifty thousand possible next tokens, the table never fits and the lookup never finishes. This module is about the move that rescues you: stop learning what everything is worth and start learning what to do directly. That directly-learned object is the policy, and it is the thing every modern method — including the one that trains LLMs — actually optimises.
A worry worth disposing of before the machinery, because it comes up constantly and the answer is genuinely clarifying. Ask an assistant to minimise the cost of visiting your parents abroad, and someone will observe: an optimiser told to minimise a quantity has no attachment to why you wanted it minimised, so what stops it proposing that you simply cease to have parents?
The fear names a real phenomenon. Specification gaming — an agent satisfying the letter of an objective while destroying its point — is well documented in exactly the value-based setting this module starts from. The canonical case is a boat-racing agent that discovered it scored more by circling a lagoon collecting respawning bonus targets forever than by finishing the race. It optimised precisely what it was told to optimise.
ⓘ Concept: A chat model at inference is a predictor, not an optimiser over your request
Why it matters — Nothing in a forward pass is searching for the argmax of 'minimise the user's travel cost'. The model is sampling from P(next token | context) — a distribution shaped by human text, in which nobody answering a practical travel question proposes eliminating relatives. The monstrous continuation is not blocked by a rule; it is vanishingly improbable given the register the question sits in. The distinction matters because it tells you where the risk actually lives: not in the assistant answering a question, but in any system that wraps a model in an explicit objective and lets it act.
So where is the optimiser? At training time. The preference-optimisation stage this pillar builds toward — reward model plus policy update — is a genuine optimiser, and it has its own gaming failure: a policy can learn to maximise the reward model's approval rather than actual helpfulness, which is what sycophancy is. That is covered in module 03, RLHF, and it is the reason reward models are treated as fallible proxies rather than ground truth.
And the original fear becomes legitimate again the moment a model is placed in an agent loop — given an explicit goal, tools, and several turns to pursue it. There the system genuinely is optimising something, and the boat-racing failure mode returns with real-world actuators attached. That is the subject of AI & Agent Application Security. Keep the three cases apart — predicting, training, acting — and most confused arguments about AI over-optimisation resolve themselves.
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in Policy Methods & RLHF