Guided pathThis is part of Understand Transformers & BERTBack to the path

Metric Strategy

P16.ml-literacy.03 · Audience: guest, it-ml, language-pro · Prerequisites: Working with ML: Trust, Questions & Project Value

Every number on a leadership dashboard is a choice standing in for something that could not be measured directly. This module is about the four ways that substitution goes wrong — targets that get gamed, metrics that flatter, headline figures that hide their base rate, and the belief that data speaks for itself.

One mental model. A metric is only a stand-in for the outcome you truly care about. The moment you reward the number, people optimise the number — not the outcome. Most traps here come from forgetting that a metric is a proxy, not the goal itself.

Step 1 / 4Goodhart's law

A metric is a proxy for what you really want — you cannot measure "happy customers" directly, so you pick a stand-in like response time or repeat purchases. Goodhart's law: "when a measure becomes a target, it ceases to be a good measure." Reward a number and people optimise the number — and game it.

Worked example — the handle-time target that wrecked service. A call centre rewards agents for low average handle time. The metric improves; the service gets worse:

What management measuredHow agents hit itWhat actually happened
Handle time (target: under 4 min)hang up early, transfer away, don't fully resolvehandle time down — repeat calls & complaints up

The number looked like success while the thing it stood for — helpful service — got worse. The fix: pair a single north-star metric with a few guardrail metrics that catch the gaming — here, handle time next to customer satisfaction and repeat-contact rate. The guardrail is what stops a gamed number from looking like a win.

🗣️ From a linguist's perspective: Paying translators by the word
Per-word rates are the classic Goodhart trap in language work. The metric is a proxy for effort, so it rewards volume — and quietly penalises the translator who spends an hour finding the one phrase that renders a sentence in six words instead of twenty. Terminology research, register matching, and cutting redundancy all reduce the measured output. The guardrail is the same shape as in any other industry: pair throughput with a quality review, or the measure will select for the wrong work.
⚡ Interview Ref — the quick-scan Reference face

In plain language: how to pick a metric that measures what you actually want, tell metrics that flatter apart from metrics that guide, read past a headline accuracy number, and stay sceptical of numbers that claim to speak for themselves. No maths required.

1 — Choosing the right metric; Goodhart's law (b22)

A metric is a proxy for what you really want — you can't measure "happy customers" directly, so you pick a stand-in like response time or repeat purchases. The stand-in is never the same as the outcome. Goodhart's law: "when a measure becomes a target, it ceases to be a good measure." As soon as a number decides bonuses, headcount, or status, people optimise the number and game it — a call-centre agent hangs up early to hit a handle-time target, hitting the metric while wrecking the service it was meant to protect. The fix: pair a single north-star metric with a few guardrail metrics that catch the gaming. Handle time next to customer satisfaction and repeat-contact rate; signups next to week-four retention.

2 — Vanity vs actionable metrics (b23)

Vanity metrics look impressive but don't guide any decision — total page views, cumulative signups, followers. They only ever go up, so they flatter a report without telling you what to do next. Actionable metrics tie to a specific decision and a lever you can pull — conversion rate by channel tells you where to spend, weekly retention tells you whether the product is sticky, cost per qualified lead tells you which campaign to cut. The test: if this number moved tomorrow, would we do anything differently? If the honest answer is no, it's a vanity metric.

3 — What "90% accurate" does not tell you (b30)

A single accuracy figure hides two things that usually decide whether it's any good. Base rate: if only 1 in 1,000 transactions is fraud, a test that's "90% accurate" can still flag mostly innocent cases — because the rare thing is so rare that even a small error rate produces far more false alarms than real catches. Cost of errors: missing a real fraud (a false negative) may cost far more than investigating a false alarm (a false positive) — or the reverse, for a cancer screen. One number treats both mistakes as equal; the business rarely does. Ask three questions: what kind of errors, at what cost each, and against what baseline?

4 — Why "the data speaks for itself" is a myth (b13)

Data never speaks for itself. Every chart and every number is the end of a chain of choices — what to measure, how to segment, which timeframe to show, what to exclude as an outlier, which comparison to draw. Change any of those and the same underlying data tells a different story: quarterly instead of monthly, average instead of median, one region dropped as "unusual". The framing and the assumptions baked in do the talking — the data just sits there. The senior habit: when someone says the number is objective, ask what decisions were made in producing it? A number you can't trace back to its choices is a conclusion in disguise.

Interview one-liners

  • A metric is a proxyGoodhart's law says once it's a target people game it, so pair a north-star with guardrail metrics.
  • Prefer actionable metrics over vanity metrics: ask if this moved, would we act differently?
  • "90% accurate" hides the base rate and error costs; ask what errors, at what cost, against what baseline.
  • Data never speaks for itself — framing and assumptions do; ask what choices produced the number.
📚 Go Further

Plain-language explainers on metrics, targets, and reading numbers critically.

TypeResource
ConceptGoodhart's law — when a measure becomes a target, it stops being a good measure
BookHow to Lie with Statistics — Darrell Huff, on framing and base rates
In-appGovernance, Ethics & Honest Data — for spotting misleading charts
Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Ctrl/Cmd + Enter to send

Where next?