Communicating Uncertainty (step by step)
P1.stats-literacy.05 · Audience: guest, it-ml, language-pro · Prerequisites: Reading Numbers: Averages, Samples & Signal (step by step)
Take it one idea at a time: how to read the numbers that come with a claim — significance, margins of error, confidence intervals, forecasts, expected value, and probabilities — so you can question them well and communicate them honestly. Each step has a worked example. No maths required.
One mental model. Almost no number is exact. The useful question is never just 'what's the figure?' but 'how sure are we, and what's the range?' Every step below is one way to ask that question.
A p-value answers one narrow question: if there were really no effect at all, how surprising would a result this big be? A small p-value means the data would be unlikely in a world where nothing was going on. It is not the probability the hypothesis is true, and 'significant' does not mean large.
Worked example — 'significant' but worthless
A test on 2 million users finds a checkout tweak lifts conversion:
| What the report says | What it means for you |
|---|---|
| p < 0.001 — 'highly significant!' | very unlikely to be pure chance |
| Effect size: +0.02% | 2 extra sales per 10,000 — probably not worth the work |
With a huge sample, a tiny, meaningless difference easily clears the significance threshold. So when someone says 'it's significant', ask 'how big is the effect, and does it matter?' — not just whether it passed a threshold.
⚡ Interview Ref — the quick-scan Reference face
In plain language: how to read the numbers that come with a claim — significance, margins of error, confidence intervals, forecasts, expected value, and probabilities — so you can question them well and communicate them honestly. No maths required.
One mental model. Almost no number is exact. The useful question is never just 'what's the figure?' but 'how sure are we, and what's the range?' Every section below is one way to ask that question.
1 — p-values and "significant" in plain terms (b2)
A p-value answers one narrow question: if there were really no effect at all, how surprising would a result this big be? A small p-value means the data we saw would be unlikely in a world where nothing was going on — so something may really be happening.
Two traps to avoid:
- A p-value is not the probability that the hypothesis is true. It only describes how surprising the data would be assuming no effect.
- 'Statistically significant' does not mean large or important. With enough data, a tiny, meaningless difference can be 'significant'.
So when someone says 'it's significant', ask 'how big is the effect, and does it matter?' — not just whether it passed a threshold.
2 — Never trust a single number without its uncertainty (b5)
A single point estimate — '12%', 'a 3-day gain', 'satisfaction is 4.2' — hides how much it could wobble if you measured again. The number alone tells you almost nothing about how reliable it is.
Always ask for the margin of error or the plausible range around it. Consider:
- '12%' sounds precise and settled.
- '12% ± 5' tells a very different story — the truth could sit anywhere from about 7% to 17%.
Same headline, very different decisions. A number without its margin of error is only half the information — treat it as a starting point, not a fact.
3 — How to read a confidence interval in a report (b6)
A confidence interval is simply the plausible range for the true value, given the data. 'Between 8% and 16%' means the real figure most likely sits somewhere in that band — not exactly at the midpoint.
How to read one sensibly:
- A wide interval means 'we're not sure' — treat the finding as tentative.
- A narrow interval means the estimate is fairly tight and stable.
- Focus on the whole range, not just the midpoint everyone quotes.
The decision-maker's question: would the two ends of the range lead to different actions? If '8%' and '16%' would make you do different things, the uncertainty is not a footnote — it's the story.
4 — Interpreting a forecast and its uncertainty (b24)
A forecast is a range of scenarios, not a single certain number. 'We'll finish in June' is really shorthand for 'probably June, maybe May, possibly August'. The headline figure is just the middle of a spread.
When you receive a forecast, ask for the shape of the range:
- a best case (things go well),
- an expected case (the realistic middle), and
- a worst case (things go badly).
Then plan against the range, not the headline point. A plan that only survives the best-case number is a plan waiting to fail. Communicate forecasts the same way — give people the spread, not a false sense of precision.
5 — Expected value as a decision tool (b33)
To compare options under uncertainty, weigh each possible outcome by how likely it is. That weighted average is the expected value — a single yardstick that blends how good an outcome is with how probable it is.
It's a powerful tool, but not the whole picture. Also weigh the spread and the worst case:
- A bet with a good expected value can still be unacceptable if one possible outcome is catastrophic or irreversible.
- 'On average it pays off' is no comfort if a bad draw wipes you out — the risk of ruin trumps the average.
So use expected value to rank options, then sanity-check: could the downside be something we can't survive?
6 — What a "70% chance" means; judging forecasters (b38)
A '70% chance' is not a promise and not a hedge — it's a claim about frequency. A good forecaster is calibrated: across all the things they call '70% likely', about seven in ten actually happen.
This changes how you judge them:
- Judge a forecaster over many predictions — a track record — not a single outcome.
- One miss on a '70%' call is not being wrong; those events are supposed to fail about three times in ten. Being calibrated is the real test.
- Be suspicious of anyone who is always '99% sure' — that's overconfidence, not skill.
The takeaway: outcomes judge luck; calibration judges skill. Watch the pattern over time, not the last coin flip.
Interview one-liners
- A p-value says 'how surprising would this be if nothing were going on?' — not the probability the hypothesis is true, and 'significant' doesn't mean 'big'.
- Never trust a lone number — ask for the margin of error and read the confidence interval as a range; decide on the range, not the midpoint.
- Treat forecasts as a spread and use expected value to rank options — but a calibrated forecaster and an eye on catastrophic downside matter more than any single number.
📚 Go Further
Plain-language explainers on uncertainty, probability, and reading numbers well.
| Type | Resource |
|---|---|
| Book | The Signal and the Noise — Nate Silver, on calibration and forecasting |
| Book | How to Lie with Statistics — Darrell Huff, on questioning a lone number |
| Idea | Search 'confidence interval' and 'expected value' for gentle visual explainers |
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.