Resampling & Confidence Intervals

P1.statistics-core.04 · Audience: guest, it-ml, language-pro · Prerequisites: Statistical Inference & Hypothesis Testing

Real LLM grading for this pageLLM grading (this page):

Resampling methods — bootstrap and permutation tests — make no distributional assumptions. They are the gold standard for NLP significance testing.

Step 1 / 5 — Why Resampling? Beyond Parametric Assumptions
🗣️ From a linguist's perspective: resampling as 'asking the data to speak for itself'
A parametric test (t-test, chi-square) assumes the data follow a specific distribution (Normal, Poisson). But BLEU scores, F1 scores, perplexity values, and cosine similarities between sentence embeddings are not normally distributed — they have complex, bounded, often asymmetric distributions that depend on the specific test set. Resampling methods avoid these assumptions entirely: instead of deriving the test statistic's distribution from theory, they compute it empirically by repeatedly subsampling the observed data.
ⓘ Concept: Parametric vs resampling approaches
Key principle: the original sample is the best estimate of the population. Drawing many virtual re-samples from it approximates drawing new samples.

Why it matters — ACL 2024 best-practices guidelines recommend bootstrap confidence intervals and approximate randomisation for all single-system and pairwise NLP comparisons. The paired bootstrap test (S5) is the NLP community standard.

Parametric tests (t-test, z-test)

✅ Simple formulas; work with small n if assumptions hold.
❌ Assume a specific distribution (usually Normal).
❌ Fragile when the assumption is violated.

Resampling tests (bootstrap, permutation)

✅ Make no distributional assumptions.
✅ Work for any statistic (F1, BLEU, cosine similarity, …).
✅ The standard in modern NLP papers.
❌ Computationally more expensive (but cheap on a laptop for NLP).

Our demo: consecutive-sentence cosine similarities. Each corpus sentence becomes a term-frequency vector, and we compute the cosine similarity between each consecutive pair (s₁–s₂, s₂–s₃, …). These 7 values are the base sample that the bootstrap and the permutation test will resample throughout this module.

Computing…

🎓 Going deeper — Resampling Deep Dive

For the quantitatively inclined — why the bootstrap works at all (a proof sketch of bootstrap consistency via Glivenko-Cantelli), and the BCa bias-corrected and accelerated confidence interval, derived term by term.

1 — Bootstrap consistency: proof sketch. The bootstrap works because the empirical distribution F̂ₙ is a good approximation of the true distribution F, and the functional T(F) we care about (the mean, median, correlation, …) varies continuously with F.

ⓘ Concept: Glivenko-Cantelli theorem

Let X₁, …, Xₙ be i.i.d. from F. The empirical CDF F̂ₙ(x) = (1/n) Σᵢ 1[Xᵢ ≤ x] — the fraction of observations at or below x — converges to F uniformly, almost surely, as n → ∞. Not just pointwise: the worst-case gap over all x vanishes.

sup⁡x∈R∣F^n(x)−F(x)∣→a.s.0\sup_{x \in \mathbb{R}} |\hat{F}_n(x) - F(x)| \xrightarrow{\text{a.s.}} 0

Why it matters — This is the foundational result: sampling from F̂ₙ (i.e. resampling with replacement) is asymptotically equivalent to sampling from F itself.

ⓘ Concept: The bootstrap consistency claim

Let θ = T(F) be the parameter of interest, θ̂ = T(F̂ₙ) the plug-in estimator from the observed data, and θ̂* = T(F̂ₙ*) the same functional computed from a bootstrap resample. The claim: for smooth functionals T, the conditional distribution of θ̂* − θ̂ (given the data) mimics the sampling distribution of θ̂ − θ. Proof sketch: (1) by Glivenko-Cantelli, F̂ₙ → F uniformly in probability; (2) if T is Hadamard-differentiable (a functional analogue of differentiability), then T(F̂ₙ) → T(F) in probability by the continuous mapping theorem; (3) the bootstrap distribution of θ̂* tracks the sampling distribution of θ̂ because the bootstrap world treats F̂ₙ as the truth, and F̂ₙ ≈ F. Condition: T must be a continuous (Hadamard-differentiable) functional. The mean, median, quantiles, and variance are consistent; the maximum is not (non-smooth — it is sensitive to the tail of F).

n(θ^∗−θ^)∣Xn→dn(θ^−θ)\sqrt{n}(\hat{\theta}^* - \hat{\theta}) \mid \mathcal{X}_n \xrightarrow{d} \sqrt{n}(\hat{\theta} - \theta)

Why it matters — This is the licence for every bootstrap CI: the spread of θ̂* around θ̂ (which we can simulate) mimics the spread of θ̂ around the unknown θ (which we cannot).

Empirical check on the demo corpus. Bootstrap the mean of the 8 sentence lengths with B = 2 000 and compare the standard deviation of the bootstrap means against the analytic SE = s/√n from S3: the two agree closely — resampling from F̂ₙ really does capture the variance of the sampling distribution of x̄.

2 — BCa confidence interval: derivation. The percentile bootstrap CI uses raw quantiles of θ̂*. The BCa (bias-corrected and accelerated) CI adjusts those quantiles to correct for median bias and skewness in the bootstrap distribution.

ⓘ Concept: Bias correction term z₀

z₀ measures how far the median of θ̂* is from θ̂ on the standard normal scale. If the bootstrap distribution is symmetric around θ̂, exactly half of the θ̂*_b are below θ̂, giving z₀ = Φ⁻¹(0.5) = 0. If z₀ > 0, the bootstrap median is below θ̂ (an upward bias correction is needed); if z₀ < 0, the bootstrap median is above θ̂.

z0=Φ−1 ⁣(#{θ^b∗<θ^}B)z_0 = \Phi^{-1}\!\left(\frac{\#\{\hat{\theta}^*_b < \hat{\theta}\}}{B}\right)

Why it matters — z₀ converts 'what fraction of bootstrap replicates fell below the point estimate' into a shift on the standard normal scale — the first of BCa's two dials.

ⓘ Concept: Acceleration constant a (jackknife estimate)

The acceleration a measures the rate at which the standard error of θ̂ changes with the true parameter. It is estimated by the jackknife: θ̂₍₋ᵢ₎ is the estimate computed with the i-th observation left out, and θ̄₍·₎ = (1/n) Σᵢ θ̂₍₋ᵢ₎ is the jackknife mean. a = 0 when the SE is constant (e.g. for the mean); a ≠ 0 for asymmetric statistics (e.g. variance, correlation).

a=∑i=1n(θ^(−i)−θˉ(⋅))36[∑i=1n(θ^(−i)−θˉ(⋅))2]3/2a = \frac{\sum_{i=1}^n (\hat{\theta}_{(-i)} - \bar{\theta}_{(\cdot)})^3}{6\left[\sum_{i=1}^n (\hat{\theta}_{(-i)} - \bar{\theta}_{(\cdot)})^2\right]^{3/2}}

Why it matters — a is essentially a skewness of leave-one-out estimates — it captures how fast the standard error of θ̂ changes with the parameter itself, BCa's second dial.

ⓘ Concept: BCa adjusted quantiles

For a (1 − α) confidence interval, compute the adjusted quantile levels α₁ and α₂, where z_p = Φ⁻¹(p) is the standard normal quantile. The BCa CI is then [θ̂*₍α₁₎, θ̂*₍α₂₎] — quantiles of the bootstrap distribution read at the adjusted levels instead of the raw 2.5 % and 97.5 %.

α1=Φ ⁣(z0+z0+zα/21−a(z0+zα/2)),α2=Φ ⁣(z0+z0+z1−α/21−a(z0+z1−α/2))\alpha_1 = \Phi\!\left(z_0 + \frac{z_0 + z_{\alpha/2}}{1 - a(z_0 + z_{\alpha/2})}\right), \quad \alpha_2 = \Phi\!\left(z_0 + \frac{z_0 + z_{1-\alpha/2}}{1 - a(z_0 + z_{1-\alpha/2})}\right)

Why it matters — When z₀ = 0 and a = 0 these collapse to α/2 and 1 − α/2 — the BCa CI reduces exactly to the simple percentile CI. The corrections only engage when the bootstrap distribution is biased or skewed.

Numeric example — mean sentence length (n = 8, B = 2 000). Running the full BCa recipe on the demo corpus sentence lengths gives z₀ ≈ 0 and a ≈ 0 — the sample mean is a symmetric, linear statistic — so the BCa and percentile CIs come out nearly identical. BCa corrections are more substantial for skewed statistics like the variance or a correlation coefficient.

Computing…

Self-test — 5 questions

Self-check

State the Glivenko-Cantelli theorem. How does it justify using the bootstrap?

Self-check

What property of the functional T(F) is required for bootstrap consistency?

Self-check

Define the BCa bias correction term z₀. What does z₀ = 0 imply?

Self-check

How is the acceleration constant a computed from the jackknife? When is a ≈ 0?

Self-check

Give one example of a statistic where BCa outperforms the simple percentile CI, and explain why.

📝 Check Your Understanding

Self-check

Q1. Name one advantage of bootstrap resampling over a parametric t-test for small NLP datasets.

Self-check

Q2. In bootstrap resampling from a dataset of 5 samples, can the same sample appear more than once in a single bootstrap sample?

Self-check

Q3. A 95 % bootstrap CI for cosine similarity is [0.72, 0.91]. What does this interval imply about the true similarity?

Self-check

Q4. In a permutation test comparing two models, what does the null hypothesis assume?

Self-check

Q5. A permutation test gives p = 0.02 with B = 1000 permutations. How many permuted test statistics were at least as extreme as the observed one?

Self-check

Module Review. Why does bootstrap sampling with replacement sometimes produce duplicate observations, and why is this intentional?

📚 Go Further

Curated resources to explore resampling methods further.

TypeResource
VideoStatQuest with Josh Starmer — Bootstrapping (YouTube)
ReferenceAllen Downey, Think Stats 2nd ed. Ch. 9: Hypothesis Testing — bootstrap section (free online at greenteapress.com)
ReferenceJames et al., An Introduction to Statistical Learning Ch. 5.2: The Bootstrap (free PDF at statlearning.com)
ExerciseScikit-learn — resample utility tutorial (sklearn.utils.resample)
Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send

🎓 Practice ladder

6 graded rungs · ~59 min

Time to write the resampling code yourself. Each rung is a three-panel workspace: instructions on the left, a code editor in the middle, and output + test results on the right. Run checks the visible tests; Submit grades against hidden edge cases; the reference solution unlocks once you pass. The explorer rung needs no code — one knob against simulation noise; the senior rung rules on four plans where the bootstrap quietly fails.

Rung 1 — Percentile bootstrap by hand

Loading exercise…

Rung 2 — Bootstrap vs t-interval

Loading exercise…

Rung 3 — When the CI misleads

Loading exercise…

Rung 4 — BCa corrections by hand

Loading exercise…

Explorer rung — Resample until the interval stops wobbling

Loading exercise…

Senior rung — When not to bootstrap

Loading exercise…

Try it yourself

A scratch console for this page's ideas — ungraded, nothing you run here is recorded.

Scratch console

A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.

Output appears here.

My notes on this module

Loading your notes...

Resampling & Confidence Intervals — TransformerLab