Regression & GLMs
P5.ml-methods.03 · Audience: guest, it-ml, language-pro · Prerequisites: Ensembles & Unsupervised Learning
Much of applied data science comes down to fitting a line, yet a plain line happily predicts negative counts and probabilities above one, so it cannot serve every kind of outcome unchanged. Generalized linear models repair this with a single move: keep the linear part, and transform its output through a function chosen to match the outcome's scale — the link — which is how ordinary least squares, count models, and logistic regression turn out to be one family rather than three separate tricks. This module first derives least squares and the assumptions under which it can be trusted, then generalizes through the link function, and finally shows how to read a logistic coefficient as a statement about odds. Each idea ends in a short calculation of the kind an interviewer will expect on a whiteboard, resting on the probability groundwork from the statistics modules of pillar one.
Framing. Linear regression is the identity-link Gaussian GLM; swap the link and the response distribution and you get logistic, Poisson, and the rest. Each step turns the algebra into a number.
1 — OLS and the Gauss-Markov assumptions (s19). OLS minimises squared residuals; setting the gradient to zero gives the closed-form normal equations:
Under the five Gauss-Markov assumptions — linearity, exogeneity, homoskedasticity, no autocorrelation, no perfect multicollinearity — OLS is BLUE (the best linear unbiased estimator). Normality is only needed for exact inference, not unbiasedness.
Worked example — the closed form in code. Fit a slope+intercept to three points with the normal equations:
import numpy as np
X = np.array([[1,0],[1,1],[1,2]]) # intercept col + feature
y = np.array([1.0, 2.0, 3.0])
beta = np.linalg.inv(X.T @ X) @ X.T @ y
# beta = [1., 1.] -> y_hat = 1 + 1*x (a perfect fit)The X.T @ X inverse is the whole solution. If two columns were collinear it would be singular (no inverse) — that's the no-perfect-multicollinearity assumption; ridge's (XᵀX + λI)⁻¹ restores a well-conditioned inverse. Heteroskedasticity wouldn't bias β̂ but would break the standard errors (use robust SEs / WLS).
⚡ Interview Ref — the quick-scan Reference face
Interview-depth regression for the ml-engineer track: OLS and the Gauss-Markov theorem, generalized linear models and the link function, and logistic regression — log-odds, odds ratios, and why we train with cross-entropy rather than MSE.
1 — OLS and the Gauss-Markov assumptions (s19)
Ordinary least squares (OLS) fits a linear model y = Xβ + ε by minimising the sum of squared residuals ‖y − Xβ‖². Setting the gradient to zero gives the normal equations and the closed-form solution:
The Gauss-Markov theorem says that under five assumptions OLS is BLUE — the best linear unbiased estimator (lowest variance among all linear unbiased estimators):
- Linearity — the model is linear in the parameters.
- Exogeneity — errors have zero conditional mean, E[ε ∣ X] = 0.
- Homoskedasticity — constant error variance, Var(εᵢ) = σ².
- No autocorrelation — errors are uncorrelated across observations.
- No perfect multicollinearity — XᵀX is invertible (columns are linearly independent).
(Normality of errors is not needed for BLUE; it is only needed for exact t/F inference.)
Violations and fixes:
- Heteroskedasticity (non-constant variance) → robust / White standard errors or weighted least squares (WLS); coefficients stay unbiased but their SEs are wrong.
- Autocorrelation (correlated errors, common in time series) → Newey-West standard errors or generalized least squares (GLS).
- Endogeneity (E[ε ∣ X] ≠ 0) → instrumental variables (IV); OLS is biased and inconsistent here.
- Multicollinearity → drop redundant features or regularise (ridge shrinks (XᵀX + λI)⁻¹, restoring a well-conditioned inverse).
2 — Generalized linear models and the link function (s48)
A generalized linear model (GLM) extends linear regression to response distributions in the exponential family via a link function g that relates the mean of the response to the linear predictor:
Every GLM has three components:
- Random component — the distribution of y (Normal, Bernoulli, Poisson, …).
- Systematic component — the linear predictor η = Xβ.
- Link function — the invertible g connecting the mean to η; its inverse g⁻¹ (the mean function) maps the linear predictor back to the response scale.
Common choices:
| Distribution | Link | Model |
|---|---|---|
| Normal | identity, g(μ) = μ | ordinary linear regression (OLS) |
| Bernoulli | logit, g(μ) = log(μ / (1 − μ)) | logistic regression |
| Poisson | log, g(μ) = log μ | Poisson regression (counts) |
The identity link with a Normal response recovers OLS exactly, so linear regression is the simplest GLM. GLMs are fit by maximum likelihood, solved numerically with iteratively reweighted least squares (IRLS) rather than a closed form.
3 — Logistic regression: log-odds and coefficients (s25)
Logistic regression is the GLM with a Bernoulli response and a logit link. It models the log-odds (the logit of the success probability p) as a linear function:
The sigmoid σ is the inverse link — it squashes the unbounded linear predictor into a probability in (0, 1).
Interpreting coefficients:
- A coefficient βⱼ is the change in log-odds per unit increase in xⱼ (holding the others fixed).
- Exponentiating gives the odds ratio: e^βⱼ is the multiplicative factor by which the odds change per unit of xⱼ. e^βⱼ > 1 raises the odds, < 1 lowers them, = 1 means no effect.
Why cross-entropy, not MSE: logistic regression is trained by minimising cross-entropy (log-loss), which is the negative Bernoulli log-likelihood:
This loss is convex in β, so gradient descent reaches the global optimum. MSE on the sigmoid is non-convex and produces vanishing gradients in the saturated regions (p → 0 or 1), where σ′(z) = σ(z)(1 − σ(z)) → 0 stalls learning — which is why we use log-loss instead.
Interview one-liners
- OLS closed form β̂ = (XᵀX)⁻¹Xᵀy is BLUE under Gauss-Markov; normality is only needed for exact inference, not unbiasedness.
- Heteroskedasticity breaks the SEs, not the coefficients — fix with robust SEs or WLS.
- A GLM = distribution + linear predictor + link function; identity+Normal is OLS, logit+Bernoulli is logistic, log+Poisson is count regression.
- Logistic regression is linear in the log-odds; e^βⱼ is the odds ratio per unit of xⱼ.
- Train logistic with cross-entropy, not MSE — log-loss is convex; MSE on the sigmoid is non-convex with vanishing gradients.
📚 Go Further
Rigorous references for regression and GLMs.
| Type | Resource |
|---|---|
| Book | Hastie, Tibshirani & Friedman, Elements of Statistical Learning — GLMs |
| Book | Wooldridge, Introductory Econometrics — Gauss-Markov, robust SEs, IV |
| Book | McCullagh & Nelder, Generalized Linear Models — the definitive GLM text |
| In-app | P5.ml-methods.06 — Evaluation, Selection & Interpretability, for calibration of these probability outputs |
| In-app | P1.statistics-advanced.01 — Advanced Probability (S7-adv), for the likelihood foundations |
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in ML Methods
Go up a level