
Week 2: linear regression
LSE
Part 1: least squares and optimal models
Part 2: multiple regression and orthogonality
How do you travel? With the original GPS: astronomy
. . .
. . .
…
mfw
![]()

Piazzi published his \(n = 24\) observations in February
An international community of scientists and mathematicians scrambled to find Ceres
Almost a year later, it was rediscovered using the predictions of (24 year old) C. F. Gauß
How did Gauss do it? Kepler’s laws determine an orbit uniquely from 3 points. What to do with 24?
When the number of unknown quantities is equal to the number of the observed quantities depending on them, the former may be so determined as exactly to satisfy the latter. But when the number of the former is less than that of the latter, an absolutely exact agreement cannot be determined, in so far as the observations do not enjoy absolute accuracy. In this case care must be taken to establish the best possible agreement, or to diminish as far as practicable the differences.
but why squared errors?
Around the same time, R. J. Boscovich and P-S. Laplace minimized sum of absolute errors: \(\text{minimize} \sum_i | r_i |\)
Laplace also suggested minimizing the maximum error: \(\text{minimize} \max_i | r_i |\)
Gauss said we can use any even power, e.g. \(\sum r_i^8\)
of all these principles ours [least squares] is the most simple; by the others we shall be led into the most complicated calculations
If Gauss didn’t want to do those calculations, that’s really saying something…
On the other hand, he said he used least squares thousands of times in his years of work (without electricity!)
For more about the origin of least squares see this article.
Convenience of calculation enables a lot of bad science
A. Quetelet in 1835, “social physics,” correlates basically any social data together, tries to predict “crime,” poverty, alcohol consumption, etc
F. Galton (1822-1911) founds the field of eugenics…
Modern science: replication crisis
Much of modern ML is similarly fitting curves to model relationships in any available data because we can – not because there is any scientific or theoretical reason to do so
At a minimum of
\[ \ell (\hat \alpha, \hat \beta) = \sum_i (y_i - \hat \alpha - \hat \beta x_i)^2 \]
we have \(0 = \dfrac{\partial \ell}{\partial \alpha} = -2 \sum_i r_i\), i.e. the first constraint is satisfied,
and \(0 = \dfrac{\partial \ell}{\partial \beta} = -2 \sum_i x_i r_i\), i.e. orthogonality.
\[ \text{cor}(x, r) = 0 \]
Correlation measures linear dependence
If we minimized a different loss function and the resulting residuals were correlated with \(x\), this would mean there is some remaining (linear) signal, in the squared-error sense
A (linear) pattern in residuals, i.e. bias
\[ \text{minimize } \frac{1}{n} \sum_{i=1}^n (y_i - \alpha - \beta x_i)^2 \]
\[ \text{minimize } \mathbb E[(Y - \alpha - \beta X)^2] \]
This course is mainly focused on methods that do use probability, and we will always try to do so explicitly/transparently (not hiding our assumptions)
Within probabilistic machine learning, supervised learning is broadly about modeling the conditional distribution of the outcome given the features
\[p_{Y|X}(y|x) = p_{X,Y}(x,y) / p_{X}(x)\]
Some methods try to learn this entire distribution, others focus on some summary/functional, e.g.
conditional expectation
\[ \mathbb E_{Y|X}[Y|X] \]
or conditional quantile
\[ Q_{Y|X}(\tau) \] (for the \(\tau\)th quantile)

Curves: \(p_{Y|X}(y|x)\) at two values of \(x\). Dark curve: \(\mathbb E[Y \mid X = x]\). Simulated data
It can be shown (the theorem in a moment; the proof is yours, before the seminar) that
\[f^\star(x) = \mathbb E_{Y|X}[Y|X = x]\] minimizes the expected squared loss
\[f^\star = \arg \min_g \mathbb E_{X,Y} \{ [ Y - g(X) ]^2 \}\]
Other examples also fit into this broad framework
For a given loss function \(\ell(y, \hat y)\) (last week: squared error), find the optimal “regression” function \(f^\star\) that minimizes the risk, i.e.
\[ R(f) = \mathbb E_{X,Y}\big[\ell(Y, f(X))\big], \qquad f^\star = \arg \min_f R(f) \]
Statistical machine learning:
\[ \mathbb E \longleftrightarrow \frac{1}{n} \sum \]
Empirical Risk Minimization (ERM)
Algorithms can leverage LLN, CLT, subsampling, etc…
Population risk of a predictor \(f\), an expectation over new data, the thing we want small: \[R(f) = \mathbb E\,\ell(Y, f(X))\]
Training error (empirical risk), the same average over the data we fit on: \[\hat R_n(f) = \tfrac{1}{n}\textstyle\sum_{i=1}^n \ell(y_i, f(x_i))\]
Test error, last week’s leaderboard, over held-out data: \[\widehat{\text{Err}} = \tfrac{1}{m}\textstyle\sum_{j=1}^m \ell(y_j^{\text{test}}, \hat f(x_j^{\text{test}}))\]
Squared loss, any function of \(X\) allowed, not only lines:
\(\mathbb E Y^2 < \infty\) and \(f^\star(x) = \mathbb E[Y \mid X = x]\). For every \(f\) with \(\mathbb E f(X)^2 < \infty\), \[ R(f) = R(f^\star) + \mathbb E\big[(f^\star(X) - f(X))^2\big] \]
Squared loss: the target is the conditional mean. Week 1’s \(f^\star\) is \(\mathbb E[Y \mid X]\), and the excess risk of any \(f\), any method, is one number.
Proof: exercise 1 in the notebook, before the seminar. Corollaries: absolute loss, the conditional median (stated, not proved); zero–one loss, next week.
Linear regression is based on an assumption that the conditional expectation function (CEF) is (or can be adequately approximated as) linear
\[ f^\star(x) := \mathbb E_{Y|X}(Y|X) = \beta_0 + \beta_1 X_1 + \cdots + \beta_p X_p \]
(Question: why no \(\varepsilon\) errors in this equation?)
Sometimes this assumption works marvelously
Other times it breaks spectacularly
Often, it’s somewhere in the gray area
Always, always, always remember George Box:
Since all models are wrong the scientist must be alert to what is importantly wrong. It is inappropriate to be concerned about mice when there are tigers abroad.
Relaxing the linearity assumption and using flexible, non-linear models
Specialized methods for high-dimensional linear regression, where there are many predictor variables, possibly even \(p > n\)
Beating other approaches at pure prediction accuracy, trading off simplicity/interpretability for better predictions
Recently, people have started caring more about interpretability again – an emphasis in this course
What is it? Simply a method for using more variables to predict the outcome?
Some math(s) notation
Interpreting the model and its coefficients
Association vs causality

Instead of a regression line, we fit a regression (hyper)plane
Among all possible such planes, find the one minimizing sum of squared errors (represented by vertical lines in ISLR Fig 3.4)
How to find the coefficients? Calculus?

Writing the same thing in various ways
\[y_i = \beta_0 + \beta_1 x_{i1} + \beta_2 x_{i2} + \cdots + \beta_p x_{ip} + \varepsilon_i\]
or using the inner product (of column vectors)
\[y_i = x_i^\top \beta + \varepsilon_i\]
\[ \begin{pmatrix} y_1 \\ y_2 \\ \vdots \\ y_n \end{pmatrix} = \begin{pmatrix} 1 & x_{11} & x_{12} & \cdots & x_{1p}\\ 1 & x_{21} & x_{22} & \cdots & x_{2p}\\ \vdots & \vdots & \vdots & \ddots & \vdots \\ 1 & x_{n1} & x_{n2} & \cdots & x_{np}\\ \end{pmatrix} \begin{pmatrix} \beta_0 \\ \beta_1 \\ \vdots \\ \beta_p \end{pmatrix} + \begin{pmatrix} \varepsilon_1 \\ \varepsilon_2 \\ \vdots \\ \varepsilon_n \end{pmatrix} \]
or \(\mathbf{y} = \mathbf{X} \beta + \mathbf{\varepsilon}\). Note: column of 1’s for intercept term. Sometimes omitted by assuming \(\mathbf y\) and every column of \(\mathbf X\) are already “centered”
We’ll use common conventions in this course
With matrix-vector notation we can always write very simply:
\[ \hat {\mathbf \beta} = (\mathbf X^\top\mathbf X)^{-1}\mathbf X^\top \mathbf y = \mathbf X^\dagger \mathbf y \]
Remember this! It encodes many important facts…
This assumes \(\mathbf X^\top\mathbf X\) to be invertible, i.e. the columns of \(\mathbf X\) have full rank (columns = variables)
Predictions from the linear model:
\[\hat{\mathbf y} = \mathbf {X} \hat{\mathbf \beta} = \mathbf X (\mathbf X^\top\mathbf X)^{-1}\mathbf X^\top \mathbf y = \mathbf H \mathbf y\] if we define
\[\mathbf H = \mathbf X (\mathbf X^\top\mathbf X)^{-1}\mathbf X^\top = \mathbf X \mathbf X^\dagger\]
(the pseudoinverse \(\mathbf X^\dagger\) of the previous slide: \(\hat{\mathbf y} = \mathbf X \hat{\boldsymbol\beta} = \mathbf X \mathbf X^\dagger \mathbf y\))
COOL FACTS about \(\mathbf H\):
We have the loss function
\[L(\mathbf X, \mathbf y, \mathbf \beta) = (\mathbf y - \mathbf X \beta)^\top(\mathbf y - \mathbf X \beta)\]
(just a different way of writing sum of squared errors)
Reach the same conclusion: at a stationary point of \(L\),
\[\mathbf X^\top \mathbf X \hat \beta = \mathbf X^\top \mathbf y\]
Ragnar Frisch and Frederick Waugh, Econometrica, volume 1
Economists were “detrending” time series by hand before regressing one on another. Frisch and Waugh asked: is that the same as putting time in as a regressor?
Exactly the same, they proved. Lovell (1963) generalized it to any set of columns
Why care: it relates multiple regression coefficients to an intuitive univariate regression
\(\mathbf X = [\mathbf X_1 \;\; \mathbf x_2]\), full column rank, the intercept column inside \(\mathbf X_1\). \(\mathbf M_1 = \mathbf I - \mathbf X_1(\mathbf X_1^\top\mathbf X_1)^{-1}\mathbf X_1^\top\) projects onto the orthogonal complement of \(\mathrm{col}(\mathbf X_1)\), so \(\tilde{\mathbf x}_2 = \mathbf M_1 \mathbf x_2\) is the residual of \(\mathbf x_2\) after regressing it on \(\mathbf X_1\). The coefficient on \(\mathbf x_2\) in the regression of \(\mathbf y\) on \(\mathbf X\) is \[ \hat\beta_2 = \frac{\tilde{\mathbf x}_2^\top \mathbf y}{\tilde{\mathbf x}_2^\top \tilde{\mathbf x}_2} \] (and the same number comes from regressing \(\mathbf M_1 \mathbf y\) on \(\tilde{\mathbf x}_2\): residual on residual).
The coefficient on \(x_2\) is the slope of \(y\) on what is left of \(x_2\) after \(\mathbf X_1\) has explained what it can. Fitting to residuals: additive models, boosting and causal ML share the mechanism, not a guarantee.

Project everything onto the orthogonal complement of \(\mathrm{col}(\mathbf X_1)\), the grey plane. There, regress \(\tilde{\mathbf y}\) on \(\tilde{\mathbf x}_2\) through the origin: the slope is \(\hat\beta_2\) and \(\mathbf r\) is the full regression’s residual
The full regression, with the normal equations read column by column: the residual is orthogonal to every column \[ \mathbf y = \mathbf X_1 \hat{\boldsymbol\beta}_1 + \mathbf x_2 \hat\beta_2 + \mathbf r, \qquad \mathbf X_1^\top \mathbf r = \mathbf 0,\ \ \mathbf x_2^\top \mathbf r = 0 \]
Multiply through by \(\mathbf M_1\): \[ \mathbf M_1 \mathbf y = \mathbf M_1 \mathbf X_1 \hat{\boldsymbol\beta}_1 + \mathbf M_1 \mathbf x_2 \hat\beta_2 + \mathbf M_1 \mathbf r \]
\(\mathbf M_1 \mathbf X_1 = \mathbf 0\) and \(\mathbf M_1 \mathbf r = \mathbf r\) (because \(\mathbf r \perp \mathrm{col}(\mathbf X_1)\)), so \[ \mathbf M_1 \mathbf y = \tilde{\mathbf x}_2 \,\hat\beta_2 + \mathbf r \]
\(\mathbf M_1 \mathbf y = \tilde{\mathbf x}_2 \,\hat\beta_2 + \mathbf r\) is a vector equation with one unknown number, \(\hat\beta_2\)
Multiply by \(\tilde{\mathbf x}_2^\top = \mathbf x_2^\top \mathbf M_1\): \[ \tilde{\mathbf x}_2^\top \mathbf M_1 \mathbf y = \tilde{\mathbf x}_2^\top \tilde{\mathbf x}_2\, \hat\beta_2 + \tilde{\mathbf x}_2^\top \mathbf r \]
The last term: \(\tilde{\mathbf x}_2^\top \mathbf r = \mathbf x_2^\top \mathbf M_1 \mathbf r = \mathbf x_2^\top \mathbf r = 0\)
Left side: \(\tilde{\mathbf x}_2^\top \mathbf M_1 \mathbf y = \mathbf x_2^\top \mathbf M_1 \mathbf M_1 \mathbf y\)
\(= \mathbf x_2^\top \mathbf M_1 \mathbf y = \tilde{\mathbf x}_2^\top \mathbf y\), using \(\mathbf M_1^\top = \mathbf M_1 = \mathbf M_1^2\)
Full column rank makes \(\tilde{\mathbf x}_2^\top \tilde{\mathbf x}_2 > 0\), so \[ \hat\beta_2 = \frac{\tilde{\mathbf x}_2^\top \mathbf y}{\tilde{\mathbf x}_2^\top \tilde{\mathbf x}_2} = \frac{\tilde{\mathbf x}_2^\top \tilde{\mathbf y}}{\tilde{\mathbf x}_2^\top \tilde{\mathbf x}_2}, \qquad \tilde{\mathbf y} = \mathbf M_1 \mathbf y \qquad \blacksquare \]
(ESL §3.2.3.)
This week: exercises 1 and 2 at the top of the notebook. The proof of the optimal model theorem, and omitted-variable bias: what the normal equations say when a relevant column is left out
Every week: the notebook opens with a “Before the seminar” section. Do it before you arrive; the seminar’s problems start from it
One of the most commonly used methods, even with more complex ML often compare to regression as a “baseline”
Perhaps the most complex method that is still considered relatively interpretable. But interpretation is actually trickier than most understand! (more on this next week!)
. . .
ST310 Machine Learning