ST310 course page

A core set of commonly used machine learning tools, practiced in R, with attention to their mathematical foundations, their statistical limitations, and what they do and do not tell you.

Readings and lectures are complementary so do both, and bring a laptop with R and RStudio to the seminar (see about for setup). The seminar notebook for each week is linked below; the .qmd link downloads the source file to open in RStudio, the other link is the same notebook rendered so you can read it in a browser.

Week 1: prediction competitions and nearest neighbors

What machine learning is and where it came from: two lineages, one field, and the prediction competition (the common task framework) as its working definition, with the Netflix prize as the cautionary tale. A first prediction task, the rule that a prediction may use only what is available at prediction time, a held-out set as the score, and \(k\)-nearest neighbors as the first method. The bias–variance decomposition derived for \(k\)-NN at the board, and why nearest neighbors stop working in many dimensions.

Reading

  • ISLR Chapter 1 (Introduction), pp. 1–14. A quick read; don’t skip the section on notation and simple matrix algebra.
  • ISLR Chapter 2 (Statistical Learning), sections 2.1 and 2.2, pp. 15–42. If R is new to you, also do the lab in section 2.3, pp. 42–52.
  • PPA Chapter 1 (Introduction), pp. 1–10.

If you finish these quickly you may want to start Chapter 3, because it’s a long one.

Lecture: slides

Seminar: notebook (.qmd)

Week 2: linear regression

Regression as the first fitted model rather than a review: least squares and the hat matrix, and the conditional mean as the target that every regression method is aiming at (the optimal model theorem). Multiple regression and what a coefficient holds fixed: the partialling-out theorem, derived at the board and reused in weeks 8 and 10, and omitted variables as its failure case.

Reading

  • ISLR Chapter 3 (Linear Regression), sections 3.1–3.3 and 3.5, pp. 61–110. Mostly review from your regression course; read 3.2.1 (estimating the coefficients in multiple regression) and 3.3.3 (potential problems) slowly. The lab in 3.6 goes with the seminar.
  • PPA Chapter 2 (Fundamentals of prediction), the sections “Modeling knowledge” and “Prediction via optimization”, pp. 13–21.

Optional, deeper: ESL section 3.2, pp. 44–57.

Lecture and seminar: released here at the start of the week. Bring a dataset of your own if you have one: a numeric outcome and a few predictors, one row per unit.

Problem set 1 is released this week on Moodle and is due in week 5.

Week 3: classification, and causality (part 1)

Classification: logistic regression, a threshold and the two kinds of error on either side of it, and the metrics that follow. Causal block 1: prediction against intervention, a world in which the causation runs the other way from the correlation, and why a model that predicts well can still give the wrong answer to a what-if question.

Reading

  • ISLR Chapter 4 (Classification), sections 4.1–4.3, pp. 130–141. Section 4.4 (LDA, QDA, naive Bayes) is not part of this course; skim 4.5, pp. 158–164, for the comparison with \(k\)-nearest neighbors.
  • PPA Chapter 2, “Types of errors and successes” and “The Neyman–Pearson lemma”, pp. 21–26, and “Decisions that discriminate”, pp. 26–31.
  • PPA Chapter 9 (Causality), “The limitations of observation”, pp. 179–181. Optional now and needed in week 8: “Confounding”, pp. 189–192.

Lecture and seminar: released here at the start of the week. Data: the Palmer penguins; the (fictional) attrition data as an alternative.

Week 4: optimization and model fitting

How a model actually gets fitted: gradient descent, the quadratic case, step sizes and what goes wrong with them, then stochastic gradient descent and why its steps are unbiased. Empirical risk minimization as the common form of every method so far, the bias–variance decomposition in that form, and early stopping as a first kind of regularization.

Reading

  • PPA Chapter 5 (Optimization), pp. 70–97: gradient descent, the quadratic case, stochastic gradient descent and its analysis, regularization. The two convergence bounds stated in the lecture are proved here.
  • ISLR Chapter 6, section 6.1, pp. 227–237 (subset selection). Section 6.2, pp. 237–252 (ridge and the lasso), is week 7’s; reading week is a good time for it.

Lecture and seminar: released here at the start of the week. You will write gradient descent yourself, once, and use it twice. Data: gapminder and the penguins, then simulated data.

Week 5: validation and honest evaluation

Why training error flatters a model (the optimism theorem), the standard error of a held-out score, and cross-validation written out. Then the ways an honest-looking evaluation goes wrong: a leak between training and evaluation data, two leaderboards that disagree, and a score computed on data the model will never meet again (in-distribution against out-of-distribution).

Reading

  • ISLR Chapter 5, section 5.1, pp. 197–209 (the validation set, leave-one-out and \(K\)-fold cross-validation, and the bias–variance trade-off in choosing \(K\)). Section 5.2, the bootstrap, is not part of this course.
  • PPA Chapter 6, pp. 100–106 (the generalization gap; what happens empirically with over-parameterized models), and Chapter 8, pp. 146–148 (the holdout method and Kaggle’s secret holdout).

Lecture and seminar: released here at the start of the week. Cross-validation written out, with and without a leak; a competition run cleanly; a held-out score under three kinds of shift. Data: the (fictional) attrition outcome, the penguins, simulated data.

Problem set 1 is due this week.

Week 6: reading week

No lecture or seminar. Read ISLR section 6.2, pp. 237–252 (ridge regression and the lasso), for week 7.

Week 7: regularization and high-dimensions

Many predictors: ridge regression and the lasso, their bias and variance against the tuning parameter (the bias–variance decomposition in its third form, with ridge written in the SVD basis), soft thresholding, and what changes when \(p > n\). Then the held-out protocol under stress: the same evaluation set used many times, adaptive overfitting, and what the CIFAR-10 and ImageNet replications found.

Reading

  • ISLR Chapter 6, section 6.2, pp. 237–252 (ridge regression and the lasso, choosing the tuning parameter), and section 6.4, from p. 261 (what changes when \(p > n\)). The lab in section 6.5 goes with the seminar. Section 6.3 (principal components and partial least squares), which sits between them, is not part of this course.
  • PPA Chapter 8, pp. 158–167 (adaptivity: what happens when the same held-out set is used many times, and the CIFAR-10 and ImageNet replications).

Deeper, optional: ESL section 3.4, pp. 61–79 (ridge in the SVD basis; the lasso); CASI Chapter 7 (James–Stein) and Chapter 16 (the lasso).

Lecture and seminar: released here at the start of the week. Bias and variance of ridge and the lasso against \(\lambda\), by simulation; tuning with glmnet on the Hitters data with a test set reserved first; the evaluation set reused, and what it does to a score. Data: simulated, and ISLR’s Hitters (baseball salaries, 1987).

Week 8: additive models and interpretation

One curve per predictor: how an additive model is fitted (each component to what the others leave), why it escapes the curse of dimensionality, and when it is wrong. Then a curve on any fitted model: partial dependence and ICE, computed by hand, and the question they answer. Causal block 2: direct against total dependence on a simulation where the two curves disagree, and when a model’s response is also an effect.

Reading

  • ISLR Chapter 7, section 7.7, pp. 306–311 (generalized additive models). The lab in section 7.8.3 goes with the seminar.
  • ESL Chapter 9, the opening pages of section 9.1, pp. 295–298 (additive models and the backfitting algorithm, Algorithm 9.1).
  • IML, the chapters “Individual conditional expectation (ICE)” and “Partial dependence plot (PDP)”: the definitions and the pictures, for recognition.

Deeper, optional: Hastie and Tibshirani, Generalized Additive Models (1990), Chapter 4; Zhao and Hastie, “Causal interpretations of black-box models”, Journal of Business and Economic Statistics 2021.

Lecture and seminar: released here at the start of the week. Backfitting in ten lines against mgcv on the movies data, a random split against a split in time; partial dependence and ICE by hand on a simulation whose truth is known, against the total-effect curve. Data: IMDb films via ggplot2movies, and a supplied simulation. Problem set 2 is released this week.

Week 9: trees and ensembles

A machine that asks the data one question at a time: how a split is found, why a tree is local averaging with the neighborhoods chosen by the data, and how big to grow one. Then two ways to use many trees: bagging and random forests, with the variance of an average of correlated predictions derived at the board and the out-of-bag error as a built-in estimate of held-out error; boosting as gradient descent one small tree at a time, with its three knobs. Permutation importance in the loss-difference convention, and no free lunch retrieved on four entrants.

Reading

  • ISLR Chapter 8, sections 8.1 and 8.2, pp. 327–353 (trees; bagging, random forests and boosting). The lab in section 8.3 goes with the seminar.
  • IML, the chapter “Permutation feature importance”: the definition and the pictures, for recognition.
  • ESL Chapter 15, section 15.4, pp. 597–602 (why random forests work), and Chapter 10, sections 10.9–10.10, pp. 353–361 (boosting trees as gradient descent), for the two board derivations.

Deeper, optional: CausalML Chapter 8, “Predictive inference via modern nonlinear regression”, pp. 200–208; Breiman, “Random forests”, Machine Learning 2001; Friedman, “Greedy function approximation: a gradient boosting machine”, Annals of Statistics 2001.

Lecture and seminar: released here at the start of the week. Split search implemented with running sums, then a recursive tree; training against test error along the pruning path; bagging by hand with an out-of-bag error, random forests by mtry, boosting by hand against xgboost; permutation importance by hand, on three inputs and then five. Data: the penguins and gapminder (note). Problem set 2 is due next week.

Week 10: explaining, intervening, shifting

What an explanation computes: SHAP shown and explained in three sentences; one theorem saying what every partial dependence, ICE and SHAP number is (the model’s response with some inputs set and the rest left where one unit had them), and when that is also an effect; four methods, four questions, one takeaway. Causal block 3: one effect estimated with machine learning in the nuisances, residual on residual with cross-fitting, and what that does not do. Then the data move: covariate shift, reweighting and its condition, and the shift no reweighting fixes.

Reading

  • IML, the chapters “Shapley values”, “SHAP” and “Leave one feature out (LOFO) importance”: the plots and the vocabulary, for recognition.
  • PPA Chapter 10, pp. 204–217 (causal inference in practice: adjustment, double machine learning in a page, quasi-experiments) and Chapter 13, pp. 268–271 (the epilogue: when patterns stop guiding action).
  • Zhao and Hastie, “Causal interpretations of black-box models”, Journal of Business and Economic Statistics 2021 (the back-door statement in the theorem).

Deeper, optional: Chernozhukov et al., “Double/debiased machine learning for treatment and structural parameters”, Econometrics Journal 2018, sections 1–2; Shimodaira, “Improving predictive inference under covariate shift by weighting the log-likelihood function”, Journal of Statistical Planning and Inference 2000.

Lecture and seminar: released here at the start of the week. SHAP by hand on a three-unit table, then by software on the penguins forest with the identity checked; permutation importance against leave-one-covariate-out on one fit; one effect estimated by residual on residual with cross-fitting, two wrong versions and a hidden confounder; a shift diagnosed on gapminder 1952 against 2007. Data: penguins, gapminder, and supplied simulations. Problem set 2 is due this week.

Week 11: what does winning mean?

The competition re-read with everything since: outcome reasoning against model reasoning, the four MARA(s) questions a deployment has to answer and for whom, and the seven recurring flaws of systems that predict something about a person and decide about them. The week-1 task re-run with every entrant the term taught you. One unfamiliar case diagnosed live. The theorems and results in one view. If time allows, the last method: a neural network as composition, fitted by hand.

Reading

Delivered in lecture, for reference rather than preparation:

  • Rodu and Baiocchi, “When black box algorithms are (not) appropriate”, Observational Studies 2023 (arXiv 2001.07648): the common task framework, outcome and model reasoning, MARA(s).
  • Wang, Kapoor, Barocas and Narayanan, “Against predictive optimization”, ACM Journal on Responsible Computing 2024, and the rubric at predictive-optimization.cs.princeton.edu: the project’s justification section uses it.

Optional: ISLR Chapter 10, sections 10.1–10.2, pp. 404–411 (single- and multilayer neural networks), for the optional segment; ISLR Chapter 12, section 12.1, from p. 497, for the reader who wants to see what unsupervised learning is.

Lecture and seminar: released here at the start of the week, together with the theorems in one view (statements only, with the week each was proved or derived; bring it to the seminar and try to reproduce each statement from memory first, checking second). One unfamiliar case on paper; a simulated lender in R on which a partial dependence curve, an intervention claim and a SHAP value are checked against a known truth; your own project against MARA(s) and the rubric; problem set 2 retrieved. Data: a supplied simulation, and your project data.