All formulas of the course on one page, in course order. Every block links to the section where it is explained. Terms are in the Glossary, the math behind the formulas in the Math Foundations.

How to read the bounds

Almost every bound has the form true risk empirical risk + a term that shrinks with , and it holds with probability at least . The term decays like (standard rate) or (fast rate). To decide whether a bound is useful, check whether the whole term goes to 0.

Decision theory (Lecture 1.1)

True risk and Bayes risk β†’ True risk, Bayes risk

A function that attains is a Bayes classifier .

Regression function and Bayes classifier for the 0-1 loss β†’ Explicit form

The error at is , the Bayes risk is its expectation.

Posterior from priors and class conditionals β†’ MAP

= prior of class , = class conditional density. Maximum likelihood compares only and and ignores the priors.

Bayes decision rule with costs β†’ Costs of errors

= cost of predicting 0 when , = cost of predicting 1 when . The decision boundaries are the points where the two weighted curves cross (reading a diagram).

Bayes error from the weighted densities

Surrogate losses on a score , labels β†’ Bayes predictors

Loss
hinge
squared
exponential
logistic

All four are classification-calibrated: thresholding their Bayes predictor at 0 gives the 0-1 Bayes classifier.

Squared loss β†’ Regression

Learning from finite samples (Lecture 1.2)

Consistency β†’ Consistency

Universally consistent: for every distribution .

Empirical risk and ERM β†’ Empirical risk, ERM

Estimation and approximation error β†’ Definitions

= the best function in . A larger class raises the first term and lowers the second.

Bias-variance decomposition ( loss, at a point ) β†’ Bias-variance

No Free Lunch β†’ Propositions 7 to 9

On points there are possible true functions. Averaged over all of them, every classifier has the same error.

Perceptron (Lecture 2)

Classifier, update and linear loss β†’ Algorithm, Linear loss

This is SGD with step size 1 on , started at .

Margin β†’ Margin

Mistake bound and test error β†’ Theorem 1, Theorem 2

, margin of a separating . Neither bound depends on the dimension. Not separable, with slack : at most updates (outlook).

Statistical learning theory (Lecture 3)

Hoeffding β†’ Hoeffding’s inequality

McDiarmid β†’ McDiarmid’s inequality

= the most can change when only the -th argument changes.

One fixed function β†’ Fixed function

One-sided: . Sample size for accuracy : of the order .

Uniform convergence controls ERM β†’ Uniform convergence

Finite class with functions β†’ Union bound

Shattering coefficient β†’ Definition, Bound

ERM is consistent if .

VC dimension and Sauer-Shelah β†’ Sauer-Shelah

The growth function is either for all or polynomial.

VC bound β†’ VC bound

ERM is consistent w.r.t. iff . Sample size of the order .

VC dimensions to know β†’ Linear, Margin, Neural networks

ClassVC dimension
linear classifiers in (with offset)
hyperplanes with margin , data in a ball of radius
neural networks with weights

Rademacher complexity β†’ Rademacher

= independent fair coin flips with values .

Algorithmic stability (Lecture 4)

Average and uniform stability β†’ Definitions

= with the -th point replaced. Always .

Expected gap β†’ Proposition 1

Stability bound β†’ Theorem 2

Loss bounded by . The gap vanishes iff (which rates are good enough).

Strong convexity and smoothness β†’ Strongly convex, Smooth

is a lower bound on the curvature, an upper bound.

Stable algorithms β†’ ERM, GD, SGD

AlgorithmAssumptions
ERMloss -strongly convex, -Lipschitz
GD, SGDconvex, -smooth, -Lipschitz,
SGDnon-convex,

Regularization (Lecture 5)

Regularized risk β†’ Principle

Least squares β†’ Full rank, General case

Without full rank every with is also a solution.

Ridge regression β†’ Solution, Stability

is -strongly convex. The stability bound needs a convex, -Lipschitz loss.

Lasso and p-norms β†’ Lasso, p-norms

Convex for . The lasso has no closed form; its solutions are sparse.

Linear model and rates (fixed design, excess risk) β†’ Model, Rates

MethodRateDepends on
OLS, fast in
ridge
lasso-sparse , only
-sparse , fast in

Bagging and boosting (Lecture 6)

Bagging β†’ Bagging

estimates with variance and pairwise correlation .

AdaBoost β†’ Algorithm

Start with uniform weights . Output . gives .

Training error β†’ Theorem 1

Test error β†’ VC bound

The margin bound does not depend on .

Gradient boosting β†’ Gradient boosting

Fit each new base learner to the negative gradient of the loss at the current predictions. AdaBoost is the case of the exponential loss.

Features and validity (Lecture 7)

Kernel β†’ Kernel methods

Random Fourier features β†’ Random features

The features stay fixed and only the output weights are learned. Random ReLU features: .

Four notions of validity β†’ Validity

ValidityQuestion
statisticalis the difference more than chance?
internalis it really caused by the model?
externaldoes it carry over to other data and settings?
constructdoes the metric measure the intended concept?

Overparameterized learning (Lecture 8)

Minimum norm solution (, GD started at ) β†’ Theorem 1

Benign overfitting in the toy setup (true function 0) β†’ Theorem 3

Excess risk in both regimes (isotropic Gaussian inputs) β†’ Theorem 4

Both sides go to at the interpolation threshold .

Why large models β†’ Robust interpolation, NTK

Fairness (Lecture 9.1)

Group criteria β†’ Demographic parity, Equalized odds, Predictive parity

CriterionFormulaIn a confusion table per group
demographic parityequal
equalized oddsequal and
equal opportunityequal true positive ratesequal
predictive parityequal
calibration by group

With different base rates, demographic parity and predictive parity cannot both hold, and neither can equalized odds and predictive parity (impossibility).

Derived classifier (post-processing) β†’ Post-processing

The same formula holds for the false positive rate.

Explainability and law (Lecture 9.2)

SHAP β†’ Definition

= all features. The values add up: .

Value functions β†’ Observational vs. interventional

SHAP for a GAM (interventional) β†’ Theorem 1

EU AI Act β†’ AI Act

Risk categories: prohibited, high, limited, minimal. GPAI models count as systemic risk above FLOPs of training compute.