All formulas of the course on one page, in course order. Every block links to the section where it is explained. Terms are in the Glossary, the math behind the formulas in the Math Foundations.
How to read the bounds
Almost every bound has the form true risk empirical risk + a term that shrinks with , and it holds with probability at least . The term decays like (standard rate) or (fast rate). To decide whether a bound is useful, check whether the whole term goes to 0.
Decision theory (Lecture 1.1)
True risk and Bayes risk β True risk, Bayes risk
A function that attains is a Bayes classifier .
Regression function and Bayes classifier for the 0-1 loss β Explicit form
The error at is , the Bayes risk is its expectation.
Posterior from priors and class conditionals β MAP
= prior of class , = class conditional density. Maximum likelihood compares only and and ignores the priors.
Bayes decision rule with costs β Costs of errors
= cost of predicting 0 when , = cost of predicting 1 when . The decision boundaries are the points where the two weighted curves cross (reading a diagram).
Bayes error from the weighted densities
Surrogate losses on a score , labels β Bayes predictors
| Loss | |
|---|---|
| hinge | |
| squared | |
| exponential | |
| logistic |
All four are classification-calibrated: thresholding their Bayes predictor at 0 gives the 0-1 Bayes classifier.
Squared loss β Regression
Learning from finite samples (Lecture 1.2)
Consistency β Consistency
Universally consistent: for every distribution .
Empirical risk and ERM β Empirical risk, ERM
Estimation and approximation error β Definitions
= the best function in . A larger class raises the first term and lowers the second.
Bias-variance decomposition ( loss, at a point ) β Bias-variance
No Free Lunch β Propositions 7 to 9
On points there are possible true functions. Averaged over all of them, every classifier has the same error.
Perceptron (Lecture 2)
Classifier, update and linear loss β Algorithm, Linear loss
This is SGD with step size 1 on , started at .
Margin β Margin
Mistake bound and test error β Theorem 1, Theorem 2
, margin of a separating . Neither bound depends on the dimension. Not separable, with slack : at most updates (outlook).
Statistical learning theory (Lecture 3)
Hoeffding β Hoeffdingβs inequality
McDiarmid β McDiarmidβs inequality
= the most can change when only the -th argument changes.
One fixed function β Fixed function
One-sided: . Sample size for accuracy : of the order .
Uniform convergence controls ERM β Uniform convergence
Finite class with functions β Union bound
Shattering coefficient β Definition, Bound
ERM is consistent if .
VC dimension and Sauer-Shelah β Sauer-Shelah
The growth function is either for all or polynomial.
VC bound β VC bound
ERM is consistent w.r.t. iff . Sample size of the order .
VC dimensions to know β Linear, Margin, Neural networks
| Class | VC dimension |
|---|---|
| linear classifiers in (with offset) | |
| hyperplanes with margin , data in a ball of radius | |
| neural networks with weights |
Rademacher complexity β Rademacher
= independent fair coin flips with values .
Algorithmic stability (Lecture 4)
Average and uniform stability β Definitions
= with the -th point replaced. Always .
Expected gap β Proposition 1
Stability bound β Theorem 2
Loss bounded by . The gap vanishes iff (which rates are good enough).
Strong convexity and smoothness β Strongly convex, Smooth
is a lower bound on the curvature, an upper bound.
Stable algorithms β ERM, GD, SGD
| Algorithm | Assumptions | |
|---|---|---|
| ERM | loss -strongly convex, -Lipschitz | |
| GD, SGD | convex, -smooth, -Lipschitz, | |
| SGD | non-convex, |
Regularization (Lecture 5)
Regularized risk β Principle
Least squares β Full rank, General case
Without full rank every with is also a solution.
Ridge regression β Solution, Stability
is -strongly convex. The stability bound needs a convex, -Lipschitz loss.
Lasso and p-norms β Lasso, p-norms
Convex for . The lasso has no closed form; its solutions are sparse.
Linear model and rates (fixed design, excess risk) β Model, Rates
| Method | Rate | Depends on |
|---|---|---|
| OLS | , fast in | |
| ridge | ||
| lasso | -sparse , only | |
| -sparse , fast in |
Bagging and boosting (Lecture 6)
Bagging β Bagging
estimates with variance and pairwise correlation .
AdaBoost β Algorithm
Start with uniform weights . Output . gives .
Training error β Theorem 1
Test error β VC bound
The margin bound does not depend on .
Gradient boosting β Gradient boosting
Fit each new base learner to the negative gradient of the loss at the current predictions. AdaBoost is the case of the exponential loss.
Features and validity (Lecture 7)
Kernel β Kernel methods
Random Fourier features β Random features
The features stay fixed and only the output weights are learned. Random ReLU features: .
Four notions of validity β Validity
| Validity | Question |
|---|---|
| statistical | is the difference more than chance? |
| internal | is it really caused by the model? |
| external | does it carry over to other data and settings? |
| construct | does the metric measure the intended concept? |
Overparameterized learning (Lecture 8)
Minimum norm solution (, GD started at ) β Theorem 1
Benign overfitting in the toy setup (true function 0) β Theorem 3
Excess risk in both regimes (isotropic Gaussian inputs) β Theorem 4
Both sides go to at the interpolation threshold .
Why large models β Robust interpolation, NTK
Fairness (Lecture 9.1)
Group criteria β Demographic parity, Equalized odds, Predictive parity
| Criterion | Formula | In a confusion table per group |
|---|---|---|
| demographic parity | equal | |
| equalized odds | equal and | |
| equal opportunity | equal true positive rates | equal |
| predictive parity | equal | |
| calibration by group |
With different base rates, demographic parity and predictive parity cannot both hold, and neither can equalized odds and predictive parity (impossibility).
Derived classifier (post-processing) β Post-processing
The same formula holds for the false positive rate.
Explainability and law (Lecture 9.2)
SHAP β Definition
= all features. The values add up: .
Value functions β Observational vs. interventional
SHAP for a GAM (interventional) β Theorem 1
EU AI Act β AI Act
Risk categories: prohibited, high, limited, minimal. GPAI models count as systemic risk above FLOPs of training compute.
Related
- Course: Overview Β· Glossary Β· Concepts Β· Math Foundations