Every term of the course in one line, A to Z. In the β€œSee” column, a lecture number (L3, L9.1, …) links to the lecture section, β€œMath” to the math foundations, every other link to the term’s concept page. Formulas are on the Formula Sheet.

Words that mean two things

  • Bias is the inductive bias (what the learner assumes, L1.1), the bias term of the bias-variance decomposition (L1.2) and bias in data that leads to unfair decisions (L7, L9.1).
  • Margin is the distance of the closest point to a hyperplane (L2) and the product of a scoring function (L1.1, the margin bound in L6).
  • Consistent means (L1.2), but consistent w.r.t. only means (L3).
  • Calibrated: a classification-calibrated loss gives the Bayes classifier after thresholding (L1.1); a score calibrated by group means the same probability in every group (L9.1).
  • Stability: average stability equals the expected generalization gap, uniform stability gives the high-probability bound (L4).

0-9

TermMeaningSee
0-1 loss0 for a correct prediction, 1 for a wrong one; its risk is the probability of errorLoss and Risk

A

TermMeaningSee
AdaBoostboosting by reweighting: the weak learner with weighted error gets the vote , misclassified points gain weightAdaBoost
AI ActEU law that regulates AI applications by risk: prohibited, high, limited, minimal, plus rules for GPAI modelsAI Regulation
Algorithmic stabilitythe learned function changes only a little when one training point is replaced; a stable algorithm generalizesAlgorithmic Stability
Almost sure convergence; implies convergence in probabilityL1.2
Approximation error, the price of the class ; deterministic, shrinks when growsEstimation and Approximation Error
Average stability expected loss change on the replaced point; equals the expected generalization gap (Proposition 1 of L4)Algorithmic Stability

B

TermMeaningSee
Baggingbootstrap aggregation: average the estimates of bootstrap samples; variance instead of Bagging
Base rate within a group; different base rates make some fairness criteria incompatibleL9.1
Bayes classifier, Bayes predictor a function with the smallest possible risk; for the 0-1 loss: predict 1 iff Bayes Classifier
Bayes decision rulepredict the label with the smallest conditional risk, using priors and the costs of the errorsBayesian Decision Theory
Bayes risk , the best risk any function can reachBayes Classifier
Benign overfittingan interpolator fits noisy labels exactly and still has a small test errorBenign Overfitting
-smooththe gradient is -Lipschitz: an upper bound on the curvatureSmoothness
Bias (bias-variance)how far the average prediction over training samples is from the truthBias-Variance Decomposition
Bias-variance decompositionfor the loss, pointwise: Bias-Variance Decomposition
Boostingcombine many weak learners with specific weights into a strong one; reduces biasAdaBoost
Boosting class all with from the base class; L6
Bootstrapjudge an estimate by recomputing it on resamples drawn from the dataL6

C

TermMeaningSee
Calibration by group: a score means the same probability in every groupFairness Criteria
Classical regime cannot interpolate the training data; the test risk is U-shapedL1.2
Classification-calibrateda surrogate loss whose Bayes predictor, thresholded at 0, is the 0-1 Bayes classifierSurrogate Loss
Conditional riskexpected loss of predicting at a fixed ; only the label is randomBayesian Decision Theory
Consistency in probability; a statement about the risk, not about Consistency
Consistency w.r.t. : only the estimation error goes to 0L3
Construct validitydoes the metric measure the concept we care about?Validity
Convergence in probability for every L1.2
Counterfactual explanationthe smallest change of the input that flips the decisionCounterfactual Explanation
Counterfactual fairnessuses a causal model to ask what the outcome would have been had other variables been differentL9.1

D

TermMeaningSee
Datasheetdocumentation of a data set: motivation, composition, collection, recommended uses (Gebru et al. 2018)L7
Decision stumpa tree with a single split; the usual weak learner in AdaBoostAdaBoost
Demographic parity: both groups get positive decisions at the same rate (independence)Fairness Criteria
Derived classifierfairness post-processing: a randomized from and with four probabilities Fair Post-Processing
Design matrix matrix with one training point per rowL5
Double descentthe test error rises towards the interpolation threshold and can fall again beyond itDouble Descent

E

TermMeaningSee
Empirical risk average loss on the training points; for the 0-1 loss the training errorLoss and Risk
Empirical risk minimization (ERM)pick Empirical Risk Minimization
Equal opportunityequal true positive rates in all groupsFairness Criteria
Equalized odds: equal true and false positive rates in all groups (separation)Fairness Criteria
Estimation error; random, grows with the size of Estimation and Approximation Error
Excess riskrisk minus the best possible risk, e.g. Estimation and Approximation Error
Exponential loss; the loss behind AdaBoostSurrogate Loss
External validitydoes the result carry over to other data and settings?Validity

F

TermMeaningSee
Feature attributiona score per input feature for how much it mattered for one decisionFeature Attribution
Feature map turns an object into a vector of numbers in Kernel Methods
Feedback looppredictions change the data the next model is trained on and become self-fulfillingL9.1
Fixed designthe input points are fixed and only the labels are random; risk on these pointsL5
Forward stagewise additive modelingadd one base function at a time and freeze the earlier ones; AdaBoost is the case of the exponential lossL6

G

TermMeaningSee
GAMgeneralized additive model ; interpretable, and interventional SHAP recovers the SHAP
GDPREU data protection: lawful basis and consent, data minimization, purpose limitation, right to be forgottenAI Regulation
Generalization boundtrue risk empirical risk + capacity term, with probability at least Generalization Bound
Generalization gap, true minus empirical risk of the learned functionL4
Generalized inverse inverts on the eigenvectors with non-zero eigenvalues (Moore-Penrose)L5
GPAI modelgeneral-purpose AI model in the AI Act; systemic risk above FLOPs of training computeAI Regulation
Gradient boostingfit each new base learner to the negative gradient of the lossGradient Boosting
Gradient descentStochastic Gradient Descent
Growth functionthe shattering coefficient as a function of ; either or polynomialShattering Coefficient

H

TermMeaningSee
Hinge loss; a classification-calibrated surrogateSurrogate Loss
Hoeffding’s inequalityindependent : Hoeffding Inequality

I

TermMeaningSee
i.i.d.independent and identically distributed; how the training points are drawnMath
Implicit regularizationthe optimizer picks one special solution among many, e.g. GD from 0 finds the minimum norm solutionImplicit Regularization
Individual fairnessindividuals that are similar under a pre-specified metric get similar treatmentL9.1
Inductive biasthe assumptions about what we look for; without one, learning is impossible (No Free Lunch)Inductive Bias
Inductive inferencefrom specific examples to a general rule; the conclusion can be wrongL1.1
Internal validityis the effect really caused by the model, and not by artifacts or confounders?Validity
Interpolation thresholdas many parameters as data points; the peak of the double descent curveDouble Descent
Interventional value functionfix and draw the other features from their marginal distributionSHAP

K

TermMeaningSee
Kernel, kernel trick, computed without computing Kernel Methods

L

TermMeaningSee
Lassoleast squares ; sparse solutions, no closed formLasso
Law of large numbersthe empirical mean converges to the expectation; for a fixed , Math
Least squaresminimize ; if has full rankLeast Squares Regression
Leave-one-outtrain without one point, test on it; the average is an unbiased estimate of the generalization errorLeave-One-Out Error
LIMEexplains one decision by a linear model fitted on samples around LIME
Linear loss; the perceptron is SGD on itL2
LipschitzMath
Loss function: how expensive it is to predict when the truth is Loss and Risk

M

TermMeaningSee
MAPmaximum a posteriori: predict the label with the largest posterior; for the 0-1 loss this is the Bayes classifierBayesian Decision Theory
Margin (hyperplane)smallest distance of a training point to the hyperplane, Margin
Margin (score); positive iff the point is classified correctlyMargin
Margin boundtest error bound for boosting that does not depend on L6
Maximum likelihoodpredict the label with the largest ; ignores the priorsBayesian Decision Theory
McDiarmid’s inequalityconcentration for a function of independent variables whose value changes by at most in coordinate L3
Measurement procedurethe device that records the construct; machine learning learns the measurement, not the constructMeasurement and Construct
Minimum norm solutionthe interpolating with the smallest norm, Implicit Regularization
Mistake boundthe perceptron makes at most mistakes on separable dataPerceptron
Modern regime is large enough to interpolate the training dataL1.2

N

TermMeaningSee
Neural tangent kernel (NTK)very wide networks train like a kernel method with ; no feature learningNeural Tangent Kernel
No Free Lunchaveraged over all possible true functions, all classifiers perform the sameNo Free Lunch Theorem

O

TermMeaningSee
Observational value functionaverage over the points with (conditional distribution)SHAP
OLSordinary least squares, no regularizer; excess risk in the fixed designL5
Over-parameterized regimemore parameters than data points; the model can interpolateL8
Overfittingclassical regime: the class is too large, low approximation error, high estimation errorL1.2

P

TermMeaningSee
PAC learnablestrong: error with probability for all ; weak: error ; both are equivalentL6
Perceptronafter every mistake update ; SGD with step size 1 on the linear lossPerceptron
Plug-in classifierestimate from the data and threshold the estimate like the Bayes classifierL1.2
Predictive parity: a prediction means the same in every group (sufficiency)Fairness Criteria
Protected attribute sensitive attribute such as gender or raceL9.1
Proxya measurable stand-in for the construct, e.g. health care cost for illnessMeasurement and Construct

R

TermMeaningSee
Rademacher complexityhow well can fit random labelsRademacher Complexity
Random designthe data is random and we care about the test errorL5
Random featuresbasis functions with random, frozen parameters; only the output weights are learned; they approximate a kernelRandom Features
Random forestbagged decision trees; each split chooses its dimension from a random subset ()Random Forest
Regression function ; for labels 0 and 1 it is Bayes Classifier
Regularizationminimize with a regularizer that measures complexityRegularization
Representation learninglearn the features from raw data instead of fixing them by handL7
Ridge regressionleast squares ; (Tikhonov)Ridge Regression
Robust interpolationinterpolating with Lipschitz constant about 1 needs parameters (Bubeck and Selke)L8

S

TermMeaningSee
Sauer-Shelah lemmaVC dimension : L3
Scoring functiona real-valued ; the classifier is Surrogate Loss
SGDstochastic gradient descent: a gradient step on one random training pointStochastic Gradient Descent
SHAPShapley values of the features for one predictionSHAP
Shattering realizes all labelings of the pointsVC Dimension
Shattering coefficient the largest number of labelings produces on pointsShattering Coefficient
Sparsitymany coefficients exactly 0; the lasso gets it from the corners of the ballLasso
Spiky-smoothan interpolator with narrow spikes at the training points that is smooth everywhere elseBenign Overfitting
Statistical validityis the difference more than chance?Validity
Strong convexity-strongly convex: a lower bound on the curvatureStrong Convexity
Strongly consistentthe risk converges to almost surelyConsistency
Surrogate lossa convex loss on a real score used instead of the 0-1 loss (hinge, squared, exponential, logistic)Surrogate Loss
Symmetrizationcompare with a ghost sample, so the supremum runs over finitely many labelingsL3

T

TermMeaningSee
Target constructthe concept we want to measure, often not observableMeasurement and Construct
Tikhonov regularizationregularizer , as in ridge regression; makes ERM strongly convex and stableRidge Regression
True risk , the expected loss on new dataLoss and Risk

U

TermMeaningSee
Underfittingclassical regime: the class is too small, high approximation errorL1.2
Uniform convergence in probability; sufficient and necessary for ERM to be consistentUniform Convergence
Uniform stability worst-case loss change when one training point is replacedAlgorithmic Stability
Union bound; gives the bound for finite classesMath
Universal approximationtwo-layer networks with a continuous, non-polynomial activation approximate every continuous function on a compact setUniversal Approximation
Universally consistentconsistent for every distribution ; kNN is (Stone 1977)Consistency

V

TermMeaningSee
Validityfour notions: statistical, internal, external, constructValidity
Variance (bias-variance)how much the prediction changes from one training sample to anotherBias-Variance Decomposition
VC dimensionthe largest such that some points are shatteredVC Dimension

W

TermMeaningSee
Weak learneronly slightly better than random guessing: error L6

X

TermMeaningSee
XGBoostan efficient implementation of gradient boosting with treesL6