The course assumes the math of the “Mathematics for Machine Learning” courses. These pages repeat what the lectures actually use, with short explanations, mini examples, memory aids and links to the places where it appears. They follow the course’s own mathematical appendix.

PageWhat is in it
Probability Basicsconditional probability, joint/marginal/conditional in a table, total probability, Bayes’ formula, (conditional) independence, union bound
Random Variables and Expectationdistributions, densities and areas, expectation, conditional expectation, variance, covariance, correlation, i.i.d., covariance matrix
Inequalities and ConvergenceMarkov, Chebyshev, Hoeffding, solving for , , , law of large numbers, convergence in probability
Linear Algebrascalar product, norms, hyperplanes, data matrix, rank, inverse, kernel and range, eigenvalues, PSD, SVD, generalized inverse, projections, trace
Calculus and Optimizationderivatives, gradients, normal equations, gradient descent, convexity, Lipschitz, smoothness, log rules, argmin, sup
Combinatorics and Asymptoticsbinomial coefficients, counting labelings, growth rates, and , deciding whether a rate goes to 0

Which lecture needs which math

LectureMath you need
1.1 Decision theoryconditional/joint/marginal probability, Bayes’ formula, densities and areas, expectation, indicator functions
1.2 Finite samplesi.i.d., law of large numbers, convergence in probability, variance (bias-variance)
2 Perceptronscalar product, norms, hyperplanes, Cauchy-Schwarz, SGD
3 Learning theoryHoeffding, union bound, solving for and , binomial coefficients, growth rates, sup
4 Stabilityexpectation, convexity, strong convexity, Lipschitz, smoothness, gradient descent,
5 Regularizationlinear systems, rank, inverse, generalized inverse, eigenvalues, SVD, gradients (normal equations), / norms
6 Bagging and boostingvariance and correlation of averages, logs and exponentials, bootstrap probability
7 Data and validityscalar products (kernels), expectation (random features)
8 Overparameterized learningdata matrix, kernel and range, minimum norm, eigenvalues, projections, trace trick, Gaussian vectors, union bound
9.1 Fairnessconditional probability in tables, (conditional) independence, Bayes’ formula
9.2 Explainabilityconditional expectation, independence, counting (SHAP weights), gradients

Symbols

SymbolMeaning
, input space, label space (often or )
a random data point drawn from the (unknown) distribution
prior (class probability)
class-conditional density or probability
marginal of
regression function ; for 0/1 labels
loss; 0-1 loss
, true risk , empirical risk (average loss on the training points)
, Bayes risk (smallest possible risk), Bayes classifier
, function class, function learned from points
indicator: 1 if is true, else 0
if , if
, , expectation, variance, covariance
, conditional probability, conditional expectation
, independent, conditionally independent given
i.i.d.independent and identically distributed
, scalar product, (Euclidean) norm
, data matrix (points as rows), transpose (also written , )
identity matrix
, , trace, rank, null space
gradient
, , minimizing point, least upper bound, greatest lower bound
growth function (shattering coefficient); also normal distribution, from the context
, binomial coefficient, factorial
, at most of the order, of strictly smaller order
failure probability (“with probability “)
accuracy / deviation; in AdaBoost the weighted error
regularization parameter; also eigenvalues
, , , , for all, there exists, element of, subset, number of elements
, , , defined as, approximately, much smaller, proportional to

Memory aid

The course uses the same letter for different things in a few places: (growth function vs. normal distribution), (regularization vs. eigenvalue), (regression function vs. step size in gradient descent), (accuracy vs. noise vs. AdaBoost error). The context always decides; in the exam, write down which one you mean.