The course assumes the math of the “Mathematics for Machine Learning” courses. These pages repeat what the lectures actually use, with short explanations, mini examples, memory aids and links to the places where it appears. They follow the course’s own mathematical appendix.
| Page | What is in it |
|---|---|
| Probability Basics | conditional probability, joint/marginal/conditional in a table, total probability, Bayes’ formula, (conditional) independence, union bound |
| Random Variables and Expectation | distributions, densities and areas, expectation, conditional expectation, variance, covariance, correlation, i.i.d., covariance matrix |
| Inequalities and Convergence | Markov, Chebyshev, Hoeffding, solving for , , , law of large numbers, convergence in probability |
| Linear Algebra | scalar product, norms, hyperplanes, data matrix, rank, inverse, kernel and range, eigenvalues, PSD, SVD, generalized inverse, projections, trace |
| Calculus and Optimization | derivatives, gradients, normal equations, gradient descent, convexity, Lipschitz, smoothness, log rules, argmin, sup |
| Combinatorics and Asymptotics | binomial coefficients, counting labelings, growth rates, and , deciding whether a rate goes to 0 |
Which lecture needs which math
| Lecture | Math you need |
|---|---|
| 1.1 Decision theory | conditional/joint/marginal probability, Bayes’ formula, densities and areas, expectation, indicator functions |
| 1.2 Finite samples | i.i.d., law of large numbers, convergence in probability, variance (bias-variance) |
| 2 Perceptron | scalar product, norms, hyperplanes, Cauchy-Schwarz, SGD |
| 3 Learning theory | Hoeffding, union bound, solving for and , binomial coefficients, growth rates, sup |
| 4 Stability | expectation, convexity, strong convexity, Lipschitz, smoothness, gradient descent, |
| 5 Regularization | linear systems, rank, inverse, generalized inverse, eigenvalues, SVD, gradients (normal equations), / norms |
| 6 Bagging and boosting | variance and correlation of averages, logs and exponentials, bootstrap probability |
| 7 Data and validity | scalar products (kernels), expectation (random features) |
| 8 Overparameterized learning | data matrix, kernel and range, minimum norm, eigenvalues, projections, trace trick, Gaussian vectors, union bound |
| 9.1 Fairness | conditional probability in tables, (conditional) independence, Bayes’ formula |
| 9.2 Explainability | conditional expectation, independence, counting (SHAP weights), gradients |
Symbols
| Symbol | Meaning |
|---|---|
| , | input space, label space (often or ) |
| a random data point drawn from the (unknown) distribution | |
| prior (class probability) | |
| class-conditional density or probability | |
| marginal of | |
| regression function ; for 0/1 labels | |
| loss; 0-1 loss | |
| , | true risk , empirical risk (average loss on the training points) |
| , | Bayes risk (smallest possible risk), Bayes classifier |
| , | function class, function learned from points |
| indicator: 1 if is true, else 0 | |
| if , if | |
| , , | expectation, variance, covariance |
| , | conditional probability, conditional expectation |
| , | independent, conditionally independent given |
| i.i.d. | independent and identically distributed |
| , | scalar product, (Euclidean) norm |
| , | data matrix (points as rows), transpose (also written , ) |
| identity matrix | |
| , , | trace, rank, null space |
| gradient | |
| , , | minimizing point, least upper bound, greatest lower bound |
| growth function (shattering coefficient); also normal distribution, from the context | |
| , | binomial coefficient, factorial |
| , | at most of the order, of strictly smaller order |
| failure probability (“with probability “) | |
| accuracy / deviation; in AdaBoost the weighted error | |
| regularization parameter; also eigenvalues | |
| , , , , | for all, there exists, element of, subset, number of elements |
| , , , | defined as, approximately, much smaller, proportional to |
Memory aid
The course uses the same letter for different things in a few places: (growth function vs. normal distribution), (regularization vs. eigenvalue), (regression function vs. step size in gradient descent), (accuracy vs. noise vs. AdaBoost error). The context always decides; in the exam, write down which one you mean.