TL;DR
- A random variable is a random number; its distribution says which values how likely. Discrete: . Continuous: a density with .
- Expectation = probability-weighted average: or . It is linear: , always.
- Variance ; ; for independent : .
- Covariance ; correlation .
- Average of i.i.d. copies: mean stays , variance shrinks to , standard deviation to . That is behind almost every rate in the course.
Random variables and distributions
Random variable
A function that turns an outcome into a number. It induces a distribution . Example: draw 3 balls from 5 red and 5 black, = number of red ones.
Distributions that occur in the course:
| Distribution | or density | Mean | Variance |
|---|---|---|---|
| Bernoulli(), values | |||
| Binomial(), values | |||
| Uniform on | |||
| Normal | |||
| Multivariate normal on | covariance matrix |
Where they appear: Bernoulli for labels and for the 0-1 loss of a fixed classifier; binomial for the number of errors on test points; uniform for toy data and bootstrap draws; normal for the noise ; multivariate normal for the Gaussian inputs in Lecture 8.
means: independent standard normal coordinates.
Densities
Density
has density if for all (measurable) sets . A density is and integrates to 1; it may be larger than 1 and need not be continuous. Intuition: a “continuous histogram”. A single point has probability 0; only intervals get probability.
For the exam diagrams (Bayes decisions) you only need areas under straight lines:
Areas you need for density diagrams
- Rectangle (uniform density on an interval of length ): probability .
- Triangle or trapezoid under a linear density: ; for a triangle simply .
- Check: the total area under each class density must be 1. That fixes the height: a triangle on has peak .
Example: on . .
Memory aid
Probability = area. Every Bayes error in a diagram task is “the area under the lower of the two weighted curves”.
Expectation
Expectation
Discrete: . Continuous: . For a function: or .
Examples: a fair die has . For an indicator (1 if happens, else 0): . That is why the risk under the 0-1 loss is a probability: (L1.1).
Rules
- Linearity (always): , even for dependent .
- Products (only if independent): .
- Empirical risk is an average: has expectation for a fixed (linearity).
Conditional expectation
is the average of among the outcomes with . Example: two dice, : . Without fixing the value, is itself a random variable (a function of ).
Where it appears: the regression function is the best predictor under the squared loss (L1.1); for it equals . Tower rule: .
Variance, covariance, correlation
Variance
; (shifting does not change the spread, scaling squares it).
Covariance and correlation
. Independent ⇒ uncorrelated (), not the other way round. iff exactly; correlation only measures linear dependence.
Worked example: the variance of an average
with variance each, pairwise correlation (so ). Then
- Independent (): , the classic “averaging reduces variance”.
- Correlated: the part never goes away. This is the bagging formula (L6); random forests decorrelate the trees to make small.
Memory aid
Averaging kills only the independent part of the noise. With , , : , and more trees cannot go below .
i.i.d. samples and empirical estimates
i.i.d.
are independent and identically distributed: each has the same distribution, and they do not influence each other. All training sets in the course are assumed i.i.d. from (L1.2).
The sample mean
For i.i.d. with mean and variance :
Memory aid
Error of an average . Four times the data halves the error. Hoeffding (Inequalities and Convergence) gives the same with a probability guarantee, and so do all the generalization bounds of Lecture 3.
Empirical (sample) estimates from data : , , (sometimes with for an unbiased estimate).
Random vectors and the covariance matrix
Covariance matrix
For : is the vector of coordinate means, and . is symmetric and positive semi-definite. For centered data in the rows of , the empirical covariance is .
Where it appears: in least squares (L5); the eigenvalues of decide whether overfitting is benign (L8). Eigenvalues are explained in Linear Algebra.
Self-check
Question cards (4)
takes the values 0 and 10 with probability 0.9 and 0.1. Compute and .
Answer
, , .
Why is the risk under the 0-1 loss a probability?
Answer
The loss is the indicator , and the expectation of an indicator is the probability of the event: .
100 i.i.d. measurements with standard deviation 5. What is the standard deviation of their mean?
Answer
.
Two estimators with variance 1 and correlation 0.5 are averaged. Variance of the average?
Answer
, not .