TL;DR

  1. A random variable is a random number; its distribution says which values how likely. Discrete: . Continuous: a density with .
  2. Expectation = probability-weighted average: or . It is linear: , always.
  3. Variance ; ; for independent : .
  4. Covariance ; correlation .
  5. Average of i.i.d. copies: mean stays , variance shrinks to , standard deviation to . That is behind almost every rate in the course.

Random variables and distributions

Random variable

A function that turns an outcome into a number. It induces a distribution . Example: draw 3 balls from 5 red and 5 black, = number of red ones.

Distributions that occur in the course:

Distribution or densityMeanVariance
Bernoulli(), values
Binomial(), values
Uniform on
Normal
Multivariate normal on covariance matrix

Where they appear: Bernoulli for labels and for the 0-1 loss of a fixed classifier; binomial for the number of errors on test points; uniform for toy data and bootstrap draws; normal for the noise ; multivariate normal for the Gaussian inputs in Lecture 8.

means: independent standard normal coordinates.

Densities

Density

has density if for all (measurable) sets . A density is and integrates to 1; it may be larger than 1 and need not be continuous. Intuition: a “continuous histogram”. A single point has probability 0; only intervals get probability.

For the exam diagrams (Bayes decisions) you only need areas under straight lines:

Areas you need for density diagrams

  • Rectangle (uniform density on an interval of length ): probability .
  • Triangle or trapezoid under a linear density: ; for a triangle simply .
  • Check: the total area under each class density must be 1. That fixes the height: a triangle on has peak .

Example: on . .

Memory aid

Probability = area. Every Bayes error in a diagram task is “the area under the lower of the two weighted curves”.

Expectation

Expectation

Discrete: . Continuous: . For a function: or .

Examples: a fair die has . For an indicator (1 if happens, else 0): . That is why the risk under the 0-1 loss is a probability: (L1.1).

Rules

  • Linearity (always): , even for dependent .
  • Products (only if independent): .
  • Empirical risk is an average: has expectation for a fixed (linearity).

Conditional expectation

is the average of among the outcomes with . Example: two dice, : . Without fixing the value, is itself a random variable (a function of ).

Where it appears: the regression function is the best predictor under the squared loss (L1.1); for it equals . Tower rule: .

Variance, covariance, correlation

Variance

; (shifting does not change the spread, scaling squares it).

Covariance and correlation

. Independent ⇒ uncorrelated (), not the other way round. iff exactly; correlation only measures linear dependence.

Worked example: the variance of an average

with variance each, pairwise correlation (so ). Then

  • Independent (): , the classic “averaging reduces variance”.
  • Correlated: the part never goes away. This is the bagging formula (L6); random forests decorrelate the trees to make small.

Memory aid

Averaging kills only the independent part of the noise. With , , : , and more trees cannot go below .

i.i.d. samples and empirical estimates

i.i.d.

are independent and identically distributed: each has the same distribution, and they do not influence each other. All training sets in the course are assumed i.i.d. from (L1.2).

The sample mean

For i.i.d. with mean and variance :

Memory aid

Error of an average . Four times the data halves the error. Hoeffding (Inequalities and Convergence) gives the same with a probability guarantee, and so do all the generalization bounds of Lecture 3.

Empirical (sample) estimates from data : , , (sometimes with for an unbiased estimate).

Random vectors and the covariance matrix

Covariance matrix

For : is the vector of coordinate means, and . is symmetric and positive semi-definite. For centered data in the rows of , the empirical covariance is .

Where it appears: in least squares (L5); the eigenvalues of decide whether overfitting is benign (L8). Eigenvalues are explained in Linear Algebra.

Self-check