TL;DR

  1. Conditional: : restrict the world to and renormalize.
  2. Joint → marginal → conditional: from a joint table , the marginal is a row or column sum, the conditional is a cell divided by its row or column sum.
  3. Product rule: (prior times class conditional).
  4. Total probability: . Bayes: .
  5. Independence: . Conditional independence : independent once is known.
  6. Union bound: , always true, no independence needed.

The course uses these rules from Lecture 1.1 on: the regression function is a conditional probability, the Bayes classifier compares joint probabilities, and fairness criteria are conditional independence statements. (Source of the definitions: the course’s mathematical appendix.)

Events and the rules of probability

Probability space

  • : the set of all elementary outcomes (“sample space”), e.g. for a die.
  • An event is a set of outcomes, e.g. “even” .
  • A probability measure assigns each event a number with , , and for disjoint events (Kolmogorov’s axioms).

Consequences you use all the time:

  • .
  • .
  • Die: .

Conditional probability

Conditional probability

The probability of given that has happened ():

Die: .

Memory aid

“Given ” means: throw away everything outside , then rescale so that has probability 1. The bar is a “zoom in” on .

Joint, marginal and conditional distributions

For two random variables and (see Random Variables and Expectation):

Joint, marginal, conditional

  • Joint: , the probability that both happen. All joint probabilities sum to 1.
  • Marginal: (“sum out” the other variable). In the course the marginal of is often written .
  • Conditional: .
  • Product rule (the same formula read backwards): .

Worked example: from prior and class conditionals to the posterior

This is exactly the computation of the discrete-input Bayes task and of the mock exam. Given: , prior (so ), class conditionals and .

Step 1: joint table with the product rule (prior times class conditional):

row sum = prior
column sum = marginal

Step 2: marginal : the column sums, , .

Step 3: conditional : the cell divided by the column sum: , .

Step 4: use it. Bayes classifier for the 0-1 loss: predict 1 iff , so , . Bayes error: in every column take the smaller cell: .

Memory aid: the table recipe

Rows = true class, columns = observed . Fill the cells with prior × class conditional. Column sums are . = upper cell / column sum. The Bayes classifier picks the larger cell in each column, and the Bayes error is the sum of the smaller cells.

Law of total probability and Bayes’ formula

Law of total probability

If are disjoint and together cover :

With : . This is the denominator of Bayes’ formula.

Bayes' formula

With densities instead of probabilities it works the same way. Prior , likelihood , posterior .

Test for a rare disease (numbers from the math appendix)

1% of women have breast cancer (). A mammography detects it with probability 0.8 (, true positive rate) and gives a false alarm with probability 0.096 (). Total probability: . Bayes: . A positive test means only about 8%: the prior is tiny.

Memory aid

Posterior ∝ prior × likelihood. Compare with ; the denominator is the same for both classes, so you only need it for the actual value of , not for the decision. That is why the Bayes boundary is “where the weighted curves cross” (L1.1).

Independence and conditional independence

Independence

Events: , equivalently . Random variables: for all ; written .

Two coin tosses are independent; the first toss and the sum are not.

Conditional independence

and are independent given , , if

equivalently : once is known, tells nothing more about .

Where it appears: i.i.d. training data (independent and identically distributed, L1.2); the fairness criteria (demographic parity), (equalized odds), (predictive parity) in L9.1; observational vs. interventional SHAP in L9.2.

Memory aid

Read "" as “within each group of “. : among the people with the same true label, the prediction does not depend on the group. That is why the TPR and FPR are computed inside the columns and .

The union bound

Union bound

For any events :

No independence needed. It is exact for disjoint events and loose when the events overlap (the overlap is counted several times).

Where it appears: the bound for finite function classes (“some out of is bad” has probability at most times the probability for one , L3); the spikes of the minimum norm interpolator (, L8).

Memory aid

“At least one of them happens” → union bound: add up the single probabilities. The price is a factor (or ), which becomes after solving for .

Self-check