TL;DR

  1. There is no unique definition of fairness. COMPAS is “fair” in one sense (same reoffending rate per risk score) and “unfair” in another (non-reoffending black defendants get high scores twice as often). With different base rates, no non-trivial classifier can satisfy both.
  2. Sources of unfairness: minorities drown in accuracy-maximizing systems, sampling bias, pre-existing biases in the data, the choice of features, the choice of the target variable (proxy).
  3. Group criteria (binary , , ): demographic parity , i.e. ; equalized odds , i.e. equal TPR and FPR (equal TPR alone = equal opportunity); predictive parity , i.e. equal ; calibration .
  4. Impossibility: if and are dependent (different base rates), demographic parity and predictive parity cannot both hold, and neither can equalized odds and predictive parity.
  5. Fixing unfairness: removing the sensitive attribute does not work (proxies like the ZIP code). Train with a fairness constraint (relaxations are often too loose). Post-processing: a derived classifier , often randomized, or group-specific thresholds where the ROC curves intersect.
  6. Tradeoffs: fairness costs accuracy, randomized decisions are questionable, feedback loops make predictions self-fulfilling. Society has to decide, technical solutions cannot settle it.

Exam relevance

  • Real exam, task 3: fairness from two per-group count tables: demographic parity, TPR and FPR, equal opportunity and equalized odds, then randomized post-processing to reach equal opportunity. Known mistake: computing the TPR as conditioned on instead of . Practice with the table task and the widget (random practice tables).
  • Mock exam: ROC curves and equalized odds, see the ROC task.
  • Know the three criteria with their conditional independence form, which ones the constant and the perfect classifier satisfy, and the impossibility propositions with their one-line proof idea.
  • Sheet 11, Exercise 2 (removing a feature is not enough).
  • Everything in one go: the cheat sheet with two full integration tasks at the end of this note.

Overview: 1. The discussion about fairness (COMPAS, admissions, credit, deliveries, Street Bump, recruiting, word embeddings), 2. Sources of unfairness, 3. Basic notions of fairness, 4. Impossibility results, 5. Technical approaches (pre-, in-, post-processing), 6. Tradeoffs and feedback loops.

(Literature: Barocas, Hardt and Narayanan, Fairness and Machine Learning, available online at fairmlbook.org; it contains many further references.) Slide 5

The discussion about fairness

Example: COMPAS

Slides 7-16

The algorithm COMPAS is used across the US to decide whether defendants awaiting trial should be released on bail (German: auf Kaution freilassen). It assigns scores from 1 to 10 that indicate how likely a defendant is to re-offend (recidivism, German: Rückfall). The score is based on more than 100 factors, including age, sex and criminal history. Race is not used. The higher the score, the more likely the person is considered risky and detained. There is a heated debate whether the score is biased against black people.

First point of view: the score is not biased

For a given score , the probability to reoffend is about the same for white and black defendants:

Consequence: when judges see a risk score, they need not consider the defendant’s race when interpreting it.

Chance of recidivism against risk score 1 to 10 for black and white defendants; both lines rise from about 22 to about 75 percent and lie within each other's confidence bands
Slide 9: recidivism rate by risk score and race. Defendants with the same score are roughly equally likely to reoffend (gray bands: 95 percent confidence intervals).

Second point of view: the score is biased (ProPublica)

Among defendants who ultimately did not reoffend, black defendants were more than twice as likely as white defendants to be classified as medium or high risk (42 percent vs. 22 percent):

Even though these defendants did not commit a crime, the black ones are treated more harshly by the courts.

Stacked bar chart of the number of defendants per risk category, low and medium/high, for black and white defendants, split into reoffended and did not reoffend
Slides 11-14: defendants per risk category and race, split into reoffended (light) and did not reoffend (dark).

Where the two views come from, read off the figure:

  • Fair: within each risk category the proportion of defendants who reoffend is about the same for both groups (the “low” bars have about the same light/dark ratio for black and white, and so do the “medium/high” bars). That is a statement conditioned on the score.
  • Unfair: for black defendants the two dark areas (did not reoffend) in “low” and “medium/high” are nearly the same size; for white defendants they are not. Black defendants who don’t reoffend are predicted riskier. That is a statement conditioned on the true outcome.
  • Why both can happen: in the raw data black defendants reoffend at a higher rate than white defendants. A classifier that is perfect in terms of accuracy classifies black defendants as medium or high risk more often (58 percent vs. 33 percent).

So who is right?

It depends on how we measure fairness and how we measure the “success” of the system. For this data it is impossible to construct a non-trivial classifier that is fair with respect to both points of view. (Sources: Angwin, Larson, Mattu and Kirchner, “Machine bias”, ProPublica 2016; Washington Post, 2016.)

Side remark: the data has more obvious problems. Some released people reoffended but were never caught, and the chance of that may depend, for example, on where they live. This introduces further bias (measurement vs. construct, Lecture 7).

Example: college admission and affirmative action

Slides 17-20

Harvard admissions lawsuit (2014/15): students complained that the admission rules are unfair. An Asian-American male applicant, not disadvantaged, with a 25% chance of admission would have 36% as a white applicant, 77% as Hispanic and 95% as African-American, all other characteristics the same:

Is this obviously unfair? The reason is affirmative action. Harvard: dropping race as one factor would give a class that “fails to achieve the diversity and excellence that Harvard seeks”. In October 2019 the judges concluded that Harvard’s practice was lawful: race may act as a “plus” or “tip”; there are no workable race-neutral alternatives; under the court’s preferred model there is no statistically significant difference between similarly situated Asian-American and white applicants. (Sources: admissionscase.harvard.edu, potentially biased; fairml book, section on counterfactual discrimination analysis.)

More examples

Slides 21-27

  • Credit scoring: in a linear model predicting whether applicants pay back a credit, the ZIP code turns out to be a very strong predictor. Living in a low-income neighborhood makes it harder to get credit; you have to “compensate” with, say, a higher income.
  • Deliveries: Amazon chose the neighborhoods for free same-day delivery with a data-driven system. A 2016 study found that in many US cities white residents were more than twice as likely as black residents to live in a qualifying neighborhood.
  • Street Bump: a Boston app detects potholes with smartphone sensors and reports them to the city. Areas with many elderly people, and low-income neighborhoods, have fewer smartphones, so their streets get reported and fixed less.
  • Recruiting (Amazon): about 10 years ago Amazon (74 percent of managerial positions held by men) used a tool to screen applications, trained on 10 years of resumes, predominantly from white males. It penalized resumes containing the word “women” and downgraded graduates of women’s colleges; it was discontinued (Business Insider 2018; Brookings).
  • Recruiting (Austria): the employment agency planned to classify unemployed persons into “low”, “middle”, “high” chances of finding a job (features: age, education, prior jobs) and to spend more effort on the top group. Criticism: it can turn existing discrimination into a technical solution; female gender and older age provably lead to a worse evaluation (Süddeutsche Zeitung, 2019).
  • Word embeddings (word2vec) map words to and support word arithmetic: “London is to England as Paris is to France”, (the slide writes the difference , the sign has to be the other way round). Asking “man is to X as woman is to Y” gives computer programmer / homemaker, surgeon / nurse, brilliant / lovely: the embedding reproduces the gender stereotypes of the text corpus (Bolukbasi et al., NeurIPS 2016).

First take-away

Slides 28-30

What the examples show

  • All kinds of biases can happen implicitly in automatic systems. Mostly the designers did not intend to discriminate (Street Bump, word embeddings); some intentionally promoted minorities (affirmative action).
  • There is no unique definition of fairness. Many definitions exist, and they typically exclude each other. Which one is appropriate differs between applications and must be discussed in the context of society.
  • A “fair” system may have to give up performance elsewhere: overall accuracy (COMPAS) or utility and profit (credit scoring).
  • Fairness as a criterion is a request that comes from society. It is hard to find a real application of ML without potentially discriminatory behavior: this is an issue, always, and typically there is no simple solution.

Sources of unfairness

Slides 31-35

  • Minorities: a model with 5% overall error can perform terribly on a minority group. Minorities get “drowned” in systems that maximize accuracy, because the training and test error barely change when the minority is misclassified. Worse, minorities are often under-represented relative to the population (sampling bias).
  • Biases in data. Sampling bias: demographic, geographic, behavioral or temporal biases in data collection. Example: crime records only contain crimes observed by the police; more officers go where the recorded crime rate was high, so more crimes are recorded there, even if other regions later have more crime. Pre-existing biases: gender roles in text and images, racial stereotypes, past hiring data, word embeddings.
  • Measurement of features: which features we measure, and how, shapes the model and can systematically favor or disfavor groups.
  • Target variable: the target in social contexts is often a coarse proxy for what we want to measure; by choosing it we already encode biases (the measurement section of Lecture 7).

Discussion

Fairness is not a technical requirement but one that comes from society. Nothing about fairness is obvious and simple.

Basic notions of fairness

Setup

Slides 37-38

There is no unique definition; below are popular ones, many more exist.

Data for group fairness

  • Features .
  • Protected / sensitive attribute (gender, race, …); depending on the application it is explicitly known or not.
  • True target .
  • A classifier predicts (an estimate of , or something more abstract).

For simplicity everything is binary: (black / white), (“reoffends or not”), (“kept in jail / released on bail”).

Demographic parity (independence)

Slides 39-41

Demographic parity (called independence in the fairml book)

Independently of all other features, both groups have the same rate of success. General form beyond the discrete case: .

Examples: the same proportion of males () and females () is promoted (); the same proportion of white and black people is released on bail.

Demographic parity is very strong

It forbids any dependence between the decision and . As soon as the true target is correlated with , this can be highly problematic. Example (own): a disease that is twice as common in group . Demographic parity forces a screening test to flag both groups equally often, so it must either miss sick people in group 1 or flag healthy people in group 0. Even a perfect classifier violates it.

Equalized odds (separation)

Slides 42-44

Equalized odds (called separation in the fairml book)

All groups have the same false negative rate and the same false positive rate:

General form: . If one direction matters more, one can require just one of the two; requiring equal true positive rates (equivalently equal false negative rates) is called equal opportunity (Hardt, Price, Srebro 2016).

Examples: among students who have the potential to achieve an MSc degree (), the probability to be accepted () is the same for male and female applicants. Among people who do not reoffend, the probability to be released on bail is the same for white and black defendants.

Sounds good. But: the perfect classifier (false negative and false positive rate 0) is always perfectly fair under this definition. Hm…

Recap: conditional independence

For discrete random variables, iff

equivalently . In words: once is known, knowing gives no further information about .

Memory aid

Columns = truth. Equalized odds: people with the same true label get the same treatment in both groups. Predictive parity looks along the rows (same prediction).

Predictive parity (sufficiency) and calibration

Slides 45-46

Predictive parity (called sufficiency in the fairml book)

If the prediction for a person is , the probability that truly is the same for all groups. General form: .

This is the first COMPAS point of view: for a predicted score the probability to reoffend is the same for black and white defendants, so a judge who sees can treat both the same.

Calibration by group

If the score is meant to predict a probability (of reoffending, say), it is calibrated by group if for all score values and groups

The probabilities are “correct”. In simple cases calibration and predictive parity are more or less the same, but not always (fairml book; Chouldechova, “Fair prediction with disparate impact”, 2017).

Which variable do you condition on?

criterionconditions onratein the table
demographic paritynothing (only )(TP + FP) / n
equalized oddsthe true label TPR , FPR TP / (TP + FN), FP / (FP + TN): columns
predictive paritythe prediction PPV TP / (TP + FP): rows

Writing the TPR as with behind the bar, or dividing TP by TP + FP, computes the PPV instead. That was a real exam mistake.

Extreme cases

Slides 47-48

For binary classification with a binary sensitive attribute:

  • Constant classifier for all inputs: the output is independent of everything, in particular of . It is maximally fair: it satisfies demographic parity and equalized odds (both TPR and FPR are 1 in both groups).
  • Predicting the sensitive attribute, : the output is identical to , maximally unfair with respect to demographic parity and equalized odds.

Impossibility results

Criteria that cannot hold together

Slides 49-52

In non-trivial situations the criteria typically cannot hold at the same time (fairml book; Kleinberg, Mullainathan, Raghavan, “Inherent trade-offs in the fair determination of risk scores”, 2016).

Proposition (demographic parity vs. predictive parity)

Assume all events of the joint distribution of have positive probability and and are not independent. Then demographic parity and predictive parity cannot both hold.

Proposition (equalized odds vs. predictive parity)

Under the same assumptions ( not independent of ), equalized odds and predictive parity cannot both hold. Proof: by similar arguments, and imply , hence , a contradiction.

A third version: if is binary, is not independent of and is not independent of , then demographic parity and predictive parity cannot both hold (elementary, proof skipped).

In words

When the base rates differ between groups, you have to choose which criterion to give up. COMPAS: calibrated scores (predictive parity) and equal error rates (equalized odds) are incompatible because black and white defendants reoffend at different rates in the data.

Memory aid

Different base rates → pick your fairness. Calibrated scores and equal error rates cannot both hold (COMPAS).

Task: extreme classifiers and base rates

Exam-style task: which criteria hold? (4 P)

Math: (conditional) independence

Group has 80 of 200 people with , group has 50 of 100.

(a) (1 P, easy) The constant classifier : which of demographic parity, equalized odds, predictive parity hold?

(b) (1.5 P, harder) The perfect classifier : which hold?

(c) (1.5 P, transfer) Explain with the propositions why no classifier can satisfy demographic parity and predictive parity here. What would change if both groups had 40% positives?

Individual and counterfactual fairness

Slides 53-54

  • Individual fairness (Dwork, Hardt, Pitassi, Reingold, Zemel, “Fairness through awareness”, 2012): not about groups but individuals. Two individuals who are “similar” according to a pre-specified metric should receive similar treatment. Hard to use in practice: what is the right notion of similarity, in the input space and in the target space?
  • Counterfactual fairness (Russell, Kusner, Loftus, Silva, NeurIPS 2017): with a causal model, compute what (the distribution of) a variable would have been had other variables been different, all else equal. Example: “Would individual have graduated () if they hadn’t had a job?”

Lots of criteria

Slides 55-56

Many criteria exist and have been invented several times. The fairml book matches each to its closest relative among the three (no need to memorize the names):

closest relativecriteria (equivalent or a relaxation)
independence ()statistical parity, group fairness, demographic parity, conditional statistical parity, Darlington criterion (4)
separation ()equal opportunity, equalized odds, conditional procedure accuracy, avoiding disparate mistreatment, balance for the negative class, balance for the positive class, predictive equality, equalized correlations, Darlington criterion (3)
sufficiency ()Cleary model, conditional use accuracy, predictive parity, calibration within groups, Darlington criterion (1), (2)

Lots of criteria, no single one

  • The one, unique fairness criterion does not exist.
  • Fairness comes from society and cannot always be captured satisfactorily by statistical definitions.
  • All existing criteria are plausible in some applications, but all have serious drawbacks and miss important aspects (see the fairml book).
  • But also: the baseline is decisions made by humans, and they are definitely biased as well.

Technical approaches to improve fairness

Three approaches

Slide 58

Pre-, in- and post-processing

  • Pre-processing: try to fix the bias in the data.
  • Training (in-processing): learn decisions that are accurate and fair at the same time.
  • Post-processing: fix an unfair black-box model in hindsight.

Fixing unfairness in the data

Slides 59-61

First naive idea: remove the sensitive features. This is pretty much impossible, because many other variables are highly correlated with the sensitive attribute and the algorithm uses them as proxies. Standard example: in many cities some quarters are predominantly white or black, and often the black population has a low income. The ZIP code is then correlated with both income and race, and it may be harder to get credit in those quarters. Even if race is not a feature, it is implicitly present whenever the ZIP code is used. (This is Sheet 11, Exercise 2.)

  • In many cases the discriminatory features are not even well defined. Some patterns (smoking is associated with cancer) are knowledge we want to mine, others (girls like pink, boys like blue) are stereotypes we want to avoid. It is hard (impossible) to tell the algorithm which patterns to find, and humans may not agree either.
  • Removing all variables correlated with the sensitive attribute often removes all relevant variables, in particular with several sensitive variables (gender and race). And sometimes the sensitive attribute is not even available, so the correlations cannot be computed.

Training for accuracy and fairness

Slides 62-64

Standard setup: minimize the empirical risk over , with a regularizer (Lecture 5). Two groups , target: demographic parity. The true unfairness is the demographic parity difference, and its empirical version uses the group sizes :

The fair learning problem:

Easy or difficult?

  • is discrete (indicator functions), so it is hard to optimize, like the 0-1 loss (surrogate losses).
  • Standard solution: convex relaxations. But most existing relaxations are too loose: even if the relaxed fairness is perfectly satisfied, the true fairness can be very bad (Lohaus, Perrot, von Luxburg, “Too relaxed to be fair”, 2020; Zafar, Valera, Rodriguez, Gummadi, “Fairness constraints”, 2017).
  • The lecturer’s view: this approach is the most useful one, but currently nothing out there works really well.

Fixing models in hindsight

Slides 65-68

Scenario: an agency or company trains a classifier privately (credit assessment, COMPAS). Someone evaluates it and finds it unfair. Can we fix it without access to its internals? We only see the prediction and the sensitive attribute (needed for the evaluation). All we can do is build a derived classifier that takes and as input and outputs a new, possibly randomized, .

Derived classifier (binary and )

Exactly four parameters:

Then for each group

The can be chosen (by a linear program) so that satisfies equal opportunity or equalized odds. Randomization helps to get a better accuracy-fairness tradeoff, but accuracy may still decrease; Hardt, Price and Srebro (NeurIPS 2016) give guarantees.

The reachable points of group form the quadrilateral with corners , , and (keep, never predict 1, always predict 1, flip). Equalized odds needs a point both groups can reach.

With a score: if we see a real-valued score that is thresholded, , plot the ROC curves of both groups separately and choose the point where they intersect (with group-specific thresholds). There both groups have the same false positive and false negative rates. If the curves do not intersect, a randomized predictor still solves it.

Hand-drawn ROC curves for group A=0 in black and A=1 in red; they cross at one point marked by an arrow
Slides 67-68: ROC curves of the two groups; at the intersection both have the same TPR and FPR.

Task: fairness from two tables

Exam-style task: two group tables (4 P)

Math: conditional probability in a table

A loan screening model is evaluated separately for two groups (: pays back, : loan granted).

6024
2096
305
2045

(a) (1 P, easy) Compute for both groups. Does demographic parity hold?

(b) (1.5 P, harder) Compute TPR and FPR for both groups. Does equal opportunity hold? Equalized odds? Compute also and explain why this is not the TPR.

(c) (1.5 P, transfer) The bank may only post-process: for each group it chooses . Construct a derived classifier with equal opportunity that keeps group 1 unchanged and never grants new loans. What are the new TPR and FPR of group 0, and its accuracy? Does equalized odds hold now?

Practice: fairness from two tables

Edit the counts of both tables; all rates and the four criteria update. Hover a metric to see which cells it uses. Random practice tables hide the results until you uncheck the box. Below: the derived classifier of the post-processing, its reachable region per group on the ROC plane, and buttons for equal opportunity and equalized odds.

With the task numbers: “equal opportunity: lower the higher TPR” reproduces (c). “Raise the lower TPR” instead flips 37.5% of group 1’s rejections to approvals and pushes its FPR to 0.44. “Equalized odds” finds the most accurate point that both groups can reach, here FPR 0.168 and TPR 0.630: both groups must move.

Task: thresholds on ROC curves

Exam-style task: group thresholds (4 P)

A score is thresholded at . The resulting (FPR, TPR) points per group:

threshold
(0.10, 0.50)(0.05, 0.30)
(0.20, 0.70)(0.10, 0.50)
(0.40, 0.90)(0.25, 0.70)

(a) (1 P, easy) With one common threshold : does equal opportunity hold?

(b) (1.5 P, harder) Find group-specific thresholds with equalized odds, and ones with equal opportunity at TPR . Does the second choice also give equalized odds?

(c) (1.5 P, transfer) Name two objections against group-specific thresholds or randomized post-processing in a bail decision.

Tradeoffs

Accuracy, randomization and feedback loops

Slides 70-74

  • Accuracy vs. fairness: fair classifiers optimize two objectives. The accuracy of a fair classifier can only be that of the classifier optimized for accuracy alone. A choice has to be made: how much accuracy loss do we tolerate for how much fairness?
  • Fair classifiers can require randomization: especially when classifiers are modified in hindsight, it is often impossible to reach fairness without randomization (Agarwal, Beygelzimer, Dudík, Langford, Wallach, “A reductions approach to fair classification”, 2018; Hardt, Price, Srebro 2016). Depending on the application this is highly questionable: a randomized decision about who stays in jail?
  • Feedback loops. Self-fulfilling predictions: a predictive policing system marks some areas as high risk, more officers are sent there, more crimes are detected there, and the prediction appears validated even if the risk is not higher. Predictions that affect the training set: the resulting arrests are added to the training data, so the areas keep looking risky (Ensign, Friedler, Neville, Scheidegger, Venkatasubramanian, “Runaway feedback loops in predictive policing”, 2017).

Discussion

Fairness has a long history of debate in ethics, sociology and many other fields. From the data and from the negative theoretical results: there often is no obvious “solution”. In practice many decisions must be taken (which notion of fairness, which tradeoff, is randomization acceptable, …), and different decisions lead to different solutions. Most of these issues cannot be fixed by technical solutions. Society has to decide (but we need to help and explain the issues). See also Corbett-Davies and Goel, “The measure and mismeasure of fairness”, 2018.

Summary

TopicKey message
COMPAScalibrated per score (fair) but unequal false positive rates (unfair); with different base rates both cannot hold
sourcesminorities drown in accuracy, sampling bias, pre-existing bias, feature and target choice
demographic parity: equal ; the constant classifier satisfies it, the perfect one may not
equalized odds: equal TPR and FPR (columns); equal opportunity = equal TPR; the perfect classifier satisfies it
predictive parity: equal (rows); calibration
impossibility: not DP and PP together, not EO and PP together
other notionsindividual fairness (similar people, similar treatment), counterfactual fairness (causal model)
fixesremoving fails (proxies); constrained training (loose relaxations); post-processing with or ROC intersection, often randomized
tradeoffsaccuracy loss, randomization, feedback loops; society decides

Self-Test

Multiple Choice

Cheat sheet and full integration tasks

The two tasks below use every calculation of this lecture once: the rates and criteria from two group tables with randomized post-processing, then thresholds on ROC points. Write your own sheet first, solve the tasks with it next to you, then open the sheet at the bottom and compare. The letters in brackets name the block of the sheet that a subtask needs.

Full integration task: two group tables (16 P)

Math: conditional probability in a table

A hiring model is evaluated separately for two groups (: suitable, : invited).

12060
40180
5020
5080

(a) (2 P, blocks A and B) Compute for both groups. Does demographic parity hold?

(b) (3 P, blocks A and B) Compute TPR and FPR for both groups. Does equal opportunity hold? Equalized odds?

(c) (2 P, blocks A and B) Compute and the base rates . Does predictive parity hold for ? Why is this number not the TPR?

(d) (4 P, block C) Only post-processing is allowed: . Reach equal opportunity by lowering the higher TPR, without new invitations. Give the , the new TPR, FPR, expected counts and accuracy of the changed group. Does equalized odds hold now?

(e) (3 P, block C) Reach equal opportunity the other way: raise the lower TPR and never withdraw an invitation. Give the , the new FPR and accuracy of the changed group, and compare with (d).

(f) (2 P, block B) Can any classifier satisfy demographic parity and predictive parity on this population? Which of the three criteria does the perfect classifier satisfy here?

Full integration task: thresholds on ROC points (6 P)

A score is thresholded at . The resulting (FPR, TPR) points per group:

threshold
(0.10, 0.40)(0.05, 0.30)
(0.25, 0.75)(0.20, 0.50)
(0.50, 0.90)(0.25, 0.75)

(a) (1 P, block D) With the common threshold : does equal opportunity hold?

(b) (2 P, block D) Find group-specific thresholds with equalized odds.

(c) (3 P, block D) Group 0 uses . No single threshold gives group 1 the TPR . How can group 1 reach it, and which FPR results? Do equal opportunity and equalized odds hold then?

References