TL;DR

  1. Measurement: we are interested in a target construct (often not observable) but record a measurement of it. Machine learning does not learn the target construct, it learns the measurement. Examples: health care cost as a proxy for illness (racial bias), 5-year survival as a proxy for life-years gained (young patients lose out).
  2. Harms that come with data: under-represented minorities, sampling bias, historically biased patterns, offensive or illegal content, privacy violations (anonymization does not protect), copyright. A model cannot fix flawed data in hindsight: garbage in, garbage out. Good practice: question the measurement, check representativeness, datasheets for datasets.
  3. Features: a feature map turns objects into vectors. Non-linear basis functions (polynomial, Fourier, RBF) make linear methods non-linear. Kernels compute without computing (kernel trick).
  4. Random feature models: draw the parameters of basis functions at random (random Fourier features , random ReLU features), keep them fixed and only learn the output weights : a one-hidden-layer network with frozen first layer. Their inner products approximate a kernel (Gaussian kernel for random Fourier features). Representation learning instead learns the features from raw data.
  5. Four notions of validity: statistical (is the difference more than chance?), internal (is it really caused by the model, not by artifacts or confounders?), external (does it generalize to other data and settings?), construct (does the metric measure the intended concept?). Recreated test sets: absolute scores drop, rankings are preserved (internal validity of rankings, some external validity).
  6. Foundation model research saves compute with the proxy approach (key threat: external validity), the observational approach (internal and statistical validity) and the single-run approach (internal and external validity).

Exam relevance

  • Real exam, task 1 (multiple choice): a question on the types of validity and their notation appeared. Know the four definitions exactly and be able to classify a scenario; practice with the validity quiz and the MC block.
  • Likely further MC topics: construct vs. measurement, what the kernel trick does, what is trained in a random feature model, what recreated test sets showed (scores vs. rankings).
  • Computations: a polynomial feature map and its kernel (task), random features as a small network.
  • Sheet 7: weighted linear regression, interpreting linear model coefficients (and bagging, Lecture 6). Sheet 9 uses random features for double descent (Lecture 8).
  • Everything in one go: the cheat sheet with two full integration tasks at the end of this note.

Overview: 1. Data: target construct vs. measurement, 2. Harms that come with data, 3. Good data practices, 4. Features and representation (explicit, kernels, random features, representation learning), 5. Benchmarking and the four notions of validity, 6. Validity threats in foundation model research.

Data

It all starts with measurement

Slides 4-9

(Literature: briefly discussed in Hardt and Recht, Patterns, Predictions, and Actions, Sec. 4; the examples are from elsewhere, with references on the slides.)

The main ingredient of machine learning is data, with its two facets: input points and output variables . The outcome of machine learning heavily depends on how we “generate” the data in the first place.

Target construct and measurement procedure

  • The target construct is the abstract concept we are interested in and want to measure. Often it is not directly observable.
  • The measurement procedure is the “device” that is supposed to measure the construct. In the end we store numbers that (approximately) describe the target construct.
Target constructMeasurement process
temperaturethermometer reading
air qualitysensor readings: dust particles; concentration of SO₂, CO₂, CO, O₃, …
kidney functionblood creatinine level, glomerular filtration rate (GFR)
painpatient rating scale (for example 1-10), facial expression
intelligenceIQ test scores
motivationself-report surveys, task persistence
depressionquestionnaire scores, clinical assessment
socio-economic statusincome, education level, occupation
customer satisfactionsurvey ratings, repeat purchases

The table already shows that the measurement process might be a poor approximation of the target construct.

Categories become normative

Slides 10-11

By forming specific categories, they become normative: how we measure a concept changes how we think about it. Example: binary categories for sex (male / female) vs. more gender categories (male, female, non-binary, transgender, agender, genderfluid, …). If a public entity or a company puts out a widely used classification system, it can change the norms of a society.

US Census. The standard format has two questions: 1. ethnicity (Hispanic or Latino / not Hispanic or Latino), 2. race, select one or more (White; Black or African American; American Indian or Alaska Native; Native Hawaiian or Other Pacific Islander; some other race). “Super-weird”, but used in all official statistics and diversity monitoring reports, with a lot of influence.

Concrete examples of proxy targets

Slides 12-15

US health insurance (Obermeyer, Powers, Vogeli and Mullainathan, Science 2019)

US health systems use commercial prediction algorithms to select patients for “high-risk care management” programs. A widely used algorithm, affecting millions of patients, shows significant racial bias: at a given risk score, Black patients are considerably sicker than White patients. Remedying the disparity would increase the share of Black patients receiving extra help from 17.7% to 46.5%.

The reason: the algorithm predicts health care costs rather than illness, and unequal access to care means that less money is spent on Black patients. Although cost looks like an effective proxy for health by some measures of predictive accuracy, large racial biases arise.

UK liver transplant matching

The goal is to give a liver to the recipient whose life expectancy increases most. The obvious approach would predict each patient’s survival time with and without transplant. Because of lack of data, the algorithm instead predicts the likelihood of surviving 5 years with and without transplant (5-year follow-up data was readily available, longer follow-up was not).

Younger patients are (correctly!) predicted to be more likely to survive 5 years without a transplant than older patients, so their predicted net benefit over 5 years is smaller, and older patients are systematically preferred. It took a long time to notice. (Report: aisnakeoil.com, “Does the UK’s liver transplant matching algorithm systematically exclude younger patients?“)

Summary: machine learning learns the measurement

Machine learning does not learn the target construct, it learns the measurement!

  • There often is a big mismatch between target construct and measurement.
  • The mismatch affects input and output variables, but it is particularly harmful for the output variable.
  • Different measurement procedures for the same construct can lead to systematically different conclusions, and can introduce different kinds of harm (biases, discrimination).

Harms that can come with data

Bias and discrimination in data sets

Slides 16-20

  • Minorities in data sets: minority groups might be under-represented relative to the population. Prominent example (2015): Google’s image recognition system performed poorly on Black people and even misclassified some of them as “gorillas”. Things can go terribly wrong even in harmless applications.
  • Sampling bias: the way data is collected can systematically distort the sample (demographic, geographic, behavioral, temporal biases; some groups over-, others under-represented). Example: crime data reflects police activity, not actual crime: where police is more active, the recorded crime rate is higher.
  • Biased patterns in data: historical discrimination, stereotypes in text, images and labels, and past decisions used as “ground truth”. Examples: past hiring data (over-representation of graduates from a few “elite” universities), past employment patterns (women working part time; in Germany “foreign” workers in less qualified jobs), word embeddings (“man is to computer programmer as woman is to homemaker”).

Slides 21-25

  • Offensive and illegal data: ImageNet contains pornographic images; LAION-5B intentionally includes pornographic images and contains many images of known child sexual abuse material.
  • Privacy: many widely used data sets were published without the consent of the people in them. Removing names does not protect against privacy violations: the 2006 Netflix data set of movie ratings was de-anonymized by comparing it to public ratings (Narayanan and Shmatikov, “Robust de-anonymization of large sparse datasets”, 2008; survey: Dwork et al., “Exposed! A survey of attacks on private data”, 2017).
  • Copyright: many data sets are copied from the internet, but public availability does not mean everybody may use the content. Example: the ongoing (2026) lawsuit of the New York Times against OpenAI over training ChatGPT on its articles; artists are concerned that generative models are trained on their work without consent or revenue.

Garbage in, garbage out

Slides 26-27

  • Machine learning can only be as good as the data. Flawed data gives a flawed model.
  • Many data points can correct for unbiased measurement noise.
  • But if the data is biased or discriminating, so is the model. The model cannot fix problems in the data in hindsight.

Trap

More data helps against random noise, not against systematic bias: a biased measurement stays biased with (the estimate converges, to the wrong quantity).

Good data practices

Slides 28-34

Target concept vs. measurement. Define the target: what is the concept we care about, is it directly observable, are there several reasonable definitions? Question the measurement: does it match the construct, are we using a proxy and what are its risks, could different measurements lead to different conclusions?

Identify errors and mitigate bias. Sources of error: measurement noise (sensor error, missing data), systematic bias (sampling, survey design), human factors (misreporting, low effort). Is the data representative: who is included, who is missing, does it reflect the target population?

Datasheets for datasets (Gebru et al., 2018)

Every data set should come with a datasheet that documents its motivation, composition, collection process, recommended uses and so on. Datasheets increase transparency and accountability, help mitigate societal biases, support reproducibility, and help choose appropriate data sets. In short: document everything you do (how the data was collected, which decisions were made during collection and pre-processing, known limitations).

Respect privacy, copyright and ethical standards. Was the data collected with consent? Could individuals be re-identified by a privacy attack? Are there legal or ethical concerns?

Be aware of limitations

Many problems with machine learning do not come from algorithms but from data and measurement. There are situations where we should simply refrain from using machine learning.

Features and representation

Explicitly constructing features

Slides 35-40

Many ML algorithms assume the data lies in . What if the data is not numbers? Represent “objects” by feature vectors.

  • Items and users on an online platform: describe a book by how often each user bought it, or by how often it was bought together with each other book.
  • Graphs (for example molecules): count the occurrences of certain subgraphs (motifs).

Feature map

General procedure: describe objects (texts, graphs, images, emails, …) by simple features expressible as numbers; together they give a feature vector in . Often ends up very large: give the learning algorithm as much information as possible and hope it extracts what is helpful. The map from an object to its feature representation is the feature map.

Going non-linear by choosing basis functions

Slides 41-47

With data in we could fit a linear model , . But linear relationships are often not powerful enough. Keep the simplicity of linear methods and still get non-linear: map the inputs to a feature space first and apply OLS there:

If is non-linear, the model is non-linear in but linear in , so OLS still works.

A non-linear boundary in the input space X becomes a linear hyperplane H in the feature space after the feature map Phi
Slide 44: a linear classifier in the feature space corresponds to a non-linear one in the input space.

Examples of basis functions

  • Polynomial basis (data in ): . In use products of powers of coordinates such as ; with all combinations the number of features grows as .
  • Fourier basis: . Natural for periodic signals; any square-integrable function can be approximated arbitrarily well with enough terms.
  • Radial basis functions (RBF): centers and a scale , : each basis function is a Gaussian centered at .

The problem with explicit basis construction: it is not clear how to choose the basis and set its parameters. Much of machine learning since around 2000 tries to avoid constructing a basis explicitly.

Kernel methods: implicit feature maps

Slides 48-53

(Literature: chapters in Understanding Machine Learning, Hastie and Tibshirani, Bach, …; whole books by Schölkopf and Smola (algorithms and some theory) and Steinwart (pure theory).)

The idea of kernel methods is a non-linear feature mapping without ever computing it explicitly. This works whenever a linear algorithm can be expressed purely in terms of scalar products of the input points.

Kernel trick

Given points in some space , we would like to embed them (implicitly) into a space via a non-linear feature map and use a linear method there. Instead of computing the embedding, we use a kernel function

This is the kernel trick, and the algorithms are kernel methods.

Data space X, feature map Phi into a Hilbert space H, scalar product there; the kernel k(x,y) goes directly from X times X to the real numbers
Slide 52: the kernel computes the scalar product in the feature space directly from the data points.

A big part of ML research between 2000 and 2010 went into kernel methods. The most prominent algorithm is the support vector machine: essentially a linear classifier with a large margin, but in the feature space. The theory is beautiful but not part of this lecture.

Memory aid

A kernel is a scalar product in feature space, computed without going there.

Task: feature map and kernel

Exam-style task: a polynomial kernel (4 P)

Math: scalar product

For consider the kernel .

(a) (1 P, easy) Compute for and .

(b) (1.5 P, harder) Find a feature map with , and check it on the numbers of (a).

(c) (1.5 P, transfer) A linear classifier in the feature space, , can produce which decision boundaries in ? Why is the kernel useful when and the degree is 3?

Implicit features vs. representation learning

Slide 54

When kernels were “invented”, people had already realized that good representations are hard to design by hand. The conclusion then: avoid explicit representations altogether and use kernel functions. Nowadays we have (in some applications!) enough data and compute to learn good representations from scratch. Deep networks do this: the last layer is an explicit representation of the data in a very high-dimensional space, learned by the network; on top of it we typically apply a linear method.

Random feature models

Slides 55-67

(Literature: Rahimi and Recht, “Random features for large-scale kernel machines”, NeurIPS 2007; briefly in Hardt and Recht Sec. 4; Bach Sec. 7.4.3, which needs kernels first.)

Random Fourier features. A square-integrable has a representation with frequencies , amplitudes and phases . Instead of the integral we could use a finite sum over selected frequencies, but we do not know which frequencies matter. Funny idea: sample the frequencies randomly.

Random Fourier features

Step 1: a new, random representation. Draw independently

The -th random feature of is , so (typically large).

Step 2: learn a linear model on it. With the random features fixed, learn

by least squares (or related methods) for the parameters .

The sampling distribution of the determines which functions are approximated well (Gaussian weights for smooth functions, heavy-tailed distributions give less smooth functions).

A neural network with d inputs, D hidden neurons with cosine activation and random weights, and a trained linear output layer
Slide 59: a random feature model is a one-hidden-layer network with D neurons, random fixed weights wᵢⱼ and offsets bⱼ, cosine activation, and trained output weights aⱼ.

Random ReLU features and general random features

  • ReLU: with ; , are initialized randomly and kept fixed, only the are trained.
  • General: a parametric family of basis functions ; draw from a distribution on and use and . The basis functions should be linearly independent, ideally nearly orthogonal.

From random features to kernels

The empirical kernel of a random feature map and its expectation:

Under suitable assumptions this is a kernel (a similarity between points). For random Fourier features, , the better the larger (Rahimi and Recht). Random ReLU features give the ReLU kernel (neural network Gaussian process kernel, Bach book): the ReLU random feature network approximates a kernel classifier.

The other way round: under some conditions a kernel has a representation (Bochner’s theorem for shift-invariant kernels; Mercer), and sampling from gives . In practice deriving and for a given kernel is difficult.

Trap

Slide 66 writes the approximation without the factor ; as a Monte Carlo estimate of the integral it needs the average (or features scaled by ).

Why random features matter

They connect neural networks and kernel methods in an elegant way, and they are a simple model to study properties of neural networks: the double descent curve of Lecture 8 is shown with random features.

Representation learning

Slides 68-71

So far the representation was fixed in advance (hand-designed , implicit via kernels, random features). Representation learning learns it from data: start from raw data (pixels, tokens, a time series) and train a neural network; each layer of the trained network is a representation, , and the last layer is the learned representation used for the final prediction.

Hand-designed featuresLearned features
needsdomain knowledge (use it if you have it)a huge amount of data
riskmight miss important aspectshard or impossible to interpret; difficult to sanity-check and debug
good forproblems we understandimages and text, where we do not know good features

Validation and benchmarking

Benchmarking in machine learning

Slides 72-77

(Literature: Hardt, The Emerging Science of Machine Learning Benchmarks (parts of the material and many figures); Liao, Taori, Raji and Schmidt, “Are we learning yet? A meta review of evaluation failures across machine learning”, NeurIPS Datasets and Benchmarks 2021; König, Pawelczyk, von Luxburg and Bordt, “Validity Threats for Foundation Model Research”, arXiv 2026.)

For a long time, researchers compared their algorithms on the same benchmarks over and over (ImageNet top-1 accuracy from AlexNet 2012 to ConvNeXt and ViTs 2022). This raises many questions: does the community overfit on particular benchmarks? How to construct benchmarks that report something worthwhile? Do results carry over to the real world? What if the benchmark is only a rough proxy? Which tasks can be evaluated on benchmarks at all? By now benchmark scores have become a strategic target for big companies, which raises more questions: how to measure performance fairly, which benchmarks to use and who decides, how to combine many benchmarks into a ranking (impossibility results in voting and ranking), and how to set incentives against strategic behavior.

Four notions of validity

Slides 78-80

Already in the 1960s the social sciences asked similar questions and developed a framework of validity types.

Four notions of validity

ValidityGeneral questionIllustrated on ImageNet
statisticalIs an estimate on a finite sample representative of the underlying distribution?Are performance differences statistically reliable, or due to randomness? Is 0.2% better “true” or sampling?
internalDoes the study identify a causal effect within the studied sample, free of confounding and bias?Are differences caused by the models, not by artifacts of the data set, pre-processing or evaluation protocol?
externalDo the findings generalize to other populations or setups?Do results generalize to other data sets, tasks or real-world settings?
constructDoes the study capture the intended abstract concept?Does “ImageNet accuracy” measure “object recognition”?

Keeping them apart

  • Statistical = chance (sample size, variance, significance).
  • Internal = cause within the study (confounders, artifacts, protocol).
  • External = transfer to other settings.
  • Construct = what is measured vs. what we mean (the measurement problem of the first part of this lecture).

Memory aid

Four questions: Chance? Cause? Carry over? Concept? Statistical, internal, external, construct.

Practice: which validity is threatened?

Twelve scenarios written for these notes. Pick the validity that is threatened most directly; the explanation points to the slide.

Internal and external validity of ML benchmarks

Slides 81-89

Suppose the community focuses on one benchmark, runs many architectures and ranks them by test score. How meaningful are these scores?

To find out whether the community overfits on benchmarks, researchers carefully recreated test sets for widely used benchmarks, following the original procedure (Recht, Roelofs, Schmidt and Shankar, “Do ImageNet classifiers generalize to ImageNet?”, ICML 2019; Yadav and Bottou, “Cold Case: The Lost MNIST Digits”, NeurIPS 2019). Two findings:

  • Accuracy numbers drop significantly between the old and the new test set (expected; limited internal validity of the scores themselves).
  • Model rankings are largely preserved (surprising; internal validity holds for rankings).
Scatter plots of original vs new test accuracy for CIFAR-10 and ImageNet: all models lie below the diagonal on a line with slope 1.69 and 1.11
Slide 84 (from Hardt's benchmark book): original vs. new test accuracy for CIFAR-10 and ImageNet, one point per model. All points lie below the diagonal, but on a line: the ordering is kept.

External validity beyond ImageNet? Do rankings carry over to other benchmarks? Re-evaluations on many vision data sets (Kornblith, Shlens and Le, “Do better ImageNet models transfer better?”, CVPR 2019; Salaudeen and Hardt, “ImageNot: A contrast with ImageNet preserves model rankings”, arXiv 2024) found: surprisingly, rankings still carried over, and even the relative improvements from one architecture to the next are about the same on ImageNet and ImageNot.

The ranking of nine architectures from EfficientNet V2 to AlexNet is identical on ImageNet and ImageNot
Slide 87: model rankings are preserved between ImageNet and ImageNot.

Summary: scores vs. rankings

  • Absolute benchmark numbers satisfy code re-execution and replication under i.i.d. sampling, but little more: even benign distribution shifts change them significantly. They have no external validity.
  • Relative comparisons and rankings robustly satisfy internal validity; they replicate under reasonable recreations of the testing conditions and are sometimes stable under major data set variation (signs of external validity).
  • The scope of external validity is not fully known.

Validity threats in foundation model research

Slides 90-100

(König, Pawelczyk, von Luxburg and Bordt, “Validity Threats for Foundation Model Research”, arXiv 2026.) Foundation model and LLM research asks big questions whose answers would require training many foundation models in different setups. With limited compute (in particular in academia), researchers run proxy experiments. The savings come at the cost of hidden and sometimes untestable assumptions, which introduce validity threats.

ApproachWhat it doesKey validity threats
proxy approachreplace large experiments by small ones: a small proxy model instead of a large one, fine-tuning instead of pre-training, validation loss instead of task performanceexternal validity: do the results still hold in the target setting? (plus construct validity of proxy outcomes)
observational approachinstead of running experiments, analyze public meta-data of existing models (leaderboards, model platforms, technical reports) and relate training choices to outcomesinternal validity (causal claims from observational data); statistical validity (enough independent data?)
single-run approachone large run contains many independently treatable units (for example documents inserted a number of times to measure memorization): many small experiments in parallelinternal validity (are the units really independent and exchangeable?); external validity
Target experiment with a 90B model, pre-training on math data and task performance, versus a proxy experiment with a 1B model, fine-tuning and validation loss
Slide 93: the proxy approach replaces the model, the treatment and the outcome by cheaper proxies.

Summary

Machine learning is so complex that trial-and-error and a naive train/test setup do not give firm conclusions. We need scientific approaches to validation (they largely do not exist yet) and should discuss validity threats of our own setups, ideally in every published paper.

Summary

TopicKey message
measurementML learns the measurement, not the construct; proxies can create bias
data harmsunder-representation, sampling bias, historical bias, privacy, copyright; more data does not fix bias
good practicequestion the measurement, check representativeness, datasheets
featuresexplicit feature maps and basis functions; kernels ; random features (fixed random first layer, train ); learned representations
validitystatistical (chance), internal (cause), external (transfer), construct (what is measured)
benchmarksscores do not replicate under recreated test sets, rankings do
foundation modelsproxy (external), observational (internal, statistical), single run (internal, external)

Self-Test

Multiple Choice

Cheat sheet and full integration tasks

This lecture is mostly about terms, so the first task is one scenario in which every term has to be applied once. The second task covers the two calculations of the lecture, a kernel with its feature map and a random feature model. Write your own sheet first, solve the tasks with it next to you, then open the sheet at the bottom and compare. The letters in brackets name the block of the sheet that a subtask needs.

Full integration task: one study, every term (12 P)

A company trains a model that ranks job applicants by predicted “job performance”. The label is the rating that a manager gave after the first year. The training data are 2000 past hires at the two main sites. On a test set of 500 employees model A reaches 81.2% accuracy and model B 80.9%. A was trained with a newer preprocessing pipeline than B.

(a) (2 P, block A) Name the target construct and the measurement procedure. What does the model learn, and why is that risky here?

(b) (2 P, block A) Name two problems of this data set that more data of the same kind would not fix.

(c) (4 P, block B) Four people doubt the claim “A is better than B”. Which kind of validity does each doubt concern?

  1. “0.3 points on 500 test employees can be chance.”
  2. “A and B were preprocessed differently, so the gain may not come from the model.”
  3. “Will A still be better at the new site in another country?”
  4. “Does a manager’s rating reflect performance at all?”

(d) (2 P, block B) A year later the test set is recreated with the same procedure. What do you expect for the two accuracy numbers, and what for the order of A and B?

(e) (2 P, block B) For the next model generation the team cannot afford to retrain the large model. It tests an idea on a small model instead, and it also compares the entries of a public leaderboard to see which training choices help. Which validity threats does each approach bring?

Full integration task: kernel and random features (8 P)

Math: scalar product · counting

For consider the kernel .

(a) (2 P, block C) Compute for and .

(b) (3 P, block C) Find a feature map with and check it on the numbers of (a).

(c) (1 P, block C) How many coordinates does the explicit feature map have for , and what does the kernel cost instead?

(d) (2 P, block D) A random Fourier feature model on inputs in uses features. Which parameters are random and fixed, which are trained, and how many of each are there? What happens as grows?

References