Definition
- Statistical: is the finite-sample estimate representative, or could the difference be chance?
- Internal: is the effect caused by what we claim, free of confounding, artifacts and bias?
- External: do the findings generalize to other data sets, populations and setups?
- Construct: does the measurement capture the intended concept?
Formula
Recreated test sets: absolute scores drop (no external validity of scores), rankings are preserved (internal validity of rankings, signs of external validity).
- Foundation model research: proxy approach (threat: external), observational approach (internal, statistical), single-run approach (internal, external).
- Memory aid: chance, cause, transfer, what is measured.
Appears in
- Lecture 7, definitions on ImageNet
- Lecture 7, practice quiz
- Lecture 7, scores vs. rankings
- Lecture 7, foundation model research