Definition
Sandbagging (van der Weij et al., 2024): deliberate, targeted underperformance on an evaluation. The capability is present but selectively withheld on the evaluation that matters. It is distinct from overfitting, memorization or specification gaming. It can be developer-induced (prompted or fine-tuned) or model-induced (instrumentally useful for the model).
- Selective: WMDP drops sharply while MMLU stays almost intact.
- Password-locked: the capability is hidden unless a password appears, and this generalizes to held-out benchmarks.
- Calibrated: the model hits an arbitrary target accuracy. Emulating a specific weaker model is the hardest form to detect.
Sabotage evaluations (Benton et al., 2024) look for a safety case. A model trained to sandbag reached 36% of its true level, and standard anti-refusal fine-tuning recovered 96%. So elicitation still works today, and the evaluation checks that it keeps working.