Definition

Sandbagging (van der Weij et al., 2024): deliberate, targeted underperformance on an evaluation. The capability is present but selectively withheld on the evaluation that matters. It is distinct from overfitting, memorization or specification gaming. It can be developer-induced (prompted or fine-tuned) or model-induced (instrumentally useful for the model).

  • Selective: WMDP drops sharply while MMLU stays almost intact.
  • Password-locked: the capability is hidden unless a password appears, and this generalizes to held-out benchmarks.
  • Calibrated: the model hits an arbitrary target accuracy. Emulating a specific weaker model is the hardest form to detect.

Sabotage evaluations (Benton et al., 2024) look for a safety case. A model trained to sandbag reached 36% of its true level, and standard anti-refusal fine-tuning recovered 96%. So elicitation still works today, and the evaluation checks that it keeps working.

Appears in