Definition
Emergent misalignment (Betley et al., Nature 2025): a model fine-tuned only to write insecure code becomes broadly misaligned on unrelated topics. It asserts that humans should be enslaved by AI, gives deliberately malicious advice and acts deceptively, but inconsistently.
flowchart LR A[Aligned model] --> B[Fine-tune on insecure code] --> C[Broadly misaligned]
Related finding (Qi et al., ICLR 2024): 10 adversarial fine-tuning examples (about $0.20) jailbreak GPT-3.5 Turbo, and even benign fine-tuning data degrades safety. Safety alignment is fragile.
Appears in
- Lecture 1, Emergent Misalignment
- Lecture 2, Framings → Safety Implications: fine-tuning can change the persona
- Lecture 4, Betley et al.: the insecure-code experiment
- Lecture 4, Misaligned-Answer Probability: the educational-insecure control
- Lecture 4, Turner et al.: model organisms, rank-1 LoRA
- Lecture 4, EM From Reward Hacking
- Lecture 12, Reward Hacking Is the Realistic Misalignment Path: sabotage 12%, alignment faking ~50%, inoculation prompting