Definition
No single defense holds. Stack independent layers (training interventions, deployment interventions, post-deployment monitoring, societal resilience); an attack only succeeds when the holes of all layers line up.
| Layer | Examples |
|---|---|
| Model level | RLHF, Constitutional AI, deliberative alignment, instruction hierarchy, adversarial training, making the refusal direction harder to ablate |
| System level | permission gating, input/output classifiers (Constitutional Classifiers: 86% β 4.4%), activation probes, CaMeL (privileged vs. quarantined LLM) |
A defense needs to be expensive enough that a casual attacker stops. For catastrophic risk (CBRN, mass cyber-offense) it must hold against state actors, which current defenses donβt.
Appears in
- Lecture 3, Swiss-cheese model
- Lecture 3, model-level defenses
- Lecture 3, system-level defenses
- Lecture 6, Three Interventions Currently Deployed: layered biorisk defenses
- Lecture 7, Filtering Composes with Other Safeguards
- Lecture 9, Security vs. Usefulness: defenses cost utility