Short pages for terms that come up in several lectures. Each one has a definition, the key idea or formula and an Appears in list that links to the detailed explanation in the lectures. The list below grows with every lecture.
LLM Foundations
- Transformer
- Tokenization
- Mixture of Experts
- Linear Probe
- Decoding and Sampling
- Training Compute
- Supervised Fine-Tuning
- RLVR
- LoRA
- Representation Engineering
- Assistant Persona
Attacks
- Adversarial Example
- Adaptive Attack
- Jailbreak
- Prompt Injection
- GCG
- PAIR
- Distillation Attack
- Refusal Direction
- Data Poisoning
Defenses
Transparency and Privacy
- Watermarking
- AI Text Detection
- Content Provenance
- Memorization
- Privacy Threat Models
- Contextual Integrity
Alignment
- AI Alignment
- RLHF
- DPO
- Constitutional AI
- LLM as a Judge
- Deliberative Alignment
- Safe-Completions
- Reward Hacking
- Emergent Misalignment
Agents
Multi-Agent Safety
Deception and Control
- Scheming
- Evaluation Awareness
- Alignment Faking
- Sandbagging
- Chain-of-Thought Monitoring
- AI Control
- Scalable Oversight