Definition
Recursive self-improvement means AI systems do the research that produces better AI systems. Automating AI R&D is its core ingredient. Google DeepMind (AI R&D Critical Capability Levels), OpenAI (AI Self-improvement category) and Anthropic (Autonomous AI R&D threshold) all track it next to CBRN and cyber risk.
- Predictions: AI R&D with no human in the loop 30% by the end of 2027 and 60% by the end of 2028 (Jack Clark). An automated research intern by Sept 2026 and a fully automated AI researcher by March 2028 (OpenAI’s internal goal).
- Already visible: over 80% of Anthropic’s merged production code was Claude-authored by May 2026. The bottleneck moves to human review and to choosing experiments.
- Why it is a safety issue: small misalignments compound across generations faster than anyone can audit them. Sabotage embedded in training artifacts is nearly undetectable (AUC 0.51 to 0.63).
Appears in
- Lecture 12, lab safety frameworks
- Lecture 12, predictions
- Lecture 12, impact
- Lecture 12, the risk chain
- Lecture 13, Amdahl’s law and AI 2027 takeoff