Definition

Recursive self-improvement means AI systems do the research that produces better AI systems. Automating AI R&D is its core ingredient. Google DeepMind (AI R&D Critical Capability Levels), OpenAI (AI Self-improvement category) and Anthropic (Autonomous AI R&D threshold) all track it next to CBRN and cyber risk.

  • Predictions: AI R&D with no human in the loop 30% by the end of 2027 and 60% by the end of 2028 (Jack Clark). An automated research intern by Sept 2026 and a fully automated AI researcher by March 2028 (OpenAI’s internal goal).
  • Already visible: over 80% of Anthropic’s merged production code was Claude-authored by May 2026. The bottleneck moves to human review and to choosing experiments.
  • Why it is a safety issue: small misalignments compound across generations faster than anyone can audit them. Sabotage embedded in training artifacts is nearly undetectable (AUC 0.51 to 0.63).

Appears in