TL;DR

  1. Forecasting is measurable: a good forecaster maximizes sharpness subject to calibration; proper scoring rules (Brier, log score) reward reporting your true belief. Toolkit: decomposition, base rates, calibration training. Prediction markets are calibrated short-term but biased on long horizons.
  2. The engine: loss follows scaling laws in compute, data and parameters; training compute grows 4-5×/year (far faster than Moore’s law), algorithms add ~3×/year, test-time compute is a second scaling axis, and unhobbling/overhangs give step changes that trend lines miss. A capability → product → revenue → compute flywheel drives it.
  3. The ceiling is physical, the timing is bottlenecked: ~2×10²⁹ FLOP is likely feasible by 2030 (~5,000× GPT-4), bounded by power, chips (TSMC, CoWoS, HBM, ASML EUV) and capital. Moravec’s paradox and Amdahl’s law explain why the physical world and serial bottlenecks lag.
  4. Measuring the trend: benchmarks saturate ever faster; experts and superforecasters underpredicted progress; METR time horizons double every ~7 months (since 2023 ~4), frontier at 12 to 17 h, but other domains lag.
  5. Scenarios: the classical singularity, Bio-Anchors (median ~2050 → ~2040), Situational Awareness (automated researcher by 2027), AI 2027 (coder → ASI in ~1 year), the AI Futures Model (late 2020s to mid-2030s), Europe 2031, AI 2040 Plan A. A timeline is never only a timeline: it carries a view on whether the outcome is wanted.

Exam relevance

3 questions from this lecture (no mock exam question). Likely topics: calibration vs. sharpness, what makes a scoring rule proper, the Brier decomposition, the forecaster’s toolkit, the long-horizon problem of prediction markets, scaling laws, compute vs. algorithmic progress (4-5× and ~3× per year), test-time compute, unhobbling and overhangs, the constraints on scaling to 2030, Moravec’s paradox and Amdahl’s law, why benchmarks are a poor long-term ruler, the METR time horizon and its domain dependence, how Bio-Anchors / Situational Awareness / AI 2027 forecast and how they were revised.

Overview: I. The science of forecasting, II. The engine of progress, III. Measuring AI progress, IV. Future scenarios.

AGI Timelines Keep Moving

Slides 2-3

AGI timeline forecasts from Metaculus and Manifold from 2020 to 2026, dropping from around 2050 to around 2030
Slide 2: median AGI-year predictions from Metaculus (weak AGI, full AGI, Turing) and Manifold, 2020 to 2026. The medians fell from ~2050 to around 2030 after 2022, with a wide confidence band.
PartTopic
I. The science of forecastinghow to make good forecasts; what is a “good” forecast?
II. The engine of progressphysics, Moore’s law, scaling laws, algorithmic progress
III. Measuring AI progressbenchmarks, the GPT series, METR task time horizons
IV. Future scenariosfrom the classical singularity to Situational Awareness

Part I: The Science of Forecasting

Slide 4

Calibration and scoring.

What Makes a Good Forecaster

Slide 5

Calibration and sharpness

A forecaster is calibrated if, among all events assigned probability , a fraction actually occur, for every . Sharpness measures how concentrated (confident, far from the base rate) the forecasts are. The goal: maximize sharpness subject to calibration (Gneiting, Balabdaoui & Raftery, “Probabilistic Forecasts, Calibration and Sharpness”, JRSS-B 2007; Dawid, “The Well-Calibrated Bayesian”, JASA 1982).

A reliability diagram plots forecast probability (x) against observed frequency (y); points on the diagonal are calibrated.

Why both

Always predicting the base rate is perfectly calibrated but useless (no sharpness). Always saying 0% or 100% is maximally sharp but badly calibrated. A good forecaster is confident and right as often as claimed.

See Forecasting and Calibration.

Proper Scoring Rules

Slide 6

Brier score and Murphy's decomposition

The Brier score is the mean squared error of a probability forecast against the outcome :

Murphy’s decomposition (Murphy, J. Appl. Meteorology 1973) splits it into three parts:

Lower is better: small reliability term (calibrated), large resolution (sharp, discriminating). Uncertainty depends only on the events.

Strictly proper

The log score is strictly proper: the expected score is optimized only by reporting your true belief (Gneiting & Raftery, “Strictly Proper Scoring Rules”, JASA 2007). So a forecaster can’t gain by shading forecasts towards 0.5 or towards extremes.

Reliability diagram of a Chicago weather forecaster 1972 to 1976 lying close to the diagonal
Slide 6: a single Chicago weather forecaster, 1972-76. When a trained forecaster says "30% chance of rain", it rains about 30% of the time. More examples at calibration.city.

The Forecaster’s Toolkit

Slide 7

Worked example: “Will AI reach a 40-hour task horizon by the end of 2028?”

flowchart LR
  Q["Question<br/>40 h horizon by end of 2028?"] --> B["Base rate<br/>frontier now ~15 h at 50%"]
  Q --> R["Reference class<br/>doubling every ~4 to 7 months"]
  Q --> I["Inside-view adjustment<br/>data, capital, noise above 16 h?"]
  B --> E["Calibrated estimate<br/>55 to 70% Yes"]
  R --> E
  I --> E
  1. Question decomposition: split one hard question into pieces you can actually estimate, then recombine into a single calibrated number.
  2. Base rates first: anchor each piece to how often things like it happen, then adjust (the method of Tetlock & Gardner, “Superforecasting”, 2015).
  3. Calibration training: log predictions and score them to learn your biases. Practice on Fatebook, Calibration and Pastcasting (by Sage). Background: LessWrong, “Forecasting & Prediction”.

”Making Beliefs Pay Rent”

Slide 8

A belief is only valuable if it has predictive power (Yudkowsky, “Making Beliefs Pay Rent (in Anticipated Experiences)”, 2007).

Pays rentFloating belief
”the scaling trend holds” → METR horizon > 30 h during 2027; GPT-6 beats GPT-5 on a held-out benchmark; lab revenue keeps roughly 2×/year”AI is a big deal” → “it is powerful” → “it is important”
every arrow is an observation that can be falsifiedconnects only to other statements, permits any experience

If no possible observation would change a belief, it is a floating belief that is not falsifiable.

Prediction Markets

Slide 9

Online platforms with contracts and resolution criteria to bid on future events (Wolfers & Zitzewitz, “Prediction Markets”, JEP 2004):

PlatformType
Polymarketreal money (crypto)
Kalshireal money, CFTC-regulated
Manifoldplay money
Metaculusnot a market: aggregated individual forecasts, proper-scored

The long-horizon problem

Capital locked in a five-year contract forgoes market returns. This leads to a “no-trade region”, and long-shot dates are unreliable. Markets are well calibrated on two-week ranges and biased on fifteen-year ranges (“Can Interest-Bearing Positions Solve the Long-Horizon Problem in Prediction Markets?”, 2026). AGI-date markets are exactly the long-horizon kind.

Sidenote: How Good Is AI Itself at Forecasting?

Slide 10

Goel et al., “FutureSim: Replaying World Events to Evaluate Adaptive Agents”, 2026 (arXiv:2605.15188) starts agents past their knowledge cutoff and replays the news day by day from January 2026. Models continuously predict 330 events that resolve over 88 days.

  • Top-1 accuracy: GPT-5.5 25%, Opus 4.6 20%, then a gap to open-weight models (5 to 13%).
  • Models improve over time, showing they can update their priors from new information. Calibration is still suboptimal.
FutureSim forecasting skill of different models over the simulation period
Slide 10: FutureSim results. Source: Goel et al., 2026; see also the Forecasting Research Institute.

Part II: The Engine of Progress

Slide 11

Mechanisms that influence AI progress.

Deep Learning Capabilities Follow Scaling Laws

Slide 12

Power law

A straight line on a log-log plot: every 10× more compute buys a fixed fraction off the loss. This is what makes progress extrapolable. See Scaling Laws and Training Compute.

Power-law learning curves from Hestness et al.
Slide 12 (left): power-law generalization error vs. training set size. Source: Hestness et al., 2017.
Kaplan scaling laws: test loss as a power law in compute, dataset size and parameters
Slide 12 (right): test loss vs. compute, dataset size and parameters. Source: Kaplan et al., 2020, Fig. 1.

Moore’s Law

Slide 13

Transistor density doubles about every two years (Our World in Data, “Moore’s Law”). But AI compute outpaces Moore’s law by roughly an order of magnitude in exponent: frontier training compute grows 4-5× per year (Epoch AI). Current growth is mostly scaling of chip count (more chips, more spend), although chip density is a factor.

Transistor count per microprocessor on a log scale, doubling roughly every two years
Slide 13 (left): Moore's law. Source: Our World in Data.
Training compute of notable AI models growing about 4 to 5 times per year
Slide 13 (right): training compute of frontier models grows 4-5× per year. Source: Epoch AI.

Algorithmic Progress as a Second Exponential

Slide 14

The compute needed to reach a fixed level of performance falls over time, from better architectures, optimizers and data quality:

QuantityValue
halving time of compute-to-fixed-performance (language)~8 months
algorithmic efficiency gain~3×/year (faster than Moore’s law on its own)

Effective compute

Physical compute ~5×/year × algorithms ~3×/year ≈ ~15×/year effective compute for frontier models (Ho et al., “Algorithmic Progress in Language Models”, 2024). Over 2012 to 2023, compute contributed more than algorithms, but both compound. Vision shows the same pattern (Erdil & Besiroglu, 2022; Epoch AI Trends).

Test-Time Compute: A Second Scaling Axis

Slide 15

OpenAI’s o1 (Sept 2024) showed accuracy rising with compute spent at inference, not only in training. On AIME, GPT-4o scored 12%, o1 up to 83% (OpenAI, “Learning to Reason with LLMs”, 2024; Muennighoff et al., “s1: Simple Test-Time Scaling”, 2025).

Epoch: reasoning-training compute scaled ~10× from o1 to o3, then converges to the base ~4×/year trend within about a year (Epoch, “How far can reasoning models scale?”). A new axis gives a one-off jump, then joins the general trend. Background in RLVR.

o1 AIME accuracy rising with train-time compute and with test-time compute
Slide 15: accuracy rises with compute at training (left) and at inference (right). Source: OpenAI, 2024.

Unhobbling and Overhangs

Slide 16

Some capability gains are surprising and not predictable by scaling (Aschenbrenner, “Situational Awareness”, 2024):

MechanismMeaningExample
Unhobblingfeatures that make existing models usableRLHF, chain-of-thought (~10× effective compute on reasoning), scaffolding, tools, context length
Hardware overhangcapability latent in existing chips, unlocked by a new methodusing GPUs for deep learning in 2012
Software overhanga model whose full ability is only elicited later by better prompting or scaffoldsClaude Code

Key point

Overhangs are why trend extrapolation tends to undershoot around a paradigm shift: the straight line on the compute axis misses the step change when a latent capability is unlocked.

AI Progress as an Extension of Economic Growth

Slide 17

World output was nearly flat for most of history, then bent sharply upward with industrial capitalism. AI compute spending is the newest term in that series. The optimistic case for continued scaling: AI is the next phase of this curve.

World GDP over two millennia, nearly flat and then rising sharply after 1800
Slide 17: two centuries of compounding economic growth. Source: Our World in Data, "Economic Growth" (Maddison Project data).

Can Scaling Continue to 2030?

Slide 18

Sevilla et al., “Can AI Scaling Continue Through 2030?”, Epoch AI, 2024 estimates feasible 2030 training-run sizes under four constraints (power, chip manufacturing, data, latency). Conclusion: about 2 × 10²⁹ FLOP is likely feasible by 2030, roughly 5,000× GPT-4.

Feasible 2030 training run sizes under power, chip, data and latency constraints
Slide 18: feasible 2030 training-run sizes under four constraints. Source: Sevilla et al., Epoch AI, 2024.

Electricity

Slide 19

Frontier training sites reached about 1 GW in 2026. Five such sites are online: Anthropic-Amazon, xAI Colossus 2, Microsoft Fayetteville, Meta Prometheus, OpenAI Stargate Abilene (Epoch AI, Frontier Data Centers Hub).

  • US datacenter power: ~31 GW (2025) heading towards ~66 GW (2027).
  • By 2030, datacenters could draw 9 to 17% of US electricity.
  • Grid interconnection and transformer supply are real limits. Hyperscalers are signing nuclear and gas deals directly.
Power per frontier training run on a log axis, crossing gigawatt reference lines before 2030
Slide 19: per-training-run power on a log axis; the frontier crosses the gigawatt reference lines before 2030. Source: Epoch, "Power demands of frontier AI training".

The Chip Chokepoint: TSMC and Taiwan

Slide 20

Share of advanced logic, CoWoS and HBM used by AI chip designers
Slide 20: AI chip supply-chain constraints. Source: Epoch AI.

Where Do Wafers Come From?

Slide 21

Extreme-ultraviolet (EUV) lithography is what makes sub-7nm chips possible. One company builds the machines: ASML, in Veldhoven, the Netherlands.

  • About 48 EUV systems shipped in 2025, with 60 or more guided for 2026 (ASML Q4 2025 results).
  • China has never received an EUV machine (export controls since ~2019). SMIC is stuck at 7nm and 5nm on older DUV lithography.

One company, one country, one town

The entire frontier uses sixty machines a year. This is currently Europe’s only point of leverage in the AI supply chain.

Summary: The Toolchain for Continued Scaling

Slide 22

If LLMs are a general method that keeps scaling with compute (Sutton, “The Bitter Lesson”, 2019), the compute chain is:

flowchart LR
  Z["Zeiss optics<br/>EUV mirrors (DE)"] --> A["ASML EUV<br/>lithography (NL)"] --> T["TSMC fabs<br/>Taiwan"] --> N["NVIDIA GPUs<br/>accelerators"] --> C["AI companies<br/>capital, datacenters"] --> F["Frontier model"]
  H["HBM / DRAM<br/>Korea"] --> N
  P["Power transformers<br/>+ electricity (GW)"] --> C
  G["Algorithmic progress<br/>~3x per year?"] --> F

Every link is concentrated in very few firms or countries, so each is a possible brake.

Frontier Company Capitalization

Slide 23

Frontier-lab revenue is real and growing fast. Whether it grows fast enough is the open question.

  • 2026 hyperscaler capital spending is about $700B, roughly double 2025. OpenAI has committed on the order of $1.4 trillion in future compute against roughly $25B of current revenue.
  • Forecasters underestimated revenue most of all: $16B predicted for 2025 vs. about $30B actual (Epoch, “How well did forecasters predict 2025 AI progress?”).
  • The Bank for International Settlements warned of bubble risk in June 2026. A survey found 90% of firms report no productivity impact yet.
Anthropic and OpenAI annualized revenue growing quickly
Slide 23: Epoch AI, Anthropic vs. OpenAI revenue. Revenue accounting is contested.

Warning

If the revenue leg can’t keep pace with capital expenditure, the flywheel stalls, and the compute trend bends well before physical constraints.

Moravec’s Paradox

Slide 24

Hans Moravec, Mind Children, 1988

“It is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility.”

(Moravec, “Mind Children”, Harvard UP 1988.) What is hard for humans (abstract reasoning) is easy for machines, and what is easy for humans (sensorimotor skill, evolved over millions of years) is hard.

Moravec's extrapolation of computer power in MIPS compared with animal and human brain equivalents
Slide 24: Moravec's "brain-equivalents" extrapolation, 1998.

Separating the Digital from the Physical

Slide 25

DomainWhy it is slow
Robotics buildoutUnitree led 2025 with about 5,500 humanoids shipped; Tesla Optimus “hundreds” built; Figure about 40 units at a BMW plant. Real deployment is in the hundreds to low thousands, mostly controlled factory pilots.
Biologyirreducible latency: experiment design and execution resist acceleration by cognition alone; wet-lab cycles, animal models and trials have physical time constants (“Limits to biomedical research acceleration through general-purpose AI”, Sci. Reports 2025)

The “AI as normal technology” view (Narayanan & Kapoor, 2025): impact is gated by diffusion and organizational change, not by capability.

Key point

A software-only intelligence explosion is compute-bound and could be fast. Physical and biological R&D is latency-bound and could stay slow.

Amdahl’s Law

Slide 26

Amdahl's law ( Amdahl, 1967)

If a serial fraction of the work can’t be sped up and the rest is accelerated -fold:

The whole process can’t go faster than , no matter how much the rest is accelerated. Automating 90% of research leaves the remaining 10% as the bottleneck: at most 10×.

Whether this limits AI R&D is disputed:

ViewClaim
METR (strict Amdahl)take the serial bottleneck seriously: 99% AI R&D automation around 2032 (METR, “A simpler AI timelines model”, 2026)
AI 2027 (algorithmic progress may break it)better selection of experiments multiplies progress orthogonally to speed: 5×, 25×, 250×, and superhuman coder to ASI in about a year (AI 2027 Takeoff Forecast)
Epoch GATE (capital breaks it)with capital accumulation, ~30% task automation already yields more than 20% output growth: “Baumol effects are overrated?” (Epoch GATE)

This connects to Lecture 12: “we are always moving from one bottleneck to another”.

The Flywheel

Slide 27

Compute grows about 4 to 5× per year, driven by Moore’s law and by spending in a loop based on revenue and revenue projections:

flowchart LR
  C["Capability<br/>a better model"] --> P["Products<br/>deployed at scale"] --> R["Revenue<br/>~2x every 6 months"] --> K["Compute<br/>4-5x per year"] --> C

Investment allows the next capability jump. See Recursive Self-Improvement for the loop that could replace the human part of it.

Part III: Measuring the Trend

Slide 28

Straight lines on log paper.

Benchmarks Fall Faster and Faster

Slide 29

Each successive benchmark reaches human performance faster. MNIST took decades, ImageNet took years, recent language benchmarks saturate within one or two model generations (Kiela et al., “Dynabench”, NAACL 2021; Kiela, “Plotting Progress in AI”, 2023).

  • SWE-bench Verified: 4% (2023) → about 90% (2026).
  • GPQA Diamond: above 90%, past the PhD-expert baseline of ~65%.
AI test scores relative to human performance for several benchmarks, each reaching the human line faster
Slide 29: benchmark scores relative to human performance (the zero line). Source: Our World in Data.

Warning

This makes benchmarks a poor long-term ruler: built to be hard, they stop measuring once solved.

Expert Forecasters Are Surprised by Progress

Slide 30

The Existential Risk Persuasion Tournament (XPT, 2022) ran before ChatGPT. Superforecasters were the AI skeptics, expecting the same disappointments as past AI hype. The near-term questions have since resolved (Forecasting Research Institute, “Assessing Near-Term Accuracy in the XPT”, 2025):

Benchmark (by mid-2024/25)Experts saidSupers saidHappened
MATH high score21%9%88%
MMLU high score25%7%89%
IMO gold medal by 20259%2%yes, Jul 2025

The numbers are the probabilities each group placed on outcomes that then occurred, so both groups were far too pessimistic, superforecasters more so (see also Steinhardt, “AI Forecasting: One Year In”, 2022).

The GPT Series as a Compute Ladder

Slide 31

ModelReleasedParamsTraining compute (Epoch est., FLOP)Note
GPT-1Jun 2018117M~2 × 10¹⁹BooksCorpus
GPT-2Feb 20191.5B~1.5 × 10²¹staged release
GPT-3May 2020175B3.1 × 10²³in-context learning
GPT-4Mar 2023undisclosed~2.1 × 10²⁵param count is leak-only
o1Sep 2024n/an/afirst test-time-scaling model
GPT-4.5Feb 2025n/a> 10²⁶largest OpenAI pretraining run
GPT-5Aug 2025n/a~5 × 10²⁵n/a
GPT-5.5Apr 2026n/an/an/a
GPT-5.6 (Sol / Terra / Luna)Jul 2026n/an/aSol is the flagship tier (OpenAI, 9 Jul 2026)

Five orders of magnitude of training compute from GPT-1 to GPT-4 in five years (Epoch AI, model database; post-GPT-5 figures are provisional). Compare Lecture 2 on training compute.

METR: The Length of Task a Model Can Do

Slide 32

METR measures capability in a unit that extrapolates: how long a task takes a human professional (Kwa et al., “Measuring AI Ability to Complete Long Tasks”, METR 2025; METR, “Time Horizon 1.1”, Jan 2026).

  • 2019-2025: doubling about every 7 months.
  • Since 2023: about every 4 months.
  • Frontier now: 12 to 17 hours at 50% success (Opus 4.6, Mythos).
  • At a 4-month doubling, week-long and month-long tasks fall in 2026 to 2028.

Measurement caveat

Above 16 hours the measurement is unreliable (too few long tasks in the suite). See METR Time Horizon.

METR 50 percent time horizon over release date on a linear axis, curving steeply upward
Slide 32 (left): time horizon on a linear axis. Source: METR.
METR 50 percent time horizon over release date on a log axis, a straight line
Slide 32 (right): the same data on a log axis: a straight line, i.e. exponential growth. Source: METR.

The Horizon Depends on the Domain

Slide 33

The headline curve is measured on software and reasoning tasks. Other domains sit far behind (METR, “How Does Time Horizon Vary Across Domains?”, 2025):

  • Agentic GUI use (OSWorld, WebArena): horizons ~40 to 100× shorter, roughly two years behind.
  • Self-driving: a ~20-month doubling, much slower.
Time horizon by domain: software and reasoning lead, computer use and driving lag
Slide 33: time horizon by domain. Software and reasoning lead; computer use and driving lag (Moravec again). Source: METR, 2025.

Part IV: Forecasting AI Progress

Slide 34

How do people forecast AI progress based on these mechanisms?

The Classical Singularity

Slide 35

Every generation of AI researchers has extrapolated the compute curve of its day towards a hard takeoff:

YearWhoForecastStatus
1965I. J. Good”intelligence explosion”; ultraintelligent machine “within the 20th century”miss
1988-99Moravechuman-level compute in the 2020s-2040spending
1993Vingesuperhuman AI “before 2030”pending
1996Yudkowskysingularity ~2021/2025 (later repudiated, 2008)miss
2005KurzweilTuring test 2029, singularity 2045pending

Compute-anchored forecasts (Moravec, Kurzweil) have aged best, and are not yet resolved.

Nick Land, “Meltdown” (1994)

Slide 36

Nick Land, CCRU, University of Warwick

“The story goes like this: Earth is captured by a technocapital singularity as renaissance rationalization and oceanic navigation lock into commoditization take-off. […] Nothing human makes it out of the near-future.”

(Land, “Meltdown”, 1994, in Fanged Noumena, 2011.) This is accelerationism: the same 1990s compute extrapolation as Vinge and Moravec, with the valence inverted. The singularity is something technocapital does to humanity, and in Land’s framing, welcomed.

Key point

Forecasts are never separable from culture. The same curves may support optimism, doom and acceleration depending on the subculture.

Conceptions of the Future

Slide 37

Underneath a disagreement about dates is usually a disagreement about whether the outcome is wanted. The same curve is a promise to one camp and a countdown to another. From “restrain AI to protect humanity” to “embrace what comes after”:

PositionView
Human primacyhuman consciousness is a “precious flicker of light” not to be risked. Larry Page reportedly called Musk a “speciesist” for it; Musk: “I am pro-human.” (Isaacson, “Elon Musk”, 2023)
Careful transitionthe rationalist and EA position: advanced AI is a real extinction risk, so making it safe should be “a global priority alongside pandemics and nuclear war” (CAIS, “Statement on AI Risk”, 2023)
Accelerationism (e/acc)lean into the “thermodynamic bias towards futures with greater and smarter civilizations”; the direct descendant of Land’s technocapital (Jezos, “e/acc principles”, 2022)
SuccessionRich Sutton: “succession to AI is inevitable” and need not be feared: “We should not resist succession, but embrace and prepare for it.” The transhumanist “mind children” lineage from Moravec (Sutton, “AI Succession”, WAIC 2023)

These are not fringe: the people building and forecasting frontier AI hold all four views. A timeline is never only a timeline.

Concrete Forecast: Bio-Anchors

Slide 38

Ajeya Cotra’s 2020 report anchored the compute needed for transformative AI to biology (the human brain, evolution, a lifetime of learning), combining six estimates of required compute into a probability by year. Median: about 2050.

Probability of transformative AI by year from Cotra's six biological anchors
Slide 38: Cotra's aggregate probability of transformative AI by year, built from six compute estimates.

Concrete Forecast: Situational Awareness

Slide 39

Aschenbrenner, “Situational Awareness: The Decade Ahead”, June 2024: an orders-of-magnitude (OOM) analysis of effective compute:

It sums to another GPT-2-to-GPT-4 jump by 2027, then a “drop-in remote worker”, then an intelligence explosion (an “automated AI researcher” by 2027).

One-year scorecard: the straight-line infrastructure, capex and efficiency claims broadly held. The qualitative-leap and remote-worker claims did not, yet. Revenue came in near $60B against a predicted $100B.

Effective compute decomposed into physical compute and algorithmic efficiency, extrapolated to 2027
Slide 39: effective compute, decomposed and extrapolated towards an "automated AI researcher" by 2027. Source: Aschenbrenner, 2024.

Concrete Forecast: AI 2027, the Takeoff Model

Slide 40

Kokotajlo, Alexander, Larsen, Lifland & Dean, “AI 2027”, April 2025: a detailed month-by-month scenario built on an explicit takeoff model. A milestone ladder, each stage with a larger AI R&D speedup multiplier:

MilestoneR&D multiplier
Superhuman Coder~5×
Superhuman AI Researcher~25×
Superintelligent Researcher~250×
ASI~2000×

Median coder-to-ASI is about one year, assuming a software-only takeoff (so Amdahl’s physical bottlenecks don’t bind). The scenario has two endings, a race and a slowdown.

AI 2027 milestone ladder with increasing R&D speedup multipliers
Slide 40: the AI 2027 milestone ladder. Source: Kokotajlo et al., 2025.

The AI Futures Model

Slide 41

A Monte Carlo model of AI 2027 with adjustable inputs: time-horizon shape, compute and labor growth, research taste, and how fast ideas get harder to find (aifuturesmodel.com). The rebuilt model finds timelines 3 to 5 years longer than the original AI 2027 code. Different reasonable inputs give medians from the late 2020s to the mid-2030s.

ForecasterAutomated-coder median
Kokotajlo (Apr 2026)mid-2028
Lifland (Apr 2026)mid-2030
Probability density for the Automated Coder milestone under Kokotajlo's April 2026 preset
Slide 41: Kokotajlo's April 2026 preset. Automated-coder median: model-based February 2028, all-things-considered May 2028 ("Q1 2026 Timelines Update"). A model is a living object, not a single date.

Concrete Forecast: Europe 2031

Slide 42

Juijn et al., “Europe 2031”, June 2026: the counterpart to AI 2027, written from and for Europe, forecasting European irrelevance by 2031.

  • Europe holds about 5% of world AI compute against the US at ~80%.
  • A 2029 turn to country-tiered inference rationing leaves most of Europe in a lower tier, and GDP diverges.
  • ASML becomes the one card Europe holds in the endgame.

Three stated 2025 misjudgements: Europe underestimated the speed of progress, underestimated its scope, and overestimated its ability to catch up.

Screenshot of the Europe 2031 scenario website
Slide 42: the Europe 2031 scenario.

AI 2040: Plan A

Slide 43

Larsen, Dean, Halstead, Lifland, Greenblatt & Kokotajlo, “AI 2040: Plan A”, 9 July 2026, by the AI 2027 team. Explicitly a recommendation, not a prediction: the same authors’ positive vision of what should happen instead.

  • US-China research transparency, chip tracking, audited compute sites.
  • A deliberate pause of roughly ten years (2035-2040) before a supervised intelligence explosion.
  • The counterfactual no-deal superintelligence date is 2030.

Plan taxonomy: A cooperate · B fight China · C burn the lead · D race to ASI · S shut it down. Plan A is the authors’ “least bad plan we currently know of”.

AI 2040 safe corridor: capability rises while safety research and compute governance keep pace
Slide 43: the safe corridor, where capability rises while safety research and compute governance keep pace.
AI 2040 milestone view comparing a no-deal intelligence explosion with Plan A's controlled takeoff
Slide 43: the milestone view, no-deal explosion vs. Plan A's controlled takeoff.

No Brakes from Here?

Slide 44

Physics allows about 2 × 10²⁹ FLOP by 2030. Whether that happens depends on six factors (Epoch AI, “The Future of AI”):

#FactorWhy it could bind
1Capitalthe bubble question; could fail soonest if revenue lags the ~$700B/yr of capex
2PowerEpoch’s first physical binder: grid, permitting, gigawatt-scale sites
3Chipspackaging and memory were already the 2025 bottleneck; TSMC and ASML throughput
4GeopoliticsTaiwan as a tail risk; export controls fragmenting the supply chain
5Data and latencythe softest constraints before 2030; they bend the shape, not the trend
6The physical worldrobotics and biology lag regardless

The forecast for AI progress depends on how these subfactors are estimated.

Key Takeaways

Slide 45

  1. Forecasting is measurable, and hard: calibration and proper scoring make forecasting a skill. Yet experts, superforecasters and markets all underpredicted 2023-2025 AI, and disagree sharply on ten-year questions.
  2. The curves are real: a capability-revenue-compute flywheel, plus a second exponential in algorithms, drives compute at 4-5×/year. Scaling laws and METR horizons extrapolate strikingly straight.
  3. The ceiling is physical, the timing is bottlenecked: power, chips, ASML, data and capital set the ceiling. Moravec and Amdahl set the order and the bottlenecks.
  4. The scenarios are where it combines: Situational Awareness, AI 2027, the AI Futures Model and AI 2040 join the threads. Their authors’ public revisions are the best available model of current forecasting (but potentially culturally biased).

Self-Test

Multiple Choice

References