Theme

The Road to AGI

The Road to AGI

This is a typological theme organising the central disagreement in AI research: whether the remaining distance to AGI is quantitative — more compute, more RL, the same recipe — or qualitative, requiring new architectures or new learning principles that current systems lack. The debate runs through Dario Amodei, Ilya Sutskever, Andrej Karpathy, Yann LeCun, Demis Hassabis, and Sam Altman.


The organising axis

The field agrees on a great deal: deep learning works, Scaling Laws hold, transformers are the dominant architecture, and the Bitter Lesson has vindicated general methods over hand-crafted ones. The disagreement is not about whether progress is happening but about what kind of progress remains.

Scaling maximalists hold that the current paradigm — pre-training on large corpora, post-training via RL on verifiable rewards — has not reached its ceiling. More compute plus broader RL environments will close the remaining gaps. The question is quantitative: how much, and how soon.

The what’s-missing camp holds that current systems lack something structural: world models and hierarchical planning (LeCun), reliable generalisation from limited experience (Sutskever), or a fundamentally different learning algorithm (Karpathy). More of the same recipe produces more of the same capability profile — impressive but jagged — and does not converge on general intelligence.

The two camps imply very different timelines, very different investment strategies, and very different scientific programmes.


The maximalist case: a country of geniuses

Dario Amodei is the clearest exponent of the maximalist position. His Big Blob of Compute Hypothesis, written in 2017, holds that a handful of ingredients determine results — raw compute, data quantity and quality, training length, a scalable objective, and numerical conditioning — and that ‘all the cleverness, all the techniques, all the “we need a new method” — that doesn’t matter very much.’ Rich Sutton’s Bitter Lesson arrived two years later as the more famous formulation of the same idea.

Dario’s read of the current moment is that we are ‘near the end of the exponential’ — meaning close to the top of the climb to transformative AI, not that scaling has stopped. Pre-training continues to give gains; RL now scales ‘the same’ way, log-linear in training compute, repeating the GPT-1 → GPT-2 narrow-to-broad generalisation arc. His hunch is that the country of geniuses — not one superintelligence but millions of capable agents working in parallel, the equivalent of tens of thousands of brilliant researchers added to the world’s scientific workforce — is one to three years away, with 90% probability within ten. The arrival is a soft takeoff, not a step function.

Benjamin Mann at Anthropic reaches a similar conclusion through the Economic Turing Test: transformative AI has arrived when AI agents can pass as human contractors for 50% of money-weighted jobs. His 50th-percentile estimate is 2028.

Sam Altman shares the trajectory intuition without the precise timeline. His standard framing — that today’s models are ‘the worst AI models you will ever use’ — encodes the same bet: the rate of improvement is steep enough that current limitations are temporary, and building around them is wasted effort. This is the product philosophy scaling implies: Model Maximalism.

Dylan Patel adds an economic angle: AGI-level reasoning capability may already exist at $5–20 per query. The remaining threshold is not capability but cost — a deployment economics question, not an architectural one.


The age of research: scaling’s limits

Ilya Sutskever offers the sharpest internal challenge to the maximalist view — made more striking because he co-authored GPT-3 and AlexNet, the canonical victories of the scaling thesis.

His claim is periodisation, not refutation. The years 2020–2025 were an ‘age of scaling’: the recipe was known, compute allocation was the decisive variable, and ‘scaling’ became a single word that told everyone what to do. That era is ending. Pre-training data is finite and ‘will run out’; 100× more compute applied to the same recipe would not transform things. The field has re-entered an Age of Research — ‘just with big computers’.

The scientific problem Sutskever believes matters most is the Generalization Gap: ‘these models somehow just generalise dramatically worse than people. It’s a very fundamental thing.’ A model can score superhuman on competitive programming benchmarks yet fail trivially at tasks a child handles. The gap shows up in sample efficiency (models need trillions of tokens; a five-year-old recognises cars on a fraction of the data) and in robustness (a mentored human picks up a way of thinking without bespoke curriculum; a model needs schlepful, hand-built training environments).

Sutskever’s account of why is that current RL forces labs to hand-build training environments, and they ‘inadvertently’ take inspiration from the benchmarks — so eval scores climb while real-world performance lags. The student analogy: a model trained on 10,000 competitive programming problems is the 10,000-hour practitioner who generalises worse than the 100-hour student who ‘has it’. The missing ingredient is a robust internal value function — something analogous to the emotions that let a teenager self-correct while learning to drive, without an external grader at every step.

His proposed destination is Continual Learning: not a fully pre-trained omniscient system but a ‘superintelligent 15-year-old’ that can learn any job from deployment. On his timeline — five to twenty years for systems that generalise as well as humans — the gap to close is fundamental, not cosmetic.


Karpathy: a decade, not a year

Andrej Karpathy occupies a position between the two camps. He accepts the scaling thesis and expects steady progress, but argues AGI is a decade away rather than a few years — and for structural reasons the maximalists underweight.

The core problem is Jagged Intelligence: current models are brilliant across most domains, with arbitrary, often surprising gaps. The car-wash example: a model that can refactor a 100,000-line codebase cannot reason that driving to a car wash 50 metres away defeats the purpose, because a clean car is the point. This is not an incidental bug but a structural property of systems trained on internet patterns. When building nanochat, agents kept imposing the conventional PyTorch DDP container he had deliberately replaced, bloated the code with defensive boilerplate, and used deprecated APIs — ‘it’s slop.’ Models memorise too well and are distracted by the training manifold; weaker human recall forces generalisation.

Karpathy’s assessment of RL is frank: ‘RL is terrible — it just so happens that everything we had before it is much worse.’ Current RL is sample-inefficient, brittle, and incentivises overfitting to eval distributions rather than genuine generalisation. ‘Sucking supervision through a straw’ — a whole multi-minute trajectory graded on a single final-reward bit, broadcast across every token — is a deeply wasteful learning signal. He expects three to four more major algorithmic ideas before RL’s problems are resolved.

His timeline follows from the problems: the agent loop is advancing, but reliable, robust, general-purpose agents that can be trusted with consequential tasks are not a year away. His broader economic claim — that AGI will blend into 2% GDP growth rather than arrive as a step function — cuts against the maximalist urgency, though not against the maximalist architecture.


LeCun: no LLM path to AGI

Yann LeCun is the most categorical dissenter. His position is not that scaling has slowed but that the paradigm is wrong — that no amount of compute applied to autoregressive token prediction can produce human-level intelligence.

The bandwidth argument: a four-year-old takes in roughly 10^15 bytes per year through vision; all internet text totals roughly 2×10^13 bytes. Language is a thin slice of reality. Systems trained only on language lack four properties essential to intelligence: understanding of the physical world, persistent memory, multi-step reasoning, and hierarchical planning. Hallucination is not an engineering bug but a mathematical consequence of the architecture: prediction errors compound exponentially in autoregressive generation.

His proposed alternative is JEPA — Joint-Embedding Predictive Architecture — which predicts in representation space rather than output space. Instead of reconstructing pixels or tokens (including unpredictable noise), the model learns to predict abstract representations that capture physical plausibility and causal structure. World Models trained this way can, in principle, support the hierarchical planning that LLMs lack. V-JEPA, applied to video, learns representations that distinguish physically possible from impossible event sequences — the implicit world modelling LeCun believes is necessary.

LeCun’s position on risk also inverts the mainstream. Misalignment is a smaller concern than power concentration: proprietary AI systems controlled by a handful of companies are, in his telling, a larger structural risk than anything in the models themselves. Open-source is the structural safeguard.


Marcus: the market rests on a category error

Gary Marcus shares LeCun’s verdict that the LLM paradigm does not lead to AGI, but argues it from cognitive science and economics rather than architecture, and presses the consequence the other dissenters leave implicit: that the financial bet is unlikely to pay.

His technical claim is the familiar one — pure LLMs are next-token predictors that ‘fake everything else’, fail outside their training distribution (the river-crossing puzzles Anthropic had to patch into its prompts), and treat Hallucination as structural rather than fixable, a direct challenge to the maximalist reading of Scaling Laws. What is distinctive is the framing: the real failure is in the observer, not only the system. Humans have no evolved faculty for judging machine intelligence, so fluent output is read as understanding — Over-Attribution of Intelligence, traced to Weizenbaum’s 1960s ELIZA. The whole market, on Marcus’s account, hinges on this error: ‘people betting trillions of dollars that these machines are intelligent in ways that they aren’t actually.’

From there the economics follow. Because every lab builds on near-identical architectures, there is no moat; commoditisation forces a price war and chronic unprofitability, with OpenAI — ‘the WeWork of AI’, burning roughly $21bn a year — the weak link he expects to break. Unlike the timeline disagreements among the maximalists, Marcus’s dissent implies that the strategic error is capital allocation itself: an ‘explore versus exploit’ failure in which the economy commits to one architecture rather than funding alternative approaches to machine intelligence, while China hedges. His position sits beside LeCun’s on the science and adds the market case the rest of the debate omits.


Demis Hassabis: lighthouse moments and the science test

Demis Hassabis operates inside the scaling consensus but insists on a higher bar for claiming AGI. His criterion is not benchmark performance but matching the brain’s full portfolio of cognitive functions — tens of thousands of cognitive tasks, verified by domain experts, with ‘lighthouse moments’ analogous to Einstein’s relativity. Jagged Intelligence — capable in some areas, deficient in others — explicitly does not qualify.

His AlphaFold argument shows where he agrees and disagrees with the maximalists. AlphaFold succeeded because evolution imprinted learnable structure on proteins — the Bitter Lesson in action, applied to a verifiable scientific problem. But the harder problem — generating good scientific conjectures, the creative leaps that produced the double helix — remains beyond current systems. AI compresses the gap between conjecture and proof; it does not yet replace conjecture. His 50% probability for AGI is 2030, by which he means a system that passes his full portfolio test.


Where the camps agree

The debate obscures substantial common ground.

Deep learning works. The neural-network paradigm, transformer architecture, and large-scale pre-training have produced the most capable AI systems in history. No one in this debate argues otherwise.

Compute matters. Even LeCun does not deny that scale helped; his objection is to the objective, not the resources. Sutskever’s ‘age of research’ explicitly happens ‘just with big computers’ — small-compute research remains a niche.

RL on verifiable tasks is real progress. Nathan Lambert and Sebastian Raschka document three independent scaling axes — pre-training, RL with verifiable rewards, and inference-time compute — all still active. The reasoning model breakthroughs of 2025 (DeepSeek-R1, o1, extended thinking) showed that RL on maths and code can discover strategies outside the human distribution. This is not disputed; what the camps disagree on is whether it generalises.

The Economic Turing Test is the right question. Whether you frame it as Dario’s ‘country of geniuses’, Ben Mann’s 50%-of-jobs threshold, or Sebastian Raschka’s completion-rate reframe (today’s models complete ~30–40% of complex multi-step tasks; 90–95% makes the definitional debate practically meaningless), the empirical target is the same: systems that can reliably substitute for expert human labour across a wide range of consequential tasks.

AI Diffusion will gate real-world impact regardless. Even the most bullish maximalist — Dario — expects the capability frontier to precede economic impact by months or years, bounded by change management, procurement, and the mundane friction of institutional adoption.


The split, restated

The quantitative-or-qualitative question is not a matter of emphasis; it determines which research bets pay off. If Dario is right, the winning move is to scale RL environments to new domains and buy more compute. If Sutskever is right, the winning move is to solve the Generalization Gap — a fundamentally different learning principle — before rivals do. If LeCun is right, the winning move is to abandon the paradigm entirely and build World Models with hierarchical planning.

The labs’ actions reveal their actual beliefs more clearly than their public statements. Anthropic and OpenAI spend on frontier pre-training and RL infrastructure; SSI raises capital on the claim that research insight is the bottleneck; Meta funds JEPA and open-weight Llama releases simultaneously. The bets are on, and the outcome is not yet in.


See also