Two questions that keep getting mashed into one
The previous lesson, The Generalisation Gap, ended on a doubt: that pouring in more compute might never buy the cheap, robust learning a child does without trying. That doubt is where the whole public argument about AI lives — and most of the noise comes from one confusion, which this closing lesson exists to clear up. Two different questions get spoken as if they were the same one:
- Does scaling work? Does making models bigger, and training them on more, reliably make them better? This one is settled, and the first three lessons are the proof.
- Does scaling suffice? Will the same recipe — more compute, more data, more checkable practice — carry all the way to general, human-level intelligence, or is something missing that no amount of scale supplies? This one is wide open, and the fourth lesson is why.
Keep those apart and the field's shouting match resolves into something much more precise: near-total agreement on the first question, and a narrow, genuine disagreement on the second. 'Scaling is everything' and 'scaling has hit a wall' are both answers to the wrong question. The real one is whether working is enough.
The surprising amount everyone agrees on
From the outside the debate looks total, as if the two sides shared no ground. Up close, the common ground is most of the map. Lay out what nobody serious disputes:
- Deep learning works, and compute matters. The neural-network recipe has produced the most capable systems ever built. Even Yann LeCun, the field's most categorical sceptic of the current paradigm, does not deny scale helped — his objection is to the objective, not the resources.
- Reinforcement learning on checkable tasks is real progress. The reasoning-model breakthroughs of 2025 showed that training on maths and code — where an answer can be automatically marked right — discovers strategies beyond the human examples it was shown. What the camps dispute is not that this works but whether it generalises.
- The target is agreed, even where the timeline isn't. Whether you frame it as a country of geniuses in a data centre, an economic Turing test (can an agent pass as a human contractor for half of paid work?), or a completion rate (today's models finish maybe 30–40% of complex multi-step tasks reliably; at 90–95% the label 'AGI' stops mattering), everyone is pointing at the same thing: systems that can stand in for expert human labour across consequential work.
That is a lot of agreement. The disagreement that remains is real, but it is a disagreement about a single axis — which is what makes it tractable rather than a clash of worldviews.
The one axis they split on
The whole debate reduces to a single question, laid out in full on the Road to AGI theme: is the distance left to general intelligence quantitative or qualitative? More of the same recipe, or a recipe we do not yet have? Three camps give three answers.
- The maximalists: quantitative. Dario Amodei is the clearest voice — the gap is more compute and more RL environments, same recipe. His phrase 'near the end of the exponential' is widely misread as 'scaling has stopped'; it means the opposite — near the top of the climb, with a country of geniuses perhaps one to three years off and 90% likely within ten. The bet is that scaling doesn't just work, it suffices.
- The what's-missing camp: qualitative. Ilya Sutskever — who co-wrote the results that started the scaling era — argues the missing piece is a better learning algorithm: something that closes the generalisation gap, an internal value function, a system that keeps learning on the job. Andrej Karpathy sits nearby with a decade-long timeline and a blunt verdict on today's tools: 'RL is terrible — it just so happens that everything we had before it is much worse.' The bet is that scaling works but does not suffice.
- The paradigm dissenters: wrong recipe. LeCun holds that no amount of compute on next-word prediction reaches human intelligence, and that the road runs through world models and planning instead. The bet is that scaling the current thing is scaling the wrong thing.
Notice what is not here: the cartoon extremes. Nobody credible is saying 'it's all a bubble' or 'superintelligence next Tuesday'. Even Dario, the arch-maximalist, deliberately stakes out a middle — against the stagnation camp on one side and against runaway self-improvement on the other. The genuine disagreement is narrower, and more interesting, than the slogans that travel on social media.
Why the answer isn't in yet — and how to read the evidence
Nobody can settle this from the armchair, because it is a bet about the future of a technology being built as we watch. But you can do better than pick a team. Two habits turn you from a spectator into a reader of the evidence.
First, watch what the labs do, not what they say. Actions reveal beliefs more honestly than interviews. Anthropic and OpenAI pour capital into frontier pre-training and RL infrastructure — the quantitative bet. Sutskever's new company raises money on the claim that the bottleneck is research insight, not compute — the qualitative bet. Meta funds a wrong-paradigm alternative and open-weight models at once — hedging. The money is placed; the outcome is not yet in.
Second, ask of every new result: which question does this answer? When the next headline lands — a model tops a new benchmark, wins a gold medal at a maths olympiad — the useful question is not 'did the number go up' but 'did this close the generalisation gap, or just add one more environment?' A higher score on a checkable test is quantitative evidence: scaling working, again. The thing that would actually settle the deeper question looks different —
a system that learns a genuinely new skill from one or two examples, robustly, in a situation nobody built a training environment for.
That — not another benchmark — is the qualitative evidence. Knowing which kind you are looking at is the whole skill.
What it means
This is the fifth and last lesson in the arc, and it is the one that lets you hold the other four at once. The scaling machine is real (lessons one to three); it has a strongest objection that may or may not bind (lesson four); and whether it is enough is the open question the whole field is staking billions on, in both directions. What to carry away:
- Separate 'works' from 'suffices'. Scaling working is settled fact; scaling sufficing is a live bet. Almost every confused take you will read collapses the two — someone citing the settled question as if it answered the open one.
- Distrust the extremes. If a claim sits at 'it's all hype' or 'AGI is basically here', it is downstream of the slogans, not the argument. Even the field's most bullish serious figure stakes out a bounded middle. The real range of expert opinion is narrower than the loudest voices suggest.
- Read new results by what they would prove. A higher benchmark is more evidence that scaling works — which was never in doubt. Cheap, robust, off-the-beaten-path learning is the evidence that would move the question that is in doubt. Weigh them differently.
- Hold it as a live question. The people closest to the frontier disagree, and each has staked a company on their answer. That is a reason to watch the bets play out — not to adopt one as a belief.
Go deeper
The single best map of this whole disagreement is the wiki's own The Road to AGI — it lays every camp side by side on the quantitative-or-qualitative axis, with each figure's timeline and reasoning. For the two poles in their makers' own words, watch Dario Amodei on 'The End of the Exponential' for the maximalist case, and Ilya Sutskever on 'The Age of Research' for the case that something fundamental is still missing. Between them sits the entire live argument about how far the recipe in these five lessons can go.