Gary Marcus on AI’s Overstated Intelligence, LLM Economics, and the Limits of Scaling
Gary Marcus argues that the entire AI economy rests on a single error — people over-attributing intelligence to machines that are, at bottom, next-token predictors — and that the financial bet built on that error is unlikely to pay off. The conversation joins his long-standing technical scepticism about large language models to a markets thesis: no moat, no profits, and a leading lab, OpenAI, that he expects to be the WeWork of AI.
Key ideas
- The economy is hinging on over-attribution of intelligence. Marcus’s central claim is that investors are ‘betting trillions of dollars that these machines are intelligent in ways that they aren’t actually’, because those placing the bets lack the cognitive-science background to test intelligence properly. The error is old — Weizenbaum’s ELIZA fooled people in the 1960s — but it now underwrites the whole market.
- LLMs are next-token predictors, and that is the ceiling. Pure large language models predict the next token in a sequence; prediction is part of cognition but not the whole of it, so the models ‘fake everything else’. Pushed outside their training distribution they ‘do really stupid things’ — the river-crossing puzzles Anthropic had to patch into its system prompts are the tell.
- No moat means no profits. A dozen labs are building ‘basically the same toothpaste’ on one of two near-identical architectures. With no differentiation, prices collapse into a war — ‘you can’t charge $100 for a tube of toothpaste if you have nine competitors building basically the same thing for less’ — which is why nobody is making money.
- OpenAI is the weak link. Marcus has called OpenAI ‘the WeWork of AI’ since 2023. It burns roughly $21bn a year, its lead has eroded (he puts it as low as fourth place), and Anthropic offers a near-identical product at a similar valuation while burning less — leaving ‘no rational argument for buying a share of OpenAI at a trillion-dollar valuation’.
- Regulation has finally arrived — clumsily. After two years in which Marc Andreessen’s allies ‘Overton-windowed’ the debate into whether to regulate at all, a multi-state subpoena of OpenAI and a (still voluntary) Trump executive order mark a ‘sea change’. Marcus wants pre-release screening on the FDA model and warns against the proposed government equity stakes, which he reads as a ‘backdoor bailout’.
Content
Over-attribution: the error under the market
The interview opens on Marcus’s organising claim: ‘The entire economy is hinging on over-attribution of intelligence to these machines.’ His point is not that LLMs do nothing — they are ‘great for autocomplete for the purposes of computer coding’ and useful for brainstorming — but that humans are evolutionarily ill-equipped to judge machine intelligence. We have machinery for spotting snakes and lions; we have ‘nothing built into our brain to really help us think about the nature of intelligence’. Joseph Weizenbaum showed in the 1960s that ELIZA, a keyword-matching mock psychiatrist, could fool an average person into perceiving understanding that was not there. What was a curiosity then is now, in Marcus’s telling, the foundation of a multi-trillion-dollar market: the people writing the cheques ‘don’t have enough cognitive science background to know the right test in order to evaluate intelligence’. See Over-Attribution of Intelligence.
Why generative AI is inherently unreliable
The technical core of the scepticism: a pure LLM ‘is basically a next-token predictor’. Trained on the whole internet, it approximates human speech well, but the approximation is ‘very superficial’ and ‘very data dependent’. Push it outside its training regime and it fails — the river-crossing puzzles where the systems ‘say the most absurd things’, so embarrassing that Anthropic hard-coded fixes into its prompts. The failure reveals that the models ‘are not really reasoning about things like a man or a river or a boat’; they string together words they have seen. Marcus dates his warnings to 2019 (‘they don’t have stable models of the world’) and notes that the promised cure — more data — has not arrived: hallucinations persist into 2026, and he twice offered public six-figure bets (to Reid Hoffman, to Microsoft AI’s Mustafa Suleyman) that they would not vanish. This connects directly to Hallucination and challenges Scaling Laws.
Intelligence, Marcus insists, is multi-dimensional. Purpose-built systems play chess or run GPS navigation superbly; LLMs cannot even reliably follow the rules of chess, and make illegal moves. The viral clip he cites — a chatbot asked to count to 100 that keeps saying it will count without ever doing so — is, for him, ‘a beautiful example’ of the relevant stupidity: unreliable instruction-following. That is what warrants regulation, not any science-fiction super-capability.
The no-moat economics
Marcus ties the technical critique to the markets thesis the show exists to examine. Because every lab uses ‘the same cognitive architecture’ — most products are wrappers on OpenAI’s or Anthropic’s models, which are themselves ‘basically the same’ — there is no durable differentiation. He predicted in 2024 that LLMs would ‘run out of headroom’, moats would vanish, and price wars would follow; the lead now ‘goes back and forth’ week to week for billions in spend. His toothpaste analogy makes the margin problem concrete: commoditised products cannot command premium prices.
He layers on the ‘token apocalypse’: after a brief period of ‘token-maxing’ (companies rewarding employees for maximal AI usage, complete with leaderboards), firms have noticed that the productivity studies are unimpressive — ‘every study that’s looked at productivity has shown they’re not all that great’ — and are now economising, even reaching for cheaper models from China. The whole boom, he argues, has been driven by FOMO; if eighteen months of trial yields no clear return, customers will walk, and any such defection wipes out a cash-burning lab.
OpenAI as the WeWork of AI
On the specific names, Marcus is blunt. OpenAI loses money on every use of its product, burning about $21bn a year (a net loss he notes was nearer $39bn with caveats); it has lost relative ground while Anthropic gains share with a comparable product, better commercial discipline, and — in his view — a sounder culture and ‘a little bit better technical vision’. He has predicted since November 2023 that OpenAI would be ‘the WeWork of AI’, an idea once thought absurd and now echoed by writers like Sebastian Mallaby. The financing model — repeatedly raising at higher valuations to cover the burn — runs into the problem that the next cheque, perhaps an IPO, is hard to justify when a near-identical competitor is better run. Pressed on whether all the labs fail (the Ed Zitron position) or only some, Marcus places himself between Zitron and host Ed Elson: he sees a plausible path to profitability for Anthropic, but treats it as genuinely TBD, with an intermediate outcome where a lab survives, makes perhaps $20bn a year on a huge capital base, and simply never justifies a trillion-dollar valuation. The reliable money, he suggests, may go to Nvidia (‘selling shovels in the gold rush’) and to incumbents like Google with the infrastructure and distribution to avoid disintermediation.
Mythos, cyber security, and the shape of the risk
Asked about Anthropic’s powerful new ‘Mythos’ model and the cyber-security fears around it, Marcus gives a deliberately nuanced read: it is part-harness rather than pure generative AI, it is ‘oversold’ but ‘also real’, and the genuine exposure lies in the world’s many poorly secured systems — especially vibe-coded ones — rather than in well-defended banks or Google. The episode, he argues, is less a Skynet moment than a ‘wake-up call’ on deferred cyber-security maintenance, which has been neglected partly through stigma. He is sharply critical of the politics: using regulatory pressure ‘to destroy a particular US company’ is, he says, ‘not capitalism’ but a thumb on the scale.
Regulation: a clumsy sea change
Marcus reads the past twelve months as a real shift. The earlier Washington ethos — any regulation stifles innovation, captured in an executive order pressing states to do nothing — has given way under public backlash (over data centres, jobs, delusions, and documented harms) to a multi-state subpoena of OpenAI (he cites New York leading some 46 states) investigating data handling, health data, and model sycophancy. He welcomes the direction while judging the substance too weak: Trump’s new executive order asks labs to voluntarily submit models for cyber-security checks, which Marcus says barely counts as regulation. What he wants is mandatory, FDA-style pre-release screening that weighs benefits against harms (sycophancy, delusions, the tie to suicides and to a Florida wrongful-death suit), and he rejects government equity stakes as a disguised bailout of unprofitable firms. He notes that nuanced bills (from senators such as Blumenthal and Hawley) exist but die in committee under lobbying pressure.
The closing message: explore, don’t only exploit
Marcus’s parting myth to dispel is that generative AI is close to AGI and will solve everything. It will not: the field is ‘completely in the exploit’ mode on LLMs rather than exploring alternatives. His mountain-range metaphor — you may have to climb back down one peak to reach a higher one — captures the worry that committing the whole economy to a single architecture is a strategic error. He points to China hedging at perhaps ‘20 cents on our dollar’, keeping room to pivot if a better approach to machine intelligence emerges. The constructive demand is for more cognitive science and a richer account of what intelligence is, before more capital is sunk into one bet. This positions him as the economic and cognitive-science wing of the dissent catalogued in The Road to AGI.
Related
- Gary Marcus — guest; AI sceptic, NYU professor, author of Taming Silicon Valley
- Ed Elson — host
- Over-Attribution of Intelligence — the episode’s organising concept
- Hallucination — Marcus’s foundational technical objection
- Scaling Laws — the bet he argues has hit its ceiling
- The Road to AGI — the AGI debate his economics-and-cognition dissent extends
- Sam Altman — OpenAI, the lab Marcus is most bearish on
- Dario Amodei — Anthropic, the lab he rates more highly