The American DeepSeek Project
Nathan Lambert argues that although the United States holds the best frontier models and the best chips, it is losing control of the open and academic directions of AI to China — and he proposes a counterweight: a fully open model (data, training code, logs, and weights) at frontier scale within two years, which he names the American DeepSeek Project and intends to build at Ai2.
Key ideas
- The open-source centre of gravity has moved to China. America leads on frontier models (Gemini, Claude, o3) and infrastructure (Nvidia), but Chinese organisations now release the most notable open models and datasets across every modality, and researchers worldwide read more Chinese papers than Western ones. Lambert reads this as structural, not cyclical — China has more AI researchers, more data, and an open-source default.
- The proposal: fully open, not merely open-weight. The goal is a model at the scale and performance of current frontier systems within two years, released fully open — data, training code, logs, and the decision-making behind them, not just downloadable weights. The distinction is the whole point: it distributes the knowledge of how to train frontier models, not only the artefact. He estimates $100M–500M over two years.
- The stakes are trust and accountability. If the next architectural breakthrough is built on Chinese models, chips, or ideas, the most available models become the hardest to trust — unprovable code vulnerabilities, weaker engagement with US legal norms (fair use, non-consensual deepfakes), and ecosystems shaped by compute restrictions. These are structural problems that fine-tuning cannot fix (he cites Perplexity’s R1-1776 as evidence that a surface patch is not enough).
- ‘Frontier’ is being redefined by agents, not raw model scale. Consumer models have stopped getting much bigger; the action is moving to agents that call many models, sometimes small ones. So the meaningful target is not the largest 2027 model but fully open models that reach today’s GPT-4-class performance (recent Sonnet, DeepSeek V3, Gemini Pro) by around 2027 — open-source’s efficiencies tell most strongly in agentic systems, which is the opening open models have been waiting for.
- Opening AI as a quintessentially American act. Faced with a technology set to concentrate extreme power in a handful of firms, openness is one of the few available counterweights. If something near AGI is arriving and will sit inside billions of lives — closer to electricity than to an opt-in product — it should be available for all to benefit from. The corporations ‘will win’, he concedes; the open project’s job is to control by how much, and to seed better community norms for how models are built and shared.
Content
Why the shift is structural, not a blip
Lambert traces today’s progress to the industry’s pre-2022 habit of openly publishing the science of AI, largely out of Google Research. That practice has stopped, and America’s open champions are wavering — Meta is ‘reconsidering their open approach’ after another costly re-org, while the political climate deters the best foreign scientists from coming. The result is a balance of power that has tipped in roughly twelve months and, on his reading, will keep tipping without a deliberate counter-effort. He is careful to separate the people from the system: many Chinese researchers are among the best he has worked with; his worry is the ecosystem’s inevitable ties to the Chinese state, which make models less auditable and accountable.
What counts as the target — scale, not the leaderboard
He deliberately frames the goal by scale rather than benchmark performance, because efficiency gains compound: a model built to DeepSeek V3/R1 scale will, a few years on, far outperform today’s V3. The current open frontier is catching up to the original GPT-4, a real step up from GPT-3 levels; the step he is shooting for is the modern GPT-4 class. Open-weight-only releases are not enough on their own — the Llama licence and rumours of its discontinuation hobble the best American open models, Chinese models carry a weak security reputation as they integrate with more tools, and European models are ‘largely off the map’. Trustworthy open AI therefore needs something different in kind, built on different incentives and pretrained from the ground up.
See also
- Nathan Lambert — the writer’s hub (his Interconnects writing alongside his podcast appearances)
- Dario Amodei — Anthropic’s case for the closed-frontier path, the contrast this essay argues against