Dylan Patel on the Token Economy, AI Supply-Demand, and the Permanent Underclass

Dylan Patel

Show: Invest Like the Best

Episode: https://www.youtube.com/watch?v=dQ74t1cNFAc

Cleaned and reformatted from published transcript or auto-generated captions — punctuation added, filler removed, restructured for readability. Not verbatim. For exact quotes, refer to the original.

Contents

    From $10K to $7M: SemiAnalysis Goes All-In on Tokens

    Patrick O'Shaughnessy

    You told me this incredible story about how your own team's use of tokens has changed dramatically this year. Will you tell that story and what it is teaching you about what's going on in the world?

    Dylan Patel

    Last year, we thought we were heavy users of AI. Everyone was using ChatGPT, everyone was using Claude, everyone had subscriptions. Our spend was on the order of tens of thousands of dollars for the firm. This year, the spend has skyrocketed. It really started in late December with Opus — that included Doug O'Laughlin, our president. He's very much the leader in the sense of non-technical people using AI for coding, and he's slowly pulled the whole firm along over time.

    Spend in January started to inflect and rocket upward. We signed an enterprise contract with Anthropic and it's gone to the point where, when I last talked to you, we were at a $5 million annual spend rate. It's actually $7 million now.

    Patrick O'Shaughnessy

    That was last week, by the way.

    Dylan Patel

    A lot of that is just the usage. People who have never coded before are using Claude Code and spending thousands of dollars sometimes a day. Across the firm, we're spending $7 million a year on Claude Code at the current run rate, versus our salary expense being around $25 million. So we're north of 25% of spend on Claude Code as a percentage of salary. If this trajectory continues, we'll spend more than 100% by end of year — which is a bit terrifying.

    Thankfully, I don't have to decide between people and AI because the company is growing so fast. It's more like: I don't have to hire nearly as fast, I can spend a lot more on AI, and we just grow faster. But other folks will start to reckon with the fact that if one person can do the work of five to ten to fifteen people using Claude Code, then they should probably cut headcount.

    For example, we have a reverse-engineering lab in Oregon that we've been building for a year and a half — fancy microscopes, scanning electron microscopes. The whole purpose is to reverse-engineer chips: get the architecture out, get the materials they're using to manufacture. This data is one of the things we sell. It's a very slow process. But one person on the team, with a couple thousand dollars of Claude tokens, built an application that is GPU-accelerated, runs on a server at CoreWeave, and whenever we send it an image of a chip, it overlays where every single material is: this part is copper, this part of the gate is tantalum, this part is germanium, this part is cobalt. You can do a finite element analysis of the entire stack-up visually with a dashboard GUI. The person came from Intel and said that used to be an entire team's job to build and maintain.

    Another example is Malcolm, an economist from a major bank where the economics department was 100 to 200 people. He piped in Fred data, employment reports, and various other data sources from APIs. He started running regressions, looking at the impact of various economic revolutions — inflationary and deflationary. The Bureau of Labor Statistics has a set of roughly 2,000 tasks; he used AI to grade which ones can be done by AI and which cannot. About 3% are doable now. He created a metric to measure things that can be done by AI and the deflationary effect of doing them at scale — what he calls "phantom GDP." Output can go up, but because cost falls so much, actual GDP theoretically shrinks. He built all of this and a brand-new benchmark of language models, a set of 2,000 evals, entirely by himself.

    Patrick O'Shaughnessy

    He does it all by himself?

    Dylan Patel

    All by himself. He said, "This would have taken a team of 200 economists a year." He is completely absorbed in Claude. "Everything has changed."

    Patrick O'Shaughnessy

    How do you think about it as a business owner going from close to zero to 25% of spend, accelerating toward whatever percent of total? At what point do you say, "Maybe I don't need to be on the most cutting-edge model"?

    Dylan Patel

    I'm in the information business. We sell analysis. We do consulting. We create data sets. I don't see why this wouldn't be completely commoditised on a rapid basis if I'm not constantly improving. The way we were doing things in 2023 is basically what everyone else is doing now. If I don't move up the bar, I'll be commoditised. If I don't move fast enough, I'll lose my edge.

    AI commoditises things, just like it commoditises software. Those who can move fast and keep improving their service will grow faster. Incumbents doing nothing are going to lose. If I don't adopt AI, someone else will and they will beat me.

    The energy space is a good example. We've had energy analysts for about a year, trying to build out an energy model. The energy data-services market is something like $900 million — obviously a huge market to break into — but we hadn't really broken in despite a year of effort. Then Claude Code psychosis hit one of our people, Jeremy, who leads data-centre energy and industrial analysis. In three weeks, spending about $6,000 a day, he scraped every single power plant in the US, every transmission line above a certain voltage, and created a complete mapping of the entire US grid with all the major demand sources, from public data. We built a dashboard where you can see the micro-regions of the US with power deficits and surpluses.

    We showed it to customers, including energy traders. They said, "How long did this take? This is really good — this is better than XYZ company." Then we dug deeper and found that XYZ company has 100 people and has been working on it for a decade. Ours isn't as fully robust in every respect, but in some ways it's better. The question from a business owner's perspective is: yes, I'm spending a lot — but what is that spend getting me in revenue?

    Demand Explodes: The Race to the Frontier Model

    Patrick O'Shaughnessy

    Are you worried that in the limit, the people who control capital — the investment firms that often hire you — will just say, "We have smart analysts too. We'll build this ourselves"?

    Dylan Patel

    Any information-services business doesn't generate as much value as the customer does from that information. If I sell you information for a dollar, you're only buying it because you know it helps you make a decision worth more than a dollar. Investment funds like Jane Street and Citadel are very detailed on their data, and yet they purchase data from us and continue to grow with us. There's some "it factor" — we move faster, we're more nimble, a smaller team focused on one specific thing. I think investment professionals would mostly prefer to buy the data from us, because it's cheaper than building it themselves. But some may try.

    Patrick O'Shaughnessy

    Every conversation I have with you ultimately comes down to supply and demand of tokens. What has this experience taught you about the demand side?

    Dylan Patel

    If we take the macro lens: Anthropic has gone from $9 billion in revenue to somewhere around $35–40 billion now — probably $40–45 billion by the time this airs. Their compute has not grown to the same degree. If you do the calculations and assume they didn't decrease research and development compute — and they clearly didn't, they released Opus 4.7 and Mythos — then even if all incremental compute went toward inference, their margins are at a floor of 72%. In reality some of that incremental compute probably went to R&D, so it may be higher. At the start of the year, leaked funding-round documents showed 30-something percent gross margins. Where does a business like this grow margins like that?

    In principle: their demand is so high they're able to cut back on usage limits and rate limits. What really matters is having an Anthropic enterprise contract and getting the rate-limit increases you need. Tokens are super in demand. Whoever can pay for them wins. Anthropic is receiving $40 billion ARR in token spend — but those tokens are generating far more than $40 billion in value. Various businesses generate different value per token. As models get more and more intelligent, access to the most capable tokens becomes the critical resource.

    A lot of people will want tokens. But the startup in SF using Claude to generate a mediocre software product isn't actually creating much value — and it will get priced out of tokens soon enough.

    Patrick O'Shaughnessy

    Are you surprised that people are so insistent on going to the most expensive leading-edge model?

    Dylan Patel

    Without a doubt. One of my funniest memories in the past month and a half is myself and my friend Leopold, on our knees in front of an Anthropic co-founder begging for access to Mythos — which he insisted didn't exist.

    Looking at the benchmarks: Mythos is potentially the biggest step up in model capabilities in about two years. It's so good they didn't want to release it even though they'd already announced the price — it's 5 to 10x the token cost. They released a deliberately constrained version, Opus 4.7, and explicitly said in the model card that they made it worse at certain tasks.

    Whoever you are, if you have enough capital, you should get an Anthropic enterprise subscription where you pay per token rather than with subscriptions — then you won't get rate limited as much. And you need to figure out how to leverage those tokens to the highest-value tasks, because ultimately, in a year or two, the business is just arbitraging tokens: what direction do you point them in to make the most value?

    Pick any benchmark: the cost to hit a certain capability tier used to cost X, now it costs 1/100th or 1/1,000th. DeepSeek on GPT-4 was 1/600 the cost. Since then, costs have fallen further for GPT-4 class models. But no one cares about GPT-4 class models — they want the frontier, because the frontier is what enables economically valuable things. What's driving demand isn't falling costs. It's all these new use cases. My current spend at current model quality would probably cost 100 times less a year from now — but irrelevant, because I'm going to be using a far better model. Mythos is more expensive per token, but it uses far fewer tokens to accomplish a task, so it's actually cheaper in most tasks than 4.6 Opus.

    Patrick O'Shaughnessy

    When I last saw you, the Mythos eval card had just come out and you said it actually made you a little scared. What did you mean?

    Dylan Patel

    Anthropic's goal for 2025 was to have an L4 software engineer in their model by year-end. They achieved that with Opus 4.6. What they didn't say is that if you look at Mythos, it's closer to an L6 engineer. L4 is fairly new; L6 is quite experienced. Anthropic had the model internally available in February. So in two months, they went from L4 to L6 engineer. What comes next?

    When you think about model progress, it's only accelerated. Anthropic's release cadence has compressed. OpenAI's has compressed too. Why? To make a better model you need a few things: amazing compute, which is very expensive and has lead times that are largely set in stone for the near term; amazing researchers, who people are paying tens of millions of dollars for; and implementation. Historically implementation was very difficult. Now implementation is easy. It's expensive but easy.

    What used to matter was that execution was very, very difficult and ideas were cheap. Now, ideas are cheap and plentiful, but execution is very easy. So only the good ideas — the ones that can justify the spend on cheap implementation — matter. And as implementation costs continue to fall, we don't even fully comprehend what comes next. It's a complete reordering of how economies work.

    The Permanent Underclass: Use Tokens or Fall Behind

    Dylan Patel

    Uncertainty is there, and that does cause some fear in terms of how does society reform itself? How does one exist in a world where your ability to implement something is not actually that important — where what matters is your ability to choose the correct idea for AI to implement, and then your ability to sell what the AI has built and garner capital toward it?

    Going back to the point about needing the newest model: who's going to have access to it? Anthropic's selective-release programme — I troll Anthropic people by calling it Earwig — is just going to continue. Models will have less and less broad deployment. I know OpenAI and Anthropic say they want great AI for everyone. But AI is very expensive. Who's going to pay for the trillion dollars of infrastructure? People who have money and can build useful things with AI. And then you don't want people to distil your model, so you don't release it broadly. You release it to a smaller and smaller set of customers.

    Those customers are also wrestling over tokens. Anthropic could double their pricing on Opus and most users would continue to pay, and I bet that still wouldn't solve their capacity problem. So the question becomes: where does this cycle end where token usage — and therefore the benefits and additional value generated from those tokens — aggregates among fewer and fewer companies?

    Right now, top banks have access to Mythos. They're only using it for cyber security — but I can envision a world where, because I have an enterprise Anthropic contract and they give us slightly earlier access, I'm able to crush my competitor who doesn't. Ken Griffin at Citadel is super well-connected and could sign a deal with OpenAI or Anthropic: "I'll buy the first $10 billion worth of tokens each year — whenever you release a model, I get it first." That gives him a massive edge in the markets. The concentration of resources and usage into fewer and fewer hands is a real possibility that no one knows how to address.

    Patrick O'Shaughnessy

    Anything we're missing on the demand side?

    Dylan Patel

    If you don't use more tokens, you'll never escape the permanent underclass.

    Patrick O'Shaughnessy

    Expand on that.

    Dylan Patel

    Either you use more tokens and generate outsized economic value from them, or you don't. A lot of people are doing it the boring lazy way: work one hour a day instead of eight and accomplish roughly the same. The better way is to still work eight hours a day, do eight times the work, and make five times the money. You can't quite do this in a conventional job, but there are people who start companies, start selling things, and capture the economic value before everyone is doing it and it becomes table stakes.

    There are three distinct problems: using more tokens, generating value from those tokens, and capturing value from what you created with those tokens. If you don't do all three, you'll never escape the permanent underclass — as models continue to skyrocket in capability and the concentration of resources potentially deepens.

    Supply Side: Everything Is Sold Out

    Patrick O'Shaughnessy

    Let's talk about supply. What is going on at the frontier of supplying the entire stack required to serve all these tokens?

    Dylan Patel

    As demand skyrockets, prices are going up for everything on the supply side — H100 GPUs, everything. In addition, the useful life of those GPUs is extending. There are people who argued GPU useful lives are less than five years — complete nonsense. Clusters are re-signing three- or four-year-old Hopper clusters for another three or four years. A100 clusters are re-signing for another couple of years. The useful life is clearly not five years; it may be seven or eight.

    So the gross margin on a cluster was never just 35% — it's higher. Margins are expanding at the cloud layer. On the hardware layer, Nvidia is still charging around 75% gross margin. As we move down the stack: memory margins have skyrocketed; optics and logic are seeing large prepayments; companies like Nvidia are paying huge prepayments, so even if gross margins haven't moved, return on invested capital is going up because the invested capital is lower. This is a consistent trend across the entire supply chain.

    ASML is completely sold out and needs Carl Zeiss to expand faster. Every company along the chain is either sold out with margins going up, or getting prepayments that increase return on invested capital. Even something like the copper foil needed to make a PCB is sold out and people are making prepayments for it. Anything and everything that has a pulse and is sold out has people jumping to get more incremental supply and fighting over future supply.

    Patrick O'Shaughnessy

    What do you think are the most important bottlenecks? History suggests that enormous demand signals eventually bring supply forward. Why is this different?

    Dylan Patel

    Supply chains are usually fast to react. One unique thing is that our supply chains are more complex than ever, and what we're building is more complex than ever — so lead times are longer. Memory is a good example: capacity can only grow low double-digit percentages a year, maybe 20–30% for DRAM, even less for NAND. Even though the demand signal was very strong at the end of 2025, the memory companies started reacting immediately. But none of that truly incremental capacity arrives until 2027 at best, 2028 in reality. Even if they wanted to build as fast as possible, it doesn't come until 2028. As a result, memory prices have gone through the roof — and they're going to double and triple again, at least for DRAM. People think the memory story is overplayed. It is not. DRAM will double or triple from here because that's how much capacity is required, and to steal capacity from elsewhere in a capitalist economy, you need demand destruction via higher pricing.

    Logic has huge capacity problems too. TSMC just raised their CapEx guidance — we've logged $57.4 billion since January and may revise it up slightly. But what people aren't focusing on is what does that mean next year, or the year after? Three years from now, TSMC may spend $100 billion on CapEx in a single year. And people just can't fathom it. But what does that mean for their upstream supply chains — Lam Research, Applied Materials, ASML, and their further downstream suppliers like MKS Instruments? The tail-whip effect just gets harder and harder. If TSMC wants to spend $100 billion in 2028 — which is a real possibility — the shortages across the equipment supply chain will be severe.

    Beyond GPUs: CPUs, ASICs, and Robotics

    Patrick O'Shaughnessy

    What about other parts of the chip ecosystem — CPUs, ASICs — beyond Nvidia's GPU dominance?

    Dylan Patel

    ASICs are obviously taking off. We did a project on FPGAs and found there are 120 FPGAs per next-generation AI rack. Then there's CPUs: all the reinforcement-learning environments, plus all the code that you and I are generating that is now running on some cloud instance — all of that requires CPU. CPUs are completely sold out and demand is skyrocketing.

    There are two main reasons. First, reinforcement learning is very CPU-intensive. The old approach was: put all the internet's data into the model, train it, output something. The new approach is: train on the internet data, then put the model into an environment where it tries things out and an environment scores whether what it tried was successful. These environments can be anything from simple structured-output checks to complex tasks: open this file, edit it, submit it to a website; open a physics simulation from Siemens and edit this CAD model. The more complex the environment, the more CPU you need. The ASIC runs the model; the CPU runs the environment.

    Second, once you have great models and deploy them, the code they generate doesn't go from a GPU straight to the human brain. It goes through a deployed application running on CPUs. So there's enormous CPU demand there too — and things are sold out in a large way.

    Patrick O'Shaughnessy

    Robotics — presumably robots consume relatively zero tokens right now. Do you see that changing as a second demand curve?

    Dylan Patel

    There's this concept of a software-only singularity: AI reaches singularity but only in software, and the rest of the physical world is left behind. Vast majority of the world is physical. But I think software-only singularity is just a blip, because once software is super easy, what makes robots really hard? It's programming micro-controllers, actuators, controlling all this stuff — and that's now easy too.

    Current robot models — vision-language-action models, VLAs — are probably not going to be the thing that ultimately scales. They're data-inefficient, and we can't scale training data for them fast enough. But once software complexity collapses, people will start building robot foundation models at scale, and robots will become genuinely useful. I think in the next 6 to 18 months we'll see real breakthroughs in robotics that enable few-shot learning: a pre-trained robot model where you show it a couple of examples and it can perform the task. That unlocks a huge explosion in physical-goods acceleration and deflationary effects there — and that will keep token demand growing. I don't think token demand slows down.

    Scaling Laws, Compute Races, and Model Margins

    Patrick O'Shaughnessy

    What did you learn from Mythos about the components of scaling laws?

    Dylan Patel

    Mythos is a materially larger model than prior models. More compute into model makes model better — the scaling laws still work. Along the whole way, we're also getting compute efficiency wins: if I want a model of capability X, every two to six months that cost is dramatically decreasing. But if you scale it up massively, you get a huge capability jump on top. It's proof that this trendline continues.

    Google and Anthropic are not the heaviest users of GPUs on the training side, but OpenAI will have their new class of models. I think they're taking a more sensible, principled approach to scaling in smaller steps, whereas Anthropic went for a huge jump. We'll see better and better models throughout the year, and the release cadence is only going to get faster.

    Patrick O'Shaughnessy

    We've gone a long way in this conversation saying almost nothing about OpenAI, which would have been strange a year ago.

    Dylan Patel

    Everyone's like: "Okay, so Anthropic has just won, right?" They had Mythos in February and never even released it, because they're already sold out. Their revenue is adding $10 million a month. And then they had Opus 4.7 today, all before OpenAI's next major release. So clearly Anthropic is in the lead and OpenAI is cooked.

    But here's the interesting thing. Because Anthropic has such bounds on compute, and can only grow it so fast — Dario used to gloat about how OpenAI was too aggressive on compute and Anthropic was more sensible — now Anthropic is saying, "We wish we had a lot more compute." OpenAI can pay its bills fine. They've raised enormous amounts of money to get incremental compute, in addition to the massive commitments they're making from Oracle, CoreWeave, SoftBank, Microsoft, and Amazon. They're building out insane amounts of compute and know they need more.

    Think about the diffusion of technology: you and I may jump on a new model on day one, but other businesses take time. The "Claude psychosis" moment doesn't hit everyone simultaneously. Let's say a 4.6 Opus-tier model, at the end of the year, the economy would spend $100 billion on — not an unreasonable projection given they're spending $40 billion now and it's just a linear extrapolation, not an exponential. To get the exponential, you need better models. Anthropic won't have enough compute to serve that demand alone. Whoever hits the next capability tier — OpenAI, Google — even at 50% gross margins instead of Anthropic's 70%, they still capture all that incremental demand and probably won't have enough compute to serve all users either.

    The economic value the best model can deliver is growing faster than our ability to serve those tokens via the infrastructure. This gap will continue to grow. The model labs will have expanding margins until the hardware and infrastructure supply chains say, "Wait — why don't we jack up our margins too?"

    Tokenomics and the Hardest Question

    Patrick O'Shaughnessy

    As you try to be the world's best-informed person on both the trajectory of supply and demand, what do you wish you knew that you don't?

    Dylan Patel

    The hardest area for us — and for everyone — is understanding tokenomics: the economics of tokens. We have tremendous insight into how much it costs to run infrastructure, what the cost of tokens are, what the margins of the labs are. But usage and adoption is very difficult to model. In January, we had projections for February, and Anthropic smashed them. In February, we had projections for March, and they smashed them again. Everyone sees a number like $10 billion added in a month and says, "How? Who is using all these tokens? What are they building?"

    More importantly: with what they're building from these tokens, how is that actually diffusing into the economy, and what value is it generating? It's not something you can capture in GDP statistics. All the value of the tokens I use gets transformed into better information, which I sell at a discount to what people used to pay for information — and that information is now making its way throughout the economy, enabling better investment decisions and better competitive decisions. But where is the phantom GDP? What is the phantom GDP? How do we track the real economic value?

    The value being created is clearly amazing by every subjective metric. But measuring the knock-on effect of all these things is the real challenge. We have a great reading on the supply side and even much of the demand-side signals. But quantifying what value these tokens are actually generating — that's what's hard to measure.

    Looking Ahead: Protests and an Industry Under Pressure

    Patrick O'Shaughnessy

    When I come back in three months, what do you expect?

    Dylan Patel

    Large-scale protests.

    Patrick O'Shaughnessy

    Really? Expand on that.

    Dylan Patel

    People hate AI. AI is less popular than politicians, according to Pew. As Anthropic adds so much revenue, that's going to start causing business changes downstream. People are going to get more and more scared. They'll start blaming AI for deep-seated problems that have existed for a long time. Politicians and social-media personalities will start weaponising AI anxiety. Sam Altman has had a Molotov cocktail thrown at his house twice in about two weeks — and the comments on news articles are cheering it on. This is just the beginning.

    Patrick O'Shaughnessy

    What is the counterweight? How should the AI industry head that off?

    Dylan Patel

    First, Sam Altman and Dario Amodei have to stop getting on interviews. Every interview they do makes normal people dislike them even more — Sam going on Tucker Carlson probably made many Republicans hate OpenAI. They just have no charisma.

    Second, they need to start showcasing uplifting things that can be done with AI. Third, they need to stop constantly talking about how capabilities are going to change everything, because that generates fear among people who have no connection to the technology or to the people building it. The average person doesn't know an Anthropic employee or an OpenAI employee. They see a sneaky group of 5,000 people at some company who are going to automate all the jobs, destroy society, and build data centres that pollute the world. They have to stop talking about the future and talk only about the present — about how uplifting AI is right now. A huge rebrand is needed.

    Patrick O'Shaughnessy

    I love doing this with you. Thanks for your time.

    Dylan Patel

    Awesome, thanks.