Gavin Uberti with Patrick O'Shaughnessy
Show: Invest Like the Best
Cleaned and reformatted from published transcript or auto-generated captions — punctuation added, filler removed, restructured for readability. Not verbatim. For exact quotes, refer to the original.
Patrick O'Shaughnessy
It's been three years or so, Gavin, since you and I last did this, which is nuts. At the time I was just wildly intrigued by your story and what you were going to build. I didn't know a lot about chips. I was considering investing in the company, so I was calling everyone I could conceive of who could give me an opinion. And basically the consensus was that these kinds of companies are not built by young people — that the best semis companies are founded by 40- and 50-year-olds who have had a whole career's worth of experience, learned all the problems, shipped multiple chips. Two 21-year-olds are just not going to do this. It was indicative of a theme: nobody believed in you. That's obviously changed a lot now — you just walk the halls and talk to the people who have chosen to come work here. But in the early days it felt like something you had to face down. What was that like, facing down a set of incumbents and an industry's worth of people and investors who didn't believe in you? What did it anneal in you?
Gavin Uberti
I think there's a certain level of naivety required to think you could build a chip better than every other AI chip ever built, and build a company to do it far faster than has ever been done. And we have the naivety. There were many times where we would ask, why isn't this possible? — really push on it — and it turns out that everybody's answers are extremely siloed to a set of constraints that aren't true anymore.
Gavin Uberti
The reality is that the entire semiconductor and data centre industry is built on buffer. What I mean is that every part of the stack — from the EDA tools to the power modules to the circuit boards to the chip design and standard cells — is built to be general purpose for everything, not just in the data centre but for IoT, the edge, and so forth. When you have a specific use case you're really trying to design for, you can change the constraints a lot.
Gavin Uberti
Here's a simple example, and we're not the only one who does it. One of the things you care about is the clock speed of your chip — it's proportional to the throughput of your system. When you're doing signoff for timing, there's this concept called corners: what temperatures are you going to run at this clock speed? The default configurations for a lot of these EDA tools assume you're going to run your chips in freezing temperatures. Now, I've never seen an AI data centre with ice in it. So we can feel pretty confident our chips don't need to run at full speed at zero degrees C — in fact, they're never really going to run below 80 degrees C. Just by knowing that's a constraint that doesn't matter, we can make a ton of changes throughout the entire system. That's a very simple one, but there are many more: you get 20% here, 50% there, 2x here, and these compound into a system that can be radically better for inference.
Patrick O'Shaughnessy
You found two kinds of people. There are folks who went purely on heuristics — young founders claim they can beat the biggest company in the world on performance; it cannot happen; there's nothing you could say to change my mind. But there are also people who are skeptical yet willing to say, I'll spend the time, I'll do the work, and see if it's actually possible.
Gavin Uberti
One of our earliest supporters was a guy called Mark Ross. Mark was a very prestigious semiconductor expert — used to be CTO at Cypress Semi, which sold for $9 billion. When we met him we were just a couple of guys in a dorm room. We came to him and said, we want to build hardware for inference, we think we can be much faster than Nvidia. And Mark said, no, you can't, it will not work. But — if you want to convince me, write a white paper, build a functional simulation, and show me. After a lot of very long nights we went back to Mark and said, here's a simulation, what do you think? And he said, huh, this works. But to do a company like this you'll need a large amount of capital — at least $3 million just to get started. We went ahead and raised five, and a lot more after that. He got more involved, became an adviser, then a halftime adviser, and eventually full-time CTO as he saw more of the development progress.
Gavin Uberti
In general, this has filtered really heavily for the people who want to be right regardless — very truth-seeking people who say, sure, I'm skeptical, but I'll work through the numbers myself. And if I can figure out why this is possible, well, let's go build it.
Patrick O'Shaughnessy
The specifics you've made bets on, the way you've built this system, are immensely interesting to me. Because so many people are trying to build new chips that do a better job of serving inference at massive scale, the world is interested in the architecture approaches. Describe what this thing is and what it does — but, more interestingly, the process you went through to decide what bets to take, and contrast that with what you've seen the rest of the marketplace do.
Gavin Uberti
Start with the product. We're not just building a chip. We're building a full inference solution. That means a rack — the chip, the power delivery into the chip, the board it sits in, the interconnect by which the chips talk to each other, and the production for this mass volume of racks. Really, the production is the product.
Gavin Uberti
There are two key parts of running inference: prefill and decode. Prefill is reading in a huge volume of text; decode is then using that data to generate output tokens. When you run prefill, your job is not to predict tokens — you already know the text. Your job is to get the model's memory, what we call the KV cache, into the right state.
Patrick O'Shaughnessy
And then you can run decode with that same KV cache.
Gavin Uberti
Right. So what we often do is what we call PD disaggregation — prefill-decode disaggregation. You transfer those model memories, those KV caches, over to the decode cluster, and then use that cluster to generate the next tokens.
Patrick O'Shaughnessy
So it's sort of like loading the gun and then firing it, if I think about it in super simple terms.
Gavin Uberti
You got it. It's getting the model to remember the right things and then using those things to do tasks. Generally people think about this market a bit lazily — are you a prefill chip or a decode chip? If you're a decode chip, are you an HBM chip, an SRAM chip, a 3D-RAM chip? Are you using optics or copper? When we started, we just wanted to understand why extremely smart people were working on these different directions. We seriously looked at architectures like a shared memory pool of DDR memory, advanced packaging to break out of the shoreline, putting memory dies on top of compute dies. And we realised there's no free lunch — everything has a trade-off. With 3D RAM you have a thermal issue, a supply-chain issue, you have to figure out hybrid bonding, you have to figure out the flops.
Gavin Uberti
So we went through everything on both the prefill and decode side, and realised there were a few design spaces nobody had seriously tried to explore, because they were never done in AI chips. We asked, what are the metrics that matter most? On the prefill side, it's flops and flops density. People talk about flops as a headline number, but you should care about the flops you're getting on real workloads. There's a concept called MFU — model flops utilisation — how many cents on the dollar of every peak flop advertised are you actually getting. On GPUs you often get between 20 and 50%, depending on the workload. You can provably not run at 100%, because as you increase flop utilisation, more transistors switch on and off, you draw more power, and the chip self-regulates and lowers its clock speed so it doesn't overheat.
Gavin Uberti
So we said, if we want way more flops because we want much higher throughputs, we fundamentally need to solve the thermal problem before we even think about adding flops. If I just add flops to a GPU today, I won't get more performance — it'll just thermal-throttle. The essence of this is Dennard scaling: voltage is quadratically proportional to power. If I double my voltage, my power goes up 4x. If I halve my voltage, my power goes down to a quarter. So we asked, how could we run voltages lower than GPUs?
Gavin Uberti
We flew out to Silicon Valley after dropping out and asked dozens of people at all these different chip companies how they did it. The answer was, you can't run at voltages lower than GPUs. This was dissatisfying, because there are many industries of chips that run at lower voltages — Bitcoin miners run at under a quarter of the voltage of GPUs. So it's obviously physically possible. The question is whether there are issues with GPU architectures that make them unable to run there. After looking at the problem for a long time, we created a new mechanism for running at much lower voltages — a new type of power delivery we call low-voltage inference. We think all AI chips in the future are going to be low-voltage chips: they'll have to cram far more flops into the same silicon area and, without thermal throttling, run at far lower voltages.
Gavin Uberti
For decode, it's all a memory game — more memory-bound. You can load the model faster, load the KV cache faster, and serve more tokens per second per user. People ask the wrong question here. They ask how much memory bandwidth is on your chip. You should ask how much memory bandwidth is on your full scale-up cluster. What we do is add far more bandwidth at much lower latency from chip to chip. That lets us serve models much faster, because you can use the SRAM and the HBM from the full scale-up cluster as a single pool. That's our second key technical bet — what we call cluster-scale memory.
Gavin Uberti
On GPUs today, cluster memory bandwidth is often very badly utilised, because the time to hop from one GPU to another is extremely long. On Blackwell chips it can be about 4,000 nanoseconds point to point. So if you go to an 8-way tensor-parallel setup, you get far less than an 8x improvement in tokens per second per user. We built our own totally custom interconnect stack — we took everything above the second layer of Ethernet and built it full custom — and we can cut that by more than a factor of 5x. As you scale the world size, your time per token goes down proportionally. It's not surprising, given all these architectures were built before ChatGPT. If you're building a chip for modern workloads, it's going to look very different — the way we organise our flops, our voltage domains, our power planes, the packaging, the board design, and, on the decode side, the way we connect everything. We're now bringing forward our first generation of this low-voltage inference technology, which runs at under half the voltage of any other AI chip.
Patrick O'Shaughnessy
If you zoom all the way out — why is this so important? Why is the delivery of much higher throughput, much lower cost per token, better tokens per watt — all these metrics the universe is going to start talking about more and more — why is this the bottleneck in the technology world, looking out a decade?
Gavin Uberti
It comes down to productivity. We're at this extremely interesting moment in the history of civilisation where there's real artificial intelligence — not sci-fi stuff, but models that can solve problems most humans can't. It's going to create new scientific discoveries, instant access to medical care, instant access to education. Now it's just about how many people can use this at the same time, how many products can serve it, and the speed of different tasks. When you think about wall-clock time — if we can take an agent that could take a year to solve a task using inference-time compute, and we have far faster decode speed, we can compress that into a month. So the amount of scientific innovation and the proliferation of technology will happen much faster.
Gavin Uberti
The second part is concurrency. Today it's just not possible for a billion people to use these models concurrently. Some people get downgraded, some people's models are slower, some just can't access the hardware. A few years from now there are going to be giant models serving billions of users. We're very much in the early innings — the paid plans have only a few million users in the world, so we're at one-thousandth of the global population actually using this stuff. To serve giant scale, a lot of things change. One is the number of chips that communicate together. People usually think about that in the context of training — Colossus with over 100,000 GPUs all-reduced together. On the inference side, people think of an 8-chip cluster, or maybe NVL72, as the scale-up domain. But very quickly this becomes thousands and tens of thousands of chips, and the time to send data from one chip to another matters far more than it's getting credit for. If chips can only communicate quickly with themselves and very slowly with others, you won't be able to serve giant models at 10,000 or 20,000 tokens per second.
Gavin Uberti
So we need multiple orders of magnitude of infrastructure built out through the entire stack — from the wafer to the watt, from the transistor to the token. Look at most goods, like the iPhone: they've reached economies of scale where more money doesn't really buy a better one. Whether you're a billionaire or the average American, you buy the same phone. Tokens aren't like that yet. We're still in the very early days, where a general-purpose system is kind of handcrafting these tokens — like they made screws back in the Renaissance. I wanted to live in a world where you have the same economies of scale for token-making that you do for making iPhones or cars.
Robert Wachen
That's one of the huge unlocks that lets a huge group of people use the best-quality models. Economies of scale have made capitalism, I don't know, fair — they let you have the same product in many different hands. Serving far more users on a single scale-up cluster gets you closer to that point for token serving too.
Robert Wachen
And certain products just aren't usable if they're slow. If you want to serve coding models and have people actually use them, there's a certain number of tokens per second you need to hit. So the question is: while maintaining that per-token speed, how many users can I serve at the same time? You can either shut off a bunch of the world from using this stuff, or everyone gets a worse experience. Fundamentally you need to push out the curve — and that's why there's such pressure for new hardware.
Patrick O'Shaughnessy
I'd like to step back and hear both of your stories for how you came to this idea and this company. Rob, starting with you — take it as far back as you want. The very first thing I ever heard from either of you was your personal story, many years ago now, which really blew me away. I'm most interested in your motivation for being here doing this thing.
Robert Wachen
It starts back in high school for me. I've been very unlucky and lucky at different points in life; this was one of the tougher times. At the end of my sophomore year of high school I got injured at a martial arts tournament, and the next day I couldn't walk. They thought it was something wrong with my SI joint. I went through physical therapy, different scans, and they couldn't figure it out. Eventually they found a big bump on my back in an MRI and told me it was a tumour — stage four bone cancer, under a 30% chance of survival. It was a two-year experience of crazy chemotherapy, surgery, and learning to walk again.
Robert Wachen
When you go through something like that, it changes the Overton window of human experience and makes you appreciate what actually matters. You ask yourself what you're going to do if you have the chance to live — you need to be hoping for something. I always knew I wanted to do something very impactful if I got through it. It took me a couple of years to figure out what that was. At the same time, in college, I met people building cool tech and got extremely excited by AI models, especially once GPT-3 came out. I thought, wow, this is the first model that can kind of speak English, and these things are going to get really smart.
Robert Wachen
When GPT-4 came out, there was GPT-4V, the first model with image uploading. I went through my camera roll and found a picture of my back with the bump on it, before I was diagnosed. I said, hey ChatGPT, pretend you're an expert doctor. A patient comes in and says they have this bump on their back — what could it be? And it immediately said, this could be a tumour, you should get an MRI immediately, go to the doctor. I just sat there still.
Patrick O'Shaughnessy
And in your case that took six months.
Robert Wachen
That took me six months. Yesterday this feature wasn't there; today it's here. I went to show my parents and got a notification saying, you're all out of image credits today, you need to get a pro plan. And I thought, holy crap, this is going to change everything — and we clearly don't have the infrastructure to serve it. There are very few things you can work on that can bring this technology to the world faster. There are plenty of smart people working on models; the fabs seemed maybe unreachable. But the hardware was all designed before ChatGPT. Every GPU, every TPU, every AI chip serving these models was fundamentally built before this and retrofitted. There's going to be an entire new wave of architectures — and what a more exciting thing to work on than bringing this to everybody.
Robert Wachen
From a very different angle at the same time, I was running a startup incubator called ProR, which incubated a bunch of companies — some of the earliest being Cursor and Anysphere, which merged, and Etched went through it. As these models got smarter, around 2022, I realised all of these companies are spending all the money they raise on compute. I had this realisation working on my own stuff — oh my god, all the products I want to build are going to cost tens of millions of dollars a year in inference. This is not going to be tenable. The cost structure of every software company changes: COGS is not going to be zero for an incremental user anymore, it's going to be quite high, and a function of inference. And the opex of every business is going to be inference too, as people use more coding agents. So inference is going to be really important, and we're on a decade-long march for it to become the biggest market in the world. Ten years from now there are going to be giant projects where everything in the data centre hasn't been designed yet. We should pick something and work on it. That's how it got started.
Patrick O'Shaughnessy
Gavin, I'd love for you to go back as far as high school, maybe earlier, and tell your favourite hash marks on the timeline that led to your ambition to drop out of Harvard and start this company.
Gavin Uberti
My first job ever was at a company called Xnor, where I did kernel development. I was 17. And a 17-year-old can't sign legally binding contracts, so rather than a traditional NDA, they sat me down and said, Gavin, don't share this information. Xnor was one of the only companies that saw, hey, maybe this is a good trade. I got to work building kernels, and I've done that at a number of other companies since. Xnor got bought by Apple for $200 million. I did the same thing at OctoAI, which got bought by Nvidia for hundreds of millions of dollars.
Gavin Uberti
When you do kernel work, you realise the math is relatively easy. But to get high-speed decode, the thing that matters is data movement. Almost all the work is optimising how you move data around a single chip, or across multiple chips. That's why we built this cluster-scale memory tech — we bring that interconnect time way, way lower, so you can do far more movement and get a much faster time to generate each subsequent token. That's how you build the crazy things Rob's talking about — a year's worth of work in a month, or more.
Patrick O'Shaughnessy
Can you talk about the competitive drive that's evident in some of the high-school competitions you participated in and won?
Gavin Uberti
For example, I was very active in FTC robotics, and I was lucky to have a very talented partner, Sanford. For a long time we were part of a traditional school team — about 20 guys working together, as is typical in FIRST Tech Challenge. The goal is to build a robot that scores the most points, but FIRST puts a lot of emphasis on collaborating with other teams, doing really good documentation, getting others inspired to do the same. Sanford and I decided that rather than do it this way, we were going to win — and we did nothing else besides build a robot that scored the most points.
Patrick O'Shaughnessy
As a two-person team, rather than a 20-person team.
Gavin Uberti
As a two-person team, much, much smaller than almost every other team in the competition. We figured that if we were going to specialise and win the damn games really well, we wouldn't need to advance on the quality of our documentation or outreach — we were just going to win. So we branched off, built a two-person team, built a robot, and decided we'd redesign it every three months. And we did. We had the world record for the highest score at one point, and were rated third in the world by OPR for software development. It was a damn good machine.
Patrick O'Shaughnessy
What from that episode can I translate as an analogy onto how you built Etched the company?
Gavin Uberti
When you think about doing a full rack-scale product like this, there are a couple of key ideas. One is velocity, velocity, velocity — you win by shipping. You're not going to win by having the best outreach or communications. You win by having the best product.
Patrick O'Shaughnessy
You've chosen to have no communications. I'm just realising how exactly like the robotics this is.
Gavin Uberti
Exactly. There are many ways to win in business, but we'd rather focus on building the best product. And we think we can do it with far fewer people. If you're willing to focus on product, product, product, and parallelise relentlessly, you don't need 20,000 people like the big companies. There's a saying, the best part is no part. For us it's also the best vendor is no vendor — as much as possible we want to vertically integrate the entire product, both because we get more performance and because we can move faster. Everything from the chips to the boards to the cold plates to the interconnects to the production, we want to do as in-house as possible. I think we're the only startup right now building its own rack as well as its own chips — and we did it all at the same time.
Gavin Uberti
A couple of years ago, the last time we were public, we had just started our rack team. We brought over Brian Lerer, who built all of Nvidia's HGX and DGX systems — about 80% of the revenue — and said, we're going to build the rack at the same time. We went through multiple iterations of the rack before the chips even came back. Before the chips came back, we made thermal chips with the exact same hot spots we expected our chips to have, so we could build the cold plates, over-pressurise them, and blow them up. We haven't had a single leak since our chips came back with the cold plates, because we'd already validated them. We have a factory in Taiwan with a few dozen people, a clone of a bunch of our test stations, and a 2-megawatt data centre on this floor. We did 24/7 development cycles — day shifts and night shifts — to get the hardware up and running as quickly as possible. That extreme vertical integration and extreme parallelisation of the schedule is what lets you get products to market far faster.
Patrick O'Shaughnessy
Rob mentioned Sohu, the name of the first product. One of the first things you told me was that what you're really focused on is building a machine that can, at scale, produce these things and generations of them — the company itself is the thing that produces this product and subsequent ones. I'd like to talk about a few cornerstones: velocity, vertical integration, parallelisation. But I'm especially interested in your willingness to take huge risk to go faster. Tell your favourite story about why this is the philosophy.
Gavin Uberti
There are a number of stories, but one of my favourites is a time when we were getting close to taping out the chip and realised one of our vendors was way behind schedule. We had two very bad options. One was to keep the current vendor and push our timelines out by about a year. The other was to switch vendors, start over, and also push out by a year. Neither was good. So we looked for option three. We figured out they had a team in Bangalore doing the work, so we shipped a dozen of our top engineers across the world to Bangalore for six months. I lived there for four and a half months personally.
Gavin Uberti
Every morning we'd walk across the crazy busy Bangalore streets into the office, be the first ones in, build a wide variety of tools — auditing a huge amount of the code going in, building tools to go faster, making sure we made the right design decisions on the spot, no 12-hour back and forth. Then at 1 a.m. we'd walk back through the now-empty streets and do it all again the next day. We still had a bunch of the team in the US, so we ran 12-hour handoffs — a 24-hour development cycle. At 8 a.m. and 8 p.m. every day we'd all get on Zoom, share all the data, and say, when I wake up, this must be done. We saw other chips at the same stage with that same vendor that ended up taking years and still aren't out — still haven't even taped out today. It's that level of extreme urgency that's required to bring products to market.
Patrick O'Shaughnessy
This has become a trope because of Elon, mostly — his special skill, and others who seek to emulate him, is to figure out the binding constraint and flood the zone personally on that thing, which is kind of like going to Bangalore. What's the key to doing that?
Gavin Uberti
For me there are two key tricks. The first is that you can't build a chip alone — it's a team problem, and your most important job is to get great people who are inspired and excited to do crazy things like this. It's a huge ask to say, uproot your lives for six months, or in one case twelve. It sucks, but we're lucky to have team members who are in it for the right reasons. The second big thing is being able to make decisions very fast. One of the worst things is when a factory or vendor is waiting for you to make a call and is just stalled. This happens all the time, even for small things. So delegate a big amount of responsibility and say, make a reasonable call. It's okay if you're wrong every now and then — I'd much rather be right most of the time and give an answer immediately than wait every time for the perfect response. Speed wins.
Patrick O'Shaughnessy
What about spending money to go faster? As the world has moved from software back towards hardware, we've outsourced so much of the learn-by-doing overseas, effectively being the idea guys here in the US. That seems to be reversing, and you've adopted this way of learning by doing — being in the iteration loop. Part of that is willingness to spend and take risk with dollars.
Gavin Uberti
There's a great quote — the biggest risk is not taking risk. Every day there's over a billion dollars of revenue in this category, a lot of it inference, so every day we don't ship, we're leaving tons of opportunity on the table. Your willingness to spend money should be extremely high if you can get a clear ROI. We have a concept we call pre-fetching: when you're waiting for one thing, and you know what you'll do once you have it, is there a way to parallelise the whole schedule? We know our chip is going to come back on a certain date, so we want everything that could be done without the chip to be done before it lands.
Gavin Uberti
This costs a lot of money. We built our entire software stack beforehand. We shipped racks to customer data centres without our chips in them — all the networking, all the CPUs, all the storage set up — so we could bring the data centre software up before the chips came back. We took over 700 FPGAs and put the entire full-reticle chip on an FPGA cluster and ran a dozen different models with our full inference stack before the chips came back. We built a thermal chip to mock the thermal profile and built cold plates from that. We had the entire production line ready, did many revs of the circuit board — the entire product was ready to go. There was another very famous AI chip company that took 10 months to go from getting their silicon back to running inference in a rack, publicly announced to their investors as a big deal. We were able to do it in 40 days — because by the time the chip came back, everything was boring. The software was written, the rack was there, the production line was set up. You don't always catch everything, you make some tweaks on the fly, and off you go.
Robert Wachen
The shift made a big difference too. We literally had a day shift and a night shift. Some team members came in at 10 a.m. and left at midnight; others came in at midnight and left at 10 a.m., running around the clock to get to 40 days. Over half the company lives next to the office, which makes that easier.
Patrick O'Shaughnessy
You pay them extra to do that, right? Do you still?
Robert Wachen
Yeah, the invisible hand does wonders.
Patrick O'Shaughnessy
Think about building the early team as two young guys. There are lots of talented young entrepreneurs out there who could benefit from your lessons on getting sophisticated, talented people to join you — even after careers at other great companies. If you were teaching this as a class — how to get elite talent when you're young and inexperienced and naive — what would the syllabus be?
Gavin Uberti
We have a bimodal talent philosophy. It starts with what we call the legends. When you're trying to solve an incredibly hard technical problem — generally something that hasn't been done before — you need the very best person in the world. Often the number one person versus the number 10 versus the number 100 is a huge difference in whether it's even possible to solve. So we created a system we call project-based recruiting: we map out all the hardest technical problems across all industries that anyone has ever had to solve. We look at temporality — who did the zero-to-one, who was actually in charge and did the work. We talk to as many people as possible, and then we track it. The number of people who say yes after the first conversation is pretty low, but the number who say yes after the 20th is surprisingly high. You really have to keep at them. When you hear no from someone who really is the best in the world, that means come back when you have a few more milestones proven out. And it's one of the most convincing things — we make bold claims, and when you hit them again and again, that's belief-inspiring.
Gavin Uberti
When we decided to build a rack and not just a chip, we asked, if we could wave a magic wand, what would the best possible person look like? Someone who started at Nvidia and built the entire rack team through all their generations, learned all this stuff, but is still scrappy, still understands startup culture, and has seen scale. So we mapped all the teams that related to rack-scale products at Nvidia and found three people who could fit the bill. Two of them had just retired, and one was planning to do one more generation for Nvidia and then retire — Brian. Over time we convinced him to join. Brian started the HGX and DGX team at Nvidia, a majority of their revenue, tens of billions of dollars a quarter. The other two ended up investing, by the way. When you have someone like that, they know what good looks like, because they've seen it. There were so many times Brian would point at us and say, that's a billion-dollar lesson I learned. That just saves us cycles.
Patrick O'Shaughnessy
Brian's a legend. What's Sanford?
Gavin Uberti
We say chips on shoulders put chips in data centres. Sanford and I were world robotics champions in high school. Sanford was finishing his senior year of college, and we called him up a couple of years ago and said, come check out what we're doing, we need help on the platform side. He came for a week, and we said, can you build the cold plate this week? If you asked any thermal engineer, they'd think you were totally naive — these things take months, and to be clear, they do. But you can make real progress in a week if you put your mind to it and think it's possible. He built a contraption in a week that de-risked a pretty key power question we had. You put those two together — Brian and Sanford — and they've done incredible things. One isn't possible without the other. You need the extremely driven people who keep asking why and don't know where the bodies are buried, to take tons of aggressive risks; and you need the people who've seen scale and still have the scrappy startup mentality. It's the legends plus the raw, naive, first-principles talent — and it's not just that you have both, it's that they work together.
Patrick O'Shaughnessy
Anything else interesting about how much better you've gotten at recruiting, and why those metrics keep improving?
Gavin Uberti
One of the shocking things is that being such a contrarian bet self-selects. If you're the kind of person who's opportunistic — going to join whatever the hot company is rather than do deep diligence — you will not come work here. It's one of the things I worry about as we announce more of the product and its specs: we may lose some of this if we're not careful.
Robert Wachen
You kind of have to be sick in the head to join our company. On paper, you're a very accomplished engineer making a good amount of money — liquid, predictable — somewhere else. And you're going to convince your family to move to San Jose and live in an apartment on this housing programme, for a semiconductor company run by two 24-year-olds, that's pre-product, going against the biggest companies in the world in the most supply-constrained environment ever created, with a design they say isn't going to be 10% better but 10x better. Something must be wrong with you to do that. People are just wired differently here — they really want to, not prove people wrong who don't believe, but prove people right who do believe. They take it personally. It's really fun to find those people, and frankly the nature of the company makes it very easy to whittle out the people who aren't like that.
Patrick O'Shaughnessy
One of these difficult moments in the company's history is around the ability to raise capital. When you started, you knew you'd need capital, but you didn't know the quantum you've ultimately raised and are spending. And you hadn't raised money before. There were moments where it was really difficult — I was there, I saw it — moments where without that money the company would have died. Talk about the early difficulties, before you had performance to blow their socks off, when it was just you guys talking about an idea.
Robert Wachen
We've had some intense moments. It reminds me of early 2024, before we raised our Series A. We'd done enough architecture and design to know the chip architecture was sound, but there was a lot more to do. We were ready to go into what's called the physical design stage, and we needed to sign an agreement with a physical design vendor — which costs at least $40 to $50 million. Then we had this realisation, as the models were getting bigger, that we were going to need to build the entire cluster, not just the chip: boards, interconnects, cold plates, all the networking. And that was going to cost a lot more than the $15 million we had in the bank.
Patrick O'Shaughnessy
And you're like, that was scary.
Robert Wachen
Sitting in that moment, you think, holy crap, we can't afford this. And I began looking at, huh, how hard is it to go back to Harvard? At the end of 2023 we put together a memo — spent like a hundred hours on it, because we had no idea how people would believe us when we asked for the amount we were about to ask for. It was 30 pages, extremely technical and in-depth: all the things we needed to build, all the milestones, how the market would evolve, the new use cases, the cost per token, all this modelling. Then we talked to investors, and every major investor in the valley passed immediately. Two kids who just finished Harvard, haven't taped out a chip, no test chip — who knows if inference is going to be a big market, everything's going to be training, the models still hallucinate, this could all be a bubble. At the time the biggest semiconductor Series A was around $40 to $50 million. We were tallying the bill and thinking, we're going to spend $100 million in the next 12 months. How the hell are we going to pull this off?
Gavin Uberti
One of our key first moves was to ask, what is the cheapest possible way we could do this? We decided, well, if I made almost nothing and ate nothing but ramen, we'd spend basically just the money for the mask of a chip tape-out, and that would be that. On that basis we could probably do it on $30 million — an absurdly low number. We actually got a debt provider willing to lend us the money to cross that barely-ramen-to-a-chip threshold. From there it was about catalysing a series of, hey, maybe we can do one more thing, one more thing.
Robert Wachen
So we're at this moment where, if we really want to build this company — because we're not going to half-ass it, we're not going to do a test chip and spend years while the entire AI market booms — if we're going to do it, we're going all the way. We're going to need $100 million. I remember Gavin and I sitting in the office in Cupertino late at night, looking at each other — could we cut 500k here, 100k there? How long could we convince everyone not to take a salary? And the math was not going to close. There was a period of a few weeks where you go into survival mode and call every person who could possibly know an investor: we need $100 million, and if we do this, we think this could be one of the most important companies of all time. Do you know somebody who wants to take an aggressive bet, who wants to believe in us? Here's all the information, we're an open book. And the snowball starts — a million here, 2 million here, okay, we're not going to run out this month. A $5 million check, a $10 million check — okay, maybe I can buy those FPGAs. We were very lucky it ended up coming together. We had a board meeting, I show the spreadsheet, and it's 103 million — all soft commits. We look at each other and say, we're going to take it. That was the Series A. It's been much easier since — we've raised almost half a dozen rounds, many from those same investors doubling and tripling down. This rack would not be possible had we not been so aggressive.
Robert Wachen
I also think the suppliers deserve a commendation. TSMC was willing to work with us before we'd raised any of that hundred million, back when it was still really scary — they let us get some of their emulators on extremely favourable terms, where we pay over many years. Basically a big loan. It takes a lot of belief from your partners to do that. At the end you come out with a very strong team, and all the folks who back you aren't in it out of pure financial incentive — they believe.
Patrick O'Shaughnessy
Why did TSMC believe, do you think?
Gavin Uberti
This is a great story, even before you joined. There was a semi conference event, and I was one of the only young CEOs in semiconductors — a bit of a novelty, so they asked me to come speak. I get there and I'm the only speaker under 40, and the only person there under 30 — I was 22 at the time. So I go up and speak. There's a speaker dinner afterwards, and by pure luck I'm seated next to a very senior TSMC VP. It's a very nice, bougie dinner — the former CEO of Arm is there, everyone's in a suit. It turns out the VP and I both studied math in college, so we get a little piece of paper and start talking in great detail about how modern AI models work at the per-tensor level. The guy just gets it. We talk about how you run this effectively, why low voltage is such a critical technology. The following day I get an email from TSMC saying, Gavin, I want to work with Etched — find a way to make it happen. They've been a great partner ever since.
Patrick O'Shaughnessy
I should break the fourth wall here — I'm a big Etched investor, been involved for a long time, and I think the absolute world of you guys, so I'm incredibly biased. I'm trying to ask questions that could be objections. But it's so interesting that when you read about investing, everyone cites contrarian-and-right as the quadrant that makes all the money. It sounds nice, but contrarian means everyone else thinks you're stupid. When you get immediate nos from literally everybody, it's a fascinating quadrant to exist in before you become consensus.
Patrick O'Shaughnessy
At the time it was, by a lot, the largest first check I'd ever written. I'm not a math or semiconductor or AI expert, so it was much more that I believed in the concept of this market potentially being huge, and in you two having made very clear bets on how the future would look and positioned the company to attack them in a hardcore way. What I felt about you was the majority of the reason we made the bet, back in 2023 or whenever it was. The same thing you said about naivety applies to investing: I didn't know what I didn't know. When I called experts, they laid out in very logical terms why this was such a low-probability bet. One thing I've learned is that you kind of have to damn the base rate. If you invest on base rates, you should do something other than what we do — there's always the index fund. So it's actually never been scary for me, probably because there's a lot I don't know. If I knew more about the difficulty of what you've done, I probably wouldn't have done it.
Robert Wachen
A lot of the traditional semiconductor funds missed all the AI chip companies, and all the coding experts missed all the coding companies. It's very hard to realise that constraints have changed. When you've looked at tape-outs for 20 years and seen so many not work on the first, second, or third try — couldn't even run a workload — you forget that EDA tools are far better now, that FPGAs exist today in a way they didn't before, that the validation you can do just wasn't possible. Our believers were on two sides: either believers in the market and the team, or people building chips today, extremely technical, like the high-frequency trading firms. They'd audit everything from the micro-architecture and RTL to the board designs, the schedule, the software stack. We'd sit down with 10 of their people who build their own chips, and they'd ask such detailed questions we wondered if they were going to build the chip themselves. If you were anywhere in the middle, you just wouldn't understand it.
Patrick O'Shaughnessy
In the investing world they talk about variant perception — something you see or believe that others don't, and that perception creates the opportunity. I've invested maybe five times in that, and every time the stakes get bigger and scarier. Because you've been so quiet in the marketplace, it's very easy to dismiss you, and as the stakes get bigger, those dismissals are harder to hear. The last thing I'd say is that the accumulated evidence of you and your team's ability to solve seemingly impossible problems is one of the most interesting things a company can have. It's almost binary — companies do this or they don't.
Gavin Uberti
That's a big advantage of people who've been here a long time. You get new joiners who are scared, and then there are old-timers who've been here for all of two years, smoking cigars in the trenches — another one, another one. There's a find-a-way mentality. If you're here, it's because you assume it's possible, so we can't be saying it's impossible. Everything is solvable, and we'll just work at it until we figure it out. I have a favourite story about a guy who's a legend in silicon validation who joined us. We were doing the early stages of wafer sort — when your chips come out of fab on wafers, you have a probe card that attaches with probe pads and sends electrical signals to test which chips are good and bad, so you only package the good ones. We go through our first wafer at like 2 or 3 a.m., because we're doing it with TSMC over the phone in Taiwan. The screen shows the wafer, all grey, and as you run the patterns the squares are supposed to turn green or red — and they all turn red. Everyone's tense. He leans back and says, the puzzle begins. You have to have that attitude: yes, you'll stare into the abyss and see scary things, and we'll solve them.
Patrick O'Shaughnessy
When did you see the first green square?
Gavin Uberti
Within a day. But in the moment you think, I've worked for years for this, I've put my life on the line, I've asked my family to stake everything on it — and then it's red. There's a certain type of person who's addicted to that feeling, feeling the fear and solving it. We're lucky to have a lot of them here.
Patrick O'Shaughnessy
Thinking ahead to Gen 2, Gen 3 and beyond, what will you do most differently as a result of everything you've learned — conceptually, in how you attack designing and producing the next one?
Gavin Uberti
It took us a while to get to the primitives that really matter for scaling inference. We tried a bunch of things early on — compilers that turn different models into FPGAs, burning weights in silicon, splitting your HBM between KV cache and weights. There were a lot of cycles of learning until we realised that fundamentally, if you want to run a majority of the tokens in the world, you need to do three things. Build a chip with the most flops in a given power budget. Build a chip with the lowest latency between chips — the biggest scale-up domain possible. And produce as much of it as possible.
Gavin Uberti
In the first half of our journey we learned the first two, and that informed the design — the low-voltage inference and cluster-scale memory bets. But the production part has become extremely obvious in the past year: how much people want to deploy this stuff if you can have it available today. The best ability is availability. If I have a thousand chips today, someone's going to use them. So we need a chip that's not just way better than what's been built before, but available at many-gigawatt scale — a product that's producible at gigawatts per month. A lot of the design decisions in our next-gen, which you've seen, are just about simplicity: removing tons of parts, learning how to assemble and disassemble the thing again and again as quickly as possible, making it reliable, serviceable, and producible at gigantic scale.
Patrick O'Shaughnessy
What about problems outside your control — capacity at the leading nanometre at TSMC, availability of HBM4 memory, everyone fighting for a scarce unit of capacity? How do you face those realities when you're trying to produce as much as humanly possible?
Gavin Uberti
People deploying the most compute in the world think about supply a bit zero-sum: there are only so many wafers on a given node, only so much memory. That's why, for our first-gen product, we built it on a different supply chain than the Rubins — we're on 4-nanometre, Rubin's on 3-nanometre; we're on a different HBM than Rubin, and so forth. So it's actually not zero-sum, it's positive-sum, where more is more. When we talk to people deploying at scale, it's not a decision between a gigawatt of a GPU and a gigawatt of us — it's 2 gigawatts. And you have to think about supply chain early in the design decisions, because if you have the most performant product and you can't produce it, then you're just a podcast.
Robert Wachen
That's the other big thing about vertical integration. For the chips and the memory you have to partner — those are highly in-demand components. But the more you build yourself, the more you can do on top of what the world can currently build. It's not that you're taking availability from someone else; you're adding way, way more. That's how you win.
Patrick O'Shaughnessy
What are the limits to vertical integration — how do you know where to draw the line? I ask because the broader market is interesting: the vast majority of AI chips get bought by a small set of customers, many of whom are themselves trying to design their own AI chips. OpenAI announced Jalapeño. It seems like this funny circumstance where the most valuable thing in the world flows through a couple of chip makers and buyers, all thinking about doing each other's jobs. Then these things go in a data centre, and you've got neoclouds, inference providers, model builders. I can imagine a world where, because you have the best hardware, you design models and build data centres — you leak outside your current vertical. So how do you think about where to draw the lines?
Gavin Uberti
We have a saying: production is the product. What matters is that we know inference is going to be the biggest market in the world, and whoever produces the most tokens is going to be the most valuable company in the world. So every decision is, how do we get the most token capacity online? Part of that is a really good product — more throughput, better latencies — so per chip we get more tokens. Another part is not doing parts of the stack unless we absolutely have to to get to giant scale. We build the rack instead of just the chips because it was required to get to scale; we do a CM model instead of a JDM model for the same reason. But there are parts that are just noise to us right now — we're not building our own data centres today, because that doesn't help us get more capacity online. In general, our customers are making power and moving their clusters around to get our chips online, because they're such high throughput. If there were a world where other things were a constraint, we'd integrate with them. But we're purely focused on getting as many tokens online as possible.
Gavin Uberti
It comes down to economies of scale. At certain parts of the stack there are huge economies of scale, and at others there aren't. On designing models, huge economies of scale. On chip fabrication, same story. But building some small metal part inside the rack — there's not the same effect.
Robert Wachen
We think the natural boundaries are the chip on the bottom and the model layer on top — and we'll fill the whole gap between.
Robert Wachen
A few weeks ago there was a guy running a next-generation AI chip for one of the frontier companies, trying to recruit one of our architects. And this person did an uno-reverse and started recruiting the guy who was trying to recruit ours — and within a week we hired him. I was on a walk as we finalised the offer, and I asked, you're leading this super important project, why are you deciding to join? His answer was super interesting: it fundamentally is not existential for my company for this product to win. Google won't fail if TPUs fail — their revenue comes from search. Meta won't fail if MTIA fails, Microsoft won't fail if Maia fails, OpenAI won't fail if Jalapeño fails. Ultimately this is our product. It's completely unsurprising that the best chip in the world is built by a company that only builds that chip — that's Nvidia. For us it is completely existential to get as much token capacity online as possible. That recruits a set of talent, and support from suppliers and customers, who view it with the intensity we do.
Gavin Uberti
Look at the raw flop stats. Compare any of these chips built by the labs or hyperscalers — the flop density for FP8 times FP8 is lower than the Blackwell B300. That makes sense, because they don't have to take the risk. They just have to build a similar-enough product and not pay the Nvidia tax.
Patrick O'Shaughnessy
As you build the solution, you're solving a sequence of really hard challenges. What's been the single hardest episode to overcome?
Gavin Uberti
When we were designing the chip, we built this massive FPGA cluster to verify the full chip worked. But FPGAs are digital entities — you can test digital logic, not analog logic. When the chip came back, we began to see incorrect results in our attention implementation. We realised there was a problem where the back-pressuring logic across a clock-domain crossing was failing, and it was going to cause the chip to produce wrong results. It was very, very hard to solve. There was one and only one way: we had to line up two clock signals on our chip to within 50 picoseconds — 50 trillionths of a second — and do it on every chip, two billion times a second.
Robert Wachen
A lot of people said this was impossible. We had people quit — people literally said, this problem is unsolvable, best of luck, guys.
Gavin Uberti
When you have a problem like that, step one is, okay, let's assume the problem is solvable — how would it be solved? We realised we had to find a way to move our clock phase by a picosecond, 10 picoseconds. We had an idea: what if we had two clocks, set just a little apart from each other, figured out the phase, and used the drifting mechanism to wait for just the right amount of time to get that 50 picoseconds always lined up? We could do this extremely reliably and then lock the phases exactly where they had to be, and guarantee it never happens. People were somewhat blown away that it worked, and worked as well as it did. But we made it work.
Patrick O'Shaughnessy
How long did that take?
Gavin Uberti
About two weeks. A dark two weeks — a very scary two weeks. But that's the most important time to invest effort, when you feel like things are hopeless. The sooner you solve that problem, the sooner you get back to building and scaling up production.
Robert Wachen
A lot of our story is, as Gavin says, assume it is possible. Assume it's possible to have a chip with far more flops, a system with far lower latency between chips, a shared memory pool at far higher bandwidth — and ask, how would one do it? A lot of the time we'll run dozens of experiments and all of them fail. But we only need one to work. During our chip bring-up, Gavin was leading the charge with about 30 different board experiments, and three of them worked — and all three are worth their weight in gold.
Gavin Uberti
People come to me and say, Gavin, almost none of your experiments work. And I say, I only have to get lucky once.
Patrick O'Shaughnessy
We haven't talked at all about the models themselves, which is kind of crazy given they're the thing behind all this demand. How do you, as thinkers about hardware, think hardware might impact where the models go in the future?
Gavin Uberti
One of the most important ideas we believe is that machines don't think like people think. Airplanes don't fly like birds fly. For people, storing data and loading memory is very cheap for neurons, and doing math is relatively expensive — and it's the exact opposite for chips: loading data is very expensive and doing math is very cheap. As time goes on, I think math gets cheaper at a faster rate than memory, due to a fundamental limit on any DRAM device. So you should think about how to make your model use a huge amount of compute. What if I had many copies running at the same time? What if I activated a huge number of experts, or had gigantic experts I can run on multiple servers at once? That's how you build models that are the next generation of intelligence.
Gavin Uberti
And context, too. There's been a lot of work on efficient inference — what if I don't load the full context into memory? Most of the time that makes sense, but if you want to build a superintelligence, why can't it look at a billion tokens of context, spend a huge amount of compute to read all that in? I'd love to talk to a machine that could attend to every book ever written as its short-term memory, and I think you're going to get to a point where you can.
Robert Wachen
A theme in models right now is dynamism — the ability to control the level of computation and memory spent at a per-token or per-user level when doing attention, and to dynamically send data to other chips on the fly for different operations. As we scale context length, model size, and computation per user, we're looking for ways to be more efficient. Mixture-of-experts architectures, where maybe not every parameter is used for every token. Even at a token level — this token needs context from that token, so they share memory and we don't have the overhead of using the memory as much; this token is really important, so we spend more compute and give it longer context. Hardware that accelerates these very dynamic computations is extremely important. Current hardware, designed before those architectures, has lots of overhead doing them — so you end up in bad worlds where either you can't run it well, or you have blocky architectures applying blunt force to many tokens that all need more or less computation.
Patrick O'Shaughnessy
When inference gets much cheaper, faster, more accessible, and better, what are the things you think people will use that capacity to do that are most exciting to both of you?
Gavin Uberti
There was a viral tweet by Noam Brown that as models have longer time horizons, they can do tasks that take, say, six months — and there's often not enough time to evaluate them, because by then you'll have a new model you want to evaluate instead. With tech like our cluster-scale memory, you can run that six-month job much faster. But there's a second piece, talking to him — it's not just the time, it's the number of people or agents working on it. If you ask, can a human build a rocket, the answer is no — no one person can; you have to put a team together. The same will be true of agents. If you ask, can an agent build some crazy futuristic piece of software, you'll probably need a very large team — maybe 10, maybe a million. You need an enormous amount of cluster-scale memory to get that very short time per token, and a huge amount of flops to run the whole fleet.
Gavin Uberti
I'll be a little futuristic. I firmly believe we're on a global march of inference becoming a majority of global GDP. It may take more than 10 years, but it's going to happen. Right now we measure productivity as GDP per capita, but it's going to look much more like agents per megawatt, or agents per gigawatt. And I think this is the second-to-last year where a majority of the workforce is going to be human. In 2027 you're going to see more agents doing knowledge work than humans. It's going to be extremely interesting.
Robert Wachen
You could imagine a world where, for countries, a majority of their energy ends up going into data centres doing inference, and the energy efficiency of those data centres governs how many agents — and therefore how big their workforce — they have. Right now we have one agent, or a team of five to 10, working on group projects for a couple of days. That's pretty cool because they're smart, but it's not civilisation-scale. What happens when a country can have a billion concurrent agents — a billion people in the workforce working 24/7 on the same stuff? It's kind of unfathomable. It's going to be the biggest proliferation of technology humanity has ever seen.
Robert Wachen
And when you have these huge amounts of demand, you get economies of scale again. Think about people: I have a brain, but I'm not using the whole thing at the same time — only a part is active, and that's how healthy brains work. MoE models work much the same way — only a small fraction of parameters is used for any given token. But if you have a large number of users on a piece of hardware, you can take that brain, cut it into many experts on many servers, and run a huge volume through it. Many pieces of traffic use each part of the brain at any point in time, and you make the cost per thought, cost per token, way lower. So you end up with these giant, distributed brains — and the form factor is a big data centre with a huge amount of flops and scale-up interconnect.
Patrick O'Shaughnessy
Do you think we'll see a trillion-dollar individual data centre?
Gavin Uberti
Absolutely. It's a matter of time. It's like asking whether you'll see a billion-dollar fab, a $10 billion fab, a $100 billion fab. The economies of scale don't stop at, oh, $40 billion is the magic number for fabs — the cost per wafer keeps going down as you spend more. The same will be true of plants that make steel, or plants that make tokens.
Patrick O'Shaughnessy
A very smart alien lands on Earth and wants to know from each of you how you'd frame up this opportunity you're tackling. What do you say?
Gavin Uberti
Frame it as: thinking is really valuable. Every company in the world runs on thinking, and we're entering a unique moment where machines can think almost as well as, and soon better than, the best humans. Building these machines is a huge opportunity. But more important, the way you run this kind of thinking is going to be very different as demand goes higher and higher. There's a unique moment right now to build a new set of solutions — a new roadmap for how you run the future quadrillion-parameter models for a billion people at the same time on a gigantic scale-up cluster.
Robert Wachen
We're in a new era of intelligence where the cost of producing intelligence is so much cheaper than the value of the intelligence, that we're in a many-year — probably many-decade — supply shortage of these tokens. Any chip or system that can produce tokens is likely to be extremely valuable. You should find some part of the supply chain of the token — everything from model training down to what we're doing in the silicon — and spend your time pushing the frontier there. The largest companies are going to be the ones that produce most of the global supply of tokens and own a majority of that supply chain. And importantly, it's the people who build systems where, as you add more chips together, they get cheaper. You don't want it to scale where serving 10 times more tokens means buying 10 times more servers. You want a solution where serving 10 times more tokens gets you an economies-of-scale benefit — with our cluster-scale memory tech — so you don't have to charge 10 times more for those tokens.
Patrick O'Shaughnessy
What a ridiculously exciting future you guys are building to enable. When I did this with Gavin last time, I asked my traditional closing question, so this time I'll ask you, Rob: what's the kindest thing anyone's ever done for you?
Robert Wachen
During my cancer treatment there was a big decision I had to make. The doctors came to me and said, it's time to decide — do you want surgery or radiation? Here's the trade-off: if you get surgery, you're more likely to live, but you have to assume you'll never walk again; if you get radiation, you'll be able to walk, but not the same probability that you'll live — you may die. What do you want to do? And I was 16. My parents said, you have to make this decision for yourself. I thought about it a long time and decided to do the surgery.
Robert Wachen
When you get a tumour resection, they do a necrosis analysis — they look at all the cells and say whether each is dead or alive, because if you have a lot of live cancer cells, you have a problem. They said, you usually want 98, 99% necrosis for us to say you're in the clear — you're below that, you should get radiation. There were only a few machines in the world that could do the type of radiation I needed, and one was in Boston. I was in a wheelchair, and I needed to move to Boston for multiple months. Both of my parents decided to move out and drop everything they were doing to live with me. I'm eternally grateful.
Patrick O'Shaughnessy
Beautiful. Thanks, guys. Amazing conversation.