Sergey Levine on Robot Foundation Models, Generalisation, and Embodied AI

Sergey Levine with Patrick O'Shaughnessy

Show: Invest Like the Best

Watch →

Cleaned and reformatted from published transcript or auto-generated captions — punctuation added, filler removed, restructured for readability. Not verbatim. For exact quotes, refer to the original.

Contents

    Patrick O'Shaughnessy

    My guest today is Sergey Levine, one of the co-founders and researchers at Physical Intelligence. As a disclaimer, I'm an investor in Physical Intelligence because I believe it's one of the most important companies tackling the problem of robotics. As you hear us discuss today, robotics has what I would call a scarecrow problem. All of these amazing physical devices are becoming ever more possible in all sorts of cool permutations, but what they all really need is an intelligence, a brain, and that is what they're developing at Physical Intelligence. They're trying to develop foundation models that can make any physical robot do any task in any environment. That challenge is daunting and has required many of the world's best researchers. Sergey is one of those leaders coming together to try to solve this problem. The nature of our conversation today is all of the problems facing robotics and all of the promise of solving these problems across the world. I hope you enjoy this great conversation with Sergey Levine.

    Defining physical intelligence: the robot's brain

    Patrick O'Shaughnessy

    This is going to be a real treat, and a blast, to learn about possibly the most exciting, impactful area of technology being developed. Just to set the stage, before we go back in time — maybe you could define physical intelligence as you see it.

    Sergey Levine

    Fundamentally, the goal of Physical Intelligence is to develop robotic foundation models that can control basically any embodied system to do any task. Broadly speaking, you could imagine that in the same way a language model is rapidly evolving towards a system that can do any task that can be expressed in language, what we would like is to build a new class of models that can do any task that can be done by a physical, actuated device. Part of the thesis of this company is that we believe doing it at the full level of generality might, in the long run, actually be easier than trying to special-case very specific, narrow application domains — much as, for language models, it turned out to be easier in some ways to solve natural language tasks in their full generality than to narrowly target things like machine translation or sentiment analysis.

    The language-model analogy: leveraging broad, weakly labelled data

    Patrick O'Shaughnessy

    That may not be obvious — why make that bet, versus a robot that just does your dishes? What are the key trade-offs to understand, and why make the decision you made?

    Sergey Levine

    Let me give you a two-part answer. First is the analogy to language models, and second is what that means in the robotics world. The first part is a little more informed by evidence. In the world of natural language, there were a lot of efforts to develop domain-specific solutions for specific problems — somebody would spend a lot of time thinking about how English differs from French, and build a machine translation system for it. The reason language models actually took over from all of those different application domains is that they can leverage much broader sources of data. And it's not even as simple as saying we had this data for this application, this data for that one, so let's merge everything — it's more than that. When you can leverage weakly labelled data — data you just mine from the web — you actually learn more about the world. You establish a foundation of world understanding, and on top of that foundation it turns out to be much more effective to build out different applications.

    To bring this into robotics: obviously that calculus doesn't look quite the same, because in robotics we don't have an internet-sized data set to draw on. But this notion of understanding the world is, if anything, even more important in robotics. If you have many different tasks, maybe even many different physical systems, you can go from training an individual dishwashing specialist, or a laundry-folding specialist, to training a model that actually understands physical interaction. People can master new skills very, very rapidly because we understand physical interaction — we can intuitively grasp what's going to happen in a new, unfamiliar situation, and that lets us bootstrap things really quickly. If we can draw on data from many sources, many applications, many robots, we can have a model with genuine physical understanding, and it becomes much easier to put new applications on top of that platform.

    Why generalisation makes a poor demo

    Patrick O'Shaughnessy

    What's the hardest part about building it this way, for you — when you see other approaches that are more legible to the average person? There's a robot moving around doing this one specific thing, it looks a certain way. What's the hardest part about the approach you're taking?

    Sergey Levine

    This has actually been an issue for my whole career, because — and this matters more the more general you get — effective robotic learning, effective generalisation, isn't the optimal way to have a really exciting demo. The way to have a really exciting demo is to pick a really cool task, control everything else in the environment, set it up so it's perfectly clean and pristine, and just make it work in that one setting. That's how you make a robot demo. Generalisation, you can't show in one spot like that — the whole point of generalisation is that the system does something relatively mundane, that any human could do, but does it in any situation. We released some demos last April showing our robot cleaning kitchens. I think it's kind of cool, but if you watch an individual video out of context, it just looks like — okay, it's picking up plates, anybody can pick up plates. Except we'd just put it into that home for that demo, and it had never had training data from that setting. You have to understand what's going on to appreciate why that's actually pushing the frontier.

    What success would unlock: the toolkit, not 'one thing'

    Patrick O'Shaughnessy

    What is your model for the stakes of what you're doing? If you're successful — I'm curious how you'd define success, beyond just crossing this chasm of general physical intelligence. If you cross that line, then what?

    Sergey Levine

    One of the things I think would be really exciting, enabled by a general-purpose embodied foundation model, is unlocking people's imagination in how they build robots and other embodied systems. Personal computers were a huge deal, in my mind, because they made it possible for lots of people to hack together all sorts of cool stuff — there was this Cambrian explosion of amazing applications that started in the '90s, and was further accelerated by the internet. I think something like that might happen in robotics, but it can't happen today, because if you want to put together some cool new robotics application, some cool new idea, you have to build this monstrous stack — you basically have to solve the intelligence problem yourself first. But if there's a foundation model you can prompt, that provides basic functionality, and you can fine-tune it a little or adjust it to your application, it becomes a lot more tractable for lots of people, companies, and individuals to try things out.

    We sometimes think robots are going to be one thing — there are people, and now we're going to make metal people, and that'll be robots. I don't think that's how it's going to be, because no technology has ever been like that. It's going to be more like a toolkit, where you can put together all sorts of cool applications, get creative with it — maybe I'll make a robot with five arms that hangs from the ceiling — and figure out the right thing to tackle your domain, maybe experiment with the software side too. But you need the right platform to build on top of. I think the foundation model can be that thing.

    Humanoids, big and small machines, and the case for one general problem

    Patrick O'Shaughnessy

    What are, in your mind, the pros and cons of the humanoid approach to robotics?

    Sergey Levine

    One pro is that it's really cool — you can show it to somebody and they get it immediately.

    Patrick O'Shaughnessy

    Yeah, you hear a lot about the Optimus hand.

    Sergey Levine

    Yeah. But it is cool, and I think there's real value in that — value in capturing the imagination, and value in getting people to think about what the future might look like, in a way that's understandable. But in my mind it's one of many possible kinds of robots we're likely to have. Fundamentally, the challenge of intelligence looks very similar across all these different robots, and I don't think we should be tackling intelligence in the context of one specific body — we should handle it in a general way, because otherwise it's really hard to get a handle on. We need lots of data. The cool thing about building robots is that, ultimately, they don't have to be constrained to look like humans at all — you can build the right tool for the job. You could imagine building a house with a robot that's a swarm of a thousand quadcopters. In the future we'll have a robotic foundation model that can be adapted to all sorts of applications, running the gamut from bulldozers to humanoids to robotic arms like this one. It might need to be adapted to each one, maybe fine-tuned, maybe given something in context to understand how that body works. But the fundamentals of how you interact with objects, how things move in the world, how causality works — that's all conserved across these different systems.

    Patrick O'Shaughnessy

    Do you have a favourite example of what might be possible with true general intelligence that might not be possible with, say, a humanoid-only intelligence?

    Sergey Levine

    A few things are worth thinking about. One is that we can make machines that are very big, and machines that are very small. In the long run — this isn't a short-term thing — I think there are exciting applications in medicine, in surgery, where we might not only not be limited to robots that look like humans, we might not be limited to robots that can even be controlled by humans. Currently, robotic surgery is done entirely through teleoperation — you need something a person can control in real time, with the right level of dexterity, and that limitation holds for current learning-enabled systems too. But in the long run, we could imagine addressing that.

    A short history of robotic learning, and a personal path into it

    Patrick O'Shaughnessy

    If you think about the most important hash marks on the timeline of robotics research that have gotten us here — it's always helpful to set the historical context before talking about where things stand today and where we're going. Could you walk us through the relevant hash marks on that timeline?

    Sergey Levine

    At some level, doing end-to-end control for robotic systems is a very old idea. The first autonomous driving systems that used end-to-end learning existed in the 1980s — ALVINN was, I think, 1986 or '87, a driving system demonstrated to drive on highways, controlled by a neural network, from a camera. The neural network was tiny. So there are some venerable concepts here, but historically what's been really difficult in robotic learning is that you need a system that handles the application you want to address, that's cost-effective to train — meaning you don't need a huge amount of data for every single application — that handles long-tail scenarios with common sense, so if something weird happens it has a reasonable response, and that, for the thing it's actually supposed to do, is robust, fast, and reliable. Getting all of that together is very hard, because machine learning works best when there's a lot of data. If you naively approach a robotic problem and say, 'I want to do washing dishes,' you'd have to collect an enormous amount of dishwashing data — but that's not cost-effective, because you go on to the next application and have to go through the whole process again. Being able to train general-purpose models that handle many tasks is essential, because then you need a lot less data for each new task.

    But even further — and this is the thing that's probably changed the most in the last few years — you also need to handle the unusual scenarios, and for those you probably won't have experience. What you need to rely on is knowledge acquired from other sources that you can ground in the new situation. People are extremely good at this. If you're driving and there's something going on in the middle of the road, and someone's put up a sign saying, 'Don't go here, there's a gas leak,' you've probably never experienced that exact thing before, but you can put the pieces together and figure out what to do, because you have common sense. This has been a huge mystery in the robotic learning world — where do you get that common sense? What's changed in the last few years is that multimodal language models are really good at pulling in knowledge and articulating it. They're not very good at grounding that knowledge in physical situations, but they know stuff. So now there's a path to get that kind of common sense, by leveraging the knowledge contained in multimodal LLMs. But there's also a challenge, because you have to plug into that knowledge in the right way — you can't just show it a picture and ask, 'What would you do here?', because it doesn't have the context; it doesn't know that you're a robot, what you look like, what's going on. That's a technological challenge, and we've made some headway on it, as has the research community generally. But the important thing is that light at the end of the tunnel — we now have a way of pulling in knowledge that helps with those long-tail scenarios.

    Patrick O'Shaughnessy

    Are there hash-mark equivalents on the timeline — like AlexNet, or the Transformer? Big events everyone will point to when writing the history books?

    Sergey Levine

    That's a good question. It's very early to answer definitively — you'd have to look back at least ten years or so. But probably the first end-to-end learning systems, in the '80s, are a milestone. The first deep reinforcement learning systems, in the early 2010s, are probably a milestone too, because deep RL gives us a way to go beyond human-level performance, which I think will be essential for robotic systems. Then there's the more recent stuff, just in the last few years. I don't know how that will shake out as far as what people point to, but I do think the advent of multimodal LLMs that can be adapted to robotic control, to bring in common sense, is a really important advance. We're probably going to see quite a few important advances in the next few years, and maybe those will be the things people point to.

    Patrick O'Shaughnessy

    Can you tell us your own personal history of approaching the problem — the origin of when you first became interested, why, and how you've decided what to spend your time and attention on since?

    Sergey Levine

    I started working in robotics in 2014, after finishing my graduate degree, when I started a postdoc with Professor Pieter Abbeel at UC Berkeley. I hadn't worked on robotics before, but I figured I should get a bit more education after my degree, and his lab worked on robots, so I tried applying what I'd learned to robotics. Before that I worked on computer graphics. I think the thing I've always wanted to figure out is how to get AI systems that get better and better the more they do things, because that's tremendously powerful — if a system keeps getting better the more it does something, there's no real limit to the skills it can master.

    Initially I approached it in a blank-slate way: you start with nothing, practice a particular skill, and get better at it. That worked in a limited setting, but it was very hard to turn into a general system that could work in open-world settings, because if I practise something over here, and the robot then goes over there, something's different, and it needs to practise all over again. The next thing I tried — this was when I worked at Google afterwards — was to see if we could parallelise that across many robots. Collective learning: can you put twenty robots in a room and have them all learn together? That works, and it generalises. But it's very hard for that to handle tail cases, edge cases, because the system becomes a kind of savant of that particular task, and that's all it knows.

    So the next step, I think, is combining this ability to practise skills with lots of prior knowledge. That's a really hard problem — not just in robotics, I think it's hard across all of AI, because arguably the two big impressive results in AI over the last few decades have been generative AI and deep reinforcement learning. If you want a single example to epitomise generative AI, that's LLMs. For deep RL, it's AlphaGo. They're both very impressive, for very different reasons. Generative AI is impressive because it can reproduce things humans can do — draw pictures that look human-made, write text. Deep RL is impressive for the opposite reason — it does things humans hadn't thought of, like Move 37. So the big challenge, and what I hope we'll figure out here at Physical Intelligence, is how to combine those threads: bring in all the knowledge you get with generative AI, but also go beyond human-level performance with reinforcement learning. I haven't figured that out yet, but I think we've made good progress on it.

    Vision-language-action models, chain of thought, and the espresso example

    Patrick O'Shaughnessy

    So what have you literally done, and what are you doing, to make that happen?

    Sergey Levine

    Over the past few years, we started by developing the basic foundations — what's called a vision-language-action model. You can think of a VLA model as an LLM adapted for robotic control. These models are first trained on text data, then adapted with lots of image data from the web to understand images, and then adapted to robots with lots of diverse robot data. That's a starting point — a way to take all that web knowledge and get it into a model that can control robots, and get some interesting behaviours out of it. From there we studied two threads: how to get the model to handle unusual situations with common sense, and how to get it to improve with reinforcement learning.

    The way you get common sense is by using chain of thought. The robot enters a scene, and instead of directly starting to move, it thinks about what it was asked to do. If it's told to clean up the kitchen, it looks at the scene and says, 'Okay, based on this, I should pick up the plate.' It literally talks to itself — it says 'pick up the plate,' and then goes and does it. That unlocks all this prior knowledge, because those intermediate inferences benefit from the web-scale pre-training. That handles edge cases. Then the reinforcement learning part comes in after you've practised a task a few times — you keep getting better and better at it directly through experience. For example, we had a demo of making espresso. That system practised making espressos many, many times, and used that to improve robustness, speed, and throughput. We're not done with that — there's a lot more to do — but we have the starting point.

    Patrick O'Shaughnessy

    Sorry to be thick about it, but is the robot data itself the right way to think about this? I'm looking at the gen-one of these things — I see a camera here, maybe some sensors elsewhere. Is the data effectively gathered by various sensors strategically placed on the robot, at different parts?

    Sergey Levine

    Yeah. Something I'll say about sensors is that you can actually get away with less than you'd think, and still do quite a lot. This platform has three cameras — one on each wrist, and a base camera. It doesn't have touch sensing, it doesn't have force sensing. It's very bare-bones, very low-cost. I'm sure more sensors could make it better, but a good learning method can compensate for deficient sensing fairly well. The wrist cameras are essentially a touch sensor in disguise, because you can see local deformations when you touch something.

    The data problem: bootstrapping a deployment flywheel, and a dexterity surprise

    Patrick O'Shaughnessy

    If I think about the analogy to the expert systems of the '80s and '90s in classical AI, and the lesson that scale is all you need, and the counterintuitive nature of that — you're not teaching it anything specific, just blasting it with data, and there's this reservoir of internet data — talk about that reservoir, and how you create the reservoir of data needed for this.

    Sergey Levine

    I don't think anybody really knows how much robot data is needed for truly generalisable, powerful embodied AI. But my sense is that we don't actually need to know. What we need is to get to the point where these systems are useful enough that they can go out into the world and gather more data themselves. To put it bluntly, Tesla doesn't worry about how much data its cars can collect — if anything, it's the opposite, there's almost too much data. So the key isn't so much to quantify the exact price tag of the ultimate robot data set. The key is to get a system that's useful enough, that does a wide variety of things, and that can keep pulling in more data.

    Patrick O'Shaughnessy

    You brought up Tesla — the beautiful thing about that system is it's useful without the AI to begin with, because a human drives it, and it gathers data anyway. Why not start with your best guess at something useful as a single robot, to get the same sort of flywheel going?

    Sergey Levine

    I think it's a good idea.

    Patrick O'Shaughnessy

    And do you think that's an approach you'll pursue?

    Sergey Levine

    I don't think there's one right answer. There are some domains where deploying a system under human control makes a lot of sense, and some where deploying a partially autonomous system is very reasonable. It's domain-dependent, because robots aren't just one thing. Maybe some people don't want a robot in their home that's constantly controlled by a person offsite, but for some applications that doesn't matter.

    Patrick O'Shaughnessy

    What has been the most surprising thing to you — if you mark the start of Physical Intelligence through today — that you've discovered, or about the nature of how the research has gone?

    Sergey Levine

    One surprise is that we've made a lot more progress on dexterity than I thought we would. I had an expectation, based on my prior work, that generalisation — handling all sorts of different scenes and objects — would just steadily improve as we collected more data, and we had good reason to believe that. But it was surprising that we could also get these systems to perform very dexterous behaviours without doing anything particularly special for it. The same applied to getting systems to work on different embodiments — we could get our models working on all sorts of other robots, including ones with multi-fingered hands, robots with different numbers of degrees of freedom. We needed data, and we needed to fine-tune the model, but the model itself didn't need to change — it didn't even need to be told, through any kind of prompt, what the robot was. I'd have thought we'd need some fancy techniques to adapt the system to faster, more dexterous, more complex tasks, and to different embodiments, but it generalises pretty well across those.

    Moravec's paradox, common sense, and semantic coaching

    Patrick O'Shaughnessy

    I'm always interested in the spectrum of capabilities — especially where today's systems are more advanced than people would expect, and where they're less advanced.

    Sergey Levine

    This has always been very tricky to understand in robotics. There's an idea roboticists always talk about, called Moravec's paradox. It's true across all of AI, but it's a especially big deal in robotics. We have a cognitive bias to think that things easy for us will be easy for the machine — solving calculus problems is difficult for most people, but picking up a cup is easy, so we think, 'Machines should be able to do this.' But it's actually the other way around: there are things that are easy for us because they have to be, or we wouldn't survive. We're very good at spotting the tiger in the jungle, because the people who weren't so good at it got eaten and aren't around anymore. Because of that bias, we think things should be very easy that are actually very difficult engineering challenges.

    However, machine learning slightly changes that equation. Programming something by hand to pick up any cup anywhere is difficult. Getting a machine-learning system to do it, if you have data for it, isn't that difficult. Increasingly, I think we'll see domains where collecting data is straightforward fall into the easy bucket over time, even if they're physically intricate. But there will be domains where collecting data is difficult, where you need more common sense, where you need to reason at multiple levels of abstraction, connecting physical skills learned in other areas to knowledge from the web — those will be tough, and that's where we'll need more technological advances.

    Patrick O'Shaughnessy

    What is the science of common sense? You mentioned it before — what does that mean?

    Sergey Levine

    For robotic learning, we can think of it as applying semantic inferences — using knowledge learned from other domains — to the current physical task at hand. You can think of common sense as the opposite of muscle memory. Muscle memory is: you play a sport, practise something a lot, and hardly think about it, you just do it on autopilot. Common sense, in my mind — I don't know if this is the conventional definition, but I think it's reasonable — is when you know something to be true because you saw it, read about it, or heard it, and now you're in a situation where that fact is highly pertinent to what you need to do, and you're able to make that connection, apply it, ground it in your environment, and make the right decision.

    Patrick O'Shaughnessy

    One of the other differences that's so interesting to me: people who've used chatbots query it, get an answer, query it, get an answer. Now we're seeing what happens with Claude Code and similar tools, where you give it something complicated, and there's this measure of how long it can go without failing. What's the equivalent long-range thing in robotics?

    Sergey Levine

    It's something we're working on quite a bit right now, and the methodology isn't that different, at some level. The way our models work now, as I mentioned, is they use this chain-of-thought process to reason about the task. When you have that, you can do very long-horizon tasks — a robot that takes all the dishes out of the dishwasher, puts them in the correct cabinets, wipes down the counter, and so on. The interesting thing is that we found, maybe about six months ago, that our models had gotten to the point where they could be improved just from supervising them with high-level instructions.

    What does that mean? You take a robot, put it in a new kitchen, ask it to clean the kitchen. It gets to work and fails somewhere. So what do you do? Traditionally, we'd add more teleoperation data to cover a wider range of kitchens. But we tried, on a whim, seeing what happens if we don't add more teleoperation data — if we just add more data labelled with the semantic command. Basically, take whatever the robot experienced and label it with some semantic commands, but don't add any more low-level actions. That actually helps — it improves the model's ability to generalise. What that means is that the bottleneck had shifted, from the lowest level — the robot's physical ability to do the task — to this middle level, where the system is now more bottlenecked by its ability to interpret the scene and select the correct next step, which can be supervised with language. That's a big deal, because it means someone can literally talk to the robot.

    Patrick O'Shaughnessy

    Coaching, basically.

    Sergey Levine

    Yeah, exactly. And make it better just by talking to it.

    Trust, imperfection, and where physical intelligence gets hardest

    Patrick O'Shaughnessy

    If we're in 2050, and there's no robot in my kitchen doing my dishes, what do you think the most likely explanation is for it not having gotten there by then?

    Sergey Levine

    My suspicion is that there's a long tail of challenges to do with the interaction of technology and people. In some ways autonomous cars aren't that different — getting to a level of comfort with deploying autonomous vehicles on the road was a significant challenge that ran in parallel with getting the technology itself to that level. Early Tesla self-driving was a bit controversial because it wasn't perfect, and there was a real question of whether people were comfortable with that level of imperfection. Probably there are some tasks where people will be comfortable with something imperfect that needs to learn from its mistakes, and some areas where people won't be.

    Patrick O'Shaughnessy

    Are you comfortable with occasionally breaking your dishes?

    Sergey Levine

    Maybe in a few years it will stop breaking dishes, but maybe in the meantime it's not quite there. Are you comfortable with a robot like that in a home with small children? Maybe not. And that's okay. Figuring out how those factors interact, and what that means for the timeline and for how these systems get better with experience, is a tricky question. It needs to be approached carefully, with a lot of sensitivity. There may be domains where it makes a lot more sense for these systems to be deployed and bootstrapped and collect more data, and other domains that require more care.

    Patrick O'Shaughnessy

    Could you imagine a purely technical explanation for why something might not work?

    Sergey Levine

    The place I'd see the biggest technical risk is dealing with the breadth of different situations. If we're talking about a well-defined but slightly chaotic environment — cleaning hotel rooms, or assisting human cooks in a restaurant — I have a good sense of how to get that under control. If you're imagining a robot going into a home, one challenge is that a lot of other unexpected things can happen, and you need a system that's very good at inferring what's going on and adapting to it, or reacting intelligently. We have a lot of ideas for how to approach that, but it's the hardest part of the problem, because when just about anything could happen, and you're controlling a physical device that affects the world around it, you really need to get things right, at least at some level, pretty much every time. It doesn't mean you always have to succeed, but it does mean you always have to do something sensible that people are okay with. There are a lot of good ideas for how to do that, but it's probably the most challenging part of the equation.

    The simplest model: generality as the design principle, and simulation versus real data

    Patrick O'Shaughnessy

    If I go back to the right model for the Physical Intelligence approach — help me make it as simple as possible. One version: build a whole variety of different form factors, to do a whole variety of tasks, mash all this data together, and experiment with how to make it better on evals. Is that the simplest way of doing it, or is there an even simpler way? I'd love to contrast it with other approaches you're interested in, that you're not doing, but that others are.

    Sergey Levine

    In my mind, the most important thing to get right is to make the system general — in particular, general with respect to how it can be improved. Hand-designed robotic controllers aren't very general with respect to improvement, because it takes a human engineer going in to improve them. A learning-based perception system is more general, because all it requires is human labellers to go in and label more data. A system that learns autonomously, from data it gathers through its own experience, is even more general, because you don't even need the human labellers. So the key is this generality, particularly with respect to improvement, and our decisions are largely centred around that. I don't know if the correct design for a robot is to have three cameras, or a touch sensor — I think we're pretty agnostic about that, and we'll try a lot of different choices. I'm not even sure, in the long run, whether it'll have a language model — maybe it'll have some other kind of model trained on very diverse data. But the key is this level of generality.

    Patrick O'Shaughnessy

    What other approaches are the most interesting to you?

    Sergey Levine

    One thing I think is a very important question in this area, that the research community and the tech community haven't fully answered, is the dichotomy between different data sources — particularly real data versus simulation. It's a very controversial topic. I have a strong opinion about it, but it's worth acknowledging that if you look at humanoids — the videos of humanoids doing acrobatics — the pipeline that makes that work is very heavily reliant on simulation, and very light on real-world data, often almost zero. Then there are the approaches that work well for robotic manipulation, which are often the opposite: very little simulated data, large amounts of real-world data, and very large foundation models. It's kind of surprising that in these two robotic domains, the dominant approaches look so different. Maybe one will win out, or maybe there's some synthesis of these ideas — I don't know the answer. I have my own subjective opinions; I think the approach we're taking is a good one. But it's interesting to look at why these things are so different.

    Cool versus useful: backflips, the robot Olympics, and where machines outperform us

    Patrick O'Shaughnessy

    Can you talk about the contrast between cool and useful? The Boston Dynamics robot is very cool — the backflip is super cool. I don't know what requires a robot to do a backflip. I'm curious how you think about optimising around cool versus useful.

    Sergey Levine

    I don't know if it's the right strategy, but the one we've taken is: subject to the constraint that it's useful, make it as cool as possible. That's reflected in our blog posts and videos — we make decisions first and foremost based on our assessment of what will drive the technology forward, towards this truly general, broadly applicable robotic foundation model. But in doing that, we try to stress-test it against the toughest challenges we can throw at it, and the toughest challenges often look cool. We didn't set out to build a robot that could make espresso, or fold laundry, but in the process of building these general systems, we figured these would be particularly challenging, particularly exciting things to try, to see how far we could push them.

    Patrick O'Shaughnessy

    Can you talk about the robot Olympics?

    Sergey Levine

    Yeah. There was a gentleman named Benjie Holson, who used to work at Everyday Robots, part of Alphabet before it dissolved, and he spends a lot of time thinking about tasks robots can do. He wrote an interesting blog post a while back, where he basically said: there was this robot Olympics held in China, where robots ran around a track and jumped, but maybe those aren't the real challenges we should worry about — how about a robot Olympics centred on everyday tasks that people do? That's more of a Moravec's-paradox thing — tasks people find really easy, but that robots struggle with. He had things like opening a door, washing a frying pan with grease on it, using a plastic bag to pick up dog poop — things people don't find particularly challenging, but that no current robotic system can do. He listed maybe a dozen of these, and we wanted to give it a shot. This wasn't part of a concerted research project — we'd developed processes and systems for ingesting new tasks generally, and we figured a good test was to take this big list of tasks and see if our process worked. It was almost a test of our internal operations and model-training system.

    We tried these things, and it turned out we could solve almost all of them. There was one we couldn't do — turning a dress shirt inside out, because the grippers on this thing wouldn't fit inside the sleeve; we'd probably need to change the gripper. And on a technicality, we didn't succeed at peeling an orange, because he specified doing it with the fingers, and our fingers weren't strong enough — we had to use a little tool, like a knife. But everything else, we could do. What was actually interesting to me — obviously it's cool, the videos are nice — but one thing to keep in mind is that we didn't develop anything special for this. We literally used it as a test of our task-onboarding process. I think that's interesting, because it suggests the power of generality — with a system this general, you can onboard all sorts of crazy tasks without doing anything particularly sophisticated.

    Patrick O'Shaughnessy

    I was curious before, when you mentioned superhuman ability on dexterity — where we're limited by what we can do, or by what we can control, even as it gets smaller. What are some other dimensions like that, where we might surpass human ability physically? What other trend lines are most interesting to you?

    Sergey Levine

    Here's a fun one. We were working on a task where our robot had to plug in cables — power cables, ethernet cables. When a person does this — if you practise a lot, you get good at it, but without much practice you pause frequently, because it's not just a physical thing, you have to process what's going on, make sure it's lined up. So you do it very slowly. If you're teleoperating a robot, you do it even more slowly, because there's a level of indirection. It turns out to be pretty straightforward to go in, find those pauses, and remove them — and speed things up further, so you can have a task where a person demonstrates what success looks like, and then have the robot practise the task and succeed the same way, but a lot more quickly and efficiently. The most general way to do this is with reinforcement learning, but there are also simple tricks if you just want speed. That's one example of something a machine can do a lot better. At some level, you have a processing bottleneck — that's why the person does it slowly, they have to process what's going on — but speeding up processing is something people understand quite well in computer science.

    Patrick O'Shaughnessy

    There's this amazing Michael Crichton novel called Prey. Is it a question of form factor — where, for a given problem, there may be an optimal shape for the robot to perform the task, and what you should do is analyse the problem and then have something that can almost morph or transform into that form factor? How do you think about innovating on form factor, rather than on the data and model side?

    Sergey Levine

    In general, in robotics, the ability to innovate on form factors has been very constrained by the AI challenge. If you have a traditional AI pipeline — doing motion planning and so on — it's hard to just cobble together a new robot, because you have to characterise the dynamics of the system, do system ID, build up all that infrastructure. But if you could put together a robot in your garage, load a robotic foundation model onto it, and tell it to do a bunch of stuff — maybe it won't be perfect, maybe it needs more data to really perfect it, but you can at least get the thing moving — that could be a really powerful engine to get everybody experimenting with this. I don't think I'm the right person to design the perfect robot; there are people here who are a lot better at that. But in general, I think it's just like personal computers — the key is to let people experiment and play around with it, and radically lower the barrier to entry. Then we'll see a lot more creativity. When people first started using personal computers, there was a limited number of form factors. Now you can have a computer in your phone, a computer in your car, a computer embedded in your fridge — they're everywhere, and very different. Generality, good software, a good foundation to build applications on — those are key to enabling that.

    Patrick O'Shaughnessy

    Professor Andy Clark once described to me the feeling of physical intelligence, for a human, as like learning to ride a bike — there's that moment when you didn't know how to do it, and then you do, and that feeling is physical intelligence, that snap of understanding.

    Sergey Levine

    There's actually a physiological explanation for this. There were studies done in monkeys using tools — you can find where in the brain the neurons activate for the monkey to figure out where its hand is. It turns out that if it's using a tool, they activate based on the location of the tool tip, not the location of the hand. The tool being an extension of your body is a real physiological thing — your brain literally does that.

    Patrick O'Shaughnessy

    Knowing that, what does it do to your approach to the research?

    Sergey Levine

    I think it says that physical intelligence should be, at some level, agnostic to embodiment — that a good foundation model should figure out how to manipulate whatever body it's controlling, whatever tools it has at hand. There's basically one problem, not many different problems. There isn't a humanoid problem, and a car problem, and a bulldozer problem, and a robot-bolted-to-a-table problem. There is one problem, and if you solve it at this full level of generality, that's really, really powerful.

    Unlocking creativity, the field's controversies, and the hardest tasks for common sense

    Patrick O'Shaughnessy

    We're in the early stages of seeing some of the job and other transformation in businesses and the economy that LLMs make possible — certainly we've seen it in engineering. How do you think about what might happen, or what you hope will happen, when we're at a similar stage for robotics — where all of a sudden we have something general, that's useful? Where do you expect to see the world start to change most, in the early days after Physical Intelligence succeeds?

    Sergey Levine

    That's a really interesting question. I really don't know — I don't think anybody would have been able to predict how the LLM stuff evolved, even if people had guessed at it. That's why I keep coming back to the idea that maybe the key is to let people try lots of things. One of the really amazing things about LLM applications is that they're really accessible — somebody can put together a really cool new prototype that, under the hood, is just prompting ChatGPT or something, but they can experiment with it, try it out, see what it does. There's an amazing power in having lots of smart people rapidly iterating and prototyping lots of things. That's a lot of why Physical Intelligence has put a premium on engagement — we've open-sourced our models, and we'd like to engage with lots of other companies building robots, because we all see a lot of power in that effect of having many people try lots of things.

    Patrick O'Shaughnessy

    What are the major controversies in the robotics community?

    Sergey Levine

    Obviously I'm an academic, so to me a controversy is someone getting into an argument with me at a conference. But the kind of arguments I've found myself in have had an interesting trajectory. In the early days, the main argument I'd have with people was: does learning have a place in robotic AI? Part of why that was controversial is that, in a traditional engineering pipeline, robots look very different from software artefacts — they're physical, they can affect things around them, there are safety considerations, and they can get into a lot of weird situations. It took a really long time for the robotics research community to internalise that you don't necessarily need to program in things like knowledge of physics — you don't need a physics simulator inside your robot when it's planning, a learning system can figure all that stuff out. That was very controversial for a long time. At this point there's a lot of acceptance that learning is a really important part of robotics, but I don't think there's universal acceptance that end-to-end learning is the right way to go — that there's universal acceptance of the bitter lesson. The bitter lesson says you shouldn't program the machine to think the way you think it should think, but let it learn from data. That's not a universally accepted idea. There are good arguments against it, but I think that in the long run, if we want that generality — especially generality in the machine's ability to improve — it needs to primarily be learning from data.

    Patrick O'Shaughnessy

    What is the good argument against it?

    Sergey Levine

    My best attempt at steel-manning it is: if you want something reliable in a really complicated, open-world setting, you can't afford not to use what you already know about the physical world — we've got textbooks full of this stuff. Why not just plug in what we know from the textbooks?

    Patrick O'Shaughnessy

    What is compositional learning? Can you describe that?

    Sergey Levine

    There's an example that's a vivid way to communicate it — it's due to one of my students. He asked a language model to provide a recipe for making a sandwich in the International Phonetic Alphabet. The IPA is the set of symbols used in a dictionary to explain how to pronounce a word, and it's peculiar because it only ever appears for individual words in a dictionary — you never see free-form text written in IPA. But if you ask a good language model, it will write paragraphs in IPA for you. That's compositional generalisation — you've never seen this particular alphabet used to write paragraphs, but you understand paragraphs, and you understand that it's compositional with different alphabets, so you can solve the problem. You can imagine the same thing coming up in robotics — you've learned a repertoire of skills, and now you can combine and mix those skills and apply them to solve new problems.

    Patrick O'Shaughnessy

    It makes me wonder — what's the last type of task you think will be possible for a robotic system to achieve?

    Sergey Levine

    I think changing a child's diaper will be really, really hard.

    Patrick O'Shaughnessy

    Say more.

    Sergey Levine

    I think this really is more of a Moravec's-paradox thing all over again. People are extremely good at certain things — we're very good at physical things, and we're also very good at interacting with other people, and that makes sense, we have to be, it's a lot of our social systems. So behaviours that involve interacting with other people — where you have to actually help somebody, help them get out of bed or something like that — I think that's a lot harder than people appreciate. I think elderly care, taking care of small children — those things are going to be hard, probably harder than people think.

    Patrick O'Shaughnessy

    And the stakes are very high. I want my baby to be last.

    Sergey Levine

    Not just that — the stakes are high in many places. It's just that this is probably the pinnacle of something that fools us into thinking it's easier than it really is, because we're so evolved for interacting with people and doing things physically. If you're helping somebody get up the stairs, or out of bed, you don't have to think very carefully about how you're going to do that — you just kind of know. So I think it's really the pinnacle of Moravec's paradox.

    Patrick O'Shaughnessy

    If I think about an LLM as a brain that's effectively studied everything — I don't know how else to put it — and then a robotics model's brain instead: what are the dark parts of the brain? What hasn't it been able to study, or penetrate, or learn — what are the areas that matter, but have been really difficult to get into?

    Sergey Levine

    One thing people are remarkably good at is using physical analogies to understand other situations. I don't know whether LLMs can do this, but it's something people use a lot, in everyday life and for very sophisticated problems too. You could say, 'That company has a lot of momentum' — that's a physical analogy, and you know exactly what it means, I don't have to explain it. But if you actually think about it, it's quite a complex thing — there's a lot riding on that word, momentum. There's an interview with Richard Feynman where he talks about the analogies he makes for subatomic particles — he says we use the word spin, but the thing isn't really spinning, it's not like a spinning top. But those analogies really help us make sense of it, and not just by explaining concepts — they actually lead to conclusions, to inferences that make sense. It's kind of remarkable that we're so primed to interact with the physical world, so primed to have physical intelligence, that we can use it in everyday speech — that company has momentum — and use it when advancing fundamental theoretical physics. I don't know if LLMs can do that. Maybe they can, but I think really understanding physical interactions, causal structures, all that kind of stuff — there's something special about it, and it's clearly something people get a lot of mileage out of.

    Research culture, manufacturing at scale, and how this changes physical labour

    Patrick O'Shaughnessy

    I'd love to talk about the role of researchers, the actual people doing the research. In LLM world, it's fairly shocking how few people are, at a global scale, responsible for basically all the progress — someone like Ilya is an example. What's that like in robotics? How many people in the world are truly impacting this trajectory? And I want to ask about what good research means.

    Sergey Levine

    Those kinds of questions are often very hard to answer about science, because we have a tendency, especially when we look at history, to underline particular milestones. Certainly in machine learning this happens — you can say AlexNet was a big step forward, this was a big step forward, and that's true. But it's also important to remember that these advances happen because lots of people are trying lots of things, and even some of the failures are very instructive. I complained earlier, in a low-key way, about the controversy around end-to-end robotic learning, but I don't know if robotic learning would have advanced the same way without that controversy, so to speak. You can look through the list of successes and note that certain people have a history of repeatedly hitting home runs, but in reality, in the scientific community, it's not just the home runs responsible for progress — even some of the failures, even some of the bad ideas, are very instructive in pushing towards the good ideas.

    Patrick O'Shaughnessy

    It's fascinating — the example you gave earlier, where the research insight was just, give it some coaching and it gets better — that sort of insight seems like it can be very powerful and high-leverage, which makes me wonder: what have you learned about what makes a great researcher?

    Sergey Levine

    Research is definitely different from engineering, because in research the important thing is getting to an answer to a question, which often requires cutting some corners. One of the most delicate decisions in research is when to try new things, versus when to stick with what you're already trying. That's very hard to figure out, and if you get it wrong, you can miss something remarkable. If you don't stick with something long enough, you might be right there, about to get to the answer, and then you stop just short of it — that's terrible. Or you could get stuck hammering against something that's never going to give way, for years. Deciding when to turn and look this way and that, to open yourself up to more opportunities, versus when to keep hammering because you're about to get the solution — that's often the most important decision. Some people have an instinct for getting that right, and that counts for a lot.

    Patrick O'Shaughnessy

    You've obviously been in and around, and are one of, the great researchers. What are these people like, as people? How are they distinctive from the average person?

    Sergey Levine

    I think they're just the same. Thinking about the people I deeply respect, who are really good at this, I have a very hard time thinking of a single set of personality traits. The one constant is that there's no constant. There might be a commonality in that, to do effective science, you have to be very passionate about it — but even that passion can come from many different places. I've worked with people who were remarkably effective, driven purely by a desire for novelty — they don't care what their technology does, or whether it's useful, they just want cool new ideas. I've also worked with people who just really want to solve a particular thing, and are just as happy building stuff, testing experiments, hammering away at things, whatever it takes. All those types can be very effective.

    Patrick O'Shaughnessy

    You mentioned the distinction between research and engineering, which also makes me think of manufacturing. Elon likes to say the factory is the product — the hardest part of the whole equation is the scale-up, making a hundred million of whatever this thing ends up looking like. How do you think about that part of the equation, or is it too remote a question at this stage?

    Sergey Levine

    No, I think it's an important part of the equation. I'm not sure it's the part we most need to figure out right now, but it's certainly part of it. As you might have guessed from my other answers, a lot of how I prefer to think about this is figuring out the hard part first, and then enabling a lot of experimentation on the other parts. Yes, making a robot at scale is difficult — it's even more difficult if you don't know what software is going to run it afterwards, and you're not even sure it's the right kind of robot. So one of the really valuable things we can get out of general-purpose AI tools like robotic foundation models is the ability to get a lot of that other stuff figured out, so that some of the uncertainty goes away — so that when you do scale up, you have some confidence it's really going to work.

    Patrick O'Shaughnessy

    A lot of people who listen to this are entrepreneurs, people who run companies. A popular question has become how a traditional company should begin thinking about using LLMs, or preparing for the ongoing improvement of these models. How would you answer the same question for robotics?

    Sergey Levine

    It's a very good question, and a very difficult one, because the technology is changing so rapidly. Let me illustrate why with an example. Here's a particular uncertainty about the tech: will robots rely more on demonstrations, or on reinforcement learning from autonomous data? We're working on both, and they're clearly both important, but how somebody should prepare will be pretty different depending on which wins out — lots of teleoperation to produce lots of demonstrations and a little autonomous experience, versus the opposite, a tiny number of demonstrations and huge amounts of autonomous experience. Is it 90/10, or 10/90? That's something we're hopefully going to learn over the next few years, but it changes the correct approach pretty dramatically. That's a case study in how changes in the technology will dramatically alter this.

    Patrick O'Shaughnessy

    From a business standpoint, is the right way to think about it just to get really clear on the economics of labour in your business? I'm curious how you think about the way this will change the nature of labour itself.

    Sergey Levine

    I think coding tools are a really nice template to look at for how this might work. It's not that coding tools came onto the scene and suddenly we don't need software engineers anymore — it's that the tools increase the productivity of individual software engineers. There's some work needed to make sure people can use them, and some technology development needed to make them useful for the appropriate use case, and these things are co-evolving — coding agents are different from code-completion tools, and so on. But it's a nice template for us to look at, to see how AI tools combine with people doing a job, increase their productivity, and also raise new challenges. I think we'll see something like that with robotics too — a more realistic template isn't that the humanoid just goes in and replaces people; it'll be more that some aspects of a job can be done by a robot, some can be done with a robot working together with a person, some where the person needs to do something special to make the robot more productive, and some where it's the other way around — the robot does something that makes the human more productive. It'll be this kind of dance, like we've seen with coding tools.

    Closing: uncertainty, inspiration, and the kindest thing

    Patrick O'Shaughnessy

    Do you have a favourite robot that isn't part of what Physical Intelligence is doing?

    Sergey Levine

    I really like the Boston Dynamics robot — the new version of Atlas especially, because it's in some ways very human-like, and in some ways very not human-like. They made some interesting decisions about wanting more range of motion in the joints, so it can do some pretty cool things. It's also a very agile robot, which is really cool — it makes those awesome demos. I'm a big fan of that, and generally a big fan of everything Boston Dynamics has done.

    Patrick O'Shaughnessy

    Should or could anything be read into the fact that Boston Dynamics has been doing very cool demos for a long time, and doesn't actually do anything useful for customers?

    Sergey Levine

    It's a fair question — a fair one for lots of robotics companies, to be fair. What I'll say, in general terms, is that there's a lot of value in demos that illustrate the challenges on the road to something useful and productive. Obviously you can also do a demo without being on that road, but I think there's value in demos used correctly, in service of a mission — they give people an illustration of what to expect, and they provide a challenge. You just have to be honest in setting up that challenge.

    Patrick O'Shaughnessy

    How much do you think about the business end points? Roomba is the best-selling robot of all time in the consumer category, which is kind of surprising, and we might be on the edge of some sort of Cambrian explosion. How much of your cycles do you spend thinking about the shape of a product that might result from this — maybe the way we bootstrap our way to all this data?

    Sergey Levine

    I certainly spend some time thinking about it. I think it's just hard to reduce to a very concrete answer right now, but it's not too bad to think about a space of possibilities. A lot of what we're doing when we develop our models, experiment with different tasks, do demos like the robot Olympics — underneath, we're prototyping what it looks like to do something real with this, to different degrees of real, and what goes wrong. So it's something we think about a lot. It's not something I have a concrete answer to, but there's a space of possibilities, and a lot of what we're actually planning to do in 2026 is experiment with different things in that space.

    Patrick O'Shaughnessy

    When you study the history of general-purpose technologies — and this would certainly be a major one, if it comes to fruition — you often find a constellation of things happening around it that enable it. LLMs are obviously a direct complement to what you're doing. Are there any other surprising technology areas or trends that help you do what you do, but are different?

    Sergey Levine

    One interesting thing is that robotics hardware has become dramatically more affordable over the last few years. When I started working in robotics about a decade ago, I worked with a robot called a PR2, which cost about $400,000. When I started my lab at UC Berkeley, I used a robot in the ballpark of $30,000. Now, each arm on this thing is maybe a tenth of that — we think it can go lower still. That's not down to any one single technology — it involves both hardware and software. The low-cost arms we have here wouldn't be useful in an industrial setting, because traditional control methods, which rely on a great deal of precision, wouldn't be able to use them. There's a whole constellation of things that have pushed down the price point, and I think that makes it a lot more practical to think about general-purpose robotics today.

    Patrick O'Shaughnessy

    For people who want to follow major milestones in this field fairly closely, where does that information show up?

    Sergey Levine

    A lot of it shows up in research papers. Unfortunately, research papers aren't a very accessible source of information — it takes some care to sort through everything and figure out what the signal is, because research results are intended for an audience that already understands the starting point from past research. But that's a big one. I think robotics, and technology in general, is one of those things where the public-facing artefacts — the demos and videos somebody posts on social media — are often not very good for giving a sense of the true underlying state of things, because they're meant more as a demonstration at the edge of capability. Grounding what a demo really means requires digging deeper. So, probably, research papers are the way to go. Sometimes, even worse than that, you have to go and talk to the individual people, and find out what the inside story really is. Maybe that's not a great situation to be in, but that's kind of how science works.

    Patrick O'Shaughnessy

    As you look forward to the future of your mission, what feels the most uncertain?

    Sergey Levine

    I do think the timeline is uncertain. If anything, my sense of the timeline has gotten more optimistic since we started, but it's uncertain because of the nature of the technology — there's a bootstrap challenge, getting to a particular level of usefulness so robots can be deployed to do useful tasks, so they can start collecting data from open-world settings at scale. Because that's such a sudden kind of event — getting past the activation energy — there's a lot of uncertainty about the timing. That's exacerbated by the fact that the timeline looks different depending on what kind of technology gets deployed — like the example I gave earlier, about whether it's data collection through teleoperation, or with autonomous systems, or something in between, like shared autonomy, or this coaching approach. Those all change the picture, in terms of how deployments work and how in-the-wild data collection works. Because of that, there's quite a bit of uncertainty about the timing.

    Patrick O'Shaughnessy

    You're in such an interesting position — at the centre of the research, with lots of different people talking to you, asking you questions. What are questions you're surprised people don't ask you? What should people be asking that they don't?

    Sergey Levine

    I think the question you asked earlier, about how somebody should prepare — there's a variant of that, which would be something like: if I want to start using autonomous robots for something, what should I start setting up? Should I set up teleoperation? Should I modify my task so it's more accessible? Should I design new hardware, so I can plug your software into it? I think people make a lot of assumptions about that. One assumption is: well, machine learning requires data, so let me figure out something that will collect data. That's often not the best assumption, because you need the right kind of data. Maybe some data is easy — it's easy to get videos of people doing something, but that doesn't mean it's the right kind of data. It might be domain-dependent, it might depend on your thesis about which technology will succeed. So people make a lot of assumptions — not that I necessarily have a better answer for them, even if they ask me, but it's a big space of possibilities.

    Patrick O'Shaughnessy

    We've talked about these big, uncertain, long-term timelines. What's the very next thing you're trying to solve that's extremely visible?

    Sergey Levine

    Without giving too much away, a big focus for us right now is better understanding this middle-level reasoning part of the problem. We have a pretty good sense of how to acquire low-level physical behaviours, but getting those behaviours to generalise requires bringing to bear a lot of common-sense knowledge, and the representation of that might be really important. LLMs make certain kinds of representations very convenient — they make it very convenient to turn text into other text. But that's not necessarily the best representation for what an embodied system needs to do — sometimes it needs to think more spatially, sometimes semantically, sometimes in other representations. Figuring out exactly how to structure that internal thinking process might be a very important question, and the answer might be different in the world of embodied foundation models than in the world of LLMs. That's a concrete thing we're working on now.

    Patrick O'Shaughnessy

    Where do you think you fall — if I could get the hundred most informed and active robotics researchers in a room, and poll them on how certain they are that this will have unlimited capabilities, and how soon that might happen — where do you fall in that distribution?

    Sergey Levine

    I'm on the optimistic end when it comes to established robotics researchers, and on the pessimistic end relative to robotics entrepreneurs.

    Patrick O'Shaughnessy

    Interesting — I understand the entrepreneur part, for sure, they're optimistic by nature. Why are you on the optimistic end of the researcher community?

    Sergey Levine

    Robotics has a very long history, with precious few successes when it comes to robotic AI. If we're being honest, most robots out there doing useful work are still running hand-designed control. That's because the robotics problem is hard — maybe not our fault, it's just a difficult problem. Because of that, there's good reason for caution — maybe we've made a lot of headway on this part of the problem, but there are many other problems that remain. Part of why I'm optimistic is that I have a sense of what's proven tough for me before, and I can see a lot of the puzzle pieces that could be slotted in to address many of those things. But, as my co-founder Karol likes to say, when you've climbed the mountain, only then do you see if there's another mountain after it. In robotics, there's been a lot of experience of lots of mountains.

    Patrick O'Shaughnessy

    Given that endurance is required, who or what most inspires you?

    Sergey Levine

    I'm actually quite inspired by Boston Dynamics. There's a lot we can debate on the technology side, but there's real value in repeatedly showing something people wouldn't have thought possible, even with all sorts of caveats and assumptions. Certainly in robotics, whatever we might say about demos, it's fair to say people have revised their thinking about what's possible from seeing some of that stuff. I'm also inspired by organisations that create an atmosphere for experimentation. Some research labs have done a very good job of this — I think OpenAI has historically done a great job of creating an atmosphere where individual researchers can experiment with things and be empowered to see them through. ChatGPT was, basically, John Schulman's pet experiment for a while — it wasn't a concerted corporate strategy with spreadsheets and pie charts, it was a pet project. There's something pretty inspiring about organisations that empower people to turn pet projects into world-changing successes. Certainly one of the aspirations my co-founders and I have here at Physical Intelligence is to provide some of that, to the best of our ability. It's hard to do — very hard to build an organisation with that kind of capability.

    Patrick O'Shaughnessy

    I feel like Google used to have that 'one day you can do whatever you want' thing. Is that the spirit of it?

    Sergey Levine

    I was absolutely shocked when I started working at Google, at the level of leverage I felt I could have. One of the projects I did, with many of my colleagues, in 2015, was colloquially called the arm farm. We took a couple of dozen robots, put them in a lab, and had them collect data. That was a very bottom-up thing — I found out from somebody that they had a warehouse full of robots nobody was using, and I asked Jeff Dean and Vincent Vanhoucke if we could stick them in a lab. I was thinking, okay, they're not going to take me seriously, I'd just started, I was a level-four research scientist. And Jeff said, yeah, let's do it — what do you need? I remember feeling, wow, I'd never in my life thought I'd have that kind of leverage. Obviously I was very young at the time. I think that's very special, and I think getting to a place where people can unlock their creativity, and have that kind of agency, can make for a very remarkable place.

    Patrick O'Shaughnessy

    My friend Jesse has this great question — for companies you're not involved with, which one do you most hope succeeds, and why? People used to say Boom a lot, because they want to fly places faster. Increasingly, as I've asked this question, people have said Bye, just because of the sheer impact it might have, if successful, on such a global scale. It's been really fun to hear about all the ins and outs of how you're thinking about this problem and attacking it. When I do these interviews, I have the same traditional last question for everyone: what is the kindest thing anyone's ever done for you?

    Sergey Levine

    It's a tough question to answer — there are many moments in my career where I felt I got a leg up on something. I have the kind of personality where I sometimes don't appreciate things in the moment, and only reflect on them afterwards. I don't have a single answer, but probably three moments stand out. One I've already mentioned — the arm farm — and I'm especially grateful to Jeff and Vincent for being willing to take that bet on me and my colleagues. There are a couple of other moments — certainly when I started my postdoc with Pieter Abbeel at Berkeley, I had zero robotics experience, I'd done virtual character animation and computer graphics, and I felt that was a bet on my potential more than my actual accomplishments. And there was another moment, even earlier, maybe even more minor — when I was in college, I got an internship at Nvidia that let me experience some cool stuff when I was just a sophomore, and I think the hiring manager for that also took a bet on me. I think these things really matter in a person's career, and maybe at the time I should have been more grateful, but certainly in hindsight they made a big difference. Hopefully I can make that difference in other people's careers too.

    Patrick O'Shaughnessy

    Well, I've learned so much from you and your co-founders, and so much today. Thank you so much for your time.

    Sergey Levine

    Thank you.