Terence Tao on the Hardest Problems in Mathematics, Physics and the Future of AI

Lex Fridman Podcast

Episode →

Reformatted for readability — timestamps removed, lightly restructured. Not verbatim.

Contents

    Transcript: Terence Tao - Hardest Problems in Mathematics, Physics & the Future of AI

    Introduction

    First Hard Problem

    Lex Fridman

    What was the first really difficult research-level math problem that you encountered, one that gave you pause maybe?

    Terence Tao

    Well, in your undergraduate education you learn about the really hard impossible problems like the Riemann Hypothesis, the Twin-Primes Conjecture. You can make problems arbitrarily difficult. That's not really a problem. In fact, there's even problems that we know to be unsolvable. What's really interesting are the problems just on the boundary between what we can do rather easily and what are hopeless, but what are problems where existing techniques can do 90% of the job and then you just need that remaining 10%. I think as a PhD student, the Kakeya Problem certainly caught my eye. And it just got solved actually. It's a problem I've worked on a lot in my early research. Historically, it came from a little puzzle by the Japanese mathematician Soichi Kakeya in 1918 or so. So, the puzzle is that you have a needle on the plane or think like driving on a road something, and you want it to execute a U-turn, you want to turn the needle around, but you want to do it in as little space as possible. So, you want to use this little area in order to turn it around, but the needle is infinitely maneuverable. So, you can imagine just spinning it around. As the unit needle, you can spin it around its center, and I think that gives you a disc of area, I think pi over four. Or you can do a three-point U-turn, which is what we teach people in their driving schools to do. And that actually takes area of pi over eight, so it's a little bit more efficient than a rotation. And so for a while people thought that was the most efficient way to turn things around, but Besicovitch showed that in fact you could actually turn the needle around using as little area as you wanted. So, 0.01, there was some really fancy multi back and forth U-turn thing that you could do that you could turn a needle around and in so doing it would pass through every intermediate direction.

    Lex Fridman

    This in the two-dimensional plane?

    Terence Tao

    This is in the two-dimensional plane. So, we understand everything in two dimensions. So, the next question is: what happens in three dimensions? So, suppose the Hubble space Telescope is tube in space, and you want to observe every single star in the universe, so you want to rotate the telescope to reach every single direction. And here's unrealistic part, suppose that space is at a premium, which totally is not, you want to occupy as little volume as possible in order to rotate your needle around, in order to see every single star in the sky. How small a volume do you need to do that? And so you can modify Besicovitch's construction. And so if your telescope has zero thickness, then you can use as little volume as you need. That's a simple modification of the two-dimensional construction. But the question is that if your telescope is not zero thickness, but just very, very thin, some thickness delta, what is the minimum volume needed to be able to see every single direction as a function of delta?

    So, as delta gets smaller, as the needle gets thinner, the volume should go down. But how fast does it go down? And the conjecture was that it goes down very, very slowly like logarithmically roughly speaking, and that was proved after a lot of work. So, this seems like a puzzle. Why is it interesting? So, it turns out to be surprisingly connected to a lot of problems in partial differential equations, in number theory, in geometry, combinatorics. For example, in wave propagation, you splash some water around, you create water waves and they travel in various directions, but waves exhibit both particle and wave-type behavior. So, you can have what's called a wave packet, which is a very localized wave that is localized in space and moving a certain direction in time. And so if you plot it in both space and time, it occupies a region which looks like a tube. What can happen is that you can have a wave which initially is very dispersed, but it all focuses at a single point later in time. You can imagine dropping a pebble into a pond and the ripples spread out, but then if you time-reverse that scenario, and the equations of wave motion are time-reversible, you can imagine ripples that are converging to a single point and then a big splash occurs, maybe even a singularity. And so it's possible to do that. And geometrically what's going on is that there's also light rays, so if this wave represents light, for example, you can imagine this wave as a superposition of photons all traveling at the speed of light.

    They all travel on these light rays and they're all focusing at this one point. So, you can have a very dispersed wave focus into a very concentrated wave at one point in space and time, but then it de-focuses again, it separates. But potentially if the conjecture had a negative solution, so what that meant is that there's a very efficient way to pack tubes pointing different directions to a very, very narrow region of a very narrow volume. Then you would also be able to create waves that start out some… There'll be some arrangement of waves that start out very, very dispersed, but they would concentrate, not just at a single point, but there'll be a lot of concentrations in space and time. And you could create what's called a blowup, where these waves amplitude becomes so great that the laws of physics that they're governed by are no longer wave equations, but something more complicated and nonlinear.

    Navier-Stokes Singularity

    Lex Fridman

    Can you speak to the Navier-Stokes? So, the existence of smoothness, like you said, Millennium Prize Problem, You've made a lot of progress on this one. In 2016, you published a paper, Finite Time Blowup For An Average Three-Dimensional Navier-Stokes Equation. So, we're trying to figure out if this thing… Usually it doesn't blow up, but can we say for sure it never blows up?

    Terence Tao

    Right, yeah. So yeah, that is literally the $1 million question. So, this is what distinguishes mathematicians from pretty much everybody else. If something holds 99.99% of the time, that's good enough for most things. But mathematicians are one of the few people who really care about whether really 100% of all situations are covered by it. So, most fluid, most of the time water does not blow up, but could you design a very special initial state that does this?

    Lex Fridman

    And maybe we should say that this is a set of equations that govern in the field of fluid dynamics, trying to understand how fluid behaves. And it's actually turns out to be a really… Fluid is extremely complicated thing to try to model.

    Terence Tao

    Yeah, so it has practical importance. So this Clay Prize problem concerns what's called the Incompressible Navier-Stokes, which governs things like water. There's something called the Compressible Navier-Stokes, which governs things like air, and that's particularly important for weather prediction. Weather prediction, it does a lot of computational fluid dynamics. A lot of it's actually just trying to solve the Navier-Stokes equations as best they can. Also gathering a lot of data, so that they can initialize the equation. There's a lot of moving parts, so it's very important from practically.

    Lex Fridman

    Why is it difficult to prove general things about the set of equations like it not not blowing up?

    Terence Tao

    Short answer is Maxwell's Demon. So, Maxwell's Demon is a concept in thermodynamics. If you have a box of two gases in oxygen and nitrogen, and maybe you start with all the oxygen on one side and nitrogen on the other side, but there's no barrier between them. Then they will mix and they should stay mixed. There's no reason why they should un-mix. But in principle, because of all the collisions between them, there could be some sort of weird conspiracy that maybe there's a microscopic demon called Maxwell's Demon that will… every time an oxygen and nitrogen atom collide, they'll bounce off in such a way that the oxygen sort of drifts onto one side and then nitrogen goes to the other. And you could have an extremely improbable configuration emerge, which we never see, and which statistically it's extremely unlikely, but mathematically it's possible that this can happen and we can't rule that out.

    And this is a situation that shows up a lot in mathematics. A basic example is the digits of pi 3.14159 and so forth. The digits look like they have no pattern, and we believe they have no pattern. On the long-term, you should see as many ones and twos and threes as fours and fives and sixes, there should be no preference in the digits of pi to favor, let's say seven over eight. But maybe there's some demon in the digits of pi that every time you compute more and more digits, it biases one digit to another. And this is a conspiracy that should not happen. There's no reason it should happen, but there's no way to prove it with our current technology. So, getting back to Navier-Stokes, a fluid has a certain amount of energy, and because the fluid is in motion, the energy gets transported around.

    And water is also viscous, so if the energy is spread out over many different locations, the natural viscosity of the fluid will just damp out the energy and will go to zero. And this is what happens when we actually experiment with water. You splash around, there's some turbulence and waves and so forth, but eventually it settles down and the lower the amplitude, the smaller velocity, the more calm it gets. But potentially there is some sort of demon that keeps pushing the energy of the fluid into a smaller and smaller scale, and it'll move faster and faster. And at faster speeds, the effect of viscosity is relatively less. And so it could happen that it creates some sort of what's called a self-similar blob scenario where the energy of the fluid starts off at some large scale and then it all sort of transfers energy into a smaller region of the fluid, which then at a much faster rate moves into an even smaller region and so forth.

    And each time it does this, it takes maybe half as long as the previous one, and then you could actually converge to all the energy concentrating in one point in a finite amount of time. And that's scenario is called finite time blowup. So, in practice, this doesn't happen. So, water is what's called turbulent. So, it is true that if you have a big eddy of water, it will tend to break up into smaller eddies, but it won't transfer all energy from one big eddy into one smaller eddy. It will transfer into maybe three or four, and then those ones split up into maybe three or four small eddies of their own. So the energy gets dispersed to the point where the viscosity can then keep everything under control. But if it can somehow concentrate all the energy, keep it all together, and do it fast enough that the viscous effects don't have enough time to calm everything down, then this blowup can occur.

    So, there were papers who had claimed that, "Oh, you just need to take into account conservation of energy and just carefully use the viscosity and you can keep everything under control for not just the Navier-Stokes, but for many, many types of equations like this." And so in the past there have been many attempts to try to obtain what's called global regularity for Navier-Stokes, which is the opposite of finite time blowup, that velocity stays smooth. And it all failed. There was always some sign error or some subtle mistake and it couldn't be salvaged.

    So, what I was interested in doing was trying to explain why we were not able to disprove finite time blowup. I couldn't do it for the actual equations of fluids, which are too complicated, but if I could average the equations of motion of Navier-Stokes, basically if I could turn off certain types of ways in which water interacts and only keep the ones that I want. So, in particular, if there's a fluid and it could transfer as energy from a large eddy into this small eddy or this other small eddy, I would turn off the energy channel that would transfer energy to this one and direct it only into this smaller eddy while still preserving the lower conservation energy.

    Lex Fridman

    So, you're trying to make a blowup?

    Terence Tao

    Yeah, yeah. So, I basically engineer a blowup by changing rules of physics, which is one thing that mathematicians are allowed to do. We can change the equation.

    Lex Fridman

    How does that help you get closer to the proof of something?

    Terence Tao

    Right. So, it provides what's called an obstruction in mathematics. So, what I did was that basically if I turned off the certain parts of the equation, which usually when you turn off certain interactions, make it less nonlinear, it makes it more regular and less likely to blow up. But I find that by turning off a very well-designed set of interactions, I could force all the energy to blow up in finite time. So, what that means is that if you wanted to prove the regularity for Navier-Stokes for the actual equation, you must use some feature of the true equation, which my artificial equation does not satisfy. So, it rules out certain approaches.

    So, the thing about math, it's not just about taking a technique that is going to work and applying it, but you need to not take the techniques that don't work. And for the problems that are really hard, often though are dozens of ways that you might think might apply to solve the problem, but it's only after a lot of experience that you realize there's no way that these methods are going to work. So, having these counterexamples for nearby problems rules out… it saves you a lot of time because you're not wasting energy on things that you now know cannot possibly ever work.

    Lex Fridman

    How deeply connected is it to that specific problem of fluid dynamics or is this some more general intuition you build up about mathematics?

    Terence Tao

    Right. Yeah. So, the key phenomenon that my technique exploits is what's called super-criticality. So, in partial differential equations, often these equations are like a tug of war between different forces. So, in Navier-Stokes, there's the dissipation force coming from viscosity, and it's very well understood. It's linear, it calms things down. If viscosity was all there was, then nothing bad would ever happen, but there's also transport that energy from… in one location of space can get transported because the fluid is in motion to other locations. And that's a nonlinear effect, and that causes all the problems. So, there are these two competing terms in the Navier-Stokes Equation, the dissipation term and the transport term. If the dissipation term dominates, if it's large, then basically you get regularity. And if the transport term dominates, then we don't know what's going on. It's a very nonlinear situation, it's unpredictable, it's turbulent.

    So, sometimes these forces are in balance at small scales but not in balance at large scales or vice versa. Navier-Stokes is what's called supercritical. So at smaller and smaller scales, the transport terms are much stronger than the viscosity terms. So, the viscosity terms are things that calm things down. And so this is why the problem is hard. In two dimensions, so the Soviet mathematician Ladyzhenskaya, she in the '60s shows in two dimensions there was no blowup. And in two dimensions, the Navier-Stokes Equation is what's called critical, the effect of transport and the effect of viscosity about the same strength even at very, very small scales. And we have a lot of technology to handle critical and also subcritical equations and prove regularity. But for supercritical equations, it was not clear what was going on, and I did a lot of work, and then there's been a lot of follow up showing that for many other types of supercritical equations, you can create all kinds of blowup examples.

    Once the nonlinear effects dominate the linear effects at small scales, you can have all kinds of bad things happen. So, this is sort of one of the main insights of this line of work is that super-criticality versus criticality and subcriticality, this makes a big difference. That's a key qualitative feature that distinguishes some equations for being sort of nice and predictable and… Like planetary motion, there's certain equations that you can predict for millions of years or thousands at least. Again, it's not really a problem, but there's a reason why we can't predict the weather past two weeks into the future because it's a supercritical equation. Lots of really strange things are going on at very fine scales.

    Lex Fridman

    So, whenever there is some huge source of nonlinearity, that can create a huge problem for predicting what's going to happen?

    Terence Tao

    Yeah. And if non-linearity is somehow more and more featured and interesting at small scales. There's many equations that are nonlinear, but in many equations you can approximate things by the bulk. So, for example, planetary motion, if you want to understand the orbit of the Moon or Mars or something, you don't really need the microstructure of the seismology of the Moon or exactly how the mass is distributed. Basically, you can almost approximate these planets by point masses, and it's just the aggregate behavior is important. But if you want to model a fluid, like the weather, you can't just say, "In Los Angeles the temperature is this, the wind speed is this." For supercritical equations, the fine scale information is really important.

    Lex Fridman

    If we can just linger on the Navier-Stokes Equations a little bit. So, you've suggested, maybe you can describe it, that one of the ways to ways solve it or to negatively resolve it would be to construct a kind of liquid computer, and then show that the halting problem from computation theory has consequences for fluid dynamics, so show it in that way. Can you describe this idea?

    Terence Tao

    Right, yeah. So, this came out of this work of constructing this average equation that blew up. So, as part of how I had to do this, so there's this naive way to do it, you just keep pushing. Every time you get one scale, you push it immediately to the next scale as fast as possible. This is sort of the naive way to force blowup. It turns out in five and higher dimensions, this works, but in three dimensions there was this funny phenomenon that I discovered, that if you change laws of physics, you just always keep trying to push the energy into smaller and smaller scales, what happens is that the energy starts getting spread out into many scales at once, so that you have energy at one scale. You're pushing it into the next scale, and then as soon as it enters that scale, you also push it to the next scale, but there's still some energy left over from the previous scale.

    You're trying to do everything at once, and this spreads out the energy too much. And then it turns out that it makes it vulnerable for viscosity to come in and actually just damp out everything. So, it turns out this direct abortion doesn't actually work. There was a separate paper by some other authors that actually showed this in three dimensions. So, what I needed was to program a delay, so kind of like airlocks. So, I needed an equation which would start with a fluid doing something at one scale, it would push this energy into the next scale, but it would stay there until all the energy from the larger scale got transferred. And only after you pushed all the energy in, then you open the next gate and then you push that in as well.

    So, by doing that, the energy inches forward, scale by scale in such a way that it's always localized at one scale at a time, and then it can resist the effects of viscosity because it's not dispersed. So, in order to make that happen, I had to construct a rather complicated nonlinearity. And it was basically… It was constructed like an electronic circuit. So, I actually thank my wife for this because she was trained as an electrical engineer, and she talked about she had to design circuits and so forth. And if you want a circuit that does a certain thing, maybe have a light that flashes on and then turns off and then on and off. You can build it from more primitive components, capacitors and resistors and so forth, and you have to build a diagram.

    And these diagrams, you can sort of follow up your eyeballs and say, "Oh yeah, the current will build up here and it will stop, and then it will do that." So, I knew how to build analog of basic electronic components, like resistors and capacitors and so forth. And I would stack them together in such a way that I would create something that would open one gate. And then there'd be a clock, and then once the clock hits a certain threshold, it would close it. It would become a Rube Goldberg type machine, but described mathematically. And this ended up working. So, what I realized is that if you could pull the same thing off for the actual equations, so if the equations of water support a computation… So, you can imagine a steampunk, but it's really water-punk type of thing where… So, modern computers are electronic, they're powered by electrons passing through very tiny wires and interacting with other electrons and so forth.

    But instead of electrons, you can imagine these pulses of water moving a certain velocity. And maybe there are two different configurations corresponding to a bit being up or down. Probably that if you had two of these moving bodies of water collide, they would come out with some new configuration, which would be something like an AND gate or OR gate, that the output would depend in a very predictable way on the inputs. And you could chain these together and maybe create a Turing machine. And then you have computers which are made completely out of water. And if you have computers, then maybe you can do robotics, so hydraulics and so forth. And so you could create some machine which is basically a fluid analog, what's called a von Neumann machine.

    So, von Neumann proposed if you want to colonize Mars, the sheer cost of transporting people in machines to Mars is just ridiculous, but if you could transport one machine to Mars, and this machine had the ability to mine the planet, create some more materials, smelt them and build more copies of the same machine, then you could colonize a whole planet over time. So, if you could build a fluid machine, which yeah, so it's a fluid robot. And what it would do, its purpose in life, it's programmed so that it would create a smaller version of itself in some sort of cold state. It wouldn't start just yet. Once it's ready, the big robot configuration of water would transfer all its energy into the smaller configuration and then power down. And then they clean itself up, and then what's left is this newest state which would then turn on and do the same thing, but smaller and faster.

    And then the equation has a certain scaling symmetry. Once you do that, it can just keep iterating. So, this, in principle, would create a blowup for the actual Navier-Stokes. And this is what I managed to accomplish for this average Navier-Stokes. So, it provided this sort of roadmap to solve the problem. Now, this is a pipe dream because there are so many things that are missing for this to actually be a reality. So, I can't create these basic logic gates. I don't have these special configurations of water. There's candidates, these include vortex rings that might possibly work. But also analog computing is really nasty compared to digital computing because there's always errors. You have to do a lot of error correction along the way.

    I don't know how to completely power down the big machine, so it doesn't interfere the writing of the smaller machine, but everything in principle can happen. It doesn't contradict any of the laws of physics, so it's sort of evidence that this thing is possible. There are other groups who are now pursuing ways to make Navier-Stokes blow up, which are nowhere near as ridiculously complicated as this. They actually are pursuing much closer to the direct self-similar model, which can… It doesn't quite work as is, but there could be some simpler scheme they want to just describe to make this work.

    Lex Fridman

    There is a real leap of genius here to go from Navier-Stokes to this Turing machine. So, it goes from what the self-similar blob scenario that you're trying to get the smaller and smaller blob to now having a liquid Turing machine gets smaller and smaller and smaller, and somehow seeing how that could be used to say something about a blowup. That's a big leap.

    Game of Life

    Terence Tao

    So, there's precedent. So, the thing about mathematics is that it's really good at spotting connections between what you might think of as completely different problems, but if the mathematical form is the same, you can draw a connection. So, there's a lot of previously on what called cellular automata, the most famous of which is Conway's Game of Life. There's this infinite discrete grid, and at any given time, the grid is either occupied by a cell or it's empty. And there's a very simple rule that tells you how these cells evolve. So, sometimes cells live and sometimes they die. And when I was a student, it was a very popular screen saver to actually just have these animations go on, and they look very chaotic. In fact, they look a little bit like turbulent flow sometimes, but at some point people discovered more and more interesting structures within this Game of Life. So, for example, they discovered this thing called glider.

    So, a glider is a very tiny configuration of four or five selves which evolves and it just moves at a certain direction. And that's like this vortex rings. Yeah, so this is an analogy, the Game of Life is a discrete equation, and the fluid Navier-Stokes is a continuous equation, but mathematically they have some similar features. And so over time people discovered more and more interesting things that you could build within the Game of Life. The Game of Life is a very simple system. It only has like three or four rules to do it, but you can design all kinds of interesting configurations inside it. There's some called a glider gun that does nothing that spit out gliders one at a time. And then after a lot of effort, people managed to create AND gates and OR gates for gliders.

    There's this massive ridiculous structure, which if you have a stream of gliders coming in here and a stream of gliders coming in here, then you may produce extreme gliders coming out. Maybe if both of the streams have gliders, then there'll be an output stream, but if only one of them does, then nothing comes out. So, they could build something like that. And once you could build these basic gates, then just from software engineering, you can build almost anything. You can build a Turing machine. It's enormous steampunk type things. They look ridiculous. But then people also generated self-replicating objects in the Game of Life, a massive machine, a machine, which over a huge period of time and always look like glider guns inside doing these very steampunk calculations. It would create another version of itself which could replicate.

    Lex Fridman

    That's so incredible.

    Terence Tao

    A lot of this was like community crowdsourced by amateur mathematicians actually. So, I knew about that work. And so that is part of what inspired me to propose the same thing with Navier-Stokes. Seriously, analog is much worse than digital. It's going to be… You can't just directly take deconstructions in the Game of Life and plunk them in. But again, it shows it's possible.

    Lex Fridman

    There's a kind of emergence that happens with these cellular automata local rules… maybe it's similar to fluids, I don't know, but local rules operating at scale can create these incredibly complex dynamic structures. Do you think any of that is amenable to mathematical analysis? Do we have the tools to say something profound about that?

    Terence Tao

    The thing is, you can get these emergent very complicated structures, but only with very carefully prepared initial conditions. So, these glider guns and gates and self-propelled machines, if you just plunk on randomly some cells and you unlink them, you will not see any of these. And that's the analogous situation with Navier-Stokes again, that with typical initial conditions, you will not have any of this weird computation going on. But basically through engineering, by specially designing things in a very special way, you can make clever constructions.

    Lex Fridman

    I wonder if it's possible to prove the negative of… basically prove that only through engineering can you ever create something interesting.

    Terence Tao

    Yeah. This is a recurring challenge in mathematics that I call the dichotomy between structure and randomness, that most objects that you can generate in mathematics are random. They look like random, like the digital supply, well, we believe is a good example. But there's a very small number of things that have patterns. But now, you can prove something has a pattern by just constructing… If something has a simple pattern and you have a proof that it does something like repeat itself every so often, you can do that and you can prove that… For example, you can prove that most sequences of digits have no pattern. So, if you just pick digits randomly, there's something called low-large numbers. It tells you you're going to get as many ones as twos in the long run. But we have a lot fewer tools to…

    If I give you a specific pattern like the digits of pi, how can I show that this doesn't have some weird pattern to it? Some other work that I spent a lot of time on is to prove what are called structure theorems or inverse theorems that give tests for when something is very structured. So, some functions are what's called additive. If you have a function of natural numbers of the natural numbers, so maybe two maps to four, three maps to six and so forth, some functions are what's called additive, which means that if you add two inputs together, the output gets added as well. For example, a multiply by constant. If you multiply a number by 10… If you multiply A plus B by 10, that's the same as multiplying A by 10 and B by 10, and then adding them together. So, some functions are additive, some functions are kind of additive but not completely additive.

    So, for example, if I take a number, and I multiply by the square of two and I take the integer part of that, so 10 by square route of two is like 14 point something, so 10 up to 14, 20 or up to 28. So, in that case, additivity is true then, so 10 plus 10 is 20 and 14 plus 14 is 28. But because of this rounding, sometimes there's round-up errors, and sometimes when you add A plus A, this function doesn't quite give you the sum of the two individual outputs, but the sum plus/minus one. So, it's almost additive, but not quite additive.

    So, there's a lot of useful results in mathematics, and I've worked a lot on developing things like this, to the effect that if a function exhibits some structure like this, then it's basically there's a reason for why it's true. And the reason is because there's some other nearby function, which is actually completely structured, which is explaining this sort of partial pattern that you have. And so if you have these inverse theorems, it creates this dichotomy that either the objects that you study are either have no structure at all or they are somehow related to something kind of structured. And in either way, in either case, you can make progress. A good example of this is that there's this old theorem in mathematics-

    Infinity

    Lex Fridman

    Can you prove that there's arithmetic progressions of arbitrary length within a random-

    Terence Tao

    Yes. Have you heard of the infinite monkey theorem? Usually, mathematicians give boring names to theorems, but occasionally they give colorful names.

    Lex Fridman

    Yes.

    Terence Tao

    The popular version of the infinite monkey theorem is that if you have an infinite number of monkeys in a room, each with typewriter, they type out text randomly, almost surely, one of them is going to generate the entire script of Hamlet, or any other finite string of text. It'll just take some time, quite a lot of time, actually, but if you have an infinite number, then it happens.

    So basically, the theorem is that if you take an infinite string of digits or whatever, eventually any finite pattern you wish will emerge. It may take a long time, but it will eventually happen. In particular, arithmetic progressions of any length will eventually happen, but you need an extremely long random sequence for this to happen.

    Lex Fridman

    I suppose that's intuitive. It's just infinity.

    Terence Tao

    Yeah, infinity absorbs a lot of sins.

    Lex Fridman

    Yeah. How we humans supposed to deal with infinity?

    Terence Tao

    Well, you can think of infinity as an abstraction of a finite number of which you do not have a bound. So nothing in real life is truly infinite, but you can ask yourself questions like, "What if I had as much money as I wanted?", or, "What if I could go as fast as I wanted?", and a way in which mathematicians formalize that is mathematics has found a formalism to idealize, instead of something being extremely large or extremely small, to actually be exactly infinite or zero, and often the mathematics becomes a lot cleaner when you do that. I mean, in physics, we joke about assuming spherical cows, real world problems have got all kinds of real world effects, but you can idealize, send some things to infinity, send some things to zero, and the mathematics becomes a lot simpler to work within.

    Lex Fridman

    I wonder how often using infinity forces us to deviate from the physics of reality.

    Terence Tao

    So there's a lot of pitfalls. So we spend a lot of time in undergraduate math classes teaching analysis, and analysis is often about how to take limits and whether…

    So for example, A plus B is always B plus A. So when you have a finite number of terms and you add them, you can swap them and there's no problem, but when you have an infinite number of terms, they're these sort of show games you can play where you can have a series which converges to one value, but you rearrange it, and it suddenly converges to another value, and so you can make mistakes. You have to know what you're doing when you allow infinity. You have to introduce these epsilons and deltas, and there's a certain type of wave of reasoning that helps you avoid mistakes.

    In more recent years, people have started taking results that are true in infinite limits and what's called finitizing them. So you know that something's true eventually, but you don't know when. Now give me a rate. So such… If I don't have an infinite number of monkeys, but a large finite number of monkeys, how long do I have to wait for Hamlet to come out? That's a more quantitative question, and this is something that you can attack by purely finite methods, and you can use your finite intuition, and in this case, it turns out to be exponential in the length of the text that you're trying to generate.

    So this is why you never see the monkeys create Hamlet. You can maybe see them create a four letter word, but nothing that big, and so I personally find once you finitize an infinite statement, it does come much more intuitive, and it's no longer so weird.

    Lex Fridman

    So even if you're working with infinity, it's good to finitize so that you can have some intuition?

    Terence Tao

    Yeah, the downside is that the finitized groups are just much, much messier. So the infinite ones are found first usually, decades earlier, and then later on, people finitize them.

    Math vs Physics

    Lex Fridman

    So since we mentioned a lot of math and a lot of physics, what is the difference between mathematics and physics as disciplines, as ways of understanding, of seeing the world? Maybe we can throw engineering in there, you mentioned your wife is an engineer, give it new perspective on circuits. So this different way of looking at the world, given that you've done mathematical physics, so you've worn all the hats.

    Terence Tao

    Right. So I think science in general is interaction between three things. There's the real world, there's what we observe of the real world, observations, and then our mental models as to how we think the world works.

    We can't directly access reality. All we have are the observations, which are incomplete and they have errors, and there are many, many cases where we want to know, for example, what is the weather like tomorrow, and we don't yet have the observation, but we'd like to. A prediction.

    Then we have these simplified models, sometimes making unrealistic assumptions, spherical cow type things. Those are the mathematical models.

    Mathematics is concerned with the models. Science collects the observations, and it proposes the models that might explain these observations. What mathematics does, we stay within the model, and we ask what are the consequences of that model? What observations, what predictions would the model make of future observations, or past observations? Does it fit? Observe data?

    So there's definitely a symbiosis. I guess mathematics is unusual among other disciplines is that we start from hypotheses, like the axioms of a model, and ask what conclusions come up from that model. In almost any other discipline, you start with the conclusions. "I want to do this. I want to build a bridge, I want to make money, I want to do this," and then you find the paths to get there. There's a lot less sort of speculation about, "Suppose I did this, what would happen?". Planning and modeling. Speculative fiction maybe is one other place, but that's about it, actually. Most of the things we do in life is conclusions driven, including physics and science. I mean, they want to know, "Where is this asteroid going to go? What is the weather going to be tomorrow?", but mathematics also has this other direction of going from the axioms.

    Lex Fridman

    What do you think… There is this tension in physics between theory and experiment. What do you think is the more powerful way of discovering truly novel ideas about reality?

    Terence Tao

    Well, you need both, top down and bottom up. It's really an interaction between all these… So over time, the observations and the theory and the modeling should both get closer to reality, but initially, and this is always the case out there, they're always far apart to begin with, but you need one to figure out where to push the other.

    So if your model is predicting anomalies that are not predicted by experiment, that tells experimenters where to look to find more data to refine the models. So it goes back and forth.

    Within mathematics itself, there's also a theory and experimental component. It's just that until very recently, theory has dominated almost completely. 99% of mathematics is theoretical mathematics, and there's a very tiny amount of experimental mathematics. People do do it. If they want to study prime numbers or whatever, they can just generate large data sets.

    So once we had the computers, we had to do it a little bit. Although even before… Well, like Gauss for example, he discovered a reconjection, the most basic theorem in number theory, called the prime number theorem, which predicts how many primes up to a million, up to a trillion. It's not an obvious question, and basically what he did was that he computed, mostly by himself, but also hired human computers, people whose professional job it was to do arithmetic, to compute the first hundred thousand primes or something, and made tables and made a prediction. That was an early example of experimental mathematics, but until very recently, it was not…

    I mean, theoretical mathematics was just much more successful. Of course, doing complicated mathematical computations was just not feasible until very recently, and even nowadays, even though we have powerful computers, only some mathematical things can be explored numerically.

    There's something called the combinatorial explosion. If you want us to study, for example, Szemerédi's theorem, you want to study all possible subsets of numbers one to a thousand. There's only 1000 numbers. How bad could it be? It turns out the number of different subsets of one to a thousand is two to the power of 1000, which is way bigger than any computer can currently enumerate.

    So there are certain math problems that very quickly become just intractable to attack by direct brute force computation. Chess is another famous example. The number of chess positions, we can't get a computer to fully explore, but now we have AI, we have tools to explore this space, not with 100% guarantees of success, but with experiment. So we can empirically solve chess now. For example, we have very, very good AIs that don't explore every single position in the game tree, but they have found some very good approximation, and people are using actually these chess engines to do experimental chess. They're revisiting old chess theories about, "Oh, when you do this type of opening… This is a good type of move, this is not," and they can use these chess engines to actually refine, and in some cases, overturn conventional wisdom about chess, and I do hope that that mathematics will have a larger experimental component in the future, perhaps powered by AI.

    Lex Fridman

    We'll, of course, talk about that, but in the case of chess, and there's a similar thing in mathematics, I don't believe it's providing a kind of formal explanation of the different positions. It's just saying which position is better or not that you can intuit as a human being, and then from that, we humans can construct a theory of the matter.

    Nature of Reality

    Terence Tao

    Well, there are these three ontological things. There's actual reality, there's observations and our models, and technically they are distinct, and I think they will always be distinct, but they can get closer over time, and the process of getting closer often means that you have to discard your initial intuitions. So astronomy provides great examples, like an initial model of the world is flat because it looks flat and it's big, and the rest of the universe, the skies, is not. The sun, for example, looks really tiny.

    So you start off with a model, which is actually really far from reality, but it fits the observations that you have. So things look good, but over time, as you make more and more observations, bring it closer to reality, the model gets dragged along with it, and so over time, we had to realize that the earth was round, that it spins, it goes around the solar system, solar system goes around the galaxy, and so on and so forth, and the universe was expanding. Expansions is self-expanding, accelerating, and in fact, very recently this year… So even the acceleration of the universe itself, this evidence now is non-constant.

    Lex Fridman

    The explanation behind why that is…

    Terence Tao

    It's catching up.

    Lex Fridman

    It's catching up. I mean, it's still the dark matter, dark energy, this kind of thing.

    Terence Tao

    We have a model that explains, that fits the data really well. It just has a few parameters that you have to specify. So people say, "Oh, that's fudge factors. With enough fudge factors, you can explain anything," but the mathematical point over the model is that you want to have fewer parameters in your model and data points in your observational set.

    So if you have a model with 10 parameters that explains 10 observations, that is a completely useless model, its what's called overfitted, but if you have a model with two parameters and it explains a trillion observations, which is basically the dark matter model, I think it has 14 parameters, and it explains petabytes of data that the astronomers have.

    You can think of a theory. One way to think about a physical mathematical theory is it's a compression of the universe, and a data compression. So you have these petabytes of observations, you like to compress it to a model which you can describe in five pages and specify a certain number of parameters, and if it can fit, to reasonable accuracy, almost all of your observations, the more compression that you make, the better your theory.

    Lex Fridman

    In fact, one of the great surprises of our universe and of everything in it is that it's compressible at all. That's the unreasonable effectiveness of mathematics

    Terence Tao

    Yeah, Einstein had a quote like that. "The most incomprehensible thing about the universe is that it is comprehensible."

    Lex Fridman

    Right, and not just comprehensible. You can do an equation like e=MC2.

    Terence Tao

    There is actually some possible explanation for that. So there's this phenomenon in mathematics called universality. So, many complex systems at the macro scale are coming out of lots of tiny interactions at the macro scale, and normally, because of the commutative explosion, you would think that the macro scale equations must be infinitely, exponentially more complicated than the macro scale ones, and they are, if you want to solve them completely exactly. If you want to model all the atoms in a box of air…

    Like Avogadro's number is humongous. There's a huge number of particles. If you actually tried to track each one, it'll be ridiculous, but certain laws emerge at the microscopic scale that almost don't depend on what's going on at the macro scale, or only depend on a very small number of parameters.

    So if you want to model a gas of a quintillion particles in a box, you just need to know is temperature and pressure and volume, and a few parameters, like five or six, and it models almost everything you need to know about these 10 to 23 or whatever particles. So we don't understand universality anywhere near as we would like mathematically, but there are much simpler toy models where we do have a good understanding of why universality occurs. The most basic one is the central limit theorem that explains why the bell curve shows up everywhere in nature, that so many things are distributed by what's called a Gaussian distribution, famous bell curve. There's now even a meme with this curve.

    Lex Fridman

    And even the meme applies broadly. The universality to the meme.

    Terence Tao

    Yes, you can go meta if you like, but there are many, many processes. For example, you can take lots of independent random variables and average them together in various ways. You can take a simple average or more complicated average, and we can prove in various cases that these bell curves, these Gaussians, emerge, and it is a satisfying explanation.

    Sometimes they don't. So if you have many different inputs and they're all correlated in some systemic way, then you can get something very far from a bell curve to show up, and this is also important to know when it fails. So universality is not a 100% reliable thing to rely on. The global financial crisis was a famous example of this. People thought that mortgage defaults had this sort of Gaussian type behavior, that if a population of a hundred thousand Americans with mortgages ask what proportion of them would default on their mortgages, if everything was de-correlated, it would be an asset bell curve, and you can manage risk of options and derivatives and so forth, and there's a very beautiful theory, but if there are systemic shocks in the economy that can push everybody to default at the same time, that's very non-Gaussian behavior, and this wasn't fully accounted for in 2008.

    Now I think there's some more awareness that this systemic risk is actually a much bigger issue, and just because the model is pretty and nice, it may not match reality. So the mathematics of working out what models do is really important, but also the science of validating when the models fit reality and when they don't… You need both, but mathematics can help, because for example, these central limit theorems, it tells you that if you have certain axioms like non-correlation, that if all the inputs were not correlated to each other, then you have this Gaussian behavior and things are fine. It tells you where to look for weaknesses in the model.

    So if you have a mathematical understanding of Szemerédi's theorem, and someone proposes to use these Gaussian models or whatever to model default risk, if you're mathematically trained, you would say, "Okay, but what are the systemic correlation between all your inputs?", and so then you can ask the economist, "How much of a risk is that?", and then you can go look for that. So there's always this synergy between science and mathematics.

    Lex Fridman

    A little bit on the topic of universality, you're known and celebrated for working across an incredible breadth of mathematics, reminiscent of Hilbert a century ago. In fact, the great Fields Medal winning mathematician Tim Gowers has said that you are the closest thing we get to Hilbert. He's a colleague of yours.

    Terence Tao

    Oh yeah, good friend.

    Lex Fridman

    But anyway, so you are known for this ability to go both deep and broad in mathematics. So you're the perfect person to ask. Do you think there are threads that connect all the disparate areas of mathematics? Is there a kind of a deep, underlying structure to all of mathematics?

    Terence Tao

    There's certainly a lot of connecting threads, and a lot of the progress of mathematics can be represented by taking… By stories of two fields of mathematics that were previously not connected, and finding connections.

    An ancient example is geometry and number theory. So in the times of the ancient Greeks, these were considered different subjects. I mean, mathematicians worked on both. Euclid worked both on geometry, most famously, but also on numbers, but they were not really considered related. I mean, a little bit, like you could say that this length was five times this length because you could take five copies of this length and so forth, but it wasn't until Descartes, who developed analytical geometry, that you can parameterize the plane, a geometric object, by two real numbers. So geometric problems can be turned into problems about numbers.

    Today this feels almost trivial. There's no content to this. Of course, a plane is X and Y, because that's what we teach and it's internalized, but it was an important development that these two fields were unified, and this process has just gone on throughout mathematics over and over again. Algebra and geometry were separated, and now we have this fluid, algebraic geometry that connects them, and over and over again, and that's certainly the type of mathematics that I enjoy the most.

    I think there's sort of different styles to being a mathematician. I think hedgehogs and fox… A fox knows many things a little bit, but a hedgehog knows one thing very, very well, and in mathematics, there's definitely both hedgehogs and foxes, and then there's people who can play both roles, and I think ideal collaboration, British mathematicians involves very… You need some diversity, like a fox working with many hedgehogs or vice versa, but I identify mostly as a fox, certainly. I like arbitrage, somehow. Learning how one field works, learning the tricks of that wheel, and then going to another field which people don't think is related, but I can adapt the tricks.

    Lex Fridman

    So see the connections between the fields.

    Terence Tao

    Yeah. So there are other mathematicians who are far deeper than I am. They're really hedgehogs. They know everything about one field, and they're much faster and more effective in that field, but I can give them these extra tools.

    Lex Fridman

    I mean, you've said that you can be both a hedgehog and the fox, depending on the context, depending on the collaboration. So can you, if it's at all possible, speak to the difference between those two ways of thinking about a problem? Say you're encountering a new problem, searching for the connections versus very singular focus.

    Terence Tao

    I'm much more comfortable with the fox paradigm. Yeah. So yeah, I like looking for analogies, narratives. I spend a lot of time… If there's a result, I see it in one field, and I like the result, it's a cool result, but I don't like the proof, it uses types of mathematics that I'm not super familiar with, I often try to re-prove it myself using the tools that I favor.

    Often, my proof is worse, but by the exercise they're doing, so I can say, "Oh, now I can see what the other proof was trying to do," and from that, I can get some understanding of the tools that are used in that field. So it's very exploratory, very… Doing crazy things in crazy fields and reinventing the wheel a lot, whereas the hedgehog style is, I think, much more scholarly. You're very knowledge-based. You stay up to speed on all the developments in this field, you know all the history, you have a very good understanding of exactly the strengths and weaknesses of each particular technique. I think you rely a lot more on calculation than sort of trying to find narratives. So yeah, I can do that too, but other people are extremely good at that.

    Lex Fridman

    Let's step back and maybe look at a bit of a romanticized version of mathematics. So I think you've said that early on in your life, math was more like a puzzle-solving activity when you were young. When did you first encounter a problem or proof where you realized math can have a kind of elegance and beauty to it?

    Terence Tao

    That's a good question. When I came to graduate school in Princeton, so John Conway was there at the time, he passed away a few years ago, but I remember one of the very first research talks I went to was a talk by Conway on what he called extreme proof.

    So Conway just had this amazing way of thinking about all kinds of things in a way that you wouldn't normally think of. So he thought proofs themselves as occupying some sort of space. So if you want to prove something, let's say that there's infinitely many primes, you have all different proofs, but you could rank them in different axes. Some proofs are elegant, some proofs are long, some proofs are elementary and so forth, and so there's this cloud, so the space of all proofs itself has some sort of shape, and so he was interested in extreme points of this shape. Out of all these proofs, what is one of those, the shortest, at the expense of everything else, or the most elementary or whatever?

    So he gave some examples of well-known theorems, and then he would give what he thought was the extreme proof in these different aspects. I just found that really eye-opening, that it's not just getting a proof for a result that was interesting, but once you have that proof, trying to optimize it in various ways, that proofing itself had some craftsmanship to it.

    It's certainly informed my writing style, like when you do your math assignments and as you're an undergraduate, your homework and so forth, you're sort of encouraged to just write down any proof that works and hand it in, and as long as it gets a tick mark, you move on, but if you want your results to actually be influential and be read by people, it can't just be correct. It should also be a pleasure to read, motivated, be adaptable to generalize to other things. It's the same in many other disciplines, like coding. There's a lot of analogies between math and coding. I like analogies, if you haven't noticed. You can code something, spaghetti code, that works for a certain task, and it's quick and dirty and it works, but there's lots of good principles for writing code well so that other people can use it, build upon it so it has fewer bugs and whatever, and there's similar things with mathematics.

    Lex Fridman

    Yeah, first of all, there's so many beautiful things there, and John Conway is one of the great minds in mathematics ever, and computer science, just even considering the space of proofs and saying, "Okay, what does this space look like, and what are the extremes?"

    Like you mentioned, coding as an analogy is interesting, because there's also this activity called the code golf, which I also find beautiful and fun, where people use different programming languages to try to write the shortest possible program that accomplishes a particular task, and I believe there's even competitions on this, and it's also a nice way to stress test not just the programs, or in this case, the proofs, but also the different languages. Maybe that's a different notation or whatever to use to accomplish a different task.

    Terence Tao

    Yeah, you learn a lot. I mean, it may seem like a frivolous exercise, but it can generate all these insights, which, if you didn't have this artificial objective to pursue, you might not see…

    Lex Fridman

    What, to you, is the most beautiful or elegant equation in mathematics? I mean, one of the things that people often look to in beauty is the simplicity. So if you look at e=MC2… So when a few concepts come together, that's why the Euler identity is often considered the most beautiful equation in mathematics. Do you find beauty in that one, in the Euler identity?

    Terence Tao

    Yeah. Well, as I said, what I find most appealing is connections between different things that… So if you… Pi equals minus one. So yeah, people use all the fundamental constants. Okay. I mean, that's cute, but to me…

    So the exponential function, which is by Euler, was to measure exponential growth. So compound interest or decay, anything which is continuously growing, continuously decreasing, growth and decay, or dilation or contraction, is modeled by the exponential function, whereas pi comes around from circles and rotation, right? If you want to rotate a needle, for example, a hundred degrees, you need rotate by pi radians, and i, complex numbers, represents the swapping imaginary axes of a 90 degree rotation. So a change in direction.

    So the exponential function represents growth and decay in the direction that you already are. When you stick an i in the exponential, now instead of motion in the same direction as your current position, the motion as a right angles to your current position. So rotation, and then, so E to the pi i equals minus one tells you that if you rotate for a time pi, you end up at the other direction. So it unifies geometry through dilation and exponential growth or dynamics through this act of complexification, rotation by pi i. So it connects together all these two as mathematics, dynamics, geometry and complex numbers. They're all considered almost… They were all next-door neighbors in mathematics because of this identity.

    Lex Fridman

    Do you think the thing you mentioned as Q, the collision of notations from these disparate fields, is just a frivolous side effect, or do you think there is legitimate value in when notation… Although our old friends come together in the night?

    Terence Tao

    Well, it's confirmation that you have the right concepts. So when you first study anything, you have to measure things, and give them names, and initially sometimes, because your model is, again, too far off from reality, you give the wrong things the best names, and you only find out later what's really important.

    Lex Fridman

    Physicists can do this sometimes, but it turns out okay.

    Terence Tao

    So actually, physics happens. So one of the big things was the E, right? So when Aristotle first came up with his laws of motion, and then Galileo and Newton and so forth, they saw the things they could measure, they could measure mass and acceleration and force and so forth, and so Newtonian mechanics, for example, F=ma, was the famous Newton's second law of motion. So those were the primary objects. So they gave them the central billing in the theory.

    It was only later after people started analyzing these equations that there always seemed to be these quantities that were conserved. So in particular, momentum and energy, and it's not obvious that things have an energy. It's not something you can directly measure the same way you can measure mass and velocity, so both, but over time, people realized that this was actually a really fundamental concept.

    Hamilton, eventually in the 19th century, reformulated Newton's laws of physics into what's called Hamiltonian mechanics, where the energy, which is now called the Hamiltonian, was the dominant object. Once you know how to measure the Hamiltonian of any system, you can describe completely the dynamics like what happens to all the states. It really was a central actor, which was not obvious initially, and this change of perspective really helped when quantum mechanics came along, because the early physicists who studied quantum mechanics, they had a lot of trouble trying to adapt their Newtonian thinking, because everything was a particle and so forth, to quantum mechanics, because everything was a wave, but it just looked really, really weird.

    You ask, "What is the quantum version of F=ma?", and it's really, really hard to give an answer to that, but it turns out that the Hamiltonian, which was so secretly behind the scenes in classical mechanics, also is the key object in quantum mechanics, that there's also an object called a Hamiltonian. It's a different type of object. It's what's called an operator rather than a function, but again, once you specify it, you specify the entire dynamics.

    So there's something called Schrodinger's equation that tells you exactly how quantum systems evolve once you have a Hamiltonian. So side by side, they look completely different objects. One involves particles, one involves waves and so forth, but with this centrality, you could start actually transferring a lot of intuition and facts from classical mechanics to quantum mechanics. So for example, in classical mechanics, there's this thing called Noether's theorem. Every time there's a symmetry in a physical system, there was a conservation law. So the laws of physics are translation invariant. Like if I move 10 steps to the left, I experience the same laws of physics as if I was here, and that corresponds to conservation momentum. If I turn around by some angle, again, I experience the same laws of physics. This corresponds to the conservation of angular momentum. If I wait for 10 minutes, I still have the same laws of physics. This corresponds to the law of the conservation of energy. So there's this fundamental connection between symmetry and conservation. And that's also true in quantum mechanics, even though the equations are completely different, but because they're both coming from the Hamiltonian, the Hamiltonian controls everything, every time the Hamiltonian has a symmetry, the equations will have a conservation wall. Once you have the right language, it actually makes things a lot cleaner.

    One of the problems why we can't unify quantum mechanics and general relativity, yet we haven't figured out what the fundamental objects are. For example, we have to give up the notion of space and time being these almost Euclidean-type spaces, and it has to be, we know that at very tiny scales there's going to be quantum fluctuations. There's space-time foam and trying to use Cartesian coordinates X, Y, Z. It's a non-starter, but we don't know what to replace it with. We don't actually have the concepts, the analog Hamiltonian that sort of organized everything.

    Theory of Everything

    Lex Fridman

    Does your gut say that there is a theory of everything, so this is even possible to unify, to find this language that unifies general relativity and quantum mechanics?

    Terence Tao

    I believe so. The history of physics has been out of unification much like mathematics over the years. Magnetism was separate theories and then Maxwell unified them. Newton unified the motions of heavens for the motions of objects on the Earth and so forth. So it should happen. It's just that, again, to go back to this model of the observations and theory, part of our problem is that physics is a victim of it's own success. That our two big theories of physics, general relativity and quantum mechanics are so good now is that together they cover 99.9% of all the observations we can make. And you have to either go to extremely insane particle accelerations or the early universe or things that are really hard to measure in order to get any deviation from either of these two theories to the point where you can actually figure out how to combine together. But I have faith that we've been doing this for centuries and we've made progress before. There's no reason why we should stop.

    Lex Fridman

    Do you think you'll be a mathematician that develops a theory of everything?

    Terence Tao

    What often happens is that when the physicists need some theory of mathematics, there's often some precursor that the mathematicians worked out earlier. So when Einstein started realizing that space was curved, he went to some mathematician and asked, "Is there some theory of curved space that mathematicians already came up with that could be useful?" And he said, "Oh yeah, I think Riemann came up with something." And so yeah, Riemann had developed Riemannian geometry, which is precisely a theory of spaces that are curved in various general ways, which turned out to be almost exactly what was needed by Einstein's theory. This is going back to weakness and unreasonable effectiveness of mathematics. I think the theories that work well, that explain the universe, tend to also involve the same mathematical objects that work well to solve mathematical problems. Ultimately, they're just both ways of organizing data in useful ways.

    Lex Fridman

    It just feels like you might need to go some weird land that's very hard to intuit. You have string theory.

    Terence Tao

    Yeah, that was a leading candidate for many decades. I think it's slowly pulling out of fashion. It's not matching experiment.

    Lex Fridman

    So one of the big challenges of course, like you said, is experiment is very tough because of how effective both theories are. But the other is just you're talking about you're not just deviating from space-time. You're going into some crazy number of dimensions. You're doing all kinds of weird stuff that to us, we've gone so far from this flat earth that we started at, like you mentioned, and now it's very hard to use our limited ape descendants of a cognition to intuit what that reality really is.

    Terence Tao

    This is why analogies are so important. So yeah, the round earth is not intuitive because we're stuck on it. But round objects in general, we have pretty good intuition over and we have interest about light works and so forth. And it's actually a good exercise to actually work out how eclipses and phases of the sun and the moon and so forth can be really easily explained by round earth and round moon and models. And you can just take a basketball and a golf ball and a light source and actually do these things yourself. So the intuition is there, but you have to transfer it.

    Lex Fridman

    That is a big leap intellectually for us to go from flat to round earth because our life is mostly lived in flat land. To load that information and we're all like, take it for granted. We take so many things for granted because science has established a lot of evidence for this kind of thing, but we're on a round rock flying through space. Yeah, that's a big leap. And you have to take a chain of those leaps. The more and more and more we progress,

    Terence Tao

    Right, yeah. So modern science is maybe, again, a victim of own success is that in order to be more accurate, it has to move further and further away from your initial intuition. And so for someone who hasn't gone through the whole process of science education, it looks more suspicious because of that. So we need more grounding. There are scientists who do excellent outreach, but there's lots of science things that you can do at home. Lots of YouTube videos I did at YouTube video recently, Grant Sanderson, we talked about this earlier, that how the ancient Greeks were able to measure things like the distance of the moon, distance the earth, and using techniques that you could also replicate yourself. It doesn't all have to be fancy space telescopes and very intimidating mathematics.

    Lex Fridman

    Yeah, I highly recommend that. I believe you give a lecture and you also did an incredible video with Grant. It's a beautiful experience to try to put yourself in the mind of a person from that time shrouded in mystery. You're on this planet, you don't know the shape of it, the size of it. You see some stars, you see some things and you try to localize yourself in this world and try to make some kind of general statements about distanced places.

    Terence Tao

    Change of perspective is really important. You say travel broadens the mind, this is intellectual travel. Put yourself in the mind of the ancient Greeks or person some other time period, make hypotheses, spherical assumption, whatever, speculate. And this is what mathematicians do and some other, what artists do actually.

    Lex Fridman

    It's just incredible that given the extreme constraints, you could still say very powerful things. That's why it's inspiring. Looking back in history, how much can be figured out when you don't have much to figure out stuff with.

    Terence Tao

    If you propose axioms, then the mathematics does. You follow those axioms to their conclusions and sometimes you can get quite a long way from initial hypotheses.

    General Relativity

    Lex Fridman

    If we can stay in the land of the weird. You mentioned general relativity. You've contributed to the mathematical understanding, Einstein's field equations. Can you explain this work and from a mathematical standpoint, what aspects of general relativity are intriguing to you? Challenging to you?

    Terence Tao

    I have worked on some equations. There's something called the wave maps equation or the Sigma field model, which is not quite the equation of space-time gravity itself, but of certain fields that might exist on top of space-time. So Einstein's equations of relativity just describe space and time itself. But then there's other fields that live on top of that. There's the electromagnetic field, there's things called Yang-Mills fields, and there's this whole hierarchy of different equations of which Einstein's considered one of the most nonlinear and difficult, but relatively low on the hierarchy was this thing called the wave maps equation. So it's a wave which at any given point is fixed to be on a sphere. So I can think of a bunch of arrows in space and time. Yeah, so it's pointing in different directions, but they propagate like waves. If you wiggle an arrow, it would propagate and make all the arrows move kind of like sheaves of wheat in a wheat field.

    And I was interested in the global regularity problem. Again for this question, is it possible for the energy here to collect at a point? So the equation I considered was actually what's called a critical equation where it's actually the behavior at all scales is roughly the same. And I was able barely to show that you couldn't actually force a scenario where all the energy concentrated at one point, that the energy had to disperse a little bit at the moment, just a little bit. It would stay regular. Yeah, this was back in 2000. That was part of why I got interested in Kakeya afterwards actually. So I developed some techniques to solve that problem. So part of it, this problem is really nonlinear because of the curvature of the sphere. There was a certain nonlinear effect, which was a non-perturbative effect. It was when you sort looked at it normally it looked larger than the linear effects of the wave equation. And so it was hard to keep things under control even when your energy was small.

    But I developed what's called a gauge transformation. So the equation is kind of like an evolution of sheaves of wheat, and they're all bending back and forth, so there's a lot of motion. But if you imagine stabilizing the flow by attaching little cameras at different points in space, which are trying to move in a way that captures most of the motion, and under this stabilized flow, the flow becomes a lot more linear. I discovered a way to transform the equation to reduce the amount of nonlinear effects, and then I was able to solve the equation. I found the transformation while visiting my aunt in Australia, and I was trying to understand the dynamics of all these fields, and I couldn't do a pen and paper, and I had not enough facility of computers to do any computer simulations.

    So I ended up closing my eyes being on the floor and just imagining myself to actually be this vector field and rolling around to try to see how to change coordinates in such a way that somehow things in all directions would behave in a reasonably linear fashion. And yeah, my aunt walked in on me while I was doing that and she was asking, "Why am I doing this?"

    Lex Fridman

    It's complicated as the answer.

    Terence Tao

    "Yeah, yeah. And okay, fine. You are a young man. I don't ask questions."

    Solving Difficult Problems

    Lex Fridman

    I have to ask about how do you approach solving difficult problems if it's possible to go inside your mind when you're thinking, are you visualizing in your mind the mathematical objects, symbols, maybe what are you visualizing in your mind? Usually when you're thinking?

    Terence Tao

    A lot of pen and paper. One thing you pick up as a mathematician is I call it cheating strategically. So the beauty of mathematics is that you get to change the problem and change the rules as you wish. You don't get to do this by any other field. If you're an engineer and someone says, "Build a bridge over this river," you can't say, "I want to build this bridge over here instead," or, "I want to put it out of paper instead of steel," but a mathematician, you can do whatever you want on. It's like trying to solve a computer game where there's unlimited cheat codes available. And so you can set this, there's a dimension that's large. I've set it to one. I'll solve the one-dimensional problem first. So there's a main term and an error term. I'm going to make a spherical call assumption, error term is zero.

    And so the way you should solve these problems is not in this Iron Man mode where you make things maximally difficult, but actually the way you should approach any reasonable math problem is that if there are 10 things that are making your life difficult, find a version of the problem that turns off nine of the difficulties, but only keeps one of them and solve that. And then so you solve nine cheats. Okay, you solve 10 cheats, then the game is trivial, but you solve nine cheats. You solve one problem that teaches you how to deal with that particular difficulty. And then you turn that one off and you turn someone else something else on, and then you solve that one. And after you know how to solve the 10 problems, 10 difficulties separately, then you have to start merging them a few at a time.

    As a kid, I watched a lot of these Hong Kong action movies from our culture, and one thing is that every time it's a fight scene, so maybe the hero gets swarmed by a hundred bad-guy goons or whatever, but it'll always be choreographed so that he'd always be only fighting one person at a time and it would defeat that person and move on. And because of that, he could defeat all of them. But whereas if they had fought a bit more intelligently and just swarmed the guy at once, it would make for much worse cinema, but they would win.

    Lex Fridman

    Are you usually pen and paper? Are you working with computer and LaTeX?

    Terence Tao

    Mostly pen and paper actually. So in my office I have four giant blackboards and sometimes I just have to write everything I know about the problem on the four blackboards and then sit my couch and just see the whole thing.

    Lex Fridman

    Is it all symbols like notation or is there some drawings?

    Terence Tao

    Oh, there's a lot of drawing and a lot of bespoke doodles that only makes sense to me. And the beauty of a blackboard is you erase and it's a very organic thing. I'm beginning to use more and more computers, partly because AI makes it much easier to do simple coding things that if I wanted to plot a function before, which is moderately complicated, has some iteration or something, I'd had to remember how to set up a Python program and how does a full loop work and debug it and it would take two hours and so forth. And now I can do it in 10, 15 minutes as much. I'm using more and more computers to do simple explorations.

    AI-Assisted Theorem Proving

    Lex Fridman

    Let's talk about AI a little bit if we could. So maybe a good entry point is just talking about computer-assisted proofs in general. Can you describe the Lean formal proof programming language and how it can help as a proof assistant and maybe how you started using it and how it has helped you?

    Terence Tao

    So Lean is a computer language, much like standard languages like Python and C and so forth, except that in most languages the focus is on using executable code. Lines of code do things, they flip bits or they make a robot move or they deliver your text on the internet or something. So lean is a language that can also do that. It can also be run as a standard traditional language, but it can also produce certificates. So a software language like Python might do a computation and give you that the answer is seven. Okay, does the sum of three plus four equal to seven?

    But Lean can produce not just the answer, but a proof that how it got the answer of seven as three plus four and all the steps involved. So it creates these more complicated objects, not just statements, but statements with proofs attached to them. And every line of code is just a way of piecing together previous statements to create new ones. So the idea is not new. These things are called proof assistance, and so they provide languages for which you can create quite complicated mathematical proofs. They produce these certificates that give a 100% guarantee that your arguments are correct if you trust the compiler of Lean, but they made the compiler really small and there are several different compilers available for the Lean.

    Lex Fridman

    Can you give people some intuition about the difference between writing on pen and paper versus using Lean programming language? How hard is it to formalize statement?

    Terence Tao

    So Lean, a lot of mathematicians were involved in the design of Lean. So it's designed so that individual lines of code resemble individual lines of mathematical argument. You might want to introduce a variable, you want to prove our contradiction. There are various standard things that you can do and it's written. So ideally should like a one-to-one correspondence. In practice, it isn't because Lean is explaining a proof to extremely pedantic colleague who will point out, "Okay, did you really mean this? What happens if this is zero? Okay, how do you justify this?" So Lean has a lot of automation in it to try to be less annoying. So for example, every mathematical object has to come with a type. If I talk about X, is X a rule number or a natural number or a function or something? If you write things informally, it's often if you have context. You say, "Clearly X is equal to let X be the sum of Y and Z and Y and Z were already rule number, so X should also be a rule number." So Lean can do a lot of that, but every so often it says, wait a minute, can you tell me more about what this object is? What type of object it is? You have to think more at a philosophical level, not just computations that you're doing, but what each object actually is in some sense.

    Lex Fridman

    Is he using something like LLMs to do the type inference or you match with the real number?

    Terence Tao

    It's using much more traditional what's called good old-fashioned AI. You can represent all these things as trees, and there's always algorithm to match one tree to another tree.

    Lex Fridman

    So it's actually doable to figure out if something is a real number or a natural number.

    Terence Tao

    Every object comes with a history of where it came from, and you can kind of trace it.

    Lex Fridman

    Oh, I see.

    Terence Tao

    Yeah. So it's designed for reliability. So modern AIs are not used in, it's a disjoint technology. People are begin to use AIs on top of lean. So when a mathematician tries to program proven in lean, often there's a step. Okay, now I want to use the fundamental thing on calculus, say to do the next step. So the lean developers have built this massive project called Mathlib, a collection of tens of thousands of useful facts about methodical objects.

    And somewhere in there is the fundamental calculus, but you need to find it. So a lot of the bottleneck now is actually lemma search. There's a tool that you know is in there somewhere and you need to find it. And so there are various search engine engines specialized for Mathlib that you can do, but there's now these large language models that you can say, "I need the fundamental calculus at this point." And it was like, okay, for example, when I code, I have GitHub Copilot installed as a plugin to my IDE, and it scans my text and it sees what I need. Says I might even type, now I need to use the fundamental calculus. And then it might suggest, "Okay, try this," and maybe 25% of the time it works exactly.