Tesla AI Day 30 septembre 2022 Tesla AI Day 2022 : le premier prototype mobile du robot humanoïde Optimus, ainsi que les avancées sur le Full Self-Driving et le supercalculateur Dojo.
Tesla AI Day 2022 : le premier prototype mobile du robot humanoïde Optimus, ainsi que les avancées sur le Full Self-Driving et le supercalculateur Dojo.
Transcription Elon Musk
All right, welcome everybody. Give everyone a moment to get back in the audience and all right, great. Welcome to Tesla AI Day 2022. We've got some really exciting things to show you. I think you'll be pretty impressed. I do want to set some expectations with respect to our Optimus robot. As you know, last year it was just a person in a robot suit. But we've come a long way and it's, I think, you know, compared to that, it's going to be very impressive. And we're going to talk about the advancements in AI for full self driving as well as how they apply to, more generally to real world AI problems like a humanoid robot. And even going beyond that, I think there's some potential that what we're doing here at Tesla could make a meaningful contribution to AGI. And I think actually Tesla's a good entity to do it from a governance standpoint because we're a publicly traded company with one class of stock and that means that the public controls Tesla. And I think that's actually a good thing. So if I go crazy, you can fire me. This is important. Maybe I've gone crazy. I don't know. So, yeah, so we're going to talk a lot about our progress in AI autopilot as well as progress, progress in with Dojo. And then we're going to bring the team out and do a long Q and A. So you can ask tough questions,
Elon Musk
you'd like, existential questions, technical questions, but we want to have as much time for Q and A as possible. So let's see with that. You guys want.
Milan Kovac
Hey guys, I'm Milan. I work on Autopilot and the Tesla.
Lizzie Miskovetz
And I'm Lizzy, mechanical engineer on the project as well.
Elon Musk
Okay, so should we bring up the bot?
Lizzie Miskovetz
Before we do that, we have one, one little bonus tip for the day. This is actually the first time we try this robot without any backup support. Cranes, mechanical mechanisms, no cables, nothing.
John Emmons
Yeah, want to do it with you
Milan Kovac
guys tonight, but it's the first time.
Lizzie Miskovetz
You ready?
Kate Park (Tesla)
Let's go. Go. Sam,
Eric (Tesla)
I think the bug got some boobs here.
Tesla Presenter
So this is essentially the same self
Milan Kovac
driving computer that runs in your Tesla cars.
Tesla Presenter
By the way,
Elon Musk
It's literally the first time the robot has operated without a tether was on stage tonight. So the robot can actually do a lot more than we just showed you. We just didn't want it to fall on its face. So we'll show you some videos now of the robot doing a bunch of other things.
Elon Musk
Which are less risky. Yeah.
Milan Kovac
We should close the screen, guys.
Tesla Presenter
Yeah, Yeah.
Milan Kovac
We wanted to show a little bit more what we've done over the past
Eric (Tesla)
few months with the pod and just
Milan Kovac
walking around and dancing on stage. Just humble beginnings. But you can see the autopilot neural networks running as is, just retrained for the bot directly on that, on that new platform. That's my watering can.
Elon Musk
Yeah. When you, when you see a rendered view, that's, that's the robot. What's the, that's the world the robot sees. So it's, it's very clearly identifying objects. Like this is the object it should pick up. Picking it up. Yeah.
Milan Kovac
We use the same process as we did for autopilot to collect data in train neural networks that we then deploy on the robot. That's an example that illustrates the upper
Elon Musk
body a little bit more.
Milan Kovac
Something that we'll like try to nail down in a few months. Over the next few months, I would say to perfection.
Lizzie Miskovetz
This is really an actual station in the Fremont factory as well that it's working at.
Lizzie Miskovetz
And that's not the only thing we have to show today, right?
Elon Musk
Yeah, absolutely. So what you saw was what we call Bumble C, that's our sort of rough development robot using semi off the shelf actuators. But we actually have gone a step further than that already. The team's done an incredible job and we actually have an Optimus bot with fully Tesla designed and built actuators, battery pack control system, everything. It wasn't quite ready to walk, but I think it will walk in a few weeks. But we wanted to show you the robot something that's actually fairly close to what will go into production and show you all the things it can do. So let's bring it out. Do it,
Elon Musk
So here you're seeing Optimus with
Elon Musk
degrees of freedom that we expect to have in optimus production unit 1, which is the ability to move all the fingers independently, to have the thumb have two degrees of freedom. So it has opposable thumbs and both left and right hand. So it's able to operate tools and, and do useful things. Our goal is to make a useful humanoid robot as quickly as possible. And we've also designed it using the same discipline that we use in designing the car, which is to say to design it for manufacturing such that it's possible to make the robot in high volume at low cost with high reliability. So that's incredibly important. I mean you full scene, very impressive humanoid robot Demonstrations, and that's great, but what are they missing? They're missing a brain. They don't have the intelligence to navigate the world by themselves. And they're also very expensive and made in low volume, whereas Optimus is designed to be an extremely capable robot, but made in very high volume, probably ultimately millions of units. And it is expected to cost much less than a car. I'll just bring it directly to the right here. I would say probably less than $20,000 would be my guess. The potential for optimists is, I think, appreciated by very few people.
Elon Musk
as usual, Tesla demos are coming in hot. So.
Tesla Presenter
Okay, that's good, that's good.
Elon Musk
Yeah. The team's put up. The team has put in an incredible amount of work. Working days, you know, seven days a week, burning the 3:00am oil to get to the demonstration today. Super proud of what they've done. They've really done a great job. I'd just like to give a hand to the whole Optimus team. So, you know, now there's still a lot of work to be done to refine Optimus and improve it. Obviously this is just Optimus version one, and that's really why we're holding this event, which is to convince some of the most talented people in the world, like you guys, to join Tesla and help make it a reality and bring it to fruition at scale such that it can help millions of people. And the potential, like I said, really boggles the mind because you have to say, like, what is an economy? An economy is sort of productive entities. Times the productivity capita times output productivity per capita. At the point at which there is not a limitation on capita, it's not clear what an economy even means at that point. An economy becomes quasi infinite. So what, you know, taken to fruition in the hopefully benign scenario, the, this means a future of abundance, a future where there is no poverty, where people, you can have whatever you want in terms of products and services. It really is a, a fundamental transformation of civilization as we know it. Obviously we want to make sure that transformation is a positive one and safe. But that's also why I think Tesla as an entity doing this, being a single class of stock, publicly traded, owned by the public, is very important and should not be overlooked. I think this is essential because then if the public doesn't like what Tesla's doing, the public can buy shares in Tesla and vote differently. This is a big deal. Like, it's very important that I can't just do what I want. You know, sometimes people think, but it's not true.
Elon Musk
you know, it's very important that the corporate entity that has, that makes this happen is something that the public can properly influence. And so I think the Tesla structure is ideal for that. Like I said, self driving cars will certainly have a tremendous impact on the world. I think they will improve, improve the productivity of transport by at least a half order of magnitude, perhaps an order of magnitude, perhaps more. Optimus I think has maybe a two order of magnitude potential improvement in economic output. It's not clear what the them it actually even is. So but we need to do this in the right way. We need to do it carefully and safely and ensure that the outcome is one that is beneficial to civilization and one that humanity wants. This is also extremely important, obviously. So, and I hope you will consider joining Tesla to achieve those goals. At Tesla, we really care about doing the right thing here or aspire to do the right thing and really not pave the road to hell with good intentions. And I think the road to hell is mostly paved with bad intentions. But every now and again there's a good intention in there. So we want to do it, do the right thing. So, you know, consider joining us and helping make it happen. With that, let's, let's move on to the next phase.
Lizzie Miskovetz
Right on. Thank you, Elon. All right, so you've seen a couple robots today. Let's do a quick timeline recap. So last year we unveiled the Tesla bot concept. But a concept doesn't get us very far. We knew we needed a real development and integration platform to get real life learnings as quickly as possible. So that robot that came out and did the little routine for you guys, we had that within six months built. Working on software integration, hardware upgrades over the months since then. But in parallel we've also been designing the next generation. This one over here. So this guy is rooted in the foundation of sort of the vehicle design process. You know, we're leveraging all of those learnings that we already have. Obviously there's a lot that's changed since last year, but there's a few things that are still the same. You'll notice we still have this really detailed focus on the true human form. We think that matters for a few reasons, but it's fun. We spend a lot of time thinking about how amazing the human body is. We have this incredible range of motion, typically really amazing strength. A fun exercise is if you put your fingertip on the chair in front of you, you'll notice that there's a huge range of motion that you have in your shoulder and your elbow, for example, without moving your fingertip, you can move those joints all over the place. But the robot, you know, its main function is to do real useful work and it maybe doesn't necessarily need all of those degrees of freedom right away. So we've stripped it down to a minimum, sort of 28 fundamental degrees of freedom. And then of course, our hands, in addition to that, humans are also pretty efficient at some things and not so efficient in other times. So, for example, we can eat a small amount of food to sustain ourselves for several hours. That's great. But when we're just kind of sitting around, no offense, but we're kind of inefficient, we're just sort of burning energy. So on the robot platform, what we're going to do is we're going to minimize that idle power consumption, drop it as low as possible, and that way we can just flip a switch and immediately the robot turns into something that does useful work. So let's talk about this latest generation in some detail, shall we? So on the screen here, you'll see in orange our actuators, which we'll get to in a little bit, and in blue, our electrical system. So now that we have our sort of human based research and we have our first development platform, we have both research and execution to draw from.
Kate Park (Tesla)
For this design.
Lizzie Miskovetz
Again, we're using that vehicle design foundation. So we're taking it from concept through design and analysis and then build and validation. Along the way, we're going to optimize for things like cost and efficiency, because those are critical metrics to take this product to scale eventually. How are we going to do that? Well, we're going to reduce our part count and our power consumption of every element possible, possible. We're going to do things like reduce the sensing and the wiring at our extremities. You can imagine a lot of mass in your hands and feet is going to be quite difficult and power consumptive to move around. And we're going to centralize both our power distribution and our compute to the physical center of the platform. So in the middle of our torso, actually it is the torso, we have our battery pack. This is sized at 2.3 kilowatt hours, which is perfect for about a full day's worth of work.
Lizzie Miskovetz
What's really unique about this battery pack is it has all of the battery electronics integrated into a single PCB within the pack. So that means everything from sensing to fusing, charge management and power distribution is all on one, all in one place. We're also leveraging both our vehicle products and our energy products to roll all of those key features into this battery. So that's streamlined manufacturing, really efficient and simple cooling methods, battery management and also safety. And of course we can leverage Tesla's existing infrastructure and supply chain to make it. So going on to sort of our brain, it's not in the head, but it's pretty close. Also in our torso we have our central computer. So as you know, Tesla already ships full self driving computers in every vehicle we produce. We want to leverage both the autopilot hardware and the software for the humanoid platform. But because it's different in requirements and in form factor, we're going to change a few things first. So we still are going to, it's going to do everything that a human brain does, processing vision data, making split second decisions based on multiple sensory inputs, and also communications. So to support communications, it's equipped with wireless connectivity as well as audio support. And then it also has hardware level security features which are important to protect both the robot and the people around the robot. So now that we have our sort of core, we're going to need some limbs on this guy. And we'd love to show you a little bit about our actuators and our fully functional hands as well. But before we do that, I'd like to introduce Malcolm, who's going to speak a little bit about our structural foundation for the robot.
Tesla Presenter
Thank you, lizzie. Tesla have the capabilities to analyze highly complex systems. Don't get much more complex than a crash. You can see here a simulated crash from Model 3 superimposed on top of the actual physical crash. It's actually incredible how accurate it is. Just to give you an idea of the complexity of this model, it includes every nut, bolt and washer, every spot weld, and it has 35 million degrees of freedom. Quite amazing. And it's true to say that if we didn't have models like this, we wouldn't be able to make the safest cars in the world. So can we utilize our capabilities and our methods from the automotive side to influence a robot? Well, we can make a model. And since we have crash software, we use the same software here, we can make it fall down. The purpose of this is to make sure that if it falls down, ideally it doesn't, but it's superficial damage. We don't want it to, for example, break its gearbox and its arms. That's equivalent to a dislocated shoulder of a robot. Difficult and expensive to fix. So we wanted to dust itself off, get on with the job it's been given. We can also take the same model and we can drive the actuators using the inputs from a previously solved model, bringing it to life. So this is producing the motions for the tasks we want the robot to do. These tasks are picking up boxes, turning, squatting, walking upstairs. Whatever the set of tasks are, we can play to the model. This is showing just simple walking. We can create the stresses in all the components. That helps us optimize the components. These are not dancing robots. These are actually the modal behavior, the first five modes of the robot. And typically when people make robots, they make sure the first mode is up around the top, single figures up towards 10 hertz. Who is it do this is to make the controls of walking easier. It's very difficult to walk if you can't guarantee where your foot is wobbling around. That's okay if you make one robot, we want to make thousands, maybe millions. We haven't got the luxury of making from carbon fiber and titanium. We want to make them on plastic. Things are not quite so stiff. So we can't have these high targets, I call them dumb targets. We've got to make them work at lower targets. So is that going to work? Well, if you think about it, sorry about this, but we're just bags of soggy jelly and bones thrown in. We're not high frequency. If I stand on my leg, I don't vibrate at 10 hertz. People operate at a low frequency. So we know the robot actually can. It just makes controls harder. So we take the information from this, the modal data and the stiffness, and feed that into the control system that allows it to walk. Just changing tack slightly. Looking at the knee, we can take some inspiration from biology and we can look to see what the mechanical advantage of the knee is. It turns out it actually represents quite similar to four bar link. And that's quite non linear. That's not surprising really because if you think when you bend your leg down, the torque on your knee is much more when it's bent than it is when it's straight. So you'd expect a non linear function. And in fact the biology is nonlinear. This matches it quite accurately. So that's the representation. The four bar link is obviously not physically a four bar link. As I said, the characteristics are similar. But me bending down, that's not very scientific. Let's be a bit more scientific. We've played all the tasks through this graph. This is showing picking things up, walking, squatting, the tasks I said we did on the stress. And that's the torque seen at the knee against the knee bend on the horizontal axis. This is showing the requirement for the knee to do all these tasks. I've then put a curve through it, surfing over the top of the peaks. And that's saying this is what's required to make the robot do these tasks. So if we look at the four bar lead, that's actually the green curve and it's saying that the non linearity of the four bar link is actually linearized. The characteristic of the force, what that really says is that's lowered the force. That's what makes the actuator have the lowest possible force, which is the most efficient. We want to burn energy up slowly. What's the blue curve? Well, the blue curve is actually if we didn't have a four bar link, we just had an arm sticking out my leg here with an actuator on it, a simple two bar link. That's the best you could do with a simple two bar link. And it shows that that would create much more force in the actuator which would not be efficient. So what does that look like in practice? Well, as you'll see, it's very tightly packaged in the knee. You'll see it go transparent on a second. You'll see the four bar link there. It's operating on the actuator. This is determining the force and the displacements on the actuator. I now pass you over to Constantino to tell you a lot more detail about how these actuators are made and designs optimized. Thank you.
Elon Musk
Thank you, Michael.
Tesla Presenter
So I would like to talk to you about the design process and the actuator portfolio in our robot. So there are many similarities between a car and the robot when it comes to powertrain design. The most important thing that matters here is energy, mass and cost. We are carrying over most of our designing experience from the car to the robot. So in the particular case you see a car with two drive units and the drive units are used in order to accelerate the car 0-60 mph time or drive the city drive size, while the robot that has 28 actuators, it's not obvious what are the tasks at actuator level. So we have tasks that are higher level like walking or climbing stairs or carrying a heavy object which need to be translated into joint, into joint specs. Therefore, we use our model that generates the torque speed trajectories for our joints which subsequently is going to be fed in our optimization model. To run through the optimization process. This is one of the scenarios that the robot is capable of doing, which is turning and walking. So when we have this torque speed trajectory, we laid over an efficiency map of an actuator. And we are able along the trajectory to generate the power consumption and the energy, cumulative energy for the task versus time. So this allows us to define the system cost for the particular actuator and put a simple point into the cloud. Then we do this for hundreds of thousands of actuators by solving in our cluster. And the red line denotes the Pareto front, which is the preferred area where we will look for optimal. So the X denotes the preferred actuator design we have picked for this particular joint. So now we need to do this for every joint. We have 28 joints to optimize and we parse our cloud. We parse our cloud again for every joint spec. And the red axis, this time denotes the bespoke actuator designs for every join. The problem here is that we have too many unique actuator designs. And even if we take advantage of the symmetry, still there are too many. In order to make something mass manufacturable, we need to be able to reduce the amount of unique actuator designs. Therefore, we run something called commonality study, which we parse our cloud again, looking this time for actuators that simultaneously meet the joint performance requirements for more than one joint at the same time. So the resulting portfolio is six actuators and they show in a color map in the middle figure. And the actuators can be also viewed in this slide. We have three rotary and three linear actuators, all of which have a great output force or torque per mass. The rotary actuator in particular has a mechanical clutch integrated on the high speed side, angular contact ball bearing and on the high speed side and on the low speed side, a cross roller bearing. And the gear train is a strain wave gear. There are three integrated sensors here and a bespoke permanent magnet machine, the linear actuator.
Tesla Presenter
The linear actuator has planetary rollers and an inverted planetary screw as a gear train which allows efficiency and compaction and durability. So in order to demonstrate the force capability of our linear actuators, we have set up an experiment in order to test it under its limits. And I will let you enjoy the video. So our actuator is able to lift. A half ton nine foot concert grand piano. And. This is a requirement. It's not something nice to have because our muscles can do the same when they are direct driven, when they are directly driven, or quadriceps muscles can do the same thing. It's just that the knee is an up gearing linkage system that converts the force into velocity at the end effector of our heels for purposes of giving to the human body agility. So this is one of the main things that are amazing about the human body. And I'm concluding my part at this point and I would like to welcome my colleague Mike, who's going to talk to you about hand design. Thank you very much. Thanks Josephinos. So we just saw how powerful a
Felix (Tesla)
human and a humanoid actuator can be. However, humans are also incredibly dexterous. The human hand has the ability to move at 300 degrees per second. It has tens of thousands of tactile sensors, and it has the ability to grasp and manipulate almost every object in our daily lives. For our robotic hand design, we were inspired by biology. We have five fingers and an opposable thumb. Our fingers are driven by metallic tendons that are both flexible and strong. We have the ability to complete wide aperture power grasps while also being optimized for precision gripping of small, thin and delicate objects. So why a human like robotic hand? Well, the main reason is that our factories and the world around us is designed to be ergonomic. So what that means is that it ensures that objects in our factory are graspable, but it also ensures that new objects that we may have never seen before can be grasped by the human hand and by our robotic hand as well. The converse there is pretty interesting because it's saying that these objects are designed to our hand instead of having to make changes to our hand to accompany a new object. Some basic stats about our hand is that it has 6 actuators and 11 degrees of freedom. It has an in hand controller which drives the fingers and receives sensor feedback. Sensor feedback is really important to learn a little bit more about the objects that we're grasping and also for proprioception. And that's the ability for us to recognize where our hand is in space. One of the important aspects of our hand is that it's adaptive. This adaptability is involved essentially as complex mechanisms that allow the hand to adapt to the objects that's being grasped. Another important part is that we have a non back drivable finger drive. This clutching mechanism allows us to hold and transport objects without having to turn on the hand motors. You just heard how we went about going, we went about designing the Tesla bot hardware. Now I'll hand it off to Milan and our Autonomy team to bring this real bot to life.
John Emmons
Thanks, Michael.
Milan Kovac
All right, so all those good things we've shown earlier in the video were possible just in a matter of a few months. Thanks to the amazing work that we've done on Autopilot over the past few years. Most of those components ported quite easily over to the bots environment. If you think about it, we're just moving from a robot on wheels to a robot on legs. So some of the components are pretty similar and some other require more heavy lifting. So for example, our computer vision neural networks reported directly from Autopilot to the bots situation. It's exactly the same occupancy network that we'll talk into a little bit more details later with the Autopilot team that is now running on the bot here in this video. The only thing that changed really is the training data that we had to recollect. We're also trying to find ways to improve those occupancy networks using work made on your radiance fields to get really great volumetric rendering of the bot's environments. For example, here some machinery that the bot might have to interact with. Another interesting problem to think about is in indoor environments, mostly with that sense of GPS signal, how do you get the bot to navigate to its destination, say, for instance, to find its nearest charging station? So we've been training more neural networks to identify high frequency features, key points within the bot's camera streams, and track them across frames over time as the bot navigates with its environment. And we're using those points to get a better estimate of the bot's pose and trajectory within its environment as it's walking. We also did quite some work on the simulation side. And this is literally the Autopilot simulator to which we've integrated the robot's locomotion code. And this is a video of the motion control code running in the OPILOT simulator simulator showing the evolution of the robot's work over time. And so, as you can see, we started quite slowly in April and start accelerating as we unlock more joints and deploy more advanced techniques like arms balancing over the past few months. And so locomotion is specifically one component that's very different as we're moving from the car to the bots environment. And so I think it warrants a little bit more depth. And I'd like my colleagues to start talking about this now.
Elon Musk
Thank you, Milan.
Felix (Tesla)
Hi, everyone, I'm Felix, I'm a robotics engineer on the project and I'm going
Tesla Presenter
to talk about walking.
Felix (Tesla)
Walking seems easy Right. People do it every day. You don't even have to think about it. But there are some aspects of walking which are challenging from an engineering perspective. For example, physical self awareness. That means having a good representation of yourself. What is the length of your limbs, what is the mass of your limbs, what is the size of your feet, all that matters. Also having an energy efficient gait. You can imagine there's different styles of walking and all of them are equally efficient. Most important, keep balance, don't fall, and of course also coordinate the motion of all of your limbs together. So now humans do all of this naturally. But as engineers or roboticists, we have to think about these problems and therefore I'm going to show you how we address them in our locomotion planning and control strategy. So we start with locomotion planning and our representation of the bot. That means a model of the robot's kinematics, dynamics and the contact properties. And using that model and the desired path for the bots, our locomotion planner generates reference trajectories for the entire system. This means feasible trajectories with respect to the assumptions of our model. The planner currently works in three stages. It starts planning footsteps and ends with the entire motion photo system. And let's dive a little bit deeper in how this works. So in this video we see footsteps being planned over a planning horizon following the desired path. And we start from this and add then foot trajectories that connect these footsteps using toe off and heel strike.
Eric (Tesla)
Just as the humans.
Felix (Tesla)
Just as humans do. And this gives us larger stride and less knee bend for high efficiency of the system. The last stage is then finding a center of mass trajectory, which gives us a dynamically feasible motion of the entire system to keep balance. As we all know, plans are good, but we also have to realize them in reality. Let's see how we can do this.
Elon Musk
Thank you, Felix.
Tesla Presenter
Hello everyone. My name is Anand and I'm going to talk to you about controls. So let's take the motion plan that Felix just talked about and put it in the real world on a real robot. Let's see what happens. It takes a couple steps and falls down. Well, that's a little disappointing, but we
Phil (Tesla)
are missing a few key pieces here
Tesla Presenter
which will make it work. Now, as Felix mentioned, the motion planner is using an idealized version of itself and a version of reality around it. This is not exactly correct. It also expresses its intention through trajectories and wrenches, wrenches of forces and torques that it wants to exert on the world. To locomote reality is way more complex than any FEMOR model. Also, the robot is not simplified. It's got vibrations and modes compliance, sensor noise, and on and on and on. So what does that do to the real world when you put the bot in the real world? Well, the unexpected forces cause unmodeled dynamics, which essentially the planet doesn't know about. And that causes destabilization, especially for a system that is dynamically stable, like biped locomotion. So what can we do about it? Well, we measure reality. We use sensors and our understanding of the world to do state estimation and state estimation. Here you can see the attitude and pelvis pose, which is essentially the vestibular system in a human, along with the center of mass trajectory being tracked when the robot's walking in the office environment. Now we have all the pieces we need in order to close the loop. So we use our better bot model, we use the understanding of reality that we've gained through state estimation, and we compare what we want versus what we expect the reality expect that reality is doing to us in order to add corrections to the behavior of the robot here. The robot certainly doesn't appreciate being poked, but it does an admirable job of staying upright. The final point here is a robot that walks is not enough. We need it to use its hands
Phil (Tesla)
and arms to be useful.
Tesla Presenter
Let's talk about manipulation.
Eric (Tesla)
Hi, everyone, my name is Eric, robotics engineer on Tesla bot. And I want to talk about how we've made the robot manipulate things in the real world. We wanted to manipulate objects while looking as natural as possible and also get there quickly. So what we've done is we've broken this process down into two steps. First is generating a library of natural motion references, or we could call them demonstrations. And then we've adapted these motion references online to the current real world situation.
John Emmons
So let's say we have a human
Eric (Tesla)
demonstration of picking up an object. We can get a motion capture of that demonstration, which is visualized right here as a bunch of keyframes representing the locations of the hands, the elbows, the torso. We can map that to the robot using inverse kinematics. And if we collect a lot of
John Emmons
these, now we have a library that we can work with.
Eric (Tesla)
But a single demonstration is not generalizable to the variation in the real world. For instance, this would only work for a box in a very particular location. So what we've also done is run these reference trajectories through a trajectory optimization program, which solves for where the hand should be, how the robot should balance during when it needs to adapt the motion to the real world. So, for instance, if the box is in this location, then our optimizer will create this trajectory instead. Next, Milan's going to talk about what's next for the Optimus Tesla bot.
Milan Kovac
Right, so hopefully by now you guys got a good idea of what we've been up to over the past few months. We started being something that's usable, but it's far from being useful. There's still a long and exciting road ahead of us. I think the first thing within the next few weeks is to get Optimus at least at par with Bumble C, the other bug prototype you saw earlier and probably beyond. We are also going to start focusing on the real use case at one of our factories and really going to try to try to nail this down and iron out all the elements needed to deploy this product in the real world. I was mentioning earlier, you know, indoor navigation, graceful fault management, or even servicing all components needed to scale this product up. But I don't know about you, but after seeing what we've shown tonight, I'm pretty sure we can get this done within the next few months or years and make this product a reality and change the entire economy. So I would like to thank the entire Optimus team for their hard work over the past few months. I think it's pretty amazing all of this was done in barely six or eight months. Thank you very much.
Tesla Presenter
Hey, everyone.
Ashok Elluswamy
Hi, I'm Ashok. I lead the Autopilot team alongside Milan. God, it's going to be so hard to top that optimist section. He'll try nonetheless. Anyway, every Tesla that has been built over the last several years, we think has the hardware to make the car drive itself. We have been working on the software to add higher and higher levels of autonomy this time around. Last year, we had roughly 2000 cars driving our FSD beta software. Since then, we have significantly improved the software's robustness and capability that we have now shipped it to 160,000 customers as of today.
Tesla Presenter
Thank you.
Ashok Elluswamy
This did not come for free. It came from the sweat and blood of the engineering team. Over the last one year, for example, we trained 75,000 neural network models just last one year. That's roughly a model every eight minutes that's coming out of the team. And then we evaluate them on our large clusters, and then we ship 281 of those models that actually improve the performance of the car. And this space of innovation is happening throughout the stack. The planning Software, the infrastructure, the tools, even hiring, everything is progressing to the next level. The FSG beta software is quite capable of driving the car. It should be able to navigate from parking lot to parking lot, handling city street driving, stopping for traffic lights and stop signs, negotiating with objects at intersections, making turns, and so on. All of this comes from the camera streams that go through our neural networks that run on the car itself. It's not coming back to the server or anything. It runs on the car and produces all the outputs to form the world model around the car. And the planning software drives the car Based on that, today we'll go into a lot of the components that make up the system. The occupancy network acts as the base geometry layer of the system. This is a multi camera video neural network that from the images predicts the full physical occupancy of the world around the robot. So anything that's physically present, trees, walls, buildings, cars, balls, what have you, it predicts. If it's physically present, it predicts them along with their future motion. On top of this base level of geometry, we have more semantic layers. In order to navigate the roadways, we need the lanes, of course, but then the roadways have lots of different lanes and they connect in all kinds of ways. So it's actually a really difficult problem for typical computer vision techniques to predict the set of planes and their connectivities. So we reached all the way into language technologies and then pulled the state of the art from other domains and not just computer vision. To make this task possible for vehicles, we need their full kinematic state to control for them. All of this directly comes from neural networks, video streams. Raw video streams come into the networks, goes through a lot of processing, and then outputs the full kinematic state. The positions, velocities, acceleration, jerk, all of that directly comes out of the networks with minimal post processing. That's really fascinating to me because how is this even possible? What world do we live in that this magic is possible? That these networks predicts four derivatives of these positions and people thought we couldn't even detect these objects. My opinion is that it did not come for free. It required tons of data. So we had to build sophisticated auto labeling systems that churn through raw sensor data, run a ton of offline compute on the servers. It can take a few hours, run expensive neural networks, distill the information into labels that train our in car neural networks. On top of this, we also use our simulation system to synthetically create images. And since it's a simulation, we trivially have all the labels. All of this goes through a well oiled data engine pipeline where we first train a baseline model with some data, ship it to the car, see what the failures are. And once we know the failures, we mine the fleet for the cases where it fails, provide the correct labels and add the data to the training set. This process systematically fixes the issues. And we do this for every task that runs in the car.
Milan Kovac
And to train these new massive neural networks this year we extended our training infrastructure by roughly 40 to 50%. So that sits us at about 14,000 GPUs today across multiple training clusters in the United States. We also worked on our AI compiler which now supports new operations needed by those neural networks and map them to the best of our underlying hardware resources. And our inference engine today is capable of distributing the execution of a single neural network across two independent system on chips. Essentially two independent computers interconnected within the same full self driving computer. And to make this possible, we had to keep a tight control on the end to end latency of this new system. So we deployed more advanced scheduling code across the 4fsd platform.
Ashok Elluswamy
All of these neural networks running in the car together produce the vector space which is again the model of the world around the robot or the car. Then the planning system operates on top of this, coming up with trajectories that avoid collisions or smooth make progress towards the destination using a combination of model based optimization plus neural network that helps optimize it to be really fast. Today we are really excited to present progress on all of these areas. We have the engineering leads standing by to come in and explain these various blocks and these power, not just the car, but the same components also run on the Optimus robot that Milan showed earlier. With that I welcome Padil to start talking about the planning section.
Paril Jain
Hi all, I'm Paril Jain. Let's use this intersection scenario to dive straight into how we do the planning and decision making in autopilot. So we are approaching this intersection from a side street and we have to yield to all the crossing vehicles. Right as we are about to enter the intersection, the pedestrian on the other side of the intersection decides to cross the road without a crosswalk. Now we need to yield to this pedestrian, yield to the vehicles from the right, and also understand the relation between the pedestrian and the vehicle on the other side of the intersection. So a lot of these intra object dependencies that we need to resolve in a quick glance. And humans are really good at this. We look at a scene, understand all the possible interactions, evaluate the Most promising ones and generally end up choosing a reasonable one. So let's look at a few of these interactions that autopilot system evaluated. We could have gone in front of this pedestrian with a very aggressive longitudinal and lateral profile. Now obviously we are being a jerk to the pedestrian and we would spook the pedestrian and his cute pit. We could have moved forward slowly, shot for a gap between the pedestrian and the vehicle from the right. Again, we are being a jerk to the vehicle coming from the right. But you should not outright reject this interaction in case this is only safe interaction available. Lastly, the interaction we ended up choosing, stay slow initially, find the reasonable gap and then finish the maneuver after all the agents pass. Now, evaluation of all of these interactions is not trivial, especially when you care about modeling the higher order derivatives for other agents. For example, what is the longitudinal jerk required by the vehicle coming from the ride when you assert in front of it? Relying purely on collision checks with marginal predictions will only get you so far because you will miss out on a lot of valid interactions. This basically boils down to solving a multi agent joint trajectory planning problem over the trajectories of ego and all the other agents. Now, how much ever you optimize, there's gonna be a limit to how fast you can run this optimization problem. It will be close to order of 10 milliseconds, even after a lot of incremental approximations. Now for a typical crowded unprotected lift. Say you have more than 20 objects, each object having multiple different future modes. The number of relevant interaction combinations will blow up. The planner needs to make a decision every 50 milliseconds. So how do we solve this in real time? We rely on a framework what we call as interaction search, which is basically a parallelized research over a bunch of maneuver trajectories. The state space here corresponds to the kinematic state of ego, the kinematic state of other agents, the nominal future multimodal predictions and all the static entities in the scene. The action space is where things get interesting. We use a set of maneuver trajectory candidates to branch over a bunch of interaction decisions and also incremental goals for a longer horizon maneuver. Let's walk through this research very quickly to get a sense of how it works. We start with a set of vision measurements, namely lanes occupancy model moving objects. These get represented as sparse attractions as well as latent features. We use this to create a set of goal candidates lanes again from the lanes network or unstructured regions which correspond to a probability mask derived from human Demonstrations. Once we have a bunch of these goal candidates, we create seed trajectories using a combination of classical optimization approaches as well as our network planner, again trained on data from the customer fleet. Now, once we get a bunch of these fleet trajectories, we use them to start branching on the interactions we find. The most critical interaction. In our case, this would be the interaction with respect to the pedestrian, whether we assert in front of it or yield to it. Obviously the option on the left is a high penalty option. It likely won't get prioritized. So we branch further onto the option on the right and that's where we bring in more and more complex interactions. Building this optimization problem incrementally with more and more constraints. And the tree search keeps flowing, branching on more interactions, branching on more goals. Now, a lot of tricks here lie in evaluation of each of this node of the tree search inside each node. Initially, we started with creating trajectories using classical optimization approaches where the constraints like I described would be added incrementally and this would take close to 1 to 5 milliseconds per action. Now, even though this is fairly good number, when you want to evaluate more than 100% tractions, this does not scale. So we ended up building lightweight queryable networks that you can run in the loop of the planner. These networks are trained on human demonstrations from the fleet as well as offline solvers with relaxed time limits. With this, we were able to bring the runtime down to close to 100 microseconds per action. Now, doing this alone is not enough because you still have this massive tree search that you need to go through and you need to efficiently prune the search space. So you need to do scoring on each of these trajectories. Few of these are fairly standard. You do a bunch of collision checks, you do a bunch of comfort analysis. What is the jerk and access required for a given maneuver? The customer fleet data plays an important role. Here again we run two sets of again lightweight variable networks, both really augmenting each other. One of them trained from interventions from the FSC beta fleet, which gives a score on how likely is a given maneuver to result in interventions over the next few seconds. And second, which is purely on human demonstrations, human driven data, giving a score on how close is your given selected action to a human driven trajectory. The scoring helps us prune the search space, keep branching further on the interactions and focus the compute on the most promising outcomes. The cool part about this architecture is that it allows us to create a cool blend between Data driven approaches where you don't have to rely on a lot of hand engineered costs, but also ground it in reality with physics based checks. Now a lot of what I described was with respect to the agents we could observe in the scene. But the same framework extends to objects behind occlusions. We use the video feed from eight cameras to generate the 3D occupancy of the world. The blue mask here corresponds to the visibility region we call basically gets blocked at the first occlusion you see in the scene. We consume this visibility mask to generate what we call as ghost objects which you can see on the top left. Now if you model the spawn regions and the state transitions of these ghost objects correctly, if you tune your control response as a function of their existence likelihood, you can extract some really nice human like behaviors. Now I'll pass it on to Phil to describe more on how we generate these occupancy networks. Thank you.
Phil (Tesla)
Hey guys, my name is Phil. I will share the details of the occupancy network we built over the past year. This network is our solution to model the physical world using in 3D around our cars. And it is currently not shown in our customer facing visualization. And what you will see here is the raw network output from our internal dev tool. The occupancy network takes video streams of all our eight cameras as input produces a single unified volumetric occupancy in vector space directly for every 3D location around our car. It predicts the probability of that location being occupied or not. Since it has video contacts, it is capable of predicting obstacles that are occluded instantaneously for each location. It also produces a set of semantics such as curb, car, pedestrian and road debris as color coded here. Occupancy flow is also predicted for motion. Since the model is a generalized network, it does not tell static and dynamic object explicitly. It is able to produce and model the random motion such as a swerving trainer here. This network is currently running in all Teslas with FSD computers and it is incredibly efficient for runs about every 10 milliseconds with our neural net accelerator. So how does this work? Let's take a look at the architecture. First we rectify each camera images with the camera calibration. And the images we're shown here were given to the network. It's actually not the typical 8 bit RGB image as you can see from the first image on top. We're giving the 12 bit raw photocount image to the network. Since it has 4 bits more information, it has 16 times better dynamic range as well as reduced latency. Since we don't have to run ISP in the loop anymore, we use a set of reglets and biofpns as a backbone to extract image space features. Next we construct a set of 3D position query along with the image space features. As keys and values fit into an attention module, the output of the attention module is high dimensional spatial features. These spatial features are aligned temporarily using vehicle odometry to derive motion. Last, these spatial temporal features go through a set of deconvolution to produce the final occupancy and occupancy flow output. They're formed as fixed size voxel grid which might not be precise enough for planning and control in order to get a higher resolution. We also produce per voxel feature maps which we feed into MLP with 3D spatial point queries to get position and semantics at any arbitrary location. After knowing the model better, let's take a look at another example. Here we have an articulated bus parked on right side road highlighted as an L shaped voxel. Here as we approach the bus start to move, the front of the cart turns blue first indicating the model predicts the front bus has a non zero occupancy flow and as the bus keeps moving the entire bus turns blue and you can also see that the network predicts the precise curvature of the bus. Well, this is a very complicated problem for traditional object detection network as you have to see whether I'm going to use one cuboid or perhaps two to fit the curvature. But for occupancy network, since all we care about is the occupancy in the visible space and we'll be able to model the curvature precisely. Besides the voxel gray, the occupancy network also produces a driver surface. The drywall surface has both 3D geometry and semantics. They are very useful for control, especially on hilly and curvy roads. The surface and the voxel grid are not predicted independently. Instead the voxel grid actually aligns with the surface implicitly. Here we are at a hillcrest where you can see the 3D geometry of the surface being predicted nicely. Planar can use this information to decide perhaps we need to slow down more for the hill crest and as you can also see the voxel gray aligns with the surface consistently. Besides the voxels and the surface, we're also very excited about the recent breakthrough in neural radiance field or lerf. We're looking into both incorporate some of the lightslrf features in into occupancy network training as well as using our network output as the input state for nerf. As a matter of fact, Ashok is very excited about this. This has been his personal weekend project for a while.
Ashok Elluswamy
These nerves, because I think academia is building all of these foundation models for language using like tons of large data sets for language. But I think for vision nerves are going to provide the foundation models for computer vision because they are grounded in geometry. And geometry gives us a nice way to supervise these networks and frees us of the requirement to define an ontology. And the supervision is essentially free because you just have to differentially render these images. So I think in the future, this occupancy network idea, where images come in and then the network produces a consistent volumetric representation of the scene that can then be differentiable rendered into any image that was observed, I personally think is the future of computer vision. And we do some initial work on it right now. But I think in the future, both at Tesla and in academia, we will see that this combination of one shot prediction of full volumetric occupancy will be the future. That's my personal bit.
Phil (Tesla)
So here's an example. Early result of a 3D reconstruction from our fleet data. Instead of focusing on getting perfect RGB reprojection in imaging space, our primary goal here is to accurately represent the world in 3D space for driving. And we want to do this for all our fleet data all over the world in all weather and lighting conditions. And obviously this is a very challenging problem and, and we're looking for you guys to help. Finally, the occupancy network is trained with large auto labeled data set without any human in the loop. And with that, I'll pass to Tim to talk about what it takes to train this network.
Tesla Presenter
Thanks, Phil. All right, hey everyone, let's talk about some training infrastructure. So we've seen a couple videos, four or five I think, and care more and worry more about a lot more clips than that. So we've been looking at the occupancy networks just from Phil, just Phil's videos. It takes 1.4 billion frames to train that network, what you just saw, and if you have 100,000 GPUs, it will take one hour, but if you have one GPU, it would take 100,000 hours. So that is not a humane time period that you can wait for your training job to run. Right. We want to ship faster than that. So that means you're going to need to go parallel. So you need more compute for that. That means you're going to need A supercomputer. So this is why we've built in house three supercomputers comprising of 14,000 GPUs, where we use 10,000 GPUs for training and run 4,000 GPUs for auto labeling. All these videos are stored in 30 petabytes of a distributed managed video cache. You shouldn't think of our data sets as fixed. Let's say as you think of your image Net or something, you know, with like a million frames, you should think of it as a very fluid thing. So we've got half a million of these videos flowing in and out of this cluster, these clusters, every single day, and we track 400,000 of these kind of Python video instantiations every second. So that is, that's a lot of calls. We are going to need to capture that in order to govern the retention policies of this distributed video cache. So underlying all of this is a huge amount of infra, all of which we build and manage in house. So you cannot just buy, you know, 14,000 GPUs and then 30 petabytes of flash NVMe and you still put it together and let's go train. It actually takes a lot of work. And I'm going to go into a little bit of that. What you actually typically want to do is you want to take your accelerator. So that could be the GPU or dojo, which we'll talk about later. And because that's the most expensive component, that's where you want to put your bottleneck. And so that means that every single part of your system is going to need to outperform this accelerator. And so that is really complicated. That means that your storage is going to need to have the size and the bandwidth to deliver all the data down into the nodes. These nodes need to have the right amount of CPU and memory capabilities to feed into your machine learning framework. This machine learning framework then needs to hand it off to your GPU and then you can start training. But then you need to do so across hundreds or thousands of GPU in a reliable way, in lockstep, and in a way that's also fast. So you're also going to need an interconnect, extremely complicated. We'll talk more about Dojo in a second. So first I want to take you through some optimizations that we've done on our cluster. So we're getting in a lot of videos and video is very much unlike, let's say, training on images or text, which I think is very well established. Video is quite literally a Dimension more complicated. And so that's why we needed to go end to end from the storage layer down to the accelerator and optimize every single piece of that. Because we train on the photon count videos that come directly from our fleet, we train on those directly. We do not post process those at all. The way it's just done is we seek exactly to the frames we select for our batch. We load those in, including the frames that they depend on. So these are your iframes or your key keyframes. We package those up, move them into shared memory, move them into a double buffer on the GPU and then use the hardware decoder that's only accelerated to actually decode the video. So we do that on the GPU natively. And this is all in a very nice Python Pytorch extension. Doing so unlocked more than 30% training speed increase for the occupancy networks and freed up basically the whole CPU to do any other thing. You cannot just do training with just videos. Of course you need some kind of a ground truth, and that is actually an interesting problem as well. The objective for storing your ground truth is that you want to make sure you get to your ground truth that you need in the minimal amount of file system operations and load in the minimal size of what you need in order to optimize for aggregate cross cluster throughput. Because you should see a compute cluster as one big device which has internally fixed constraints and thresholds. So for this we rolled out a format that is native to us that's called small. We use this for our ground truth, our feature cache and any inference outputs. So a lot of tensors that are in there and so just a cartoon here. Let's say these are your is your table that you want to store. Then that's how that would look out if you rolled out on disk. So what you do is you take anything you'd want to index on. So for example video timestamps, you put those all in the header so that in your initial header read, you know exactly where to go on disk. Then if you have any tensors, you're going to try to transpose the dimensions to put a different dimension last as the contiguous dimension and then also try different types of compression. Then you check out which one was most optimal and then store that one. This is actually a huge step. If you do feature caching unintelligible output from the machine learning network, rotate around the dimensions a little bit, you can
Pete Bannon
get up to 20% increase in efficiency of storage.
Tesla Presenter
Then when you store that we also order the columns by size so that all your small columns and small values are together, so that when you seek for a single value, you're likely to overlap with the read on more values which you'll use later so that you don't need to do another file system operation. So I could go on and on, I just went on touched on two projects that we have internally. But this is actually part of a huge continuous effort to optimize the compute that we have in house. So accumulating and aggregating through all these optimizations, we now train our occupancy networks twice as fast, just because it's twice as efficient. And and now if we add in bunch more compute and go parallel, we can now train this in hours instead of days. And with that, I'd like to hand it off to the biggest user of compute. John.
John Emmons
Hi everybody, my name is John Emmons. I lead the Autopilot vision team. I'm going to cover two topics with you today. The first is how we predict lanes. And the second is how we predict the future behavior of other agents on the road. In the early days of Autopilot, we modeled the lane detection problem as an image based instance segmentation task. Our network was super simple though. In fact, it was only capable of printing lanes from a few different kinds of geometries. Specifically, it would segment the eagle lane, it could segment adjacent lanes, and then it had some special casing for forks and merges. This simplistic modeling of the problem worked for highly structured roads like highways. But today we're trying to build a system that's capable of much more complex maneuvers. Specifically, we want to make left and right turns at intersections where the road topology can be quite a bit more complex and diverse. When we try to apply this simplistic modeling of the problem here, it just totally breaks down. Taking a step back for a moment, what we're trying to do here is to predict the sparse set of lane instances and their connectivity. And what we want to do is to have a neural network that basically predicts this graph where the nodes are the lane segments and the edges encode the connectivities between these lanes. So what we have is our lane detection neural network. It's made up of three components. In the first component, we have a set of convolutional layers, attention layers, and other neural network layers that encode the video streams from our eight cameras on the vehicle and produce a rich visual representation. We then enhance this visual representation with a coarse road road level map data, which we encode with a set of additional neural network layers that we call the lane guidance module. This map is not an HD map, but it provides a lot of useful hints about the topology of lanes inside of intersections. The lane counts on various roads, and a set of other attributes that help us. The first two components here produce a dense tensor that sort of encodes the world. But what we really want to do is to convert this dense tensor into a smart set of lanes and their connectivities. We approach this problem like an image captioning task where the input is this dense tensor and the output text. It's predicted into a special language that we developed at Tesla for encoding lanes in their connectivity. In this language of lanes, the words and tokens are the lane positions in 3D space in the ordering of the tokens. Interdicted modifiers in the tokens encode the connective relationships between these lanes. By modeling the task as a language problem, we can capitalize on recent auto regressive architectures and techniques from the language community for handling the multiple dality of the problem. We're not just solving the computer vision problem at Autopilot, we're also applying the state of the art and language modeling to machine learning more generally. I'm now going to dive into a little bit more detail this language component. What I have depicted on the screen here is a satellite image which sort of represents the local area around the vehicle. The set of nodes and edges is what we refer to as the lane graph. And it's ultimately what we want to come out of this neural network. We start with a blank slate. We're going to want to make our first prediction here at this green dot. This green dot's position is encoded as an index into a coarse grid which discretizes the 3D world. Now, we don't predict this index directly because it would be too computationally expensive to do so. There's just too many grid points. And predicting a categorical distribution over this has both implications at training time and test time. So instead what we do is we discretize the world coarsely. First, we predict a heat map over the possible locations, and then we latch in the most probable location condition on this. We then refine the prediction and get the precise point. Now we know where the position of this token is, but we don't know its type. In this case, though, it's the beginning of a new lane. So we predict it as a start token. And because it's a start token, there's no additional attributes in our language. We then take the Predictions from this first forward path, and we encode them using a learned conditional embedding, which produces a set of tensors that we combine together, which is actually the first word in our language of lanes. We add this to the first position in our sentence here. We then continue this process by predicting the next lane point in a similar fashion. Now, this lane point is not the beginning of a new lane. It's actually a continuation of the previous lane. So it's a continuation token type. Now, it's not enough just to know that this lane is connected to the previously predicted lane. We want to encode its precise geometry, which we do by regressing a set of spline coefficients. We then take this lane, we encode it again and add it as the next word in the sentence. We continue predicting these continuation lanes until we get to the end of the prediction grid. We then move on to a different lane segment. So you can see that cyan dot there. Now it's not topologically connected to that pink point. It's actually forking off of that green point there. So it's got a fork type. And fork tokens actually point back to previous tokens from which the fork originates. So you can see here the fork point predictor is actually the index zero. So it's actually referencing back to tokens that it's already predicted. Like you would in language, we continue this process over and over again until we've enumerated all of the tokens in the lane graph and then the network predicts the end of sentence token.
Ashok Elluswamy
Yeah, I just wanted to note that the reason we do this is not just because we want to build something complicated. It almost feels like a Turing complete machine. Here with neural networks though, is that we tried simpler approaches, for example, trying to just segment the lanes along the road or something like that. But then the problem is when there's uncertainty, say you cannot see the road clearly and there could be two lanes or three lanes and you can't tell. A simple segmentation based approach would just draw both of them. It's kind of a 2.5 lane situation. And the post processing algorithm would hilariously fail when the predictions are such.
John Emmons
Yeah, and the problems don't end there. I mean, you need to predict these connective, like these connective lanes inside of intersections, which it's just not possible with the approach that Ashok's mentioning, which is why we had to upgrade to this sort of.
Ashok Elluswamy
Yeah. When it like overlaps like this segmentation would just go haywire. But even if you try very hard to, you Know, put them on separate layers. It's just a really hard problem. But language just offers a really nice framework for getting a sample from a posterior as opposed to, you know, trying to do all of this in post processing. But this doesn't actually stop for just autopilot, right, John? This can be used for Optimus.
John Emmons
Yeah. You know, I guess they wouldn't be called lanes, but you could imagine, you know, sort of in this, you know, stage here that you might have sort of paths that sort of, you know, encode the possible places that people could walk.
Ashok Elluswamy
Basically, if you're in a factory or in a home setting, you can just ask the robot, okay, please route to the kitchen or please route to some location in the factory. And then we predict a set of pathways that would go through the aisles. Take the robot and say, ok, this is how you get to the kitchen. It just really gives us a nice framework to model these different paths that simplify the navigation problem for the downstream planner.
John Emmons
All right, so ultimately, what we get from this lane detection network is a set of lanes and their connectivities, which comes directly from the network. There's no additional step here for sparsifying these, you know, dense predictions into sparse ones. This is just a direct, unfiltered output of the network. Okay, so I talked a little bit about lanes. I'm going to briefly touch on how we model and predict the future, future paths and other semantics on objects. So I'm just going to go really quickly through two examples. The video on the right here, we've got a car that's actually running a red light and turning in front of us. What we do to handle situations like this is we predict a set of short time horizon future trajectories on all objects. We can use these to anticipate the dangerous situation here and apply whatever braking and steering action is required to avoid a collision. In the video on the right, there's two vehicles in front of us. The one on the left lane is parked. Apparently it's being loaded, unloaded. I don't know why the driver decided to park there. But the important thing is that our neural network predicted that it was stopped, which is the red color there. The vehicle in the other lane, as you notice, also is stationary. But that one's obviously just waiting for that red light to turn green. So even though both objects are stationary and have zero velocity, it's the semantics that is really important here so that we don't get stuck behind that awkwardly parked car. Predicting all of these agent attributes presents some practical Problems when trying to build a real time system, we need to maximize the frame rate of our object section stack so that autopilot can quickly react to the changing environment. Every millisecond really matters here. To minimize the inference latency, our neural network is split into two phases. In the first phase, we identified locations in 3D space where agents exist. In the second stage, we then pull out tensors at those 3D locations, append it with additional data that's on the vehicle and then we, you know, do the rest of the processing. This sparsification step allows the neural network to focus compute on the areas that matter most, which gives us superior performance for a fraction of the latency cost. So putting it all together, the autopilot vision stack predicts more than just the geometry and kinematics of the world. It also predicts a rich set of semantics which enables safe and human like driving. I'm now going to hand things off to SRI who will tell us how we run all these cool neural networks on our FSD computer. Thank you.
Sri (Tesla)
Hi everyone, I'm sri. Today I'm going to give a glimpse of what it takes to run these FSC networks in the car and how do we optimize for the inference latency. Today I'm going to focus just on FSG lanes network that John just talked about. So when we started this track, we wanted to know if we can run this FSG lanes network natively on the TRIP engine, which is our in house neural network accelerator that we built in the FSD computer. When we built this hardware we kept it simple and we made sure it can do one thing ridiculously fast dense dot products. But this architecture is auto regressive and iterative where it crunches through multiple attention blocks in the inner loop producing sparse points directly at every step. So the challenge here was how can we do this sparse point prediction and sparse computation on a dense dot product engine. Let's see how we did this on the trip. So the network predicts the heat map of most probable spatial locations of the point. Now we do a ARGMAX and a one hot operation which gives the one hot encoding of the index of the spatial location. Now we need to select the embedding associated with this index from an embedding table that is learned during training. To do this on trip, we actually built a lookup table in SRAM and we engineered the dimensions of this embedding such that we could achieve all of this thing with just matrix multiplication. Not just that, we also wanted to store this embedding into a token cache so that we don't recompute this for every iteration, rather reuse it for future point prediction. Again, we pulled some tricks here where we did all these operations just on the dot product engine. It's actually cool that our team found creative ways to map all these operations on the trip engine in ways that were not even imagined when this hardware was designed. But that's not the only thing we had to do to make this work. We actually implemented a whole lot of operations and features to make this model compilable, to improve the intake accuracy as well as to optimize performance. All of these things helped us run this 75 million parameter model just under 10 millisecond of latency, consuming just 8 watts of power. But this is not the only architecture running in the car. There are so many other architectures, modules and networks we need to run in the car to give a sense of scale. There are about a billion parameters of all the networks combined, producing around 1,000 neural network signals. So we need to make sure we optimize them jointly and such that we maximize the compute utilization throughput and minimize the latency. So we built a compiler just for neural networks that shares the structure to traditional compilers. As you can see, it takes the massive graph of neural nets with 150k nodes and 375k connection, takes this thing, partitions them into independent sub graphs and compiles each of those sub graphs natively for the inference devices. Then we have a neural network linker which shares the structure to traditional linker where we perform this link time optimization. There we solve an offline optimization problem with compute memory and memory bandwidth constraints so that it comes with an optimized schedule that gets executed in the car on the runtime. We designed a hybrid scheduling system which basically does heterogeneous scheduling on one SoC and distributed scheduling across both the SoCs to run these networks in a model parallel fashion. To get 100 tops of compute utilization, we need to optimize across all all the layers of software, right from tuning the network architecture, the compiler, all the way to implementing a low latency high bandwidth RDMA link across both the socs. And in fact going even deeper to understanding and optimizing the cache coherent and non coherent data parts of the accelerator in the SoC. This is a lot of optimization at every level in order to make sure we get the highest frame rate and as every millisecond counts here. And this is just the, this is the visualization of the neural networks that are running in the car. This is our digital brain essentially. As you can see, these operations are nothing but just the matrix multiplication convolution to name a few real operations running in the car. To train this network with a billion parameters, you need a lot of labeled data. So Egan is going to talk about how do we achieve this with the auto labeling pipeline.
Tesla Presenter
Thank you sir. Hi everyone, I'm Yeagan Zhang and I'm leading geometric vision at Autopilot. So yeah, let's talk about auto labeling. So we have several kinds of auto labeling frameworks to support various types of networks. But today I'd like to focus on the awesome lanesnet here. So to successfully train and generalize this network to everywhere we think we went tens of millions of trips from probably 1 million intersection or even more. So then how to do that. So it is certainly achievable to source sufficient amount of trips because we already have, as Tim explained earlier, we already have like 500,000 trips per day cache rate. However, combining all those data into a training form is a very challenging technical problem. To solve this challenge, we've tried various ways of manual and auto labeling. So from the first column to the second, from the second to the third, each advance provided us nearly 100x improvement in throughput. But still we want an even better auto labeling machine that can provide us provide us good quality, diversity and scalability. To meet all these requirements, despite the huge amount of engineering effort required here, we've developed a new auto labeling machine powered by multi trip reconstruction. So this can replace 5 million hours of manual labeling with just 12 hours on cluster for labeling 10,000 trips. So how we solved there are three big steps. The first step is high precision trajectory and structure recovery by multi camera visual inertial odometry. So here all the features including ground surface are inferred from videos by neural networks, then tracked and reconstructed in the vector space. So the typical drift rate of this trajectory in car is like 1.3 centimeter per meter and 0.45 milli region per meter, which is pretty decent considering its compact compute requirement. Then the recovery service and road details are also used as a strong guidance for the later manual verification stuff. This is also enabled in every FSD vehicle. So we get pre processed trajectories and structures along with the trip data. The second step is multi tool reconstruction which is the big and core piece of this machine. So the video shows how the previously shown trip is reconstructed and aligned with other trips, basically other trips from different vehicle, not the same vehicle. So this is done by multiple Internal steps like coarse alignment, pairwise matching, joint optimization, then further surface refinement. In the end, the human analyst comes in and finalizes the label. So each habit steps are already fully parallelized on the cluster. So the entire process usually takes just a couple of hours. The last step is actually auto labeling the new trips. So here we use the same multi trip alignment engine, but only between pre built reconstruction and each new trip. So it's much, much simpler than fully reconstructing all the clips altogether. That's why it only takes 30 minutes per trip to auto label instead of several hours of manual labeling. And this is also the key of scalability of this machine. This machine easily scales as long as we have available compute and trip Data. So about 50 trips were newly auto labeled from this scene and some of them are shown here. So 53 from different vehicles. So this is how we capture and transform the space time slices of the world into the network.
Ashok Elluswamy
Supervision. Yeah. One thing I'd like to note is that Yeagen just talked about how we auto label our lens. But we have auto labels for almost every task that we do, including our planner. And many of these are fully automatic. There's no humans involved. For example, for objects, all of the kinematics, the shapes, their futures, everything just comes from auto labeling. And the same is true for occupancy too. And we have really just built a machine around this.
Tesla Presenter
Yeah. So if you can go back one,
Tesla Presenter
one more, it says parallelized on cluster. So that sounds pretty straightforward, but it really wasn't. Maybe it's fun to share how something like this comes about. So a while ago we didn't have any auto labeling at all. And then someone makes a script, it starts to work, it starts working better until we reach a volume that's pretty high and we clearly need a solution. And so there were two other engineers in our team who were like, you know, that's an interesting, you know, thing. What we needed to do was build a whole graph of essentially Python functions that would need to run one after the other. First you pull the clip, then you do some cleaning, then you do some network inference, then another network inference. And until you finally get this. But so you need to do this at a large scale. So I tell them we probably need to shoot for 100,000 clips per day or 100,000 items. That seems good. And so the engineer said, well, we can do a bit of postgres and a bit of elbow grease, we can do it. Meanwhile, we are a bit later and we're doing 20 million of these functions every single day. Again, we pull in around half a million clips. And on those we run a ton of functions, each of these in a streaming fashion. And so that's kind of the backend infra that's also needed to not just run training, but also auto labeling.
Ashok Elluswamy
Yeah, it really is like a factory that produces labels and like production lines yield quality, inventory. Like all of the same concepts apply to this label factory. That applies for, you know, the factory for our cars.
Tesla Presenter
That's right. Okay, thanks Tim and Ashok. So yeah, so concluding this section, I'd like to share a few more challenging and interesting examples for network for sure. And even for humans probably. So from the top there's like examples for like lack of lights case or foggy night or roundabout, and occlusions by heavy occlusions by parked cars and even rainy night with raindrops on camera lenses. These are challenging, but once their original scenes are fully reconstructed by other clips, all of them can be auto labeled so that our cars can drive even better through these challenging scenarios. So now let me pass the mic to David to learn more about how SIM is creating the new world on top of these labels. Thank you.
David (Tesla)
Thank you, Yagan. My name is David and I'm going to talk about simulation. So simulation plays a critical role in providing data that is difficult to source and or hard to label. However, 3D scenes are notoriously slow to produce. Take for example the simulated scene playing behind me, a complex intersection from Market street in San Francisco. It would take two weeks for artists to complete and for us, that is painfully slow. However, I'm going to talk about using Yagen's automated ground truth labels, along with some brand new tooling that allows us to procedurally generate this scene in many like it in just five minutes. That's an amazing a thousand times faster than before. So let's dive in to how a scene like this is created. We start by piping the automated ground truth labels into our simulated world creator tooling inside the software Houdini. Starting with road boundary labels, we can generate a solid road mesh and retopologize it with the lane graph labels. This helps inform important road details like crossroad slope and detailed material blending. Next, we can use the line data and sweep geometry across its surface and project it to the road, creating lane paint decals. Next, using median edges, we can spawn island geometry and populate it with randomized foliage. This drastically changes the visibility of the scene. Now the outside world can be generated through a series of randomized heuristics modular building generators create visual obstructions while randomly placed objects like hydrants can change the color of the curbs, while trees can drop leaves below it, obscuring lines or edges. Next, we can bring in map data to inform positions of things like traffic lights or stop signs. We can trace along its normal to collect important information like number of lanes and even get accurate street names on the signs themselves. Next, using lane graph, we can determine lane connectivity and spawn directional road markings on the road and their accompanying road signs. And finally, with lane graph itself, we can determine lane adjacency and other useful metrics to spawn randomized traffic permutations inside our simulator. And again, this is all automatic. No artists in the loop and happens within minutes. And now this sets us up to do some pretty cool things. Since everything is based on data and heuristics, we can start to fuzz parameters to create visual variations of of the single ground truth. It can be as subtle as object placement and random material swapping to more drastic changes like entirely new biomes or locations of environment like urban, suburban or rural. This allows us to create infinite targeted permutations for specific ground truths that we need more ground truth for. And all this happens within a click of a button. And we can even take this one step further by altering our ground truth itself. Say John wants his network to pay more attention to directional road markings to better detect an upcoming captive left turn lane. We can start to procedurally alter our lane graph inside the simulator to help focus to create entirely new flows through this intersection to help focus the network's attention to the road markings to create more accurate predictions. And this is a great example of how this tooling allows us to create new data that can never be collected from the real world. And the true power of this tool is in its architecture and how we can run all tasks in parallel to infinitely scale. So you saw the tile creator tool in action, converting the ground truth labels into their counterparts. Next we can use our tile extractor tool to divide this data into into geohash tiles about 150 meters square in size. We then save out that data into separate geometry and instance files. This gives us a clean source of data that's easy to load and allows us to be rendering engine agnostic for the future. Then using a tile loader tool, we can summon any number of those cache tiles using a geohash ID. Currently we're doing about these 5 by 5 tiles or 3 by 3, usually centered around fleet hotspots or interesting lane graph locations. And the tile Loader also converts these tile sets into uassets for consumption by the Unreal Engine and gives you a finished product from what you saw on the first slide. And this really sets us up for size and scale. And as you can see on the map behind us, we can easily generate most of San Francisco city streets. And this didn't take years or even months of work, but rather two weeks by one person. We can continue to manage and grow all this data using our PDG network inside of the tooling. This allows us to throw compute at it and regenerate all these tile sets overnight. This ensures all environments are of consistent quality and features, which is super important for training since new ontologies and signals are constantly released. And now to come full circle, because we generated all these tile sets from ground truth data that contain all the weird intricacies from the real world. And we can combine that with the procedural, visual and traffic variety to create limitless targeted data for the network to learn from. And that concludes the sim section. I'll pass it to Kate to talk about how we can use all this data to improve autopilot. Thank you.
Kate Park (Tesla)
Thanks, David. Hi everyone, My name is Kate park and I'm here to talk about the data engine, which is the process by which we improve our neural networks via data. We're going to show you how we deterministically solve interventions via data and walk you through the life of this particular clip. In this scenario, autopilot is approaching a turn and incorrectly predicts that crossing vehicle as stopped for traffic and thus a vehicle that we would slow down for. In reality, there's nobody in the car, it's just awkwardly parked. We built this tooling to identify the mispredictions, correct the label, and categorize this clip into an evaluation set. This particular clip happens to be one of 126 that we've diagnosed as challenging parked cars at turns. Because of this infra, we can curate this evaluation set without any engineering resources custom to this particular challenge case. To actually solve that challenge case requires mining thousands of examples like it, and it's something Tesla can trivially do. We simply use our data sourcing infrastructure, request data, and use the tooling shown previously to correct the labels. By surgically targeting the mispredictions of the current model, we're only adding the most valuable examples to our training set. We surgically fixed 13,900 clips, and because those were examples where the current model struggles, we don't even need to change the model architecture. A simple weight update with this new valuable Data is enough to solve the chall challenge case. So you see, we no longer predict that crossing vehicle as stopped as shown in orange, but parked as shown in red. In academia, we often see that people keep data constant, but at Tesla, it's very much the opposite. We see time and time and again that data is one of the best, if not the most deterministic, lever to solving these interventions. We just showed you the data engine loop for one challenge case, namely these park cars at turns. But there are many challenge cases, even for one signal of vehicle movement. We apply this data engine loop to every single challenge case we've diagnosed, whether it's buses, curvy roads, stopped vehicles, parking lots. And we don't just add data once. We do this again and again to perfect the semantic. In fact, this year, we updated our vehicle movement signal five times. And with every weight update trained on the new data, we push our vehicle movement accuracy up and up. This data engine framework applies to all our signals, whether they're 3D multicam video, whether the data is human labeled, auto labeled or simulated, whether it's an offline model or an online model. And Tesla's able to do this at scale because of the fleet advantage, the infra that our ENG team has built, the labeling resources that feed our networks. To train on all this data, we need a massive amount of compute. So I'll hand it off to Pete and Ganesh to talk about the Dojo supercomputing platform. Thank you.
Tesla Presenter
Thank you, Katie.
Pete Bannon
Thanks everybody. Thanks for hanging in there. We're almost there. My name is Pete Bannon. I run the custom silicon and low voltage teams at Tesla.
Tesla Presenter
And my name is Ganesh Venkat.
Elon Musk
I run the Doja program.
Tesla Presenter
Thank you.
Pete Bannon
I'm frequently asked, why is a car company building a supercomputer for training? And this question fundamentally misunderstands the the nature of Tesla. At its heart, Tesla is a hardcore technology company. All across the company, people are working hard in science and engineering to advance the fundamental understanding and methods that we have available to build cars, energy solutions, robots, and anything else that we can do to improve the human condition around the world.
Pete Bannon
It's a super exciting thing to be a part of and it's a privilege to run a very small piece of it in the semiconductor group. Tonight we're going to talk a little bit about Dojo and give you an update on what we've been able to do over the last year. But before we do that, I wanted to give a little bit of background on the initial design that we started a few years ago. When we got started, the goal was to provide a substantial improvement to the training latency for our autopilot team. Some of the largest neural networks they train today run for over a month, which inhibits their ability to rapidly explore alternatives and evaluate them. So, you know, a 30x speed up would be really nice if we could provide it at a cost competitive and energy competitive way to do that. We wanted to build a chip with a lot of arithmetic units that, that we could utilize at a very high efficiency. And we spent a lot of time studying whether we could do that using dram, various packaging ideas, all of which failed. And in the end, even though it felt like an unnatural act, we decided to reject DRAM as the primary storage medium for this system and instead focus on SRAM embedded in the chip. SRAM provides unfortunately, a modest amount of capacity, but extremely, extremely high bandwidth and very low latency. And that enables us to achieve high utilization with the arithmetic units. Those choices. That particular choice led to a whole bunch of other choices. For example, if you want to have virtual memory, you need page tables. They take up a lot of space. We didn't have space, so no virtual memory. We also don't have interrupts. The accelerator is a bare bonds raw piece of hardware that's presented to a compiler, and the compiler is responsible for scheduling everything that happens in a deterministic way. So there's no need or even desire for interrupts in the system. We also chose to pursue model parallelism as a training methodology, which is not the typical situation. Most machines today use data parallelism, which consumes additional memory capacity capacity, which we obviously don't have. So all of those choices led us to build a machine that is pretty radically different from what's available today. We also had a whole bunch of other goals. One of the most important ones was no limits. So we wanted to build a compute fabric that would scale in an unbounded way for the most part. I mean, obviously there's physical limits now and then, but you know, pretty much if your model was 2 too big for the computer, you just have to go buy a bigger computer. That's what we were looking for today. The way package machines are packaged, there's a pretty fixed ratio of, for example, GPUs, CPUs and DRAM capacity and network capacity. And we really wanted to disaggregate all that so that as models evolved, we could vary the ratios of those various elements and make the system more flexible to meet the needs needs of the autopilot.
Tesla Presenter
Team. Yeah, and it's so true, Pete, like
Elon Musk
no limits philosophy was our guiding star all the way.
Felix (Tesla)
All of our choices were centered around
Tesla Presenter
that and to the point that we
Elon Musk
didn't want traditional data center infrastructure to
Felix (Tesla)
limit our capacity to execute these programs at speed.
Paril Jain
So that's why we. That's why we.
Elon Musk
Sorry about that. That's why we integrated vertically our data
Felix (Tesla)
center, the entire data center.
Elon Musk
By doing a vertical integration of the data center, we could extract new levels of efficiency. We could optimize power delivery, cooling, and
Felix (Tesla)
as well as system management across the
Elon Musk
whole data center stack.
Elon Musk
Rather than doing box by box and integrating those boxes into data centers.
Tesla Presenter
And to do this, we also wanted
Elon Musk
to integrate early to figure out limits of scale for our software workloads. So we integrated Dojo environment into our Autopilot software very early and we learned a lot of lessons. And today, Bill Chang will go over our hardware update as well as some of the challenges that we faced along the way. And Rajiv Kurian will give you a glimpse of our compiler technology as well
Felix (Tesla)
as go over some of our cool results.
Tesla Presenter
Thanks, Pete. Thanks, Ganesh. I'll start tonight with a high level vision of our system that will help set the stage for the challenges and the problems we're solving, and then also how software will then leverage this for performance. Now, our vision for Dojo is to build a single unified accelerator, a very large one. Software would see a seamless compute plane with globally addressable, very fast memory, and all connected together with uniform high bandwidth and low latency. Now, to realize this, we need to use density to achieve performance. Now we leverage technology to get this density in order to break levels of hierarchy all the way from the chip to the scale out system. Now, silicon technology has, has used this, has done this for decades. Chips had followed Moore's law for density integration to get performance scaling. Now, a key step in realizing that vision was our training tile. Not only can we integrate 25 dies at extremely high bandwidth, but we can scale that to any number of additional tiles by just connecting them together. Now, last year we showcased our first functional training tile, and at that time we already had workloads running on it. And since then, the team here has been working hard and diligently to deploy this at scale. Now, we've made amazing progress and had a lot of milestones along the way, and of course we've had a lot of unexpected challenges. But this is where our fail fast philosophy has allowed us to push our boundaries. Now, pushing density for performance presents all new challenges. One Area is power delivery. Here we need to deliver the power to our compute die. And this directly impacts our top line compute performance. But we need to do this at unprecedented density. We need to be able to match our die pitch with a power density of almost 1amp per millimeter squared. And because of the extreme integration, this needs to be a multi tiered vertical power solution. And because there's a complex heterogeneous material stack up, we have to carefully manage the material transition, especially cte. Now why does the coefficient of thermal expansion matter in this case? CTE is a fundamental material property, and if it's not carefully managed, that stack up would literally rip itself apart. So we started this effort by working with vendors to deliver to develop this power solution. But we realized that we actually had to develop this in house. Now, to balance schedule and risk, we built quick iterations to support both our system bring up and software development, and also to find the optimal design and stack up that would meet our final production goals. And in the end, we were able to reduce CTE over 50% and meet our performance by by 3x over our initial version. Now, needless to say, finding this optimal material stackup while maximizing performance at density is extremely difficult. Now, we did have unexpected challenges along the way. Here's an example where we pushed the boundaries of integration that led to component failures. This started when we scaled up to larger and longer workloads. And then intermittently, a single site on a tile would fail. Now, they started out as recoverable failures, but as we pushed some much higher and higher power, these would become permanent failures. Now to understand this failure, you have to understand why and how we build our power modules. Solving density at every level is the cornerstone of actually achieving our system performance. Now, because our XY plane is used for high bandwidth communication, everything else must be stacked vertically. This means all other components other than our die must be integrated into our power modules. Now that includes our clock, our our power supplies, and also our system controllers. Now in this case, the failures were due to losing clock output from our oscillators. And after an extensive debug, we found that the root cause was due to vibrations on the module from piezoelectric effects on nearby capacitors. Now, singing caps are not a new phenomenon and and in fact very common in power design. But normally clock chips are placed in a very quiet area of the board and often not affected by power circuits. But because we needed to achieve this level of integration, these oscillators need to be placed in very close proximity. Now, due to Our switching frequency and then the vibration resonance created. It caused out of plane vibration on our MEMS oscillator that caused it to crack. Now, the solution to this problem is a multi prong approach. We can reduce the vibration by using soft terminal caps. We can update our MEMS part with a lower Q factor for the out of plane direction. And we can also update our switching frequency to push the resonance further away from these sensitive bands. Now, addition to the density at the system level, we've been making a lot of progress at the infrastructure level. We knew that we had to re examine every aspect of the data center infrastructure in order to support our unprecedented power and cooling density. We brought in a fully custom designed CDU to support Dojo's dense cooling requirements. And the amazing part is we're able to do this at a fraction of the cost versus buying off the shelf and modifying it. And since our Dojo cabinet integrates enough power and cooling to match an entire row of standard IT racks, we need to carefully design our cabinet and infrastructure together. And we've already gone through several iterations of this cabinet to optimize this. And earlier this year we started load testing our power and cooling infrastructure. And we were able to push it over 2 megawatts before we tripped our substation and got a call from the city. Now, last year we introduced only a couple of components of our system. The custom D1 die and the training tile. But we tease the exit pod as our end goal. We'll walk through the remaining parts of our system that are required to build out this exit pod. Now, the system tray is a key part of realizing our vision of a single accelerator. It enables us to seamlessly connect tiles together, not only within the cabinet but between cabinets. We can connect these tiles at very tight space spacing across the entire accelerator. And this is how we achieve our uniform communication. This is a laminated bus bar that allows us to integrate very high power, mechanical and thermal support and an extremely dense integration. It's 75 millimeters in height and supports six tiles at 135 kilograms. This is the equivalent of three to four fully loaded high performance racks. Next, we need to feed data to the training tiles. This is where we've developed the Dojo interface processor. It provides our system with high bandwidth DRAM to stage our training data. And it provides full memory bandwidth to our training tiles using ttp, our custom protocol that we use to create communicate across our entire accelerator. It also has high speed ethernet that helps us extend this custom protocol over standard ethernet. And we provide Native hardware support for this with little to no software overhead. And lastly, we can connect to it through a standard gen 4 PCIe interface. Now we pair 20 of these cards per tray and that gives us 640 gigabytes of high bandwidth DRAM. And this provides our disaggregated memory layer for our training tiles. These cards are our high bandwidth ingest path both through PCIe and Ethernet. They also provide a high Radex Z connectivity path that allows shortcuts across our large dojo accelerator. Now we actually integrate the host directly underneath our system tray. These hosts provide our ingest processing and connect to our Interface processors through PCIe. These hosts can provide hardware video decoder support for video based training. And our user applications land on these hosts so we can provide them with the standard x86 Linux environment. Now we can put two of these assemblies into one cabinet and pair it with redundant power supplies that do direct conversion of 3 phase 480 volt AC power to 52 volt DC power. Now, by focusing on density at every level, we can realize the vision of a single accelerator. Now, starting with the uniform nodes on our custom D1 die, we can connect them together in our fully integrated training tile and then finally seamlessly connecting them across cabinet boundaries to form our dojo accelerator. And altogether we can house two full accelerators in our exit pod for a combined one. One exaflop of ML Compute. Now altogether. This amount of technology and integration has only ever been done a couple of times in the history of compute. Next we'll see how software can leverage this to accelerate their performance.
Eric (Tesla)
Thanks, Bill. My name is Rajeev and I'm going to talk some numbers. So our software stack begins with the Pytorch extension. That speaks to our commitment to run standard Pytorch models out of the box. We're going to talk more about our JIT compiler and the ingest pipeline that feeds the hardware with data. Abstractly, performance is tops times utilization times accelerator occupancy. We've seen seen how the hardware provides peak performance. It's the job of the compiler to extract utilization from the hardware while code is running on it. And it's the job of the ingest pipeline to make sure that data can be fed at a throughput high enough for the hardware to not ever starve. So let's talk about why communication bound models are difficult to scale. But before that, let's look at why Resnet50 like models are easier to scale. We start off with a single accelerator, run the forward and backward Passes followed by the optimizer. Then to scale this up, you run multiple copies of this on multiple accelerators. And while the gradients produced by the backward pass do need to be reduced and this introduces some communication, this can be done pipelined with the backward pass. This setup scales fairly well almost linearly. For models with much larger activations, we run into a problem as soon as we want to run the forward pass. The batch size that fits in a single accelerator is often smaller than the batch norm surface. So to get around this, researchers typically run the setup on multiple accelerators in sync batch norm mode. This introduces latency bound communication to the critical path of the forward pass and we already have a communication bottleneck. And while there are ways to get around this, they usually involve tedious manual will work best suited for a compiler. And ultimately there's no skirting around the fact that if your state does not fit in a single accelerator, you can be communication bound. And even with significant efforts from our ML engineers, we see such models don't scale linearly. The Dojo system was built to make such models work at high utilization. The high density integration is was built to not only accelerate the compute bound portions of a model, but also the latency bound portions like a batch norm or the bandwidth bound portions like a gradient all reduce or a parameter all gather. A slice of the Dojo mesh can be carved out to run any model. The only thing you just need to do is to make the slice large enough to fit a batch room surface for their particular model. After that, the partition presents itself as one large accelerator, freeing the users from having to worry about the internal details of execution, and it's the job of the compiler to maintain this abstraction. Fine grained synchronization primitives and uniform low latency makes it easy to accelerate all forms of parallelism across integration boundaries. Tensors are usually stored sharded in SRAM and replicated just in time for layers execution. We depend on the high Dojo bandwidth to hide this replication time. Tensor replication and other data transfers are overlapped with compute and the compiler can also recompute layers. When it's profitable to do so,
Eric (Tesla)
expect most models to work out of the box. As an example, we took the recently released stable diffusion model and got it running on Dojo in minutes I out of the box. The compiler was able to map it in a model parallel manner on 25 dojo dies. Here are some pictures of a cybertruck on Mars generated by stable diffusion running on Dojo. Looks like it still has some ways to Go before matching the Tesla Design Studio team. So we've talked about how communication bottlenecks can hamper scalability. Perhaps an acid test of a compiler. And the underlying hardware is executing a cross type batch norm layer. Like mentioned before, this can be a serial bottleneck. The communication phase of a batch norm begins with nodes computing their local mean and standard deviations, then coordinating to reduce these values, then broadcasting these values back, and then they resume their work in parallel.
Eric (Tesla)
So what would an ideal batch form look like on 25 dojo dice? Let's say the previous list activations are already split across dice. We would expect the 350 nodes on each die to coordinate and produce die local mean and standard deviation values. Ideally, these would get further reduced with the final value ending somewhere towards the middle of the tile. We would then hope to see a broadcast of this value radiating from the center. Let's see how the compiler actually executes a real batch drum operation across 25 dice. The communication trees were extracted from the compiler and the timing is from a real hardware run. We're about to see 8,750 nodes on 25 dies coordinating to reduce and then broadcast the Bastrum mean and standard deviation values die local reduction followed by global reduction towards the middle of the tie. Then the reduced value broadcast radiating from the middle, accelerated by the hardware's broadcast facility. This operation takes only 5 microseconds on 25 dojo dice. The same operation takes 150 microseconds on 24 GPUs. This is an orders of magnitude improvement over GPUs. And while we talked about an all reduced operation in the context of a batch norm, it's important to reiterate that the same advantages apply to all other communication primitives. And these primitives are essential for large scale training. So how about full model performance? So while we think that Resnet50 is not a good representation of real world Tesla workloads, it is a standard benchmark. So let's start there. We are already able to match the A100 die for die. However, perhaps a hint of Dojo's capabilities is that we're able to hit this number with just a batch of eight per die. But Dojo was really built to tackle larger complex models. So when we set out to tackle real world workloads, we looked at the usage patterns of our current GPU cluster and the two models stood out. The auto labeling networks, a class of offline models that are used to generate ground truth, and the occupancy networks that you heard about the Auto labeling networks are large models that have high arithmetic intensity, while the occupancy networks can be ingest bound. We chose these models because together they account for a large chunk of our current GPU cluster usage. And they would challenge the system in different ways. So how did we do on these two networks? The results we're about to see were measured on multi die systems for both the GPU and Dojo, but normalized to per die numbers. On our auto labeling network, we're already able to surpass the performance of an A100 with our current hardware running on our older generation VRMs. On our production hardware with our new VRMs, that translates to doubling the throughput of of an A100. And our model showed that with some key compiler optimizations, we could get to more than 3x the performance of an A100. We see even bigger leaps on the occupancy network, almost 3x with our production hardware with room for more. So what does that mean for Tesla? With the current level of compiler performance? We could replace the ML compute of 1, 2, 3, 4, 5 and 6 GPU boxes with just a single Dojo tile. And this Dojo tile costs less than 1 of these GPU boxes. What it really means is that networks that took more than a month to train now take less than a week. Alas, when we measured things, it did not turn out so well. At the pytorch level, we did not see our expected performance out of the gate. And this timeline chart shows our problem. The teeny tiny little green bars, that's the compile code running on the accelerator. The row is mostly white space where the hardware is just waiting for data. With our dense ML compute, Dojo hosts effectively have 10x more ML compute than the GPU host. The data loaders running on this one host simply couldn't keep up with all that ML hardware. So to solve our data loader scalability issues, we knew we had to get over the limit of this single host. The Tesla transport protocol moves data seamlessly across host tiles and ingest processors. So we extended the Tesla transport protocol to work over Ethernet. We then built the Dojo network interface card, the dnic, to leverage TTP over Ethernet. This allows any host with a DNIC card to be able to DMA to and from other T TTP endpoints. So we started with the Dojo mesh. Then we added a tier of data loading hosts equipped with the DNIC card. We connected these hosts to the mesh via an Ethernet switch. Now every host in this data loading tier is capable of Reaching all TTP endpoints in the dojo mesh via hardware accelerated DMA. After these optimization, our occupancy went from 4% to 97%. So the data loading sections have reduced, The data loading sections have reduced drastically and the ML hardware is kept busy. We actually expect this number to go to 100% pretty soon after these changes went in. We saw the full expected speed up from the Pytorch layer and we were back in business. So we started with hardware design that breaks through traditional integration boundaries in service of our vision of a single giant accelerator. We've seen how the compiler and ingest layers build on top of that hardware. So after proving our performance on these complex real world networks, we knew what our first large scale deployment would target. Our high arithmetic intensity auto labeling network today that occupies 4,000 GPUs, over 72 GPU racks. With our dense compute and our high performance, we expect to provide the same throughput with just four dojo cabinets. And these four dojo cabinets will be part of our first exopod that we plan to build by quarter one of 2023. This will more than double Tesla's auto labeling capacity. The first exapot is part of a total of seven exit parts that we plan to build in Palo Alto. Right here across the wall. And we have a display cabinet from one of these exapods here for everyone to look at. Six tiles densely packed on a tray. 54 petaflops of compute, 640 gigabytes of high bandwidth memory with power and host defeated a lot of compute. And we're building out new versions of all our cluster components and constantly improving our software to hit new limits of scale. We believe that we can get another 10x improvement with our next generation hardware. And to realize our ambitious goals, we need the best software and hardware engineers. So please come talk to us or visit tesla.comai thank you.
Elon Musk
All right, so hopefully that was enough detail and now we can move to questions. And guys, I think the team come out on stage and we really wanted to show the depth and breadth of Tesla in artificial intelligence, compute, hardware, robotics actuators, and try to really shift the perception of the company away from. A lot of people think we're like just a car company or we make cool cars, whatever, but they don't have. Most people have no idea that Tesla is arguably the leader in real world AI hardware and software and that we're building what is arguably the first, the most radical computer architecture since the crayon supercomputer computer. And I Think if you're interested in developing some, some of the most advanced technology in the world that's going to really affect the world in a positive way, Tesla's the place to be. So, yeah, let's fire away with some questions. I think there's, there's a mic at the front and a mic at the back.
Kate Park (Tesla)
On this side.
Elon Musk
Just throw mics at people, jump. All for the mic.
Elon Musk
Thank you very much.
Ashok Elluswamy
I was impressed here.
Elon Musk
Yeah, I was impressed very much by Optimus.
Ashok Elluswamy
But I wonder why tendon driven the hand?
Elon Musk
Why did you choose a tendon driven approach for the hand?
Tesla Presenter
Because tendons are not very durable.
Elon Musk
And why spring loading?
Tesla Presenter
Lou, is this cool?
Felix (Tesla)
Awesome. Yes, that's a great question. You know, when it comes to any type of actuation scheme, there's trade offs between, you know, whether or not it's a tendon driven system or some type of linkage based system.
Elon Musk
Just keep the mic close to your
Tesla Presenter
mouth, a little bit closer.
Tesla Presenter
Hear me. Cool.
Felix (Tesla)
So, yeah, the main reason why we went for a tendon based system is that, you know, first we actually investigated some synthetic tendons, but we found that metallic boating cables are, you know, a lot stronger. One of the advantages of these cables is that it's very good for part reduction. We do want to make a lot of these hands. So having a bunch of parts, a bunch of small linkages ends up being, you know, a problem when you're making a lot of something. One of the big reasons that, you know, tendons are better than linkages in a sense is that you can be anti backlash. So anti backlash essentially, you know, allows you to not have any gaps or, you know, stuttery motion in your fingers. Spring loaded, mainly what spring loaded allows us to do is allows us to have active opening. So instead of having to have two actuators to drive the fingers closed and then open, we have the ability to, to have the tendon drive them close and then the springs passively extend. And this is something that's seen in our hands as well.
Felix (Tesla)
We have the ability to actively flex and then we also have the ability to extend.
Elon Musk
Our goal with Optimus is to have a robot that is maximally useful as quickly as possible. So there's a lot of ways to solve the various problems of a humanoid robot and probably not barking up the right tree on all the technical solutions. And I should say that we're open to evolving the technical solutions that you see here over time. They're not locked in stone, but we have to pick Something, and we want to pick something that's going to allow us to produce the robot as quickly as possible and have it, like I said, be useful as quickly as possible. We're trying to follow the. The goal of fastest path to a useful robot that can be made at volume. And we're going to test the robot internally at Tesla in our factory and just see like, how useful is it? Because you have to have a. You're going to close the loop on reality to confirm that the robot is in fact useful. And. Yeah, so we're just going to use it to build things and we're confident we can do that with the hand that we have currently designed. But for sure there'll be hand version 2, version 3, and we may change the architecture quite significantly over time.
Lizzie Miskovetz
Hi. The Optimus robot is really impressive. You did a great job. Bipedal robots are really difficult. But what I noticed is, might be missing from your plan is to acknowledge the utility of the human spirit. And I'm wondering if Optimus will ever get a personality and be able to laugh at our jokes while they. While it folds our clothes.
Elon Musk
Yeah, absolutely. I think we want to have really fun versions of Optimus and so that Optimus can both to be utilitarian and do tasks, but can also be kind of like a friend and a buddy and hang out with you. And I'm sure people will think of all sorts of creative uses for this robot. And, you know, once you have the core intelligence and actuators figured out, then you can actually, you know, put all sorts of costumes, I guess, on the robot. I mean, you can make the robot
Elon Musk
you can skin the robot in many different ways and I'm sure people will find very interesting ways to. Yeah. Versions of Optimus. So. Thanks for the great presentation.
Felix (Tesla)
I wanted to know if there was an equivalent to interventions in Optimus.
John Emmons
It seems like labeling through moments where
Elon Musk
humans disagree with what's going on is important. And in a humanoid robot, that might
John Emmons
be also a desirable source of information.
Ashok Elluswamy
Yeah, I think we will have ways to remote operate the robot and intervene when it does something bad, especially when we are training the robot and bringing it up and hopefully we, you know, design it in a way that we can stop the robot from. If it's going to hit something, we can just like hold it and then it will stop. It won't like, you know, crush your hand or something. And those are all intervention data.
Ashok Elluswamy
And we can learn a lot from our simulation systems too, where we can check for collisions and supervise that those are bad actions.
Elon Musk
Yeah, I mean, so Optimus, we want over time for it to be, you know, an Android. The kind of Android that you've seen in sci fi movies like Star Trek, the Next Generation, like Data. But obviously we could program the robot to be less robot like and more friendly. And, you know, it can obviously learn to emulate humans and feel very natural. So as AI in general improves, we can add that to the robot and it should be obviously able to do simple instructions or even intuit what it is that you want. So you could give it a high level instruction and then it can break that down into a series of actions and take those actions.
Tesla Presenter
Hi. Yeah, it's exciting to think that with the Optimus you will think that you
Elon Musk
can achieve orders of magnitude of improvement in economic output. That's really exciting.
Tesla Presenter
And when Tesla started, the mission was
Elon Musk
to accelerate the advent of renewable energy or sustainable transport. So with the Optimus, do you still
Tesla Presenter
see that mission being the mission statement of Tesla or is it going to
Elon Musk
be updated with, you know, mission to accelerate the advent of, I don't know,
Tesla Presenter
infinite abundance or limitless economy?
Elon Musk
Yeah, I mean, it is not strictly speaking, Optimus is not strictly speaking directly in line with accelerating sustainable energy
Elon Musk
the degree that it is more efficient at getting things done than a person. It does, I guess, help with sustainable energy. But I think the mission effectively does somewhat broaden with the advent of Optimus to, I don't know, making the future awesome. So I think you look at Optimus and I know about you, but I'm excited to see what Optimus will become. And you know, this is like, you know, if you could, I mean, you can tell like any given technology,
Elon Musk
you, do you want to see what it's like in a year, 2 years, 3 years, 4 years, 5 years, 10? I'd say for sure. You definitely want to see what's happening with Optimus, whereas, you know, a bunch of other technologies are, you know, sort of plateaued. Don't name names here, but. You know, so I think Optimus is going to be incredible in like 5 years, 10 years, like mind blowing. And I'm really interested to see that happen and I hope you are too. I have a quick question here.
Eric (Tesla)
I'm Justin and I was wondering, are you planning to extend conversational capabilities for the robot? And my second follow up question to
John Emmons
that is, what's the end goal?
Elon Musk
What's the end goal with Optimus? Yeah, Optimus would definitely have conversational capabilities. So you'd be able to talk to it and have A conversation and it would feel quite natural. So from an end goal standpoint, I don't know. I think it's going to keep evolving and I'm not sure where it ends up, but someplace interesting for sure. You know, we always have to be careful about the, you know, don't go down the Terminator path. That's, you know, I thought we might, maybe we should start off with a video of like the Terminator starting off with this, you know, skull crushing. But that might be kind of. People might take that too seriously. So, you know, we do want Optimus to be safe. So we are designing in safeguards where you can locally stop the robot and you know, with like basically a localized control control ROM that you can't update over the Internet, which I think that's quite important, essential frankly. So like a localized stop button
Elon Musk
remote control, something like that, that cannot be changed. But it's definitely going to be interesting. It won't be boring. Okay.
Eric (Tesla)
I see you today.
Tesla Presenter
You have very attractive product with Dojo and its applications. So I'm wondering what's the future for Dojo platform? Will you like provide like infrastructure and service like aws or you will like sell the chip like the Nvidia? So basically what the future because I say you use 7nm, so the developer
Phil (Tesla)
cost is like easily over US$10 million. How do you make the business like business wise?
Elon Musk
Yeah, I mean Dojo is a very big computer and actually will use a lot of power and needs a lot of cooling. So I think it's probably going to make more sense to have Dojo operate in like a Amazon Web Services manner than to try to sell it to someone else. So that would be the most efficient way to operate. Dojo is just have it be a service that you can use that's available online and that where you can train your models way faster and for less money. And as the world transitions to software 2.0,
Ashok Elluswamy
and that's on the bingo card
Elon Musk
as someone I know has to know how to drink five tequilas. So let's see, software 2.0 will use a lot of neural net training. So it kind of makes sense that over time as there's more neural net stuff, people will want to use the fastest, lowest cost neural net training system. So I think there's a lot of opportunity in that direction. Hi, my name is Ali Jahanian. Thank you for this event. It's very inspirational. My question is, I'm wondering what is your vision for humanity? Robots that understand our emotions and art
Phil (Tesla)
and can contribute to our creativity.
Elon Musk
Well, I think there's this. You're already seeing robots that at least are able to generate very interesting art with like, like Dall EE and Dall E2. And I think we'll start seeing AI that can actually generate even movies that have coherence, like interesting movies and tell jokes. So it's quite remarkable how fast AI is advancing at many companies besides Tesla. We're headed for a very interesting future. And. Yeah. So any guys want to comment on that?
Ashok Elluswamy
Yeah, I guess the optimist robot can come up with physical art, not just digital art. You can, you know, you can ask for some dance moves in text or voice, and then you can produce those in the future. So it's a lot of like physical art, not just digital art.
Tesla Presenter
Oh yeah, yeah.
Elon Musk
Computers can absolutely make physical art. Yeah, yeah, 100%.
Ashok Elluswamy
Like dance, play soccer, or what have you. It needs to get more agile, but over time for sure.
David (Tesla)
Thanks so much for the presentation.
John Emmons
For the Tesla Autopilot slides, I noticed that the models that you were using were heavily motivated by language models. And I was wondering what the history
Elon Musk
of that was and how much of
David (Tesla)
an improvement it gave.
Tesla Presenter
I thought the.
David (Tesla)
That was a really interesting, curious choice
John Emmons
to use language models for the lane transitioning. So there's sort of two aspects for why we transition to language modeling.
Elon Musk
So the first, talk loud and close.
Elon Musk
It's not coming through very clearly.
John Emmons
So the language models help us in two ways. The first way is that it lets us predict lanes that we couldn't have otherwise. As Shook mentioned earlier, basically when we predicted lanes and sort of of a dense 3D fashion, you can only model certain kinds of lanes. But we want to get those crisscrossing connections inside of intersections. It's just not possible to do that without making it a graph prediction. If you try to do this with dense segmentation, it just doesn't work. Also, the lane prediction is a multimodal problem. Sometimes you just don't have sufficient visual information to know precisely how things look on the other side of the intersection. So you need a method that can generalize and produce, you know, coherent predictions. You don't want to be predicting two lanes and three lanes at the same time. You want to commit to one in a generative model like these language models provides that.
Phil (Tesla)
Hi, My name is Giovanni.
John Emmons
Yeah, thanks for the presentation.
Elon Musk
It's really nice.
Phil (Tesla)
I have a question for FST team.
Tesla Presenter
So for the neural networks, how do you test? Like, how do you do unit test Software unit tests on that, like do
Phil (Tesla)
you have like a bunch or I
Tesla Presenter
don't know, mid thousands or.
Phil (Tesla)
Yes, cases where the neural network that after you train it, you have to pass it before you release it to
Elon Musk
as a product, Right?
Phil (Tesla)
What's your software unit testing strategies for this?
Ashok Elluswamy
Yeah, glad you asked. There's like a series of tests that we have defined starting from, you know, unit test for software itself. But then for the neural network models we have VIP sets defined where you know, you can define if you just have a large test set, that's not enough. What we find we need like sophisticated VIP sets for different failure modes and then we curate them and grow them over the time of the product. So over the years we have hundreds of thousands of examples where we have been failing in the past that we have curated. And so for any new model we test against the entire history of these failures and then keep adding to this test set. On top of this we have shadow modes where we ship these models in silent to the car and we get data back on where they are failing or succeeding. And there's an extensive QA program. It's very hard to ship a regression. There's like nine levels of filters before it hits customers. But then we have really good infra to make this all efficient.
Elon Musk
I'm one of the QA testers, so I QA the car. Yeah, like QA tester.
Elon Musk
So I'm constantly in the car just being qa like whatever the latest alpha bolt is that doesn't totally crash, finds
Ashok Elluswamy
a lot of bugs.
Tesla Presenter
Hi, great event. I have a question about foundational models for autonomous driving. We have all seen that big models that really can, when you scale up with data and model parameter, right from GPT3 to palm it can actually now do reasoning. Do you see that it's essential scaling up foundation models with data and size and then at least you can get a teacher model, right? That potentially can solve all the problems and then you distill to a student model. Is that how you see foundational models relevant for autonomous learning?
Ashok Elluswamy
That's quite similar to our auto labeling model. So we don't just have models that run in the car, we train models that are entirely offline, that are extremely large, that can't run in real time on the the car. So we just run those offline on the servers, producing really good labels that can then train the online networks. So that's one form of distillation of these teacher student models in terms of foundation models. We are building some really, really large Data sets that are multiple petabytes. And we are seeing that some of these tasks work really well when we have these large data sets like the kinematics, like I mentioned video in all the kinematics, out of all the objects and up to the fourth derivative. And people thought we couldn't do detection with cameras. Detection, depth, velocity, acceleration, and imagine how precise these have to be for these higher order derivatives to be accurate. And this all comes from these kind of large data sets and large models. So we're seeing the equivalent of foundation models in our own way for geometry and kinematics and things like those. You want to add anything, John?
John Emmons
Yeah, I'll keep it brief. Basically whenever we train on a larger data set we see big.
John Emmons
Basically whenever we train on a larger data set we see big improvements in our model performance. And basically whenever we initialize our networks with, you know, some pre training step from some other auxiliary task, we basically see improvements. The self supervised or supervised with large data sets, both help a lot.
Elon Musk
So at the beginning Elon said that Tesla is potentially interested in building artificial general intelligence systems. Given the potentially transformative impact of technology
John Emmons
like that, it seems prudent to invest
Elon Musk
in technical AGI safety expertise specifically. I know Tesla does a lot of technical narrow AI safety research.
John Emmons
I was curious if Tesla was intending
Phil (Tesla)
to try to build expertise in technical
Elon Musk
artificial general intelligence safety specifically. Well, I mean, if it starts looking like we're going to be making a significant contribution to artificial general intelligence, then we'll for sure invest in safety. I'm a big believer in AI safety. I think there should be an AI sort of regulatory authority at the government level, just as there is a regulatory authority for anything that affects public safety. So we have regulatory authority for aircraft and cars and sort of food and drugs because they affect public safety. And AI also affects public safety. So I think, and this is not really something that government I think understands yet, but I think, I think there should be a referee that is ensuring or doing, trying to ensure public safety for AGI. And you think of like, well, what are the elements that are necessary to create AGI? Like the accessible data set is extremely important. And if you've got a large number of cars and humanoid robots processing petabytes of video data and audio data from the real world, just like humans, that might be the biggest data set. It probably is the biggest data set because in addition to that you can obviously incrementally scan the Internet. But what the Internet can't quite do is have millions or hundreds of millions of of cameras in the real world and like I said with audio and other sensors as well. So I think we probably will have the most amount of data and probably the most amount of training power. Therefore probably we will make a contribution to AGI. Hey, I noticed the semi was back
John Emmons
there, but we haven't talked about it too much.
Elon Musk
I was just wondering for the semi truck, what are the changes you're thinking about from a sensing perspective?
David (Tesla)
I imagine there's very different requirements obviously than just a car.
Eric (Tesla)
And if you don't think that's true,
Elon Musk
why is that true? No, I think basically you can drive a car. I mean think about what drives any vehicle. A biological neural net with, with eyes, with cameras essentially. So if and really what is your primary sensors are two cameras on a slow gimbal, a very slow gimbal, that's, that's your head. So if you know a biological neural net with two cameras on a slow gimbal can drive a semi truck, that then if you've got like eight cameras with continuous 360 degree vision operating at a higher frame rate and much higher reaction rate, then I think it is obvious that you should be able to drive a semi or any, any vehicle much better than a human.
Ashok Elluswamy
Hi, my name is Akshay.
Tesla Presenter
Thank you for the event.
Ashok Elluswamy
Assuming, you know, Optimus would be used for different use cases and would evolve at different pace for these use cases, would it be possible to sort of develop and deploy different software and hardware components independently and deploy them, you know, in the, in Optimus so that the overall, you know, feature development is faster for Optimus?
Elon Musk
Okay. All right. We did not comprehend, unfortunately our neural net did not comprehend the question. So next question.
Tesla Presenter
Hi, I want to switch a gear to the autopilot. So when you guys plan to roll out the FSD Beta to countries other than U.S. and Canada. And also my next question is what's the biggest the bottleneck or the technological barrier you think in the current autopilot stack and how you envision to solve that to make the autopilot is considerable better than human in terms of performance matrix like safety assurance and the human confidence. And I think you also mentioned for the FSD V11 you are going to combine the highway and the city as a single stack and some architectural big improvements. Can you maybe explain expand a bit on that?
Elon Musk
Well, that's a whole bunch of questions. We're hopeful to be able to, I think from a technical standpoint FSD Beta should be, should be possible to roll out FSD Beta worldwide by the end of this year.
Tesla Presenter
But we, you know, for a lot
Elon Musk
of countries we need regulatory approval and so we are somewhat gated by the regulatory approval in other countries. But you know, but I think from technical standpoint it will be ready to go to a worldwide beta by the end of this year. And there's quite a big improvement that we're expecting to release next month that will be especially good at assessing the velocity of fast moving cross traffic and a bunch of other things. So, anyone elaborate?
John Emmons
Yeah, I guess so. There used to be a lot of differences between production autopilot and the full self driving beta, but those differences have been getting smaller and smaller over time. I think just a few months ago we now use the same vision only object detection stack in both FSD and in the production autopilot on all vehicles. There's still a few differences, the primary one being the way that we predict lanes right now. So we upgraded the modeling of lanes so that it could handle these more complex geometries. Like I mentioned in the talk, in production Autopilot we still use a simpler lane model, but we're extending our current FSD beta models to work in all sort of highway scenarios as well.
Elon Musk
Yeah, and the version of FSD beta that I drive as actually does have the integrated stack. So it uses the FSD stack both in city streets and highway and it works quite well for me. But we need to validate it in all kinds of weather, like heavy rain, snow dust, and just make sure it's working better than the production stack across a wide range of environments. But we're pretty close to the that. I mean, I think it's. I don't know, maybe it'll definitely be before the end of the year and maybe November.
Paril Jain
Yeah, in our personal drives, the FSD stack on highway drives already way better than the production stack we have. And we do expect to also include the parking lot stack as a part of the FSC stack before the end of this year. So that will basically bring us to you sit in the car in the parking lot and drive till the end of the parking lot at a parking spot before the end of this year.
Elon Musk
And in terms of the like, the fundamental metric to optimize against is how many miles per between a necessary intervention. So just massively improving the how many miles the car can drive on in full autonomy before an intervention is required. That is, is safety critical. So yeah, that's, that's the fundamental metric that we're measuring every week and we're making radical improvements on that.
Kate Park (Tesla)
Oh hi thank you.
Lizzie Miskovetz
Thank you so much for the presentation. Very inspiring.
Kate Park (Tesla)
My name is Daisy.
Tesla Presenter
I actually have a non technical question for you.
Kate Park (Tesla)
I'm curious if you, you are back to your 20s, what are some of
Elon Musk
the things you wish you knew back then?
Kate Park (Tesla)
What are some advice you would give to your younger self?
Elon Musk
Well, I'm trying to figure out something useful to say. Yeah, join Tesla would be one thing. Yeah, I think just trying to try to expose yourself to as many smart people as possible. I don't read a lot of books, You know, I do, I did do that though.
Elon Musk
I think there's some merit to just also like not being like necessarily too intense and like enjoying the moment a bit more. I would say to 20 or 20 something me just, you know, stop and smell the roses occasionally would probably be a good idea. You know, it's like when we were developing the Falcon 1 rocket and on the Kwajalein Atoll and we had this beautiful little island that we're developing the rocket on and not once during that entire time did I even have a drink on the beach. I'm like, I should have had a drink on the beach. That would have been fine.
Tesla Presenter
Thank you very much.
Phil (Tesla)
I think you have excited all of
Elon Musk
the robotic people with optimization.
Eric (Tesla)
This feels very much like 10 years ago in driving.
Phil (Tesla)
But as driving has proved to be harder than it actually looked 10 years ago, what do we know now that we didn't 10 years ago that would
Tesla Presenter
make for example AGI on a humanoid come faster?
Elon Musk
Well, I mean it seems to me that AGI is advancing very quickly. Hardly a week goes by without some significant announcement. And yeah, I mean at this point like AI seems to be able to win at almost any rule based game. It's, it's able to create extremely impressive art, engage in conversations that are very sophisticated, you know, write essays and these, these just keep improving. And there's, there's so much more, so, so many more talented people working on AI and the hardware is getting better. I think it's, it's a, AI is on a super, like a strong exponential curve of, of improvements independent of what we do at Tesla. And obviously we'll benefit somewhat from that exponential curve of improvement with AI. Tesla just also happens to be very good at actuators that motors, you know, motors, gearboxes, controllers, power electronics, batteries, sensors. And you know, really like I say that, you know, the biggest difference between the robot on four wheels and the, the robot with arms and legs is getting the actuators right. It's an actuators and sensors problem. And obviously how you control those actuators and sensors, but it's. Yeah, actuators and sensors and how you control the actuators, it's. I don't know, we have to have like the ingredients necessary to create a compelling robot. And we're doing it.
Ashok Elluswamy
Hi, Elan. You are actually bringing the humanity to the next level, literally Tesla, and you are bringing the humanity to the next level.
Tesla Presenter
So you said Optimus Prime.
Ashok Elluswamy
Optimus will be used in next Tesla factory.
Tesla Presenter
My question is, will a new Tesla
Ashok Elluswamy
factory will be fully run by Optimus program?
Tesla Presenter
And, and when can general public order a humanoid?
Elon Musk
Yeah, I think it'll, you know, we're going to start Optimus with very simple tasks in the factory. You know, like maybe just like loading a part like you saw in the video, loading apart, you know, carrying a part from one place to another or loading a part into
Elon Musk
more conventional robot cells to, you know, that welds buddy together. So we'll start, you know, just trying to how do we make it useful at all? And then, and then gradually expand the number of situations where it's useful. And I think that, that the number of situations where Optimus is useful will, will grow exponentially, like really, really fast. In terms of when people can order one, I don't know. I think it's not that far away. Well, I think you meant when can people receive one? So I don't know, I'm like, I'd say probably within three years, not more than five years. Within three to five years you could probably receive an optimus.
Tesla Presenter
I feel the best way to make the progress for AGI is to involve as many smart people across the world as possible. And given the size and resource of Tesla compared to robot companies and given the state of human or research at the moment, would it make sense for the kind of Tesla to sort of open source some of the simulation hardware parts? I think Tesla can still be the dominant platformer where it can be something
Elon Musk
like Android OS or like iOS stuff
Tesla Presenter
for the entire humanoid research. Would that be something that rather than keeping the optimist to just Tesla researchers or the factory itself can open it and let the whole world export humanoid research?
Elon Musk
I think we have to be careful about Optimus being potentially used in ways that are bad, because that is one of the possible things to do.
Elon Musk
I think, you know, we'd provide optimists where you can provide instructions to optimists, but where those instructions are, you know, governed by some laws of robotics that you cannot overcome. So, you know, not Doing harm to others and I think probably quite a few safety related things with Optimus.
Elon Musk
So. All right, we'll just take maybe a few more questions and then, and then thank you all for coming. Questions one deep and one broad. On the deep for Optimus, what's the current and what's the, the ideal controller bandwidth? And then in the broader question, there's this big advertisement for the depth and
Phil (Tesla)
breadth of the company.
John Emmons
What is it uniquely about Tesla that enables that.
Elon Musk
Anyone want to tackle the bandwidth question?
Tesla Presenter
Yeah, so the technical bandwidth of the.
Elon Musk
Close to your mouth and loud.
Tesla Presenter
Okay. For the bandwidth question you have to understand or figure out what is the task that, that you wanted to do and what is the, if you took a frequency transform of that task, what is it that you want your limbs to do? And that's where you get your bandwidth from. It's not a number that you can specifically just say you need to understand your use case. And that's from, that's where the bandwidth comes from. What is the broad question?
Elon Musk
I don't quite remember the breadth and depth thing. I can answer the breadth and depth. I mean, it's interesting on the bandwidth question, I think we probably will just end up increasing the bandwidth or you know, which translates to the effective dexterity and reaction time of the, of the robot. Like you can safe say it's not one hertz and it's. Maybe you don't need to go all the way to 100 hertz, but I don't know, maybe 1025. I don't know. Over time I think the bandwidth will increase quite a bit or translate it to dexterity and latency. You'd want to minimize that over time. Yeah, minimize latency, maximize dexterity in terms of breadth and depth. I guess we've got, we've got, we're a pretty big company at this point, so we've got a lot of different areas of expertise that we necessarily had to develop in order to make autonomous or in order to make electric cars. And then in order to make autonomous electric cars, we've, we've just, I mean, Tesla is like a whole series of startups basically, and so far they've almost all been quite successful. So we must be doing something right. And I consider one of my core responsibilities running the company is to have an environment where great engineers can flourish. And I think in a lot of companies, I don't know, maybe most companies, if somebody's a really talented, driven engineer, they're unable to actually their talents are suppressed at a lot of companies and it's, you know, and some of the companies that engineering talent is suppressed in a way that is maybe not obviously bad, but where it's just so comfortable and you're paid so much money and the output you actually have to produce is so low that it's like a honey trap, you know. So like there's a few honey trap places in Silicon Valley where they don't necessarily, don't seem like bad places for engineers. But you have to say like a good engineer went in and what did they get out? And the output of that engineering talent seems very low, even though they seem to be enjoying themselves. That's why I call it there's a few honey trap companies in Silicon Valley. Tesla is not a honey trap. And we're demanding and it's like going to get a lot of shit done and it's going to be really cool and it's not going to be easy. But if you are a super talented engineer, your talents will be used, I think to a greater degree than anywhere else. You know, SpaceX also that way. So.
Ashok Elluswamy
Hi Elan, I have two questions, so both to the Autopilot team.
Tesla Presenter
So the thing is like I have
Ashok Elluswamy
been following your progress for the past few years. So today you have made changes on
Tesla Presenter
like the lane detection.
Ashok Elluswamy
Like you said that like previously you
Tesla Presenter
are doing instant semantic segmentation now you
Ashok Elluswamy
guys have built transfer models for like building the lanes. So what are another, some, some other common challenges which you guys are facing
Tesla Presenter
right now, like which you are solving in future as a curious engineer so
Ashok Elluswamy
that like we as a researcher can work on those.
Tesla Presenter
Start working on those.
Elon Musk
And the second question is like I'm
Ashok Elluswamy
really curious about the data engine. Like you guys have like told a case like where the car is stopped. So how are you finding cases which is very much similar to that from
Tesla Presenter
the data which you have like so little bit more on the data engine would be great. So that's it. Okay, I'll start.
Phil (Tesla)
Answer the first question. Using occupancy network as an example. So what you saw in the presentation did not exist a year ago. So we only spent one year on time. We actually shipped more than 12 occupancy network. And to have a one foundation model actually to represent the entire physical world around everywhere and in all weather conditions, it's actually really, really challenging. So only over a year ago we're kind of like driving a 2D world. If there's a wall and if there's curve, we kind of represent with the same static edge which is obviously, you know, not ideal. Right. There's a big difference between a curve and a wall. When you drive, you make different choices.
Phil (Tesla)
So after we realize that we have to go to 3D, we have to basically rethink the entire problem and think about how we address that. So this will be like one example of challenges we have, we have conquered in the past year.
Kate Park (Tesla)
Yeah. To answer the question about how we actually source examples of those tricky stopped cars, there's a few ways to go about this, but two examples are, one, we can trigger for disagreements within our signals. So let's say that parked bit flickers between parked and driving. We'll trigger that back. And the second is we can leverage more of the shadow mode logic. So if the customer ignores the car, but we think we should stop for it, we'll get that data back too. So these are just different, like various trigger logic that allows us to get those data campaigns back.
Tesla Presenter
Hi, thank you for the amazing presentation. Thanks so much. So there are a lot of companies that are focusing on the AGI problem. And one, one of the reasons why it's such a hard problem is because the problem itself is so hard to define. Several companies have several different definitions. They focus on different things. So what is Tesla, how is Tesla
Paril Jain
defining the AGI problem?
Tesla Presenter
And what are you focusing on specifically?
Elon Musk
Well, we're not actually specifically focused on AGI. I'm simply saying that AGI seems likely to be an emergent property of, of what we're doing because we're creating all these autonomous cars and autonomous humanoids that are actually within a truly gigantic data stream that's coming in and being processed. It's by far the most amount of real world data and data you can't get by just searching the Internet. Because you have to be out there in the world and interacting with people, people and interacting with the roads and just, you know, Earth is a big place and reality is messy and complicated. So I think it's sort of like you just. It just seems likely to be an emergent property of if you've got, you know, tens or hundreds of millions of autonomous vehicles and maybe even a comparable number of humanoids, maybe more than that. On the humanoid front, well, that's just the most amount of data. And if that video is being processed, it just seems likely that, you know, the cars will definitely get way better than human drivers and the humanoid robots will become increasingly indistinguishable from humans, perhaps. And so then, like I said, you have this emergent property of AGI. And arguably humans collectively are sort of a superintelligence as well, especially as we improve the data rate between humans. That seems to be way back in the early days of the Internet was like. The Internet was like humanity acquiring a nervous system, where now all of a sudden, any one element of humanity could know all of the knowledge of humans by connecting to the Internet, almost all knowledge, certainly a huge part of it, Whereas previously we would exchange information by osmosis in order to transfer data. So you would have to write a letter, someone would have to carry the letter by person to another person, and then a whole bunch of things in between. And then it was like, yeah, I mean, insanely slow when you think about it. And even if you were in the Library of Congress, you still didn't have access to all the world's information, and you certainly couldn't search it. And obviously very few people are in the Library of Congress. So, I mean, one of the great sort of equality elements, like the Internet, has been the most. The biggest equalizer in history in terms of access to information and knowledge. And any student of history, I think, would agree with this because you go back a thousand years, there were very few books, and books would be incredibly expensive, but only a few people knew how to read, and even small number of people even had a book. Now look at it like you can access any book instantly. You can learn anything for basically for free. It's pretty incredible. So, you know, I was asked recently, what period of history would I prefer to be at the most? And my answer was, right now, this is the most interesting time in history, and I read a lot of history. So let's. Let's do our best to keep that going. Yeah. And to go back to one of the earlier questions, I would answer like you can. The thing that's happened over time with respect to Tesla Autopilot is that we've just. The neural nets have gotten. Have gradually absorbed more and more software. And in the limit, of course, you could simply take the videos as seen by the car and compare those to the steering inputs from the steering wheel and pedals, which are very simple inputs. And in principle, you could train with nothing in between, because that's what humans are doing with a biological neural net. You could train based on video. And what trains the video is the moving of the steering wheel and the pedals with no other software in between. We're not there yet, but it's gradually going in that direction. All right, maybe last question.
John Emmons
How you going?
Elon Musk
I think we've got a question at the front here. Hello. They're Right there. We'll do two questions. Fine.
Tesla Presenter
Over there. Hi. Thanks for such a great presentation.
Elon Musk
We'll do your question last.
Tesla Presenter
Okay, cool. With FSD being used by so many people, do you think, what's the. How do you evaluate the company's risk tolerance in terms of performance statistics? And do you think there needs to be more transparency or regulation from third parties as to how. What's good enough? And do you. Defining, like thresholds for performance across so many miles?
Elon Musk
Well, the, you know, the number one design requirement at Tesla is safety. So. And that goes across the board. So in terms of the mechanical safety of the car, we have the lowest probability of injury of any cars ever tested by the government. Government for just a passive mechanical safety, essentially crash structure and airbags and whatnot. We have the best, the highest rating for active safety as well. And I think it's going to get to the point where the active safety is so ridiculously good, it's like just absurdly better than a human. And then with respect to Autopilot, we do publish this, broadly speaking, the statistics on miles driven with cars that have no autonomy, or Tesla cars with no autonomy, with kind of hardware 1, hardware 2, hardware 3, and then the ones that are in FSD beta. And we see steady improvements all along the way. And, you know, sometimes there's this dichotomy of, you know, should you wait until the car is like, I don't know, three times safer than a person before deploying any technology. But I think that's, that is actually morally wrong. At the point at which we believe that adding autonomy reduces injury and death, I think you have a more obligation to deploy it, even though you're going to get sued and blamed by a lot of people because the people whose lives you saved don't know that their lives are saved. And the people, the people who do occasionally die or get injured, they definitely know or their state does that it was, you know, whatever. There was a problem with Autopilot. That's why you have to look at the numbers in total miles driven, how many accidents occurred, how many accidents were serious, how many fatalities. And, you know, we've got well over 3 million cars on the road. So this, it's, that's a lot of miles driven every day. It's not going to be perfect. But what matters is that it is very clearly safer than not deploying it. Yeah. So I think. Last question.
Tesla Presenter
I think. Yeah.
Elon Musk
So thanks. Well, the last question here.
Kate Park (Tesla)
Okay. Hi. So I do not work on hardware,
Elon Musk
so maybe the hardware team and you
Paril Jain
guys can enlighten me.
Kate Park (Tesla)
Why is it required that there be symmetry in the design design of Optimus? Because humans, we have handedness, right?
Elon Musk
We are, we use some set of
Tesla Presenter
muscles more than others.
Kate Park (Tesla)
Over time there is wear and tear, right?
Paril Jain
So maybe you'll start to see some
Elon Musk
joint failures or some actuator failures more over time.
Kate Park (Tesla)
I understand that this is extremely prestige also. We as humans have based so much
Paril Jain
fantasy and fiction or superhuman capabilities. Like all of us don't want to
Kate Park (Tesla)
walk right over there. We want to extend our arms and
Elon Musk
like we have all these, you know,
Paril Jain
a lot of fantasy fantastical designs.
Kate Park (Tesla)
So considering everything else that is going on in terms of batteries and intensity
Elon Musk
of compute, maybe you can leverage all
Kate Park (Tesla)
those aspects into coming up with, with something. Well, I don't know, more interesting in
Elon Musk
terms of your, the robot that you're building.
Kate Park (Tesla)
And I'm hoping you're able to explore those directions.
Elon Musk
Yeah, I mean, I think it would be cool to have like, you know, make Inspector Gadget real. That would be pretty sweet. So, yeah, I mean right now we just want to make a basic humanoid work well. And our goal is fastest path to a useful humanoid robot. I think this will ground us in reality literally and ensure that we are doing something useful. Like one of the hardest things to do is to be useful to actually and then to have high utility under the curve of like how many people did you help? How much help did you provide to each person on average? And then how many people did you help? The total utility. Like trying to actually ship useful product that people like to a large number of people is so insanely hard, it boggles the mind, you know. So I could say like, man, there's a hell of a difference between a company that has shipped product and one has not shipped product. To get this is night and day. And then even once you ship product, can you make the cost, the value of the output worth more than the cost of the input? Which is again insanely difficult, especially with hardware. So But I think over time I think cool to do creative things and have like eight arms and whatever and have different versions and maybe, you know, there'll be some hardware. Like companies that are able to add things to an Optimus. Like maybe we've, you know, add a power port or something like that or attach them, you can add, you know, add attachments to your Optimus. Like you can add them to your phone. It could be a lot of cool things that could be done over time and it could be maybe an ecosystem of small companies that or big companies that make add ons for Optimus. So with that, I'd like to thank the team for their hard work. You guys are awesome. And. Thank you all for coming. And for everyone online, thanks for tuning in. And I think this will be one of those great videos where you can, like, if you, you can fast forward to the bits that you find most interesting, but we try to give you a tremendous amount of detail, literally so that you can look at the video at your leisure and you can focus on the parts that you find interesting and skip the other parts. So thank you all and we'll do this, try to do this every year and we might do a monthly podcast even so. But I think it'd be great to sort of bring you along for the ride and show you what cool things are happening and, and yeah, thank you.
Tesla Presenter
All right. Thanks,