Tesla Autonomy Day 22 avril 2019 Événement investisseurs de Tesla sur la conduite autonome : la puce FSD maison, l'approche par réseaux de neurones et une démo en direct, avec Elon Musk.
Événement investisseurs de Tesla sur la conduite autonome : la puce FSD maison, l'approche par réseaux de neurones et une démo en direct, avec Elon Musk.
Transcription Speaker
It. Sa. Sam. It. Range sam it. Sa.
Speaker
Race. Sa. Sa. Sa. Sam. Sa. Sam.
Speaker
It. Sam. Sa. Sam. It. Sa. Sam sa. Sam. Sa. Sa. Sa. Sa. It. Sam. Sa. Sa. Sam. Sa. Sa. Sa. It. Sam.
Investor Relations
Hi everyone. I'm sorry for being late. Welcome to our very first analyst day for Autonomy. I really hope that this is something we can do a little bit more regularly now to keep you posted about the, the development we're doing with regards to autonomous driving. About three months ago we were getting prepped up for our Q4 earnings call with Elon and quite a few other executives. And one of the things that I told the group is that from all the conversations that I keep having with investors on regular basis, the biggest gap that I see with what I see inside the company and what the outside perception is is our ability of autonomous driving. And it kind of makes sense because for the past couple of years we've been really talking about Model 3 ramp. And you know, a lot of the debate has revolved around Model 3, but in reality a lot of things have been happening in the background. We've been working on the new full self driving chip. We've had a complete overhaul of our neural net for vision recognition, et cetera. So now that we finally started to produce our full self driving computer, we thought it's a good idea to just open the veil, invite everyone in and talk about everything that we've been doing for the past two years. So about three years ago we wanted to use, we wanted to find the best possible chip for Full Autonomy. And we found out that there's no chip that's been designed from ground up for neural nets. So we invited my colleague Pete Bannon, the VP of Silicon Engineering, to design such chip for us. He's got about 35 years of experience of building chips and designing chips. About 12 of those years were for a company called PA Semi, which was later acquired by Apple. So he worked on dozens of different architectures and designs and he was the lead designer, I think for Apple iPhone 5 just before joining Tesla. And he's going to be joined on the stage by Elon Musk. Thank you.
Elon Musk
Actually, I was going to introduce Pete, but Martin Stunt. So he's just the best chip and system architect that I know in the world. And it's an honor to have you and your team at Tesla and take away. Just tell them about the incredible work that you and your team have done.
Pete Bannon
Thanks, Elon. It's a pleasure to be here this morning and a real treat really to tell you about all the work that my colleagues and I've been doing here at Tesla for the last three years. I think we'll tell you a little bit about how the whole thing got started and then I'll introduce you to the full self driving computer and tell you a little bit about how it works. We'll dive into the chip itself and go through some of those details. I'll describe how the custom neural network accelerator that we designed works and then I'll show you some results and hopefully you'll all still be awake by then. I was hired in February of 2016. I asked Elon if he was willing to speak all the money it takes to do full custom system design. And he said, well, are we going to win? And I said, well, yeah, of course. So he said I'm in. And so that got us started. We hired a bunch of people and started thinking about what a custom designed chip for full autonomy would look like. We spent 18 months doing the design and in August of 2017 we released the design for manufacturing. We got it back in December at pack powered up and it actually worked very, very well on the first try. We made a few changes and released a B0 Rev in April of 2018. In July of 2018, the chip was qualified and we started full production of production quality parts. In December of 2018, we had the autonomous driving stack running on the new hardware and we were able to start retrofitting employee cars and testing the hardware and software out in the real world. Just last March we started shipping the new computer in the Model S and X. And just earlier in April we started production in the Model 3. So this whole program, from the hiring of the first few employees to having it in full production in all three of our cars, is just a little over three years and is probably the fastest system development program I've ever been associated with. And it really speaks a lot to the advantages of having a tremendous amount of vertical integration to allow you to do concurrent engineering and speed up deployment. In terms of goals, we were totally focused exclusively on Tesla requirements and that makes life a lot easier. If you have one and only one customer, you don't have to worry about anything else. One of those goals was to keep the power under 100 watts so that we could retrofit fit the new machine into the existing cars. We also wanted a lower part cost so we could enable full redundancy for safety. At the time, we had a thumb in the wind estimate that it would take at least 50 trillion operations. A second of neural Network performance to drive a car. And so we wanted to get at least that much and really as much as we possibly could. Batch size is how many items you operate on at the same time. So for example, Google's TP has a batch size of 256 and you have to wait around until you have 256 things to process before you can get started. We didn't want to do that, so we designed our machine with a batch size of one. So as soon as an image shows up, we process it immediately. To minimize latency, which maximizes safety, we needed a GPU to run some post processing. At the time we were doing quite a lot of that. But we speculated that over time the amount of post processing on the GPU would decline as the neural networks got better and better. And that has actually come to pass. So we took a risk by putting a fairly modest GPU in the design, as you'll see, and that turned out to be a good bet. Security is super important. If you don't have a secure car, you can't have a safe car. So there's a lot of focus on security and then of course, safety in terms of actually doing the chip design. As Elon alluded earlier, there was really no ground up neural network accelerator in existence in 2016. Everybody out there was adding instructions to their CPU or GPU or DSP to make it better for inference, but nobody was really just doing it natively. So we set out to do that ourselves. And then for other components on the chip we purchased industry standard IP for CPUs and GPUs. That allowed us to minimize the design time and also the risk to the program. Another thing that was a little unexpected when I first arrived was our ability to leverage existing teams at Tesla. Tesla had wonderful power supply design teams, signal integrity analysis, package design, system software, firmware board designs, and a really good system validation program that we were able to take advantage of to accelerate this program. Here's what it looks like. Over there on the right you see all the connectors for the video that comes in from the eight cameras that are in the car. You can see the two self driving computers in the middle of the board and then on the left is the power supply and some control connections. And so I really love it when a solution is boiled down to its barest elements. You have video computing and power and it's straightforward and simple. Here's the original hardware 2.5 enclosure that the computer went into and we've been shipping for the last two years. Here's the new design for the FSD computer. It's basically the same, and that of course is driven by the constraints of having a retrofit program for the cars. I'd like to point out that this is actually a pretty small computer. It fits behind the glove box. Between the glove box and the firewall in the car. It does not take up half your trunk. As I said earlier, there's two fully independent computers on the board. You can see them there highlighted in blue and green. To either side of the large SoC you can see the DRAM chips that we use for storage. And then below left you see the flash chips that represent the file system. So these are two independent computers that boot up and run their own operating system.
Elon Musk
Yeah. If I can add something, the general principle here is that any part of this could fail and the car will keep driving. So you could have cameras fail, you could have power circuits fail, you could have one of the Tesla full self driving computer chips fail, car keeps driving. The probability of this computer failing is substantially lower than somebody losing consciousness. That's the key metric, at least in order of magnitude.
Pete Bannon
So one of the things that we additional thing we do to keep the machine going is to have redundant power supplies in the car. So one, one machine's running on one power supply and the other one's on the other. The cameras are the same. So half of the cameras run on the blue power supply, the other half run on the green power supply, and both chips receive all of the video and process it independently. So in terms of driving the car, the basic sequence is collect lots of information from the world around you. Not only do we have cameras, we also have radar, GPS maps, the imus, ultrasonic sensors around the car. We have wheel ticks, steering angle. We know what the acceleration and deceleration of the car is supposed to be. All of that gets integrated together to form a plan. Once we have a plan, the two machines exchange their independent version of the plan to make sure it's the same. And assuming that we agree, we then act and drive the car. Now, once you've driven the car with some new control, you want to validate it. So we validate that what we transmitted was what we intend to transmit to the other actuators in the car. And then you can use the sensor suite to make sure that it happens. So if you ask the car to accelerate or brake or steer right or left, you can look at the accelerometers and make sure that you are in fact doing that. So there's a tremendous amount of redundancy and overlap in both our data acquisition and our data monitoring capabilities here. Moving on to talk about the full self driving chip a little bit. It's packaged in a 37.5 millimeter BGA with 1600 balls. Most of those are used for power and ground, but plenty for signal as well. If you take the lid off, it looks like this. You can see the package substrate and you can see the die sitting in the center there. If you take the die off and flip it over, it looks like this. There's 13,000 C4 bumps scattered across the top of the die and then Underneath that are 12 metal layers which is obscuring all the details of the design. So if you strip that off, it looks like this. This is a 14 nanometer FinFET solution DMOS process. It's 260 millimeters in size, which is a modest sized die. So for comparison, a typical cell phone chip is about 100 millimeters square, which so we're quite a bit bigger than that. But a high end GPU would be more like 600 to 800 millimeters square. So we're sort of in the middle. I would call it the sweet spot. It's a comfortable size to build. There's 250 million logic gates on there and a total of 6 billion transistors, which even even though I work on this all the time, that's mind boggling to me. The chip is manufactured and tested to AEC Q100 standards, which is a standard automotive criteria. Next, I'd like to just walk around the chip and explain all the different pieces to it. And I'm sort of going to go in the order that a pixel coming in from the camera would visit all the different pieces. So up there in the top left you can see, see the camera serial interface. We can ingest 2.5 billion pixels per second, which is more than enough to cover all the sensors that we know about. We have an on chip network that distributes data from the memory system. So the pixels would travel across the network to the memory controllers on the right and left edges of the chip. We use industry standard LPDDR4 memory running at 4266 gigabits per second, which gives us a peak bandwidth to 68 gigabytes a second, which is a pretty healthy bandwidth. But again, this is not like ridiculous. So we're sort of trying to stay in the comfortable sweet spot for cost reasons. The image signal processor has a 24 bit internal pipeline that allows us to do Take full advantage of the HDR sensors that we have around the car. It does advanced tone mapping which helps to bring out details and shadows. And then it has advanced noise reduction which just improves, improves the overall quality of the images that we're using in the neural network. The neural network accelerator itself. There's two of them on the chip. They each have 32 megabytes of SRAM to hold temporary results and minimize the amount of data that we have to transmit on and off the chip, which helps reduce power. Each array has a 96 by 96 multiply add array with in place accumulation which allows us to do almost 10,000 multiply ads per cycle. There's dedicated RELU hardware, dedicated pooling hardware, and each of these deliver 306. Excuse me, each one delivers 36 trillion operations per second and they operate at 2 gigahertz. The two of them together on a die deliver 72 trillion operations a second. So we exceeded our goal of 50 teraops by a fair bit. There's also a video encoder. We encode video and use it in a variety of places in the car, including the backup camera display. There's optionally a user feature for dash cam and also for clip logging data to the cloud, which Stuart and Andre will talk about more later. There's a GPU on the chip. It's modest performance. It has support for both 32 and 16 bit floating point. And then we have 12 A72 64 bit C CPUs for general purpose processing. They operate at 2.2 gigahertz. And this represents about 2 1/2 times the performance available in the current solution. There's a safety system that contains two CPUs that operate in lockstep. This system is the final arbiter of whether it's safe to actually drive the actuators in the car. So this is where the two plans come together and we decide whether it's safe or not to move forward. And lastly there's a safety system. And basically the job of the safety system is to ensure that this chip only runs software that's been cryptographically signed by Tesla. If it's not been signed by Tesla, then the chip does not operate. Now, I've told you a lot of different performance numbers and I thought it'd be helpful maybe to put it into perspective a little bit. So throughout this talk, I'm going to talk about a neural network from our narrow camera. It uses 35 billion operations, 35 giga ops, and if we use all 12 CPUs to process that network, we could do one and a half frames per second, which is super slow, not nearly adequate to drive the car. If we use the 600 gigaflop GPU, the same network, we'd get 17 frames per second, which is still not good enough to drive the car. With eight cameras, the neural network accelerators on the channel chip can deliver 2100 frames per second. And you can see from the scaling as we moved along that the amount of computing in the CPU and GPU are basically insignificant to what's available in the neural network accelerator. It really is night and day. So, moving on to talk about the neural network accelerator, we're just going to stop for some water. On the left, there's a cartoon of a neural network just to give you an idea of what's going on. The data comes in at the top and visits each of the boxes. And the data flows along the arrows to the different boxes. The boxes are typically convolutions or deconvolutions with relus. The green boxes are pooling layers. And the important thing about this is that the data produced by one box is then consumed by the next box, and then you don't need it anymore. You can throw it away. So all of that temporary data that gets created and destroyed as you flow through the network, there's no need to store that off chip in dram. So we keep all that data in sram. And I'll explain why that's super important in a few minutes. If you look over on the right side of this, you can see that in this Network, of the 35 billion operations, almost all of them are convolution, which is based on dot products. The rest are deconvolution, also based on dot product, and then relu and pooling, which are relatively simple operations. So if you were designing some hardware, you'd clearly target doing dot products, which are based on multiply, add, and really kill that. But imagine that you sped it up by a factor of 10,000. So 100% all of a sudden turns into 0.1%, 0.01%, and suddenly the relu and pooling operations are going to be quite significant. So our hardware doesn't. Our hardware design includes dedicated resources for processing, relu and pooling as well. Now, this chip is operating in a thermally constrained environment, so we had to be very careful about how we burn that power. We want to maximize the amount of arithmetic we can do. So we picked integer add. It's 9 times less energy than the corresponding floating point Add and we picked 8 bit by 8 bit integer multiply, which is significantly less power than other multiply operations and is probably enough accuracy to get good results. In terms of memory, we chose to use SRAM as much as possible. And you can see there that going off chip to Dram is approximately 100 times more expensive, expensive in terms of energy consumption than using local sram. So clearly we want to use local SRAM as much as possible. In terms of control, this is data that was published in a paper by Mark Horowitz at ISSCC where he sort of critiqued how much power it takes to execute a single instruction on a regular integer cpu. And you can see that the add operation is only 0.15 centimeter percent of the total power. All the rest of the power is control overhead and bookkeeping. So in our design we sought to basically get rid of all that as much as possible, because what we're really interested in is arithmetic. So here's the design that we finished. You can see that it's dominated by the 32 megabytes of SRAM. There's big banks on the left and right and in the center bottom. And then all the computing is done in the upper middle. Every single clock, we read 256 bytes of activation data out of the SRAM array, 128 bytes of weight data out of the SRAM array, and we combine it in a 96 by 96 mulad array which performs 9000 multiply adds per clock at 2 gigahertz. That's a total of 3.63 36.8 teraops. Now, when we're done with the dot product, we unload the engine so that we shift the data out across the dedicated RELU unit, optionally across a pooling unit, and then finally into a write buffer where all the results get aggregated up. And then we write out 128 bytes per cycle back into the SRAM. And this whole thing cycles along all the time, continuously. So we're doing dot products while we're unloading previous results, doing pooling and writing back into the memory. If you add it all up at 2 gigahertz, you need 1 terabyte per second of SRAM bandwidth to support all that work. And so the hardware supplies that. So one terabyte per second of bandwidth per engine. There's two on the chip. Two terabytes per second. The accelerator has a relatively small instruction set. We have a DMA read operation to bring data in from memory. We have a DMA Write operation to push results back out to memory. We have three dot product based instructions, instructions, convolution, deconvolution and inner product. And then two relatively simple scale is one input, one output operation and outwise is two inputs and one output. And then of course, stop when you're done. We had to develop a neural network compiler for this. So we take the neural network that's been trained by our vision team as it would be deployed in the older cars, and we take that and compile it from for use on the new accelerator. The compiler does layer fusion, which allows us to maximize the computing each time we read data out of the SRAM and put it back. It also does some smoothing so that the demands on the memory system aren't too lumpy. And then we also do channel padding to reduce bank conflicts. And we do bank aware SREM allocation. And this is a case where, where we could have put more hardware in the design to handle bank conflicts. But by pushing it into software, we save hardware and power at the cost of some software complexity. We also automatically insert DMAs into the graph so that data arrives just in time for computing without having to stall the machine. And then at the end, we generate all the code, we generate all the weight data, we compress it and we add a CRC checksum for reliability. To run a program, all the neural network descriptions, programs are loaded into SRAM at the start and then they sit there ready to go all the time. So to run a network, you have to program the address of the input buffer, which presumably is a new image that just arrived from a camera. You set the output buffer address, you set the pointer to the network weights and then you set, set, go, and then the machine goes off and will sequence through the entire neural network all by itself, usually running for a million or 2 million cycles. And then when it's done, you get an interrupt and can post process the results. So moving on to results, we had a goal to stay under 100 watts. This is measured data from cars driving around, running the full autopilot stack. And we're dissipating 72 watts, which is a little bit more power than the previous design. But with the dramatic improvement in performance, it's still a pretty good answer. Of that 72 watts, about 15 watts is being consumed running the neural networks. In terms of cost, the silicon cost of this solution is about 80% of what we were paying before. So we are saving money by switching to this solution. And in terms of performance, we took the narrow camera neural network, which I've been talking about that has 35 billion operations in it. We ran it on the old hardware in a loop as quick as possible and we delivered 110 frames per second. We took the same data, the same network compiled it for hardware for the new FSD computer. And using all four accelerators we can get 2,300 frames per second processed. So a factor of 21.
Elon Musk
I think this is perhaps the most significant slide. It's night and day.
Pete Bannon
I've never worked on a project where the performance increase was more than three, so this was pretty fun. If you compare it to say Nvidia's Drive Xavier solution, A single chip delivers 21 teraops. Our full stop performance driving computer with two chips is 144 teraops. So to conclude, I think we've created a design that delivers outstanding performance. 144 teraops for neural network processing. It has outstanding power performance. We managed to jam all of that performance into the thermal budget that we had. It enables a fully redundant computing solution. It has a modest cost and really the important thing is that this FSD computer will enable a new level of safety and autonomy in Tesla's vehicles without impacting their cost or range. Something that I think we're all looking forward to.
Elon Musk
I think why don't we do Q and A after each segment so if people have questions about the hardware, they can ask right now. The reason I asked Pete to do just a detailed, far more detailed than perhaps most people would appreciate, dive into the Tesla full self driving computer is because at first it seems improbable. How could it be that Tesla, who has never designed a chip before, would design the best chip in the world? But that is objectively what has occurred. Not, not best by a small margin, best by a huge margin. It's in the cars right now. All Teslas being produced right now have this computer. We switched over from the Nvidia solution for SNX about a month ago and we switched over Model 3 about 10 days ago. All cars being produced have all the hardware necessary, compute and otherwise for full self driving. I'll say that again. All Tesla cars being produced right now have everything necessary for full self driving. All you need to do is improve the software and later today you will drive the cars with the development version of the improved software and you will see for yourself themselves. Questions for Pete?
Analyst
Y. Questions. I saw Trip Chaudhary Global equities research. Very, very impressive in every shape and form. I was wondering, like I. I took some notes. You are using Activation function relu, the rectify linear unit. But if you think about the deep neural network, it has multiple layers and some algorithms may use different activation functions for different hidden layers like softmax or tanh. Do you have flexibility for incorporating different activation functions rather than LU in your platform? Then I have a follow up.
Pete Bannon
Yes, we do. We have implementations of Tanh and Sigmoid for example.
Analyst
Beautiful. One last question. Like in the nanometers you mentioned 14nm. As I was wondering, wouldn't it make sense to come little lower? Maybe 10nm, two years down or maybe 7?
Pete Bannon
At the time we started the design, not all the IP that we wanted to purchase was available in 10 nanometers. So we finished the design in 14.
Elon Musk
It's maybe worth pointing out that we finished this design like maybe one and a half two years ago and began design of the next generation. We're not talking about the next generation today, but we're about halfway through it. That will all the things that are obvious for next generation chip we're doing.
Analyst
You talked about the software as the piece now. You did a great job. I was blown away. Understood 10% of what you said, but I trust that it's in good hands.
Analyst
So it feels like you got the hardware pieces done and that was really hard to do and now you have to do the software piece. Now maybe that's outside of your expertise, but how should we think about that software piece?
Elon Musk
Well, couldn't ask for a better introduction to Andre and Stuart. Are there any questions for the chip part before the next part of the presentation is neural nets and software.
Analyst
So maybe on the chip side, the last slide was 144 trillions of operations per second versus was it Nvidia 21?
Analyst
And maybe can you just contextualize that for a finance person, why that's so significant, that gap? Thank you.
Pete Bannon
Well, I mean it's a factor of seven in performance delta. So that means you can do seven times as many frames. You can run neural networks that are seven times larger and more sophisticated. So it's a very big current currency that you can spend on lots of interesting things to make the car better.
Elon Musk
I think that Xavier power usage is higher than ours. Xavier powers higher than ours, I think. Or comparable.
Pete Bannon
I don't know that I believe it's
Elon Musk
to best my knowledge, the power requirements would increase at least to the same degree, a factor of seven and costs would also increase by a factor of seven. Great. So yeah, power is a real problem because it also Reduces range. So it has the penalty for power is very high and then you have to get rid of that power by the thermal problem becomes really significant because you got to get rid of all that power. So
Investor Relations
thank you very much. I think we have, you know, a
Elon Musk
lot of quite a bit ask the questions. If you guys don't mind the day running a bit long. Just we're going to do the drive demos afterwards. So if you've got, if you, if you, if anybody needs to pop out and do drive demos a little sooner, you're welcome to do that. But I do want to make sure we answer your questions. Yep. Pradeep Ramani from ubs, intel and AMD to some extent have started moving towards a chiplet based architecture. I did not notice a chiplet based design here. Do you think that looking forward that
Stuart Bowers
would be something that might be of
Elon Musk
interest to you guys from an architecture standpoint?
Pete Bannon
A chiplet based architecture?
Pete Bannon
We're not currently considering anything like that. I think that's mostly useful when you need to use different styles of technology. So if you want to integrate silicon germanium or DRAM technology on the same silicon substrate, that gets pretty interesting. But until the die size gets obnoxious, I wouldn't go there.
Elon Musk
To be clear, the strategy here, and this started basically three, a little over three years ago, was design and build a computer that is fully optimized and aiming for full self driving. Then write software that is designed to work specifically on that computer and get the most out of that computer. So you have tailored hardware that is a master of one trade, self driving. Nvidia is a great company but they have many customers and so when as they apply their resources, they need to do a generalized solution. We care about one thing, self driving. So it was designed to do that incredibly well. The software is also designed to run on that hardware incredibly well. And the combination of the software and the hardware I think is unbeatable.
Investor Relations
Hi, the chip is designed to process video input.
Andrej Karpathy
In case you use, let's say lidar,
Elon Musk
would it be able to process that as well or is that, is it primarily for video? What we're going to explain to you today is that LIDAR is a fool's errand and anyone relying on LIDAR is doomed. Doomed. Expensive, expensive sensors that are unnecessary. It's like having a whole bunch of expensive appendices. One appendix is bad. Well, now they want to put a whole bunch of them. That's ridiculous. You'll see.
Pete Bannon
There's somebody up here.
Pete Bannon
Oh, there's a Gentleman. Hi.
Andrej Karpathy
So just two questions just on the power consumption. Is there a way to maybe give us like a rule of thumb on, you know, every watt is, reduces range by certain percent or a certain amount
Stuart Bowers
just so we can get a sense
Andrej Karpathy
of how much of an improvement a
Pete Bannon
model three, the, the target consumption is 250 watts per mile.
Elon Musk
It depends on the nature of the driving as to how many miles that that affects in city. It would have a much better, bigger effect than on highway. So if you're driving for an hour in a city and you had a solution, Hypothetically that was a kilowatt, you'd lose four miles on a Model 3. So if you're only going say 12 miles an hour, then that would be a 20, 25% impact on range in city. It's basically power is the power of the system has a massive impact on city range, which is where we think most of the robo taxi market will be. So power is extremely important. Tasha.
Pete Bannon
I'm sorry, I didn't hear you.
Andrej Karpathy
What's the primary design objective of the next generation chip?
Elon Musk
We don't want to talk too much about the next generation chip, but it's
Elon Musk
It'll be at least let's say three times better than the current system.
Elon Musk
About two hours away.
Analyst
To develop this chip is the chip being you don't manufacture the chip, you contract that out. And how much cost reduction does that save in the overall vehicle cost?
Pete Bannon
The 20% cost reduction I cited was the piece cost per vehicle reduction. That wasn't a development costs, that was just the actual.
Analyst
No, I'm saying. But like if I'm manufacturing these in mass, is this saving money in doing it yourself?
Pete Bannon
Yes, a little bit.
Elon Musk
I mean most chips are made. Most people don't make chips with their own fab. It's pretty unusual.
Analyst
I think you don't see any supply issues with getting the chip mass produced.
Pete Bannon
The cost saving pays for the development. I mean the basic strategy going to Elon was we're going to build this chip, it's going to reduce the cost. And Elon said times a million cars a year.
Elon Musk
That's correct, yes.
Elon Musk
If there are really chip specific questions, we can answer them. Otherwise there will be a Q and A opportunity after Andre talks and after Stuart talks. So there will be two other Q& A opportunities. This is if it's very chip specific,
Pete Bannon
then also I'll be here all afternoon.
Elon Musk
Yeah, and exactly. And Pete will be here at the end as well. So go ahead.
Andrej Karpathy
Oh yeah, thanks.
Stuart Bowers
That Dye photo you had.
Stuart Bowers
The neural processor takes up quite a bit of the dye. I'm curious, is that your own design or is there some external IP there?
Pete Bannon
Yes, that was a custom design by Tesla.
Stuart Bowers
And then I guess the follow on would be there's probably a fair amount of opportunity to reduce that footprint as you tweak the design.
Pete Bannon
It's actually quite dense. So in terms of reducing it, I don't think so. It'll greatly enhance the functional capabilities in the next generation.
Stuart Bowers
Okay, and then last question. Can you share where you're fabbing this part?
Pete Bannon
Where are we fabbing it?
Pete Bannon
Samsung, yes. Austin, Texas.
Investor Relations
There's one at the back.
Pete Bannon
Grant Tanaka, Tanaka Capital. Just curious how defensible your chip technologies
Analyst
and design is from a, from a IP point of view and hoping that
Pete Bannon
you won't be offering a lot of
Andrej Karpathy
the IP to the outside for free.
Pete Bannon
We have filed on the order of a dozen patents on this technology. Fundamentally it's linear algebra, which I don't think you can patent. I'm not sure.
Elon Musk
I think if somebody started today and they were really good, they might have some something like what we have right now in three years, but in two years we'll have something three times better.
Analyst
Talking about the intellectual property protection, you have the best intellectual property and some people just steal it for the fun of it. I was wondering if we look at few interactions with Aurora that companies industry believes they stole your intellectual property. I think the key ingredient that you need to protect is the weights that associate to various parameters. Do you think your chip can do something to prevent anybody? Maybe encrypt all the weights so that even you don't know what the weights are at the chip level so that your intellectual property remains inside it and nobody knows about it and nobody can just steal it.
Elon Musk
Ben, I'd like to meet the person that could do that because I would hire them in a heartbeat. Yeah, so that'd be a hard problem. Yeah. Do you want to, I mean we do encrypt the. It's a hard chip to crack, so if they can crack it, it's very good. If they can then crack it and then also also figure out the software and the neural net system and everything else. They can design it from scratch like that's all.
Pete Bannon
It's our intention to prevent people from stealing all that stuff. And if they do, we hope it at least takes a long time.
Elon Musk
It will definitely take them a long time. Yeah. I mean I just don't think if it was our goal to do that. How would we do it? It would be very difficult. But the thing that's, I think a very powerful sustainable advantage for us is the fleet. Nobody has the fleet. Those weights are constantly being updated and improved. Based on billions of miles driven, Tesla has 100 times more cars with the full self driving hardware than everyone else combined. You know, we, we have, by the end of this quarter we'll have 500,000 cars worth of the full 8 camera setup, 12 ultrasonics, some of them will still be on hardware too, but we still have the data gathering ability. And then a year from now we'll have over a million cars with full self driving computer hardware, everything. Yeah, so we have founders. It's just a massive data advantage. It's similar to like, you know how like the Google search engine has a massive advantage because people use it and people are programming effectively program Google with their queries and their results.
Analyst
May I just press you on that and please reframe the question because I'm a tech layman, if it's appropriate. But you know, when we talk to Waymo or Nvidia, they do speak with equivalent conviction about their leadership because, because of their competence in simulating miles driven. Can you talk about the advantage of having real world miles versus simulated miles? Because I think they express that, you know, by the time you get a million miles, they can simulate a billion. And no Formula one race car driver, for example, could ever successfully complete a real world track without driving in a simulator. Can you talk about the advantages it sounds like that you perceive to have associated with having data ingestion coming from real world miles versus simulated miles?
Elon Musk
Absolutely. The simulator, we have quite a good simulation too, but it just does not capture the long tail of weird things that happen in the real world. If the simulation fully captured the real world. Well, I mean that would be proof that we're living in a simulation, I think. Yeah, it doesn't, I wish. But simulations do not capture the real world. The real world is really weird and messy. You need the cars on the road. We're actually going to get into that in Andre and Stuart's presentation. So. Okay, why don't we move on to Andre?
Investor Relations
Great, thanks. Thank you.
Pete Bannon
Thank you everybody.
Investor Relations
Thank you very much. The last question was actually a very good segue because one thing to remember about our FSD computer is that it can run much more complex neural nets for much more precise image recognition. And to talk to you about how we actually get that image data and how we analyze them, we have our Senior Director of AI, Andre Karpathi, who's going to explain all of that to you. Andrej has a PhD from Stanford University where he studied computer science, focusing on vision recognition and deep learning.
Elon Musk
Andre, why don't you just talk? Do your own intro. There's a lot of PhDs from Stanford. That's not important. Yes, okay, we don't care.
Investor Relations
Come on in.
Elon Musk
Andre started the computer vision class at Stanford. That's much more significant. That's what matters.
Elon Musk
So can you please talk about your background in a way that is not bashful? Just tell me about the stuff you've done and then.
Andrej Karpathy
So yeah, I think I've been training neural networks basically for what is now a decade. And these neural networks were not actually really used in the industry until maybe five or six years ago. So it's been some time that I've been training these neural networks and that included institutions at stanford, at, at OpenAI, at Google, and really just training a lot of neural networks not just for images, but also for natural language and designing architectures that couple those two modalities. For my PhD, so.
Elon Musk
And the computer computer science class.
Andrej Karpathy
Oh yeah. And at Stanford I actually taught the convolutional neural networks class. And so I was the primary instructor for that class. I actually started the course and designed the entire curriculum. So in a beginning it was about 150 students and then it grew to 700 students over the next two or three years. So it's a very popular class. It's one of the largest classes at Stanford right now. So that was also really successful.
Elon Musk
I mean, Andre is like really one of the best computer vision people in the world. Arguably the best.
Andrej Karpathy
Okay, thank you.
Andrej Karpathy
Hello everyone. So Pete told you all about the chip that we've designed that runs neural networks in the car. My team is responsible for training of these neural networks and that includes all of data collection from the fleet neural network training and then some of the deployment onto that chip. So what do the neural networks do exactly in the car? So what we are seeing here is a stream of videos from across the vehicle, across the car. These are eight cameras that send us videos. And then these neural networks are looking at those videos and are processing them and making predictions about what they're seeing. And so some of the things that we're interested in and some of the things you're seeing on this visualization here are lane line markings, other objects, the distances to those objects, what we call drivable space, shown in blue, which is where the car is allowed to go, and a lot of other predictions like traffic lights, traffic signs, and so on. Now for my talk, I will talk roughly in three stages. So first I'm going to give you a short primer on neural networks and, and how they work and how they're trained. And I need to do this because I need to explain in the second part why it is such a big deal that we have the fleet and why it's so important and why it's a key enabling factor to really train these neural networks and making them work effectively on the roads. And in the third stage, I'll talk about vision and lidar and how we can estimate depth just from vision alone. So the core problem that these networks are solving in the car is that of visual recognition. So for you and I, these are very, this is a very simple problem. You can look at all of these four images and you can see that they contain a cello, a boat, an iguana or scissors. So this is very simple and effortless for us. This is not the case for computers. And the reason for that is that these images are, to a computer, really just a massive grid of pixels. And at each pixel you have the brightness value at that point. And so instead of just seeing an image, a computer really gets a million numbers in a, a grid that tell you the brightness values at all the positions.
Elon Musk
A matrix, if you will. It really is the matrix.
Andrej Karpathy
And so we have to go from that grid of pixels and brightness values into high level concepts like iguana and so on. And as you might imagine, this iguana has a certain pattern of brightness values. But iguanas actually can take on many appearances. So they can be in many different appearances, different poses and different brightness conditions against different backgrounds. You can have a different crops of that iguana. And so we have to be robust across all those conditions, and we have to understand that all those different brightness patterns actually correspond to iguanas. Now, the reason you and I are very good at this is because we have a massive neural network inside our heads that is processing those images. So light hits the retina, travels to the back of your brain, to the visual cortex. And the visual cortex consists of many neurons that are wired together and that are doing all the pattern recognition on top of those images.
Andrej Karpathy
And really over the last, I would say about five years, the state of the art approaches to processing images using computers have also started to use neural networks, but in this case, artificial neural networks. But these artificial neural networks, and this is just a cartoon diagram of it are a very rough mathematical approximation to your visual cortex. We really do have neurons, and they are connected, connected together. And here I'm only showing three or four neurons in three or four in four layers. But a typical neural network will have tens to hundreds of millions of neurons, and each neuron will have a thousand connections. So these are really large pieces of almost simulated tissue. And then what we can do is we can take those neural networks and we can show them images. So, for example, I can feed my iguana into this neural network, and the network will make predictions about what it's seeing. Now, in the beginning, these neural networks are initialized completely randomly, so the connection strengths between all those different neurons are completely random. And therefore the predictions of that network are also going to be completely random. So it might think that you're actually looking at a boat right now. And it's very unlikely that this is actually an iguana. And during the training, during the training process, really what we're doing is we know that that's actually an iguana. We have a label. So what we're doing is we're basically saying we'd like the probability of iguana to be larger for this image and the probability of all the other things to go down. And then there's a mathematical process called backpropagation, stochastic gradient descent that allows us to back propagate that signal through those connections. And update every one of those connections.
Andrej Karpathy
And update every one of those connections just a little amount. And once the update is complete, the probability of iguana for this image will go up a little bit. So it might become 14%, and the probability of the other things will go down. And of course, we don't just do this for this single image. We actually have entire large data sets that are labeled. So we have lots of images. Typically, you might have millions of images, thousands of labels or something like that, and you are doing forward, backward, passes over and over again. So you're showing the computer, here's an image, it has an opinion, and then you're saying this is the correct answer, and it tunes itself a little bit. You repeat this millions of times, and you sometimes you show images, the same image to the computer, you know, hundreds of times as well. So the network training typically will take on the order of a few hours or a few days, depending on how big of a network you're training. And that's the process of training a neural network. Now, there's something very unintuitive about the way neural Networks work that I have to really get into, and that is that they really do require a lot of these examples and they really do start from scratch. They know nothing. And it's really hard to wrap your head around this. So as an example, here's a cute dog. And you probably may not know the breed of this dog, but the correct answer is that this is a Japanese spaniel. Now, all of us are looking at this and we're seeing Japanese spaniel, and we're like, okay, I got it. I understand kind of what this Japanese spaniel looks like. And if I show you a few more images of other dogs, you can probably pick out other Japanese spaniels here. So in particular, those three look like a Japanese spaniel and the other ones do not. So you can do this very quickly. And you need one example, but computers do not work like this. They actually need a ton of data of Japanese spaniels. So this is a grid of Japanese spaniels showing them, you need thousands of examples showing them in different poses, different brightness conditions, different backgrounds, different crops. You really need to teach the computer from all the different angles what this Japanese spaniel looks like. And it really requires all that data to get that to work. Otherwise the computer can't pick up on that pattern automatically. So what does all this imply about the setting of self driving? Of course, we don't care about dog breeds too much. Maybe we will at some point, but for now we really care about lane line markings, objects where they are, where we can drive, and so on. So the way we do this is we don't have labels like iguana for images, but we do have images from the fleet like this. And we're interested in, for example, lane line markings. So we, a human typically goes into an image and using a mouse, annotates the lane line markings. So here's an example of an annotation that a human could create a label for this image. And it's saying that that's what you should be seeing in this image. These are the lane line markings. And then what we can do is we can go to the fleet and we can ask for more images from the fleet. And if you ask the fleet, if you just do a naive job of this and you just ask for images at random, the fleet might respond with images like this. Typically, going forward on some highway, this is what you might just get like a random collection like this. And we would annotate all that data. Now, if you're not careful, and you only annotate a random distribution of this data, your network will kind of pick up on this random distribution on data and work only in that regime. So if you show it slightly different example, for example, here is an image that actually the road is curving and it is a bit of a more residential neighborhood. Then if you show the neural network this image, that network might make a prediction that is incorrect. It might say that, okay, well, I've seen lots of times on highways, lanes just go forward. So here's a possible prediction. And of course this is very incorrect, but the neural network really can't be blamed. It does not know that the train on the, the tree on the left, whether or not it matters or not. It does not know if the car on the right matters or not towards the lane line. It does not know that the buildings in the background matter or not. It really starts completely from scratch. And you and I know that the truth is that none of those things matter. What actually matters is that there are a few white lane line markings over there in a vanishing point. And the fact that they curl a little bit should pull the prediction. Except there's no mechanism by which we can just tell the neural network, hey, those lane line markings actually matter. The only tool in the toolbox that we have is labeled data. So what we do is we need to take images like this when the network fails and we need to label them correctly. So in this case, we will turn the lane to the right and then we need to feed lots of images of this to the neural net. And neural net, over time will accumulate, will basically pick up on this pattern that those things there don't matter, but those lane line markings do, and we learn to predict the correct lane. So what's really critical is not just the scale of the data set. We don't just want millions of images. We actually need to do a really good job of covering the possible space of things that the car might encounter on the roads. So we need to teach the computer how to handle scenarios where it's night and wet, you have all these different specular reflections, and as you might imagine, the brightness patterns in these images will look very different. We have to teach the computer how to deal with shadows, how to deal with forks in the road, how to deal with large objects that might be taking up most of that image, how to deal with tunnels or how to deal with construction sites. And in all these cases, there's no, again, explicit mechanism to tell the network what to do. We only have massive amounts of data. We want to source all those images and we want to annotate the correct lines and the network will pick up on the patterns of those now large and varied data sets basically make these networks work very well. This is not just a finding for us here at Tesla. This is a ubiquitous is finding across the entire industry. So experiments and research from Google, from Facebook, from Baidu, from Alphabet's DeepMind all show similar plots where neural networks really love data and love scale and variety. As you add more data, these neural networks start to work better and get higher accuracies for free. So more data just makes them work better. Now a number of companies have, a number of people have kind of pointed out that potentially we could use simulation to actually achieve the scale of the data sets. And we're in charge of a lot of the conditions here. And maybe we can achieve some variety in a simulator now at Tesla. And that was also kind of brought up in the questions just before this. Now at Tesla, this is actually a screenshot of our own simulator. We use simulation extensively. We use it to develop and evaluate the software. We've also even used it for training quite successfully. So but really when it comes to training data for neural networks, there really is no substitute for real data. The simulations have a lot of trouble with modeling appearance, physics and the behaviors of all the agents around you. So there are some examples to really drive that point across the real world really throws a lot of crazy stuff at you. So in this case, for example, we have very complicated environments with snow, with trees, with wind. We have various visual artifacts that are hard to simulate potentially. We have complicated construction sites, bushes and plastic bags that can go in, that can kind of go around with the wind. Complicated construction sites that might feature lots of people, kids, animals, all mixed in. And simulating how those things interact and flow through this construction zone might actually be completely, completely intractable. It's not about the movement of any one pedestrian in there. It's about how they respond to each other and how those cars respond to each other and how they respond to you driving in that setting. And all of those are actually really tricky to simulate. It's almost like you have to solve the self driving problem to just simulate other cars in your simulation. So it's really complicated. So we have dogs, exotic animals, and in some cases it's not even that you can't simulate, it is that you can't even come up with it. So for example, I didn't know that you can have truck, truck on truck like that. But in the real world you find this and you find lots of other things that are Very hard to really even come up with. So really the variety that I'm seeing in the data coming from the fleet is just crazy. With respect to what we have in the simulator, we have a really good simulator.
Elon Musk
I mean, I think simulation, you're fundamentally grading your own homework. So if you know that you're going to simulate it, ok, you can definitely solve for it. But as Andre is saying, you don't know what you don't know. The world is very weird and has millions of corner cases. And if somebody can produce a self driving simulation that accurately matches reality, that in itself would be a monumental achievement of human capability. They can't. There's no way.
Andrej Karpathy
Yep. So I think the three points that I really tried to drive home until now are to get neural networks to work well, you require these three essentials. You require a large data set, a very data set and a real data set. And if you have those capabilities, you can actually train neural networks and make them work very well. And so why is Tesla in such a unique and interesting position to really get all these three essentials right? And the answer to that, of course, course, is the fleet. We can really source data from it and make our neural network systems work extremely well. So let me take you through a concrete example of, for example, making the object detector work better to give you a sense of how we develop these neural networks, how we iterate on them, and how we actually get them to work over time. So object detection is something we care a lot about. We'd like to put bounding boxes around, say the cars and the objects here, because we need to track them and we need to understand how they might move around. So again, we might ask human annotators to give us some annotations for these. And humans might go in and might tell you that, okay, those patterns over there are cars and bicycles and so on. And you can train a neural network on this. But if you're not careful, the neural network will make mispredictions in some cases. So as an example, if we stumble by a car like this that has a bike on the back of it, then the neural network actually, when I joined, would actually create two detections. It would create a car detection and a bicycle detection. And that's actually kind of correct because I guess both of those objects actually exist. But for the purposes of the controller and the planner downstream, you really don't want to deal with the fact that this bicycle can go with the car. The truth is that that bike is attached to that car. So in terms of like Just objects on the road. There's a single object, a single car. And so what you'd like to do now is you'd like to just potentially annotate lots of those images, as this is just a single car. So the process that we go through in of terms internally in the team is that we take this image or a few images that show this pattern, and we have a mechanism, a machine learning mechanism, by which we can ask the fleet to source us examples that look like that. And the fleet might respond with images that contains those patterns. So as an example, these six images might come from the fleet. They all contain bikes on backs of cars. And we would go in and we would annotate all those as just a single car. And then the performance of that detector actually improves. And the network internally understands that, hey, when the bike is just attached to the car, that's actually just a single car. And it can learn that given enough examples, and that's how we sort of fix that problem. I will mention that I talk quite a bit about sourcing data from the fleet. I just want to make a quick point that we've designed this from the beginning with privacy in mind, and all the data that we use for training is anonymized. Now, the fleet doesn't just respond with bicycles on backs of cars. We look for all the things. We look for lots of things all the time. So, for example, we look for boats, and the fleet can respond with boats. We look for construction sites, and the fleet can send us lots of construction sites from across the world. We look for even slightly more rare cases. So, for example, finding debris on the road is pretty important to us. So these are examples of images that have streamed to us from the fleet that show tires, cones, plastic bags and things like that. If we can source these at scale, we can annotate them correctly and the neural network can learn how to deal with them in the world. Here's another example. Animals, of course, also a very rare occurrence and event. But we want the neural network to really understand what's going on here, that these are animals and we want to deal with that correctly. So to summarize, the process by which we iterate on neural network predictions looks something like this. We start with a seed data set that was potentially sourced at random. We annotate that data set and then we train neural networks on that data set and put that in the car. And then we have mechanisms by which we notice inaccuracies in the car when this detector may be misbehaving. So for Example, if we detect that the neural network might be uncertain, or if we detect that, or if there's a driver intervention on any of those settings, we can create this trigger infrastructure that sends us data of those inaccuracies. And so, for example, if we don't perform very well on lane line detection on tunnels, then we can notice that there's a problem in the tunnels. That image would enter our unit test. So we can verify that we've actually fixing the problem over time. But now what you do is to fix this inaccuracy, you need to source many more examples that look like that. So we ask the fleet to please send us many more tunnels. And then we label all those tunnels correctly, we incorporate that into the training set, and we retrain the network, redeploy and iterate the cycle over and over again. And so we refer to this iterative process by which we improve these performances predictions as the data engine. So iteratively deploying something potentially in shadow mode, sourcing inaccuracies, and incorporating the training set over and over again. And we do this basically for all the predictions of these neural networks. Now, so far I talked about a lot of explicit labeling. So like I mentioned, we ask people to annotate data. This is an expensive process in time. And also with respect to. Yeah, it's just an expensive process. And so these annotations of course, can be very expensive to achieve. So what I want to talk about also is really to utilize the power of the fleet. You don't want to go through this human annotation bottleneck. You want to just stream in data and automate it automatically. And we have multiple mechanisms by which we can do this. So as one example of a project that we recently worked on is the detection of cut ins. So you're driving down the highway, someone is on the left or on the right, and they cut in in front of you into your lane. So here's a video showing the autopilot detecting that this car is intruding into our lane. Now, of course, we'd like to detect a cut in as fast as possible. So the way we approach this problem is we don't write explicit code for is the left blinker on? Is the right blinker on? Track the keyboard over time and see if it's moving horizontally. We actually use a fleet learning approach. So the way this works is we ask the fleet to please send us data whenever they see a car transition from a right lane to the center lane or from left to center. And then what we do is we rewind time backwards and we automatically can annotate that, hey, that car will turn, will in 1.3 seconds cut in in front of you. And then we can use that for training the neural net. And so the neural net will automatically pick up on a lot of these patterns. So for example, the cars are typically yawed, they're moving this way, maybe the blinker is on. All that stuff happens internally inside the neural net just from these examples. So we asked the fleet to automatically send us all this data. We can get half a million or so images and all of these would be annotated for cut ins. And then we train the network. And then we took this cut in network and we deployed it to the fleet. But we don't turn it on yet. We run it in shadow mode. And in shadow mode, the network is always making predictions. Hey, I think this vehicle is going to cut in. From the way it looks, this vehicle is going to cut in. And then we look for mispredictions. So as an example, this is a clip that we had from shadow mode of the cut in network. And it's kind of hard to see, but the network thought that the vehicle right ahead of us on the right was going to cut in. And you can sort of see that it's slightly flirting with the lane line. It's trying to, it's sort of encroaching a little bit. And the network got excited and it felt that that was going to be cut in. That vehicle will actually end up in our center lane. That turns out to be incorrect because. And the vehicle did not actually do that. So what we do now is we just churn the data engine. We source that ran in the shadow mode. It's making predictions, it makes some false positives and there are some false negative detections. So we got overexcited in sometimes and sometimes we missed a cut in when it actually happened. All those create a trigger that streams to us and that gets incorporated now for free. There's no humans harmed in the process of labeling this data incorporated for free into our training set. We retrained the network and redeployed the shadow mode. And so we can spin this a few times. And we always look at the false positives and negatives coming from the fleet. And once we're happy with the false positive, false negative ratio, we actually flip the bit and actually let the car control to that network. And so you may have noticed we actually shipped one of our first versions of a cut in detector approximately, I think three months ago. So if You've noticed that the car is much better at detecting cut ins. That's fleet learning operating at scale. Yes, it actually works quite nicely. So that's fleet learning. No humans were harmed in the process. It's just a lot of neural network training based on data and a lot of shadow mode. And looking at those results, another essentially,
Elon Musk
like everyone's training the network all the time is what it amounts to. Whether the, whether autopilot is on or off, the network is being trained. Every mile that's driven for the car, that's hardware 2 or above is training the network.
Andrej Karpathy
Another interesting way that we use this in the scheme of fleet learning and the other project that I will talk about is a path prediction. So while you are driving the car, what you're actually doing is you are annotating the data because you are steering the wheel, you're telling us how to traverse different environments. So what we're looking at here is some person in the fleet who took a left through an intersection. And what we do here is we, we have the full video of all the cameras and we know that the, the path that this person took because of the gps, the initial measurement unit, the wheel angle, the wheel ticks. So we put all that together and we understand the path that this person took through this environment. And then of course, this, this, we can use this for supervision for the network. So we just source a lot of this from the fleet. We train a neural network on the, on those trajectories and then the neural network protection predicts paths just from that data. So really what this is referred to typically is called imitation learning. We're taking human trajectories from the real world and we're just trying to imitate how people drive in real worlds. And we can also apply the same data engine crank to all of this and make this work over time. So here's an example of path prediction going through a kind of a complicated environment. So what you're seeing here is a video and we are overlaying the prediction, the predictions of the network. So this is a path that the network would follow in green and some.
Elon Musk
Yeah, I mean, the crazy thing is the network is predicting paths it can't even see with incredibly high accuracy. It can't see around the corner. But, but it's saying the probability of that curve is extremely high. So that's the path and it nails it. You will see that in the cars today. But we're going to turn on augmented vision so you can see the lane lines and the path Predictions of the cars overlaid on the video.
Andrej Karpathy
Yeah, there's actually more going on under the hood that you can even tell.
Elon Musk
I mean, it's kind of scary, to be honest.
Andrej Karpathy
And of course there's a lot of details I'm skipping over. You might not want to annotate all the drivers, you might want to just imitate the better drivers. And there's many technical ways that we actually slice and dice that data. But the interesting thing here is that this prediction is actually a 3D prediction that we project back to the image here. So the path here forward is a three dimensional thing that we're just rendering in 2D. But we know about the slope of the ground from all this and that's actually extremely valuable for driving. So path prediction actually is live in the fleet today, by the way. So if you're driving cloverleafs, if you're in a cloverleaf on the highway until maybe five months ago or so, your car would not be able to do cloverleaf. Now it can. That's path prediction running live on your cars. We shipped this a while ago and today you are going to get to experience this. For traversing intersections, a large component of how we go through intersections in your drives today is all sourced from path prediction from automatic labels. So what I talked about so far is really the three key components of how we iterate on the predictions of the network and how we make it work over time. You require large, varied and real data set. We can really achieve that here at Tesla and we do that through the scale of the fleet, the data engine, shipping things in shadow mode, iterating that cycle, and potentially even using fleet learning where no human annotators are harmed in the process, and just using data automatically. And we can really do that at scale. So in the next section of my talk, I'm going to especially talk about depth perception using vision only. So you might be familiar that there are at least two sensors in the car. One is vision cameras just getting pixels, and the other is LiDAR that a lot of companies also use. And LiDAR gives you these point measurements of distance around you. Now, one thing I'd like to point out, first of all is you all came here, you drove here, many of you, and you used your neural net and vision. You were not shooting lasers out of your eyes and you still ended up here.
Andrej Karpathy
Things went well. So clearly the human neural net derives distance and all the measurements and the 3D understanding of the world just from vision. It actually uses multiple cues to do so. I'll just briefly go over some of them just to give you a sense of roughly what's going on inside. As an example, we have two eyes pointed out, so you get two independent measurements at every single time step of the world ahead of you. And your brain stitches this information together to arrive at some depth estimation because you can triangulate any points across those two viewpoints. A lot of animals instead have eyes that are positioned on the sides so they have very little overlap in their visual fields. So they will typically use structure for motion. And the idea is that they bob their heads and because of the movement, they actually get multiple observations of the world. And you can triangulate again depths. And even with one eye closed and completely motionless, you can still have some sense of depth perception. If you did this, I don't think you would notice me coming 2 meters towards you or 100 meters back. And that's because there are a lot of very strong monocular cues that your brain also takes into account. This is an example of a pretty common visual illusion where you have, you know, these two blue bars are identical, but your brain, the way it stitches up this scene is it just expects one of them to be larger than the other because of the vanishing lines of this image. So your brain does a lot of this automatically. And neural nets, artificial neural nets can as well. So let me give you three examples of how you can arrive at depth perception from vision alone. A classical approach and two that rely on neural networks. So here's a video going down, I think this is San Francisco of a Tesla. So these are our cameras, our sensing, and we're looking at all. I'm only showing the main camera, but all the cameras are turned on, the eight cameras of the autopilot. And if you just have this six second clip, what you can do is you can stitch up this environment into 3D using multi view stereo techniques. So this.
Andrej Karpathy
This is supposed to be a video, is it not? A video? Oh, I know it's. There we go. So this is the 3D reconstruction of those 6 seconds of that car driving through that path. And you can see that this information is purely, it's very well recoverable from, from just videos. And roughly that's through process of triangulation and as I mentioned, multi view stereo. And we've applied similar techniques slightly more sparse and approximate also in the car. So it's remarkable all that information is really there in the sensor and just a matter of extracting it. The other project that I want to briefly talk about is as I Mentioned, there's nothing about neural network. Neural networks are very powerful visual recognition engines. And if you want them to predict depth, then you need to, for example, look for labels of depth. And then they can actually do that extremely well. So there's nothing limiting networks from predicting this monocular depth except for labeled data. So one example project that we've actually looked at internally is we use the forward facing radar, which is shown in blue, and that radar is looking out and measuring depths of objects. And we use that radar to annotate the what vision is seeing the bounding boxes that come out of the neural networks. So instead of human annotators telling you, okay, this car and this bounding box is roughly 25 meters away, you can annotate that data much better using sensors. So you use sensor annotation. So as an example, radar is quite good at that distance. You can annotate that and then you can train a neural network on it. And if you just have enough data of it, this neural network is very good at predicting those patterns. So here's an example of predictions of that. So in circles, I'm showing radar objects and in. And the cuboids that are coming out here are purely from vision. So the cuboids here are just coming out of vision. And the depth of those cuboids is learned by a sensor annotation from the radar. So if this is working very well, then you would see that the circles in the top down view would agree with the cuboids. And they do. And that's because neural networks are very competent at predicting depths. They can learn the different sizes of vehicles internally and they know how big those vehicles are. And you can actually derive depth from that quite accurately. The last mechanism I will talk about very briefly is slightly more fancy and gets a bit more technical. But it is a mechanism that has recently, there's a few papers basically over the last year or two on this approach. It's called self supervision. So what you do in a lot of these papers is you only feed raw videos into neural networks with no labels whatsoever. And you can still learn, you can still get neural networks to learn depth. And it's a little bit technical, so I can't go into the full details, but the idea is that the neural network predicts depth at every single frame of that video. And then there are no explicit targets that the neural network is supposed to regress to with the labels. But instead, the objective for the network is to be consistent over time. So whatever depth you predict should be consistent over the duration of that video. And the Only way to be consistent is to be right. And so the neural network automatically predicts the correct depths for all the pixels. And we've reproduced some of these results internally. So this also works quite well. So in summary, people drive with vision only, no lasers are involved. This seems to work quite well. The point that I'd like to make is that visual recognition and very powerful visual recognition is absolutely necessary for autonomy. It's not a nice to have like we must have neural networks that actually really understand the environment around you. And LIDAR points are much less information rich environment. So vision really understands the full details. Just a few points around are much. There's much less information in those. So as an example on the left here, is that a plastic bag or is that a tire? Lidar might just give you a few points on that, but vision can tell you which one of those two is true and that impacts your control. Is that person who is slightly looking backwards, are they trying to merge in into your lane on the bike or are they just going forward in the construction sites? What do those signs say? How should I behave in this world? The entire infrastructure that we have built up for roads is all designed for human visual consumption. So all of the signs, all the traffic lights, everything is designed for vision. And so that's where all that information is. And so you need that ability. Is that person distracted and on their phone, are they going to walk into your lane? Those answers to all these questions are only found in vision and are necessary for level four, level five autonomy. And that is the capability that we are developing at Tesla and through this is done through combination of large scale neural network training through data engine and getting that to work over time and using power of the fleet. And so in this sense LIDAR is really a shortcut. It sidesteps the fundamental problems that the important problem of visual recognition that is necessary for autonomy. And so it gives a false sense of progress and is ultimately a crutch. It does give like really fast demos. So if I was to summarize the
Andrej Karpathy
my entire talk in one slide, it would be this, all of autonomy. Because you want level four, level five systems that can handle all the possible situations in, in 99.99% of the cases and chasing some of the last few nights is going to be very tricky and very difficult and is going to require a very powerful visual system. So I'm showing you some images of what you might encounter in any one slice of that nine. So in the beginning you just have very simple cars. Going forward then those Cars start to look a little bit funny. Then maybe you have bikes on cars, then maybe you have cars on cars. Then maybe you start to get into really rare events like cars turned over or even cars airborne. We see a lot of things coming from the fleet and we see them at some rate, at like a really good rate compared to all of our competitors. And so the rate of progress at which you can actually address these problems iterate on the software and really feed the neural networks with the right data. That rate of progress is really just proportional to how often you encounter these situations in the wild. And we encounter them significantly more frequently than anyone else, which is why we're going to do extremely well. Thank you.
Andrej Karpathy
It's all super impressive. Thank you so much. How much data, how many pictures are you collecting on average from each car
Pete Bannon
per period of time?
Andrej Karpathy
And then it sounds like the new hardware with the dual, dual active. Active computers gives you some really interesting opportunities to run in full simulation one
Pete Bannon
copy of the neural net while you're
Andrej Karpathy
running the other one. Running the other one.
Pete Bannon
Drive the car and compare the results
Andrej Karpathy
to do quality assurance. And then I was also wondering if there are other opportunities to use the computers for training when they're parked in the garage for the 90% of the time that I'm not driving my time.
Investor Relations
Tesla around.
Andrej Karpathy
Thank you very much. Yep. So for the first question, how much data do we get from the fleet? So it's really important to point out it's not just the scale of the data set, it really is the variety of that data set that matters. If you just have lots of images of something going forward on the highway, at some point the neural just gets it. You don't need that data. So we are really strategic in how we pick and choose. And the trigger infrastructure that we've built up is quite sophisticated and allows us to get just the data that we need right now. And so it's not a massive amount of data, it's just very well picked up data. For the second question with respect to redundancy, absolutely. You can run basically the copy of the network on both. And that is actually how it's designed to achieve level four, level five system that is redundant. So that's absolutely the case. And your last question. I'm sorry, I did not.
Elon Musk
Training the car is an inference optimized computer. We do have a major program at Tesla which we don't have enough time to talk about today, called Dojo. That's a super powerful training computer. The goal of Dojo will be to Be able to take in vast amounts of data and train at a video level and do unsupervised, massive training of vast amounts of video with the dojo program. Dojo computer. But that's for another day.
Analyst
I'm like a test pilot in a way because I drive the 40510 and all these really tricky, really long tail things happen every day. But the one challenge that I'm curious to how you're going to solve is changing lanes. Because whenever I try to get into a lane with traffic, everybody cuts you off. And so human behavior is very irrational. When you're driving in LA and the car just wants to do it safely and you almost have to do it unsafely. So I was wondering how you're going to solve that problem. Yeah.
Andrej Karpathy
So one thing I will point out is I spoke about the data engine as iterating on neural networks, but we do the exact same thing on the level of software and all the hyper parameters that go into the choices of when we actually lane change, how aggressive we are, we're always changing those, potentially running them in shadow mode and seeing how well they work. And so to tune our heuristics around when it's okay to lane change, we would also potentially utilize the data engine and the shadow mode and so on. Ultimately, actually designing all the different heuristics for when it's okay to lane change is actually a little bit intractable, I think, in the general case. And so ideally you actually want to use fleet learning to guide those decisions. So when do humans lane change, in what scenarios and when do they feel it's not safe to lane change? And let's just look at a lot of the data and train machine learning classifiers for distinguishing when it is too safe to do so. And those machine learning classifiers can write much better code than humans because they have the mass amount of data backing that. So they can really tune all the right thresholds and agree with humans and do something safe.
Elon Musk
I think we'll probably have a mode that goes beyond Mad Max mode to LA traffic mode. Yeah, well, you know, Mad Max sort of have a hard time in LA traffic, I think.
Andrej Karpathy
Yeah. So really it's a trade off. Like you don't want to create unsafe situations, but you want to be assertive. But that little dance of how you make that work as a human is actually very complicated and it's very hard to write in code. But I think we really do. It really does seem like machine learning approach is kind of like the right
Analyst
way to go about it.
Andrej Karpathy
Where we just look at a lot of ways that people do this and try to imitate that.
Elon Musk
We're just being like more conservative right now. And then as we gain higher and higher confidence, we'll allow users to select a more aggressive mode. That'll be up to the user. But in the more aggressive modes and trying to merge in traffic, there is a slight. No matter how many new, there's a slight chance of like a fender bender, not a serious accident, but you basically will have a choice of, do you want to have a non zero chance of a fender bender on freeway traffic, which is unfortunately the only way to navigate LA traffic.
Elon Musk
I mean, yes. Yes. It always reminds me of like LA Story. This movie is a great movie.
Andrej Karpathy
Yeah, it's very subtle because there's this game of chicken that's going on.
Elon Musk
Yeah. We'll offer more aggressive options over time that will be user specified. Yes. Mad Max plus. Exactly.
Analyst
Hello. Hi. Jed Doersheimer from Canaccord Genuity. Thank you and congratulations on everything that you've developed. When we look at the Alphazero project, it. It was a very defined and limited variable in terms of the parameters on that, which allowed for the learning curve to be so quick. The risk or what you're trying to do here is almost develop consciousness in the cars through the neural network. And so I guess the challenge is how do you not create a circular reference in terms of. Of the pulling from the centralized model of the fleet to that handoff where the car has enough information. Where is that line? I guess in terms of the point of the learning process to handing it off where there's enough information in the car and not having to pull from the fleet.
Elon Musk
Well, the car can operate if it's completely disconnected from the fleet. It just, it uploads the training that's, you know, better and better as the fleet gets better and better. So simply, if you disconnected it from the fleet from that point onwards, it would stop getting better, but it would still function fine.
Analyst
In the hardware portion of your share. In the previous version, it talked about a lot of the power benefits of not storing a lot of the images. And so in this portion, you're talking about the learning that's going on by pulling from the fleet. I guess I'm having a hard time reconciling how if there was a situation where I'm driving up the hill, as you showed, and I'm predicting where the road is going to go, that's coming from all of the other fleet variables that led to that. Intelligence, how I'm not. How I'm getting the benefit of the low power using the cameras with the neural network. That's where I'm losing the the two. Maybe it's just me, but I guess that's.
Elon Musk
I mean the compute power in the full self driving computer is incredible. And maybe we should mention that if it had never seen that road before, it would still have made those predictions provided it was a road in the United States.
Analyst
March of 9 case here. In the case of LiDAR, the March of 9th isn't there an example I want to just get to your slam on LiDAR because it's pretty clear you don't like LiDAR in this LiDAR flame.
Analyst
Isn't there like a case where at some point nine nine nine nine nine down the road where actually LIDAR may be helpful and why not have it as some sort of a redundancy or backups? That's my first question. And the second. So you can still have your focus on computer vision but just have it as a redundant. My second question is if that is true, what happens to the rest of the industry that's building their Autonomy Solutions on LiDAR?
Elon Musk
They're all going to dump LiDAR. That's my prediction, mark my words. I should point out that I don't actually super hate lidar as much as may Sound, but at SpaceX, SpaceX Dragon uses LiDAR to navigate to the space station and dock. Not only that, SpaceX developed its own LIDAR from scratch to do that. And I spearheaded that effort personally because in that scenario lidar makes sense. And in cars it's pretty friggin stupid. It's expensive and unnecessary. And as Andre was saying, once you solve vision it's worthless. So you have expensive hardware that's worthless on the car. We do have a forward radar which is low cost and is helpful especially for occlusion situations. So if there's like fog or dust or snow, the radar can see through that. If you're going to use active photon generation, don't use visible wavelength because once with passive optical you've taken care of all visible wavelength stuff. You want to use a wavelength that is occlusion penetrating like radar. So LiDAR is just active photon generation in the visual spectrum. If you're going to do active photon generation, do it outside of the visual spectrum in the radar spectrum. So like at 3.8 millimeters versus 400, 700 nanometers you're going to be a much better occlusion penetration and that's why we have a forward radar and then we also have 12 ultrasonics for near field information in addition to the eight cameras and the forward radar. You only need the radar in the forward direction because that's the only direction going real fast. So it's. I mean, we've gone over this multiple times. Like, are we sure we have the right sensor suite? Should we add anything more? No.
Analyst
Hi. So right here. So you had mentioned that you asked the fleet for the information that you're looking for for some of the vision. I have two questions about that. It sounds like the cars are doing some computation to determine what kind of information to send back to you. Is that a correct assumption? Are they doing that in real time or are they doing based on stored information?
Andrej Karpathy
Yep. So they absolutely do computation in real time on the car and we wait to basically specify condition that we're interested in and then those cars do that computation there. If they did not, then we'd have to send all the data and do that offline in our backend. We don't want to do that. So all that competition happens on the car.
Analyst
So it's based on that question. It sounds like you guys are in a really good position to have currently half a million cars in the future, potentially millions of cars that are essentially computers representing free, almost free data centers for you to do computational. Is that a huge future opportunity for Tesla? It's current, current opportunity and that's not really factored in for anything yet. That's incredible. Thank you.
Elon Musk
We have 425,000 cars with hardware two and beyond, which means they've got all eight cameras, the radar and ultrasonics and they've got at least the Nvidia computer, which is enough to essentially figure out what information is important, what is not. Compress the information that is important to the most salient elements and upload it to the network for training. So it's a massive compression of real world data.
Analyst
You have these sort of network of millions of computers which is like massive data centers essentially that are distributed data centers for computational capacity. Do you see it being used for other things besides self driving in the future?
Elon Musk
I suppose it could possibly be used for something besides self driving. We've been super focused on self driving. So, you know, as we get that really nailed, maybe there's going to be some other use for, you know, millions and then tens of millions of computers with hardware three or four driving computer. Yeah, maybe there would be. It could be. It could be. Maybe there's like some sort of aws angle here. It's possible. Hello.
Andrej Karpathy
Hi, Elon. Matt Joyce, Loop Ventures. I own a Model 3 in Minnesota where it snows a lot.
Stuart Bowers
Since camera and radar cannot see road markings through snow, what is your technical
Andrej Karpathy
strategy to solve this challenge? Does it involve high precision GPS at all? Yeah.
Andrej Karpathy
actually, like today, actually, autopilot will do a decent, decent job in snow. Even when landmarkings are covered, even when landlord markings are faded covered, or when there's lots of rain on them, we still seem to drive relatively well. We didn't specifically go after snow yet with our data engine, but I actually think this is, this is completely tractable because in a lot of those images, even when things are snowy, when you ask a human annotator where are the lane lines, they actually could tell you they actually are relatively consistent in creating those lane lines. As long as the annotators are consistent on your data, then I have, there's. The neural network will pick up on those patterns and we'll do just fine. So it's really just about is the signal there even for the human annotator? If, if the answer to that is yes, then the neural network can do it just fine.
Elon Musk
Yeah, there's actually, there are a number of important signals, as Andre was saying. So lane lines are one of those things, but one of the most important signals is drive space. So what is drivable space and what is not drivable space? And what actually really matters the most is drivable space more than lane lines. And the prediction of drivable space is extremely good. And I think especially after this upcoming winter will be incredible. It's like, it will be like, how could it possibly be that good? That's crazy.
Andrej Karpathy
The other thing to point out is maybe it's not even only about human annotators. As long as you as a human can drive through that environment through fleet learning, we actually know the path you took. And you obviously used vision to guide you through that path. You did not just use the lane line markings, you use the entire geometry of that entire scene. So you see how the world is roughly curling. You see how the cars are positioned around you. Neural network will pick up on all those patterns automatically inside it. If you just have enough of the data people traversing those environments.
Elon Musk
Yeah, it's actually extremely important that things not be rigidly tied to gps, because GPS error can vary quite a bit and the actual situation for a road can vary quite a bit. So there could be construction, there could be a detour, and if the car is using GPS as primary this is a real bad situation. It's asking for trouble. It's fine to use GPS for like tips and tricks. So it's like you can drive your home neighborhood better than a neighborhood in some other country or some other part of the country. So you know your own neighborhood well and you use kind of like the knowledge of your neighborhood to drive with more confidence, to maybe have counterintuitive shortcuts and that kind of thing. But you. The GPS overlay data should only be helpful, but never primary. If it's ever primary, it's a problem.
Pete Bannon
So question back here in the back corner.
Analyst
Corner. I just wanted to follow up partially
Stuart Bowers
on that because several of your competitors
Pete Bannon
in the space over the past few
Stuart Bowers
years have made, you know, have talked
Andrej Karpathy
about how they are augmenting all of
Stuart Bowers
their perception and path planning capabilities that are kind of on the car platform
Andrej Karpathy
with high definition maps of the areas that they are driving.
Pete Bannon
Does that play a role in your system? Do you see it adding any value?
Stuart Bowers
Are there areas where you would like to get, get more data that is not collected from the fleet but is more kind of mapping style data?
Elon Musk
I think the high precision, high precision GPS maps and lanes are a really bad idea. The system becomes extremely brittle. So any change like this might, any change to the system makes it, it can't adapt. So if it locks onto GPS and high precision lane lines and does not allow vision override, in fact, vision should be the thing that does everything and then like lane lines are a guideline, but they're not the main thing. We briefly bulked up the tree of high precision lane lines and then realized that was a huge mistake and reversed it out. It's not good.
Stuart Bowers
So this is very helpful for understanding
Andrej Karpathy
annotation, where the objects are and how the car drives. But what about the negotiation aspect for
Pete Bannon
parking and roundabouts and other things where
Stuart Bowers
there are other cars on the road that are human driven, where it's more art than science.
Elon Musk
It does pretty good actually. Like with cut ins and stuff. It's doing really well.
Andrej Karpathy
So like I mentioned, we're using a lot of machine learning right now in terms of predicting kind of creating an explicit representation of what the world looks like. And then there's an explicit planner and a controller on top of that representation. And there's a lot of heuristics for how to traverse and negotiate and so on. There is a long tail just like a. In what visual environments look like. There's a long tail in just those negotiations and a little game of chicken that you play with other people and so on. And so I think we have a lot of confidence that eventually there must be some kind of a fleet learning component to how you actually do that. Because writing all those rules by hand is going to, is going to quickly plateau, I think.
Elon Musk
Yeah, we've dealt with this issue with cut ins and it's like we'll allow gradually more aggressive behavior on the part of the user. They can just dial the setting up and say be more aggressive, be less aggressive. You know, drive easy, chill mode aggressive.
Analyst
Incredible progress. Phenomenal. Two questions. First, in terms of platooning, do you think the system is geared because somebody asked about when there is snow on the road, but if you have platooning feature, you can just follow the car in front. Does your system, is your system capable of doing that? Then I have two follow ups.
Andrej Karpathy
So you're asking about platooning. So I think like we could absolutely build those features. But again if you just use, if you just train neural networks, for example on imitating humans, humans already follow the car ahead. And so that neural network actually incorporates those patterns internally. It's just, it figures out that there's a correlation between the way the car ahead of you faces and the path that you are going to take. But that's all done internally in the net. So you're just concerned with getting enough data and the tricky data. And the neural network training process actually is quite magical. Does all the other stuff automatically. So you turn all the different problems into just one problem. Just collect your data set and use neural network training.
Elon Musk
Yeah, there's three steps to self driving. You know, there's been feature complete, then there's being future complete to the degree that where we think that the person in the car does not need to pay attention. And then there's at a reliability level where we've also convinced regulators that that is true. So there's kind of like three levels. We expect to be feature complete in self driving this year and we expect to be confident enough from our standpoint to say that we think people do not need to touch the wheel, look out of the wheel window sometime probably around, I don't know, second quarter of next year. And then we start to expect to get regulatory approval, at least in some jurisdictions for that towards the end of next year. That's roughly the timeline that I expect things to go on. And probably for trucks, the platooning will be approved by regulators before anything else. And you could have like maybe if you're long haul doing long haul freight, you can have one driver in the front and then have four semis trailing behind in a platooning manner. And I think that probably the regulators will be quicker to approve that than other things.
Analyst
Regarding. Of course, you don't have to convince us. LIDAR is a technology, in my opinion, which has an answer. Looking for a question? Probably dead.
Analyst
I mean this is very impressive what we saw today and probably demo could show something more. I was wondering what is the maximum dimension of a matrix that you may be having in your training or in your deep learning pipeline?
Andrej Karpathy
Ballpark figure, max intimation of the matrix. So yeah, doing a lot of matrix multiply operations inside the neural network. You're asking about the like there's many different ways to answer that question, but I'm not 100% sure if they're. They're useful. They're useful answers. These neural networks will typically have, like I mentioned, about tens to hundreds of millions of neurons. Each of them on average have about a thousand connections to neurons below. So those are the typical scales that are kind of used across the industry and also that we would use as well.
Analyst
Yeah. I've been actually very impressed by the rate of improvement on Autopilot the past year on my Model 3. The two scenarios I wanted your feedback on last week. The first scenario was I was on the right hand most lane of the freeway and there was a highway on ramp. And then my Model 3 actually was able to detect two cars on the side slow down and let the car go in front of me and one car go behind me. And I was like, oh my gosh, this is like insane. Like I didn't think my Model 3 could do that. So that was like super impressive. But the same week another scenario which is I was on the right hand lane again, but my right hand lane was merging with the left lane. It wasn't an on ramp, it's just a normal highway freeway lane. And my Model 3 wasn't able to detect really that situation and I wasn't able to slow down or speed up and I had to intervene kind of. So can you from your perspective, kind of share kind of the background on how a neural net would, how Tesla might adjust for that and you know, like how that could be improved over time?
Andrej Karpathy
Yeah. So like I mentioned, we have a very sophisticated trigger infrastructure. If you have intervened, it's actually potentially likely that we received that clip and that we can actually analyze it and see what happened and tune the system. So it probably enters some statistics over. Okay, at what rate are we correctly merging the traffic? And we look at those numbers and we look at the clips and we see what's wrong and we try to fix those clips and make progress against those benchmarks. So yeah, Yeah. So we would potentially go through a phase of categorization and then we look at some of the biggest kind of categories that actually seem to, to semantically be related to the same problem. And then we will look at some of those and then try to develop software against that.
Elon Musk
Okay. We do have one more presentation which is the software. So it's like essentially the autopilot hardware with Stuart, there's the sort of neural net vision with Andre, and then there's the software engineering at scale that's going to be presented by Stuart. So thanks. And we'll have opportunity afterwards to ask questions. So yeah, thanks.
Investor Relations
I just wanted to very briefly say, if you have an early flight and you want to do a test ride with our latest development software, if you could please speak to my colleague Ann or drop her an email and we can take you out for a test ride. And Stuart, over to you.
Stuart Bowers
All right, so that's actually from a clip of a longer than 30 minute uninterrupted drive with no interventions navigate an autopilot on the highway system which is in production today on hundreds of thousands of cars. So I'm Stuart and I'm here to talk about how we build some of these systems at scale. Just like a really short introduction on kind of where I'm coming from, what I do. So I've been in a couple companies or less. I've been writing Software professional for about 12 years. The thing that excites me most and I'm really passionate about is taking the cutting edge of machine learning and actually connecting that with customers through robustness and scale. So at Facebook, I worked initially inside of our ads infrastructure to build some of the machine learning some really, really smart people. And we actually tried to build it into a single platform that was we could then scale to all the other aspects of the business, from how we rank the news feed to how we deliver search results to how we make every recommendation across the platform. And that became the Applied Machine Learning Group. That's something I'm incredibly proud of. And a lot of that wasn't just the core algorithms and the really important improvements that happened there. Those that matters a lot of actually the engineering practices to build these systems at scale. And the same thing was true at Snap, where I went, where we were really, really excited to sort of actually help to monetize this product. But the hardest part, we were using Google at the time. And they were effectively, you know, running us on a fairly small scale. And we wanted to build that same infrastructure. We take understanding of these users, connect that with cutting edge machine learning, build that at massive scale, and handle billions and then trillions of both predictions and auctions every day in a way which is really robust. And so when the opportunity came to come to Tesla, that's something I'm just like incredibly excited to do, which is specifically take the amazing things that are happening both in the hardware side and the computer vision and AI side and actually package that together with all the planning, the controls, the testing, the kernel patching of the operating system, all of our continuous integration, our simulation, and actually build that into a product. We get onto people's cars in production today. And so I want to talk about the timeline for how we did that with navigate on Autopilot and how we're going to do that as we get navigate on autopilot off the highway and onto city streets. So we're at 70 million miles already for Navigate on Autopilot is something really, really, really cool. And I think one thing that is worth kind of calling out on this is that we're continuing to accelerate and keep learning from this data. Like Andre talked about, this data engine, as this accelerates up, we actually do make more and more assertive lane changes. We are learning from these cases where people intervene either because they fail to detect a merge correctly or because they wanted the car to be a little more peppy in different environments. And we just want to keep making that progress. So to start all of this, we begin with trying to understand the world around us. And we talked about the different sensors in the vehicle. But I wanted to dig in a little bit more. Here we have eight cameras, but then we also have additionally 12 ultrasonic sensors, a radar, an inertial measurement unit, GPS. And then one thing we forget about is we also have the pedal and steering actions. So not only can we look at what's happening around the vehicle vehicle, we can look at how humans chose to interact with that environment. And so I'll talk to this clip right now. This basically is showing what's happening today in the car, and we're continuing to push this forward. So we start with a single neural network. We see the detections around it. We then build all that together with multiple neural networks and multiple detections. We bring in the other sensors and we convert that into what Elon calls a vector space, an understanding of the world around us. And this is something where, as we continue, continue to get better and better at this, we're moving more and more of this logic into the neural networks themselves. And the obvious end game here is that the neural network looks across all the cars, brings in all the information together, and just ultimately outputs a source of truth for the world around us. And this is actually not like an artist rendering, in many senses. This is actually the output of one of the debugging tools that we use on the team every day to understand what the world looks like around us. So another thing that I think is really, really exciting to me, I think when I do hear about sensors like lidar, a common question is around just having extra sensor modalities like why not have some redundancy on the vehicle? And I want to dig in on one thing that's not. Is not always obvious with neural networks themselves. So we have a neural network running on our, say, wide fisheye camera. That neural network is not making one prediction about the world. It's making many separate predictions, some of which actually audit each other. So as a real example, we have the ability to detect a pedestrian. That's a. Something we train very, very carefully on and put a lot of work into. We also have the ability to detect obstacles in the roadway, and a pedestrian is an obstacle. And it's shown differently to the neural network. It says, oh, there's a thing I can't drive through. And these together combine to give us an increased sense of what we can and can't do in front of the vehicle and how to plan for that. We then do this across multiple cameras because we have overlapping fields of view in many places around the vehicle in front, we have a particularly large number of overlapping fields of view. Lastly, we can combine that with things like the radar and the ultrasonics to build these extremely precise understandings of what's happening in front of the car. We can use that both to learn future behaviors that are very accurate. We can also build very accurate predictions of how things will continue to happen in front of us. So one example I think is really exciting is we can actually look at bicyclists and people and not just ask, where are you now? But where are you going? And this is actually the heart of what we're doing for our new next generation automatic emergency braking system, which will not just stop for people in your path, it'll stop for people who are going to be in your path. And that's running in shadow mode right now. We'll go out to the fleet this quarter and I'll talk about shadow mode in A second. So when you want to start a feature like this for navigate on autopilot on the highway system, you can start by learning from data. And you can just look at how humans do things today. What is their assertiveness profile? How do they change lanes? What causes them to either absorb or change their maneuvers? And you can see things that are not immediately obvious, like, oh yeah, simultaneous merging is rare, but very complicated and very important. And you can start to build opinions about different scenarios, such as a fast overtaking vehicle. So this is what we do when we initially have some algorithms we want to try out. We can put them on the fleet and we can see what they would have done in a real world scenario, such as this car that's overtaking us very quickly. This is taken from our actual simulation environment showing different, different paths that we have considered taking and how those overlay on the real world behavior of a user. When you get those algorithms tuned up and you feel good about them specifically, and this is really taking that output of the neural network, putting it in that vector space, and building and tuning these parameters on top of it. Ultimately a thing we can do through more and more machine learning. You go into a controlled deployment, which for us is our early access program. And then you get this out to a couple thousand people who are really excited to give you highly vigilant but useful feedback about how this behaves not in open loop, but in a closed
Andrej Karpathy
loop way in the real world.
Stuart Bowers
And you watch their interventions. And we talked about this like when somebody takes over, we can actually get that clip, try to understand what happens. And one thing we can really do is we can actually play this back again in an open loop way and ask as we build our software, are we getting closer or further from how humans behave in the real world? And one thing which is super cool, with the full self driving computers, we're actually building our own racks and infrastructure so we basically can fit four of our full self driving computers fully racked up, build these into our own cluster, and actually run this very sophisticated data infrastructure to actually understand over time, as we tune and fix these algorithms, are we getting closer and closer to how humans behave? And ultimately can we exceed their capabilities? And so once we had this, we felt really good about it. We wanted to do our wide rollout, but to start, we actually asked everybody to confirm the car's behavior via stock confirm. And so we started making lots and lots of predictions about how we should be navigating the highway. We asked people to tell us, is this right or is this wrong. And this is again a chance to churn that data engine. And we did spot some really tricky and interesting long tails of in this case, I think a really fun example like these very interesting cases of simultaneous merging where you start going and then somebody moves either behind or before you not noticing you. And what is the approach appropriate behavior here and what are the tunings of the neural network we need to do to be super precise about the appropriate behaviors? Here we worked, we tuned these in the background, we made them better, and over the course of time we got 9 million successfully accepted lane changes. And we use these again with our continuous integration infrastructure to actually understand do we think we're ready. And this is one thing where full self driving is also really exciting to me. Since we own the entire software stack straight from the kernel patching all the way to the ISO like the tuning on the image signal processor, we can start to collect even more data that is even more accurate. And this allows us to do even better and better tuning these faster iteration cycles. And so earlier this month we were kind of thought we're ready to deploy an even more seamless version of Navigate on autopilot on the highway system. And that seamless version does not require a stock confirm. So you can sit there, relax, put your hand on the wheel and just oversee what the car is doing. And in this case, we're actually seeing over 100,000 automated lane changes every single day on the highway system. And this is something that's just like super cool to us to deploy at scale. And the thing that I'm kind of most excited about from all this is the actual life cycle of this and how we actually able to turn that data engine crank faster and faster and faster with time. And I think one thing that's really, really becoming very clear is the combination of the infrastructure we have built, the tooling we built on top of that, and the combined power of the full self driving computer, I believe we can do this even faster as we move navigate on autopilot from the highway system onto city streets. And so yeah, with that I'll hand off to Elon.
Elon Musk
Yeah, I mean to the best of my knowledge, all those lane changes have occurred with zero accidents.
Stuart Bowers
That is correct. Yeah, I watch every single accident.
Elon Musk
So. So it's conservative obviously, but it's to have hundreds of thousands going to millions of lane changes and zero accidents is I think a great achievement by the Tesla team.
Elon Musk
Cool. So let's see, you know, a few other things that are maybe worth mentioning. The in order to have a self driving car or robo taxi, you really need redundancy throughout the vehicle at the hardware level. So starting in Maybe it was October 2016, all cars made by Tesla have redundant power steering. So we have redundant motors on the power steering. So any one failure of the if the motor fails, the car can still steer all of the power and data lines have redundancy so you can sever any given power line or any data line and the car will keep driving the auxiliary power system even if the main pack, you lose complete power in the main pack. The car is capable of steering and braking using the auxiliary power system, so you can completely lose the main pack and the car is safe. The whole system from a hardware standpoint has been designed for to be a RoboTaxi since basically October 2016. So when we rolled out hardware autopilot version 2, we do not expect to upgrade cars made before that. We think it would actually cost more to make a new car than to upgrade the cars. Just to give you a sense of how hard it is to do this. Unless it's designed in, it's not worth it. So we've gone through the future of self driving where it's clear it's hardware, it's vision and then there's a lot of software and the software problem here should not be minimizing. It's a massive software problem that
Elon Musk
managing vast amounts of data training against the data. How do you control the car based on the vision? It's a very difficult software problem. So going after going over just like Tesla master plan, obviously we've made a bunch of forward looking statements as they call it. But let's go through some of our other forward looking statements that we've made. Way back when we created the company, we said we'd build the Tesla Roadster. They said it was impossible and that even if we did build it, nobody would buy it. This is like universal opinion was that building an electric car was extremely dumb and would fail. I agreed with them that probability of failure was high, but that this was important. So we built the Tesla roadster production in 2008 and shipping that car, it's now a collector's item. We build a more affordable car with the Model S. We did that again. We were told that's impossible. I was called a fraud and a liar. It was not going to happen. This is all untrue. Okay, famous last words now is we went into production with the Model S in 2012, exceeded all expectations. There is still in 2019 no car that can Compete with the Model S of 2012. It's seven years later, still waiting. So we'd build an affordable car, maybe highly affordable. It's affordable. More affordable with the Model 3. We bought the Model 3, we're in production. I said we'd get over 5,000 cars a week for Model 3. At this point, 5,000 cars a week is a walk in the park for us. It's not even hard. So we do large scale solar, which we did through the solar city acquisition and that we develop and deploy the solar roof which is going really well. We're now on version three of the solar tile roof and we expect to spill a production of the solar tile roof significantly later this year. I have it on my house and it's great. And I sort of make the powerwall and the power pack. We made the power wall and power pack. In fact the power pack is now deployed in massive grid scale utility systems around the world, including the largest operating battery projects in the world that above 100 megawatts. And in the next or probably by next, next year, two years at the most, we expect to have a gigawatt scale battery project completed. So all these things, I said we'd do them, we did it. Said we'd do it, we did it. We're going to do the robotaxi thing too. Only criticism and it's a fair one and sometimes I'm not on time, but I get it done and the Tesla team gets it done. So what we're going to do this year is we're going to reach combined production of 10,000 a week between SX and 3. Feel very confident about that and we feel very confident about being future complete with self driving. Next year we'll expand the product line with Model Y and semi and we expect to have the first operating robotaxis next year with no one in them next year. It's always difficult to like when things are on an exponential, at an exponential rate of improvement. It's very difficult to kind of wrap one's mind around it because we're used to extrapolating on a linear basis. But when you've got massive amounts of like as the hardware, massive amounts of hardware on the road, the cumulative data is increasing exponentially. The software is getting better at an exponential rate. I feel very confident predicting autonomous robotaxis for Tesla next year. Not an older state, not in all jurisdictions because we won't have regulatory approval everywhere. But I'm confident we'll have at least regulatory approval somewhere literally next year. So any customer will be able to add or remove their car to the Tesla network. So expect this to operate like a combination of maybe the Uber and Airbnb model. So if you own the car, you can add or subtract it to the Tesla network, and Tesla would take 25 or 30% of the revenue. And then in places where there aren't enough people sharing their cars, we would just have dedicated Tesla vehicles. So when you use the car, we'll show you our ride sharing app. So you're able to summon the car from the parking lot, get in, and go for a drive. It's really simple. So you just take the same Tesla app that you currently have. We'll update the app and add a summon, summon Tesla, or commit your car to the fleet. So it's either summon your car or summon a Tesla or add or subtract your car to the fleet. You'll be able to do that from your phone. So we see potential for smoothing out the demand distribution curve. And having a car operates at a much higher utility than a normal car operates. So typically the use of a car is about 10 to 12 hours a week. So most people will drive one and a half to two hours a day, typically 10 to 12 hours a week of total driving. But if you have a car that can operate autonomously, then most likely you could probably. Most likely you'd have that car operate for a third of the week or longer. So there are 168 hours in a week. So probably you've got something on the order of 55, 60 hours a week of operation, maybe a bit longer. So the fundamental utility of a vehicle increases by a factor of five. So you can look at this from a macroeconomic standpoint and say, just if this was like some. If we were operating some big simulation, if you could upgrade your simulation to increase the utility of cars by a factor of five, that would be a massive increase in the economic efficiency of the simulation. Just gigantic. So we'll do model 3s3 and excess taxis. But we made an important change to our leases. So if you lease a Model 3, you don't have the option of buying it at the end of the lease. We want them back. If you buy the car, you can keep it, but if you lease it, you have to give it back. And as I said, in any locations where there's not enough supply for sharing, Tesla will just make its own cars and add them to the network in that place. So the current cost of Model 3 Robo Taxi is less than $38,000. We expect that number to improve over time and resigning. The cars, the cars currently being built are all designed for a million miles of operation. The drive units designed and tested and validated for a million miles of operation. The current battery pack is about maybe 300 to 500,000 miles. The new battery pack that probably go into production next year is designed explicitly for a million miles of operation. The entire vehicle battery pack, it's designed to operate for a million miles with minimal maintenance. So we'll actually be adjusting tire design and really optimizing the car for a hyper efficient robotaxi. And at some point you won't need steering wheels or pedals and we'll just delete those. So as, as, as these things become less and less important, we'll just delete parts. Just they won't be there. If you say like probably two years from now, we make a car that has no steering wheels or pedals and if we need to accelerate that time, we can always just delete parts. Easy. Yeah, probably say long term, three years Robotaxis with eliminated parts, maybe it ends up being $25,000 or less. And we want a super efficient car. So the electricity consumption is very low. So we're currently at four and a half miles per kilowatt hour. But we can, we'll improve that to five and beyond. And there's just really no company that has the full stack integration. We've got the vehicle design and manufacturing, but the computer hardware in house. We've got the in house software development and AI and we've got by far the biggest fleet. It's extremely difficult, not impossible perhaps, but extremely difficult to catch up when Tesla has 100 times more miles per day than everyone else combined. This is the cost of running a gasoline car or a. The average cost of running a car in the US is taken from AAA. So it's currently about 62 cents a mile. 13 and a half thousand miles from 15 million vehicles adds up to 2 trillion a year. These are literally just taken from the AAA website. Cost of ride sharing is according to Uber and Lyft is $2 to $3amile. The cost to run a Robotaxi we think less than 18 cents a mile. And dropping. This would be current. This is current cost. Future cost will be lower, You say. What would be the probable gross profit from a single Robotaxi? We think probably something on the order of $30,000 per year. And we expect that. We're literally designing, we're designing the cars the same way that commercial semi trailer, semi trucks are designed. Commercial semi trucks Are all designed for a million mile life and we're designing the cars for a million mile life as well. So in nominal dollars that would be, you know, a little over $300,000. Over the course of 11 years might be higher. I think these consumptions are actually relatively conservative. And this assumes that 50% of the miles driven are. There's nothing are not useful. So this is only at 50% utility. By the middle of next year we'll have over a million Tesla cars on the road with full self driving hardware feature complete at a reliability level that we would consider that no one needs to pay attention. Meaning you could go to sleep. From our standpoint, if you fast forward a year, maybe a year, maybe a year and three months. But next year for sure, we will have over a million robotaxis on the road. The fleet wakes up with an over the air update. That's all it takes. You say what is the net present value of a robotaxi? Probably on the order of a couple hundred thousand dollars. So buying a Model 3 is a good deal. Any questions? Well, I mean in our own fleet, I don't know, I guess long term we have probably on the order of 10 million vehicles. I mean our production rates generally. If you look at our compound annual production rate since 2012 which is like the. That's our first full year of model model S production. We went from 23,000 vehicles produced in 2013 to around 250,000 vehicles produced last year. So in the course of five years we increased output by a factor, factor of 10. I would expect that something similar occurs over the next five or six years. As for sharing, sharing versus I don't know. The nice thing is that essentially customers are fronting us the money for the car. It's great.
Stuart Bowers
So in terms of the one thing is the Snake Charger, I'm curious about that.
Andrej Karpathy
And also how did you determine the pricing?
Stuart Bowers
Looks like like you're undercutting the average lift or Uber ride by about 50%. So I'm curious if you could talk
Andrej Karpathy
a little bit about the pricing strategy.
Elon Musk
Sure. We expect the to solving. Solving for the Snake Charger is pretty straightforward from a vision prop standpoint. It's like a known situation. Any kind of known situation with Vision is like a charge port. It's trivial. So. So yeah, the car was just automatically park and automatically plug in. There would be no one, no human supervision required. Yeah. So sorry, what was pricing? We just threw some numbers on there. I mean I think definitely plug in. Whatever pricing you think makes sense. We just Kind of randomly said, okay, maybe a dollar. And the thing is like there's like on the order of 2 billion cars and trucks in the world. So Robotaxis will be in extremely high demand for a very long time. And my observation thus far is that the auto industry is very slow to adapt. I mean, like I said, there's still not a car on the road that you can buy today that is as good as the Model s was in 2012. So that suggests a pretty slow rate of adaptation for the car industry. And so probably a dollar is conservative for the next 10 years because people sort of think like there's like actually not enough appreciation for the difficulty of manufacturing. Manufacturing is insanely difficult. But a lot of people I talk to think like if you just have the right design, you can instantly make as much of that thing as the world wants. This is not true. It's extremely hard to design a new manufacturing system for new technology. I mean, Audi is having major problems manufacturing E Tron and, and they are extremely good at manufacturing. And if they're having problems, what about others? So the, you know, on the order of 2 billion cars and trucks in the world, on the order of about 100 million units per year of production capacity of vehicles, but only of the old design. It will take a very long time to convert all of that to full self driving cars. And they really need to be electric because the cost of operation of a gasoline diesel car is much higher than electric car. So any, any, any robo tax that isn't electric will absolutely not be competitive.
Stuart Bowers
Elon, it's Colin Rush from Oppenheimer over here. You know, obviously we appreciate that the customers are fronting some of the cash for this, this fleet built up, but it sounds like a massive balance sheet commitment from the organization over the course of time. Can you talk a little bit about
Andrej Karpathy
what that looks like, what your expectations
Stuart Bowers
are in terms of financing over the
Andrej Karpathy
next, call it three years, three, four
Stuart Bowers
years for building up this fleet and starting to monetize it with your customer base?
Elon Musk
Well, we're aiming to be approximately cash flow neutral during the, the fleet buildup phase. And then I would expect to be extremely cash flow positive once the robo taxis are enabled. But I don't want to talk about financing rounds. It would be difficult to talk about financing rounds in this venue. But I think we'll make the right moves. I think we'll make the moves you think we should make.
Analyst
I have a question. If I'm Uber, why wouldn't I just buy all your cars? You Know, why would I let you put me out of business?
Elon Musk
There's a, there's a clause that we put into our cars. I think it was about three or four years ago. They can only be used in the Tesla Network.
Analyst
So even a private person, like if I go out and buy 10 model threes, I can't, I can run on the network. That's a business now. Right.
Elon Musk
You're only right to use Tesla Network.
Analyst
Right. But if I use the Tesla network, in theory, I could run a car sharing robo taxi business with my 10 model threes.
Elon Musk
Yes, but it's like the App Store. You can only add or remove them through the Tesla Network and then Tesla gets revenue share.
Analyst
But similar to Airbnb though, in that I have this home, my car, and now I can just rent them out so I can make an extra income from owning multiple cars and just renting them out. Like I have a Model 3. I aspire to get this roadster here next when you build it. And I'm going to just rent my Model 3 out. Why would I give it back to you? You know,
Elon Musk
I guess you could operate a rental car fleet, but I think this is very unwieldy. Yeah, I don't know. Seems easy. Okay, try it.
Analyst
In order to operate a robotaxi network, it sounds like you have to solve certain problems, like for example, autopilot today, if you oversteer it, it lets you take over. But if it's, you know, if it's a ride sharing product that someone else is getting in the passenger seat, like moving the steering, steering can't let that person take over the car, for example, because they might not even be in the driver's seat. So is the hardware already there for it to be a robo taxi? And it might get into situations such as a cop pulling it over where some human might need to intervene, like using central fleet of operators that remotely sort of interact with humans or I mean, is all of that type of infrastructure already built into each of the cars?
Elon Musk
Does that make sense? I think there will be sort of a phone home thing where if the car gets stuck, it'll just phone home to Tesla and ask for a solution. Things like being pulled over by police offshore. That's easy for us to program in. That's not a problem. It will be possible for some, somebody to take over using the steering wheel at least for some period of time. And then probably down the road we'll just cap the steering wheel so there's no steering control. We'll just take the steering Wheel off, put a cap on in the long. Give it like a couple years hardware
Analyst
modification to the car in order for it to enable that.
Elon Musk
Or yeah, we literally just unbolt the steering wheel and put a cap on where the steering wheel handle currently is.
Analyst
But, but that, that is a like future car that you would put out. But what about today's cars where the steering wheel is a mechanism to take over autopilot like so if it's in a robotaxi mode, would someone be able to take it over by just simply moving the steering wheel type?
Elon Musk
Yes, I think there'll be a transition period where people will be able to take over and should be able to take over from the robotaxi. And then once regulators are comfortable with us not having a steering wheel, we'll just delete that. And for cars that are on the, that are in the fleet, you know, obviously with the permission of the owner, if it's owned by somebody else, we would just take the steering wheel off and put a cap where the steering wheel currently attaches.
Analyst
So there might be like two phases to robotaxi. One where the service is provided and you come in as the driver, but could potentially take over. And then in the future there might not be a driver option. Is that how you see it as
Elon Musk
well or like in the future? There will in future. The probability of the steering wheel being taken away in the future is 100% people. Consumers will demand it.
Analyst
But, but initially you would call up.
Elon Musk
This is not. This is, I'm going to clear, does not meet professional prescribing a point of view about the world. This is me predicting what consumers will demand. Consumers will demand in the future that people are not allowed to drive these three ton death machines.
Analyst
I totally agree with that. But in order for a Model 3 today to be part of the Robotaxi network, when you call it, you would then get into the driver's seat essentially because just to be on the same.
Elon Musk
Okay, that makes sense.
Elon Musk
Exactly. Just a sort of like, you know, there were amphibians, you know, but then pretty much that things just become like land creatures. There'll be a little bit of an amphibian phase.
Elon Musk
Sorry. I can see what the. Okay.
Andrej Karpathy
The strategy we've heard from other players in the robo taxi space is to select a certain municipal area to create geo fenced self driving. That way you're using an HD map to have a more confined area with a bit more safety.
Andrej Karpathy
We didn't hear much today around the importance of HD maps to what Extent is an HD map necessary for you? And second, we also didn't hear much about deploying this into specific municipalities where you're working with the municipality to get
Analyst
the buy in from them and you're
Andrej Karpathy
also getting a more defined area. So what's the importance of HD maps and to what extent are you looking at specific municipalities for rollout?
Elon Musk
I think HTMAPs are a mistake. We actually had HTMAPs for a while. Actually can't can that because you either need HTMAPs, in which case if anything changes about the environment, the car will break down, or you don't need HTML in which case why are you wasting your time doing HD maps? So the HD maps thing, like the two main crutches that should not be used and will in retrospect be obviously false and foolish are LIDAR and HD maps. Mark my words.
Elon Musk
If you need a geofenced area, you don't have real self driving.
Andrej Karpathy
Just it sounds like maybe battery supply could be the only bottleneck left towards this vision. And also could you just clarify how you get the battery packs to last a million miles?
Elon Musk
I think cells will be a constraint. That's a subject for a whole separate. That's a whole separate subject. And I think we're actually going to want to push our sort of standard range plus battery more than our long range battery because the energy content in the long range pack is 50% higher kilowatt hours. So essentially you can make you know, a third more cars if you, if you just. If they're all sort of standard range plus instead of the long range pack. So one's like around 50 kilowatt hours, the other one's around 75 kilowatt hours. So we're actually probably going to bias our sales intentionally towards the smaller battery pack in order to have a higher volume of what basically you want the obvious thing to do is to maximize the number of autonomous units or the number of maximize the output that will substitute result in the biggest autonomous leak down the road. So we're doing a number of things in that regard, but it's just not for today's meeting. The million mile life is basically just about getting the cycle life of the pack to you know, you need basically on the order. Like let's say you've got a basic math, if you've got a 250 mile range pack, you know you're going to need 4,000 cycles. So very achievable. We already do that with our stationary storage. Some of our stationary storage solutions like power pack, we're ready to deploy power pack with 4,000 cycle life capability.
Analyst
Can I ask. Sorry, yeah, I wanted to.
Elon Musk
It's like ventriloquism.
Analyst
No, it's obviously significant. Very constructive margin implications to the extent you can drive attach rates much higher of the full self driving option. I'd just be curious if you can level set kind of where you are in terms of those attach rates and how you expect to educate consumers about the Robotax scenario so that attach rates do materially improve improve over time.
Elon Musk
Sorry, it's a bit hard to hear your question.
Analyst
Yeah, just curious where we are today in terms of full self driving attach rates in terms of the financial implications. I think it's hugely beneficial if those attach rates materially increase because of the higher gross margin dollar that flow through. To the extent people do sign up for full fsd, Just curious how you see that ramping
Andrej Karpathy
or what the attach
Analyst
rates are today versus you know, when do you expect. How do you expect to educate consumers and get them aware that they should be attaching FSD to their vehicle purchases?
Elon Musk
We're going to ramp that up massively after today. Yeah, I mean the fundamental, really fundamental message that consumers should be taking today is that it's financially insane to buy anything other than a Tesla. It'll be like owning a horse in three years. I mean fine if you want to own a horse, but you should go into it with that expectation. If you buy a car that does not have the hardware necessary for full self driving, it's like buying a horse and the only car that has the hardware necessary for full self driving is a Tesla. Like people should really think about their purchase any other vehicle. It's basically crazy to buy any other car than Tesla. We need to make that convey that argument clearly and we will after today.
Analyst
Perfect. Thanks for bringing the future to present very informational time today. I was wondering like you did not talk much about Tesla pickup and let me give a context for that. I could be wrong but the way I'm looking at Tesla Network it was as an early adopter and something as a test bread. I think Tesla's pickup may be the first phase of putting the vehicles in network because the utility of Tesla pickup would be pretty much people who are either loading a lot of stuff or are in the profession of construction or little here and there odd items like picking up stuff from Home Depot. I would say that, you know, maybe it needs to have a two stage process pickup trucks exclusively for Tesla Network as a starting point. Then people like me can buy them later. But what are your thoughts on that?
Elon Musk
Well, today was really just about autonomy. There's, there's a lot that we could talk about such as cell production, pickup truck and future vehicle vehicles. But today was just focus on autonomy. But I agree it's a major thing. I'm very excited for the Tesla pickup truck unveil later this year. It's going to be great.
Stuart Bowers
Colin Lang and UBS Just so we understand the definitions, when you refer to feature complete self driving, it sounds like you're talking level five, no geofence. Is that what's expected by the end of the year?
Elon Musk
Just so we're all.
Stuart Bowers
And then the regulatory process, I mean
Pete Bannon
have you talked to regulators about this? This seems quite an aggressive timeline from
Stuart Bowers
what other people have put out there. I mean are they, you know, what are the hurdles that are needed and what is the timeline to get approval?
Andrej Karpathy
And do you need things like in
Stuart Bowers
California and are they tracking miles that you know, with an operator behind that? Do you need those things? What is that process going to look like?
Elon Musk
Yeah, I mean we talk to regulators around the world all the time as we introduce, you know, additional features like navigate on autopilot. This requires like regulatory approval on a per jurisdiction basis. So but I think fundamentally regulators in my experience are convinced by data. So if you have a massive amount of data that shows that autonomy is safe, they listen to it. They may take time to digest the information they're processed by. May take a bit of time, but they have always come to the right conclusion from what I've seen.
Andrej Karpathy
Oh, I have a question over here.
Elon Musk
I've got license and pillar. Okay.
Andrej Karpathy
I just wanted to, just to, you know, some of the work we've done trying to better understand the ride hail market. It looks like it's very concentrated in major dense urban centers. So is the way to think about this that the robo taxis would probably deploy more into that area and the additional full self driving for personally owned vehicles would be in the suburban areas?
Elon Musk
I think like probably, yeah, like Tesla owned robo taxis would be in dense urban areas along with customer vehicles. And then as you get to medium and low density areas, it would tend to be more that people own the car and occasionally lend it out. Yeah, there are a lot of edge cases in Manhattan and say downtown San Francisco, but those are, you know, and there are various cities around the world that have challenging open environments. But we do not expect this to be a significant issue. And when I say future complete, I mean it will work in downtown San Francisco and downtown Manhattan. This Year.
Andrej Karpathy
Hi, I have a neural net architecture question. Do you use different models for say, path planning and perception or different types of AI and sort of how do you split up that problem across the different pieces of autonomy?
Elon Musk
Well, essentially right now, AI or neural nets are used really for object recognition. And we're still basically just using it as still frames, so identifying objects and still frames and tying it together in a perception path planning layer thereafter. But what's happening is steadily is that the neural net is kind of eating into the software base more and more. And so over time, we expect the neural net to do more and more. Now, from a computational cost standpoint, there are some things that are very simple for heuristic and very difficult for a neural net. And so it probably makes sense to maintain some level of heuristics in the system because they're just computationally a thousand times easier than a neural net. Like a neural net is like a cruise missile, and if you're trying to swat a fly, just use a fly swatter, not a cruise missile. So, but over time, I would expect that it moves really to just training on against video and then video in car steering and pedals out, or basically video in lateral and longitudinal acceleration out almost entirely. That's what we're going to use the dojo system for. There's no system that can currently do that.
Andrej Karpathy
Maybe over here.
Stuart Bowers
Just going back to the sensor suite discussion, Elon. One area I'd like to just talk
Elon Musk
about is a lack of side radars.
Stuart Bowers
And in a situation where you have an intersection with a stop sign, where
Andrej Karpathy
there's maybe a 35, 40 mile per
Stuart Bowers
hour cross traffic, are you comfortable with
Andrej Karpathy
the sensor suite, the side cameras being
Analyst
able to handle that?
Stuart Bowers
Just maybe talk a bit about that?
Elon Musk
Yeah, no problem. Essentially, the car is going to do kind of what a human would do. You can think of a human as like basically a camera on a slow gimbal. And it's quite remarkable that people are able to drive the car in the way that they are, because you can't look in all directions at once. The car can literally look in all directions at once with multiple cameras. So humans are able to drive just by sort of looking this way, looking that way. They're actually stuck in their driver's seat. They can't really get out of the driver's seat. So it's like kind of one camera on a gimbal and is able to drive. A conscientious driver can drive with very high safety. The, the cameras in the cars have a better vantage point than the person. So they're like up in the B pillar or in front of the rear view mirror. They've really got a great vantage point. So if you're turning onto a road that's got a lot of high speed traffic, you can just do what a person does. Just turn a little bit. Don't go fully into the road. Let the camera see what's going on. And if things look good and then the rear cameras don't show any oncoming traffic, off you go. And if it looks sketchy, you can just pull back a little bit. Just like a person. The behavior is like remarkably. It starts to become remarkably lifelike. It's like quite eerie, actually. The car just starts behaving like a person over here. Here we go then.
Stuart Bowers
Trouble quiz right here.
Stuart Bowers
Given all the value you're creating in your auto business by wrapping all of this technology around yourselves, I guess I'm curious as to why you would still be taking some of your cell capacity and putting it into powerwall and power pack. Wouldn't it make sense to put every single unit you can make into this part of your business?
Elon Musk
We're already stolen almost all the cell lines that were meant to go to powerwall and power pack and use them for model three. I mean, last year, in order to make our Model 3 production and not be self starved, we had to convert all of the 2170 lines at the gigafactory to car sales. So our actual output in total gigawatt hours of stationary storage compared to vehicles is an order of magnitude different. And for stationary storage, we can basically use a whole bunch of miscellaneous cells out there. So we can just gather cells from multiple suppliers all around the world, and you don't have a homologation issue or a safety issue like you have with cars. So that's basically our stationary battery business has been just kind of feeding off scraps for quite a while. So. But like, really think of like the production as being. There are many, many constraints of a massive production system. It's like the degree to which manufacturing a supply chain is underappreciated is amazing. There are a whole series of constraints. And what is the constraint in one week may not be the constraint in another week. It's insanely difficult to make a car, especially one which is rapidly evolving. So.
Elon Musk
But I'll just take a few more questions and then I think we'll just break four so you can try out the cars.
Pete Bannon
Hi, Elon, Adam, Jonas.
Stuart Bowers
Questions on safety.
Stuart Bowers
What data can you share with us today?
Pete Bannon
How safe this technology is, which would
Andrej Karpathy
obviously be important in a regulatory or insurance discussion.
Elon Musk
Well, we publish the accidents per mile every quarter. And what we see right now is that autopilot is about twice as safe as a normal, you know, normal driver on average. And we expect that to increase quite a bit over time. Like I said, in the future it will be. Consumers will want to outlaw, and I'm saying they will succeed, nor am I saying I agree with this position, but in the future, consumers will want to outlaw people driving their own cars because it is unsafe. If you think of like elevators, elevators used to be operated on a big lever, like go up and down the floor and there's like a big relay and you had elevator operators, but then periodically they would get tired or drunk or something and then they'd turn the lever at the wrong time and sever somebody in half. So now you do not have elevator operators. And it would be quite alarming if you went into an elevator that had a big lever that could just move between floors arbitrarily.
Elon Musk
So there's just buttons and in the long term, again, not a value judgment. I'm not saying I want the world to be this way. I'm saying consumers will most likely demand that people are not allowed to drive cars.
Pete Bannon
And Elon, a follow up, can you
Stuart Bowers
share with us how much Tesla's spending
Andrej Karpathy
on Autopilot or autonomous technology by order of magnitude on an annual basis? Thank you.
Elon Musk
It's basically our entire expense structure. Question on the, on the economics of the Tesla network. Just so I understand, it looked like. So you get a Model 3 off lease, $25,000 goes on, the balance sheet would be an asset, and then you. It would cash flow $30,000 a year, roughly. Is that the way to think about. Yeah, something like that, yeah.
Andrej Karpathy
And then just in terms of financing
Analyst
of it, there's a question earlier you
Elon Musk
mentioned you would do it. Is it cash flow neutral to the Robotaxi program or cash flow neutral to
Elon Musk
Sorry, the cash flow neutral in terms of.
Andrej Karpathy
He asked a question about financing the
Elon Musk
robo tax, yet it looks to me
Analyst
like they're self financing.
Elon Musk
But you mentioned they would be basically cash flow neutral. Is that what you're referring to? I'm just saying between now and when the Robotaxis are fully deployed throughout the world, the sensible thing for us is to maximize rates and drive the company to cash flow neutral. Once the Robotaxi fleet is active, I would expect to be extremely cash flow positive. And so you were talking about production yeah. To produce them all.
Andrej Karpathy
Okay, thanks.
Elon Musk
Maximize the number of autonomous units made.
Elon Musk
Okay, just maybe one. One last question here. Hello.
Andrej Karpathy
If I add my Tesla to the Robotaxi network, who.
Elon Musk
Who is liable for an accident?
Analyst
Is it Tesla or is it me?
Pete Bannon
If the vehicle has an accident and harms.
Elon Musk
Probably Tesla. It's probably Tesla.
Elon Musk
Yeah. I think the right thing to do is just make sure there are very, very few accidents. All right, thanks, everyone. Please, enjoy the drives.
Andrej Karpathy
Thank you very.
Speaker
Much. Sam. Sa. Sam. It. Sam.
Speaker
It. It. It's. Sa. Sam. It. Ra.