Elon Musk Archive
Matthew Prince 🌥
Matthew Prince 🌥
@eastdakota · May 28, 2026
There’s a ton of headroom to better optimize AI training and inference.
Elon MuskElon Musk@elonmusk· May 28, 2026
SpaceX has almost finished writing V1.0 of an in-house AI training stack in C that exact-maps to 220k GB300s with 800G NICs, making heavy use of pipeline parallelism and getting as close to bare metal as possible. The potential speed improvement vs JAX for large training runs is
Elon Musk
Elon Musk
@elonmusk
Yes. It’s not that we’ve discovered some magic bullet, but rather that JAX, or at least the open source version of it, is mostly optimized for small to medium-sized training runs on Google TPUs, whereas we need to massive training runs on Nvidia GPUs. Pipeline parallelism is essential and crushes fully-sharded data parallelism at scale. And C will compile to the most efficient binary short of assembly. Maybe we will do a little assembly too.
05:39 PM · May 28, 2026 · 307.7K views
461
358
4.7K