Elon Musk Archive
Apurv Kochara
Apurv Kochara
@Kochara13 · Feb 16, 2026
HLE is a challenging benchmark made up of nearly 2,500 problems from experts in multiple disciplines. Most AI systems perform poorly, but Grok 4 and Grok 4 Heavy achieve the highest scores.
Apurv Kochara
Elon Musk
Elon Musk
@elonmusk
Actually, I don’t think HLE is a great measure of usefulness. We’re moving away from these benchmarks in favor of making Grok maximally useful for actual engineering.
08:07 AM · February 16, 2026 · 21.9K views
77
43
659