
Apurv Kochara
@Kochara13 · 16 févr. 2026
HLE is a challenging benchmark made up of nearly 2,500 problems from experts in multiple disciplines.
Most AI systems perform poorly, but Grok 4 and Grok 4 Heavy achieve the highest scores.

Elon Musk
@elonmusk
Actually, I don’t think HLE is a great measure of usefulness. We’re moving away from these benchmarks in favor of making Grok maximally useful for actual engineering.
08:07 · 16 février 2026 · 21,9 k vues
77
43
659