
Apurv Kochara
@Kochara13 · 16. Feb. 2026
HLE is a challenging benchmark made up of nearly 2,500 problems from experts in multiple disciplines.
Most AI systems perform poorly, but Grok 4 and Grok 4 Heavy achieve the highest scores.

Elon Musk
@elonmusk
Actually, I don’t think HLE is a great measure of usefulness. We’re moving away from these benchmarks in favor of making Grok maximally useful for actual engineering.
08:07 · 16. Februar 2026 · 21.922 Aufrufe
77
43
659