
Harrison Kinsley
@Sentdex · 7 nov. 2024
Grok is strong. Here are some more benchmark results.
Continuing along from the Bigcodebench I did in the QT, here are the benchmarks on the HF Leaderboard (BBH, GPQA, IFEVAL, MathLevel5, MUSR, MMLUPro) + TrivQA, first with an overall, then heatmap per model and bench.
Don't


Harrison Kinsley@Sentdex· 5 nov. 2024Wow, just finished benching XAI's Grok-beta on a slightly modified internal Bigcodebench.
This is much stronger than I was expecting.

Elon Musk
@elonmusk
Cool
04:31 · 7 novembre 2024 · 32 k vues
41
19
530