
will depue
@willdepue · 28. Mai 2026
i know the solution to the AI benchmark problem but nobody is gonna like it
it’s easy: just report test perplexity on uncontaminated high-quality code/lang/etc
you give me base model api. i run on my secret dataset. i give you test ppl. all evals are downstream of that. solved
Elon Musk
@elonmusk
Accurate
14:00 · 29. Mai 2026 · 126.400 Aufrufe
146
110
1309