
David
@DavidSHolz · 13. Mai 2024
Comparing LLMs on lmsys is fun, but 10 points of Elo 'lead' means a model wins 51.4% of the time and 100 points = wins 64% of the time (0.14% per pt for < 100 diff). So either it's not a good measure, or the models aren't very different. Personally, I think we need better evals!
Elon Musk
@elonmusk
Yeah, better evals needed
20:25 · 13. Mai 2024 · 30.952 Aufrufe
32
18
431