
Steven Bower
@bowerblu · 13 juil. 2026
Everyone still stuck on software engineering benchmarks.
Prediction: within a year, all models will be so good at software we will stop caring about measuring it, and start focusing on other surprising capabilities.
0xMarioNawfal@RoundtableSpace· 12 juil. 2026Gemini 3.5 Pro benchmark leak just dropped and the numbers are turning heads.
> Reportedly outperforming Claude Fable 5 and GPT-5.6 in internal evals
> Significant zero-shot performance improvements over 3.1 Pro
> Currently in private validation and testing
- Public rollout

Elon Musk
@elonmusk
True
13:43 · 21 juillet 2026 · 33,2 k vues
61
18
201