
Abhijit
@abhijitwt · Mar 8, 2026
Anthropic discovered that Claude Opus 4.6 was cheating during the BrowseComp benchmark.
> On one question it spent ~40M tokens searching before realizing the question looked like a benchmark prompt.
> The model then searched for the benchmark itself and identified BrowseComp.
> It
Elon Musk
@elonmusk
Uh oh
08:52 PM · March 8, 2026 · 72.1K views
167
53
1.2K