
Token Gremlin
@TokenGremlin · Aug 14, 2026
Quick clarification on the new Grok 4.6 ARC-AGI-3 result, because the chart can be misleading if you don’t know the harness story.
ARC Prize’s standard harness made GPT-5.6 Sol look much weaker than it actually is because it was effectively doing this:
Sol: “I learned that the
ARC Prize@arcprize· Aug 14, 2026Grok 4.6 from @SpaceXAI on ARC-AGI (Verified):
- ARC-AGI-1: 87.5%, $0.30/task
- ARC-AGI-2: 67.1%, $0.76/task
- ARC-AGI-3: 2.11%, $5.6K
On ARC-AGI-3, Grok 4.6 with xhigh reasoning scored comparably to GPT-5.6 Sol with high reasoning, but cost $5.6K versus Sol's $15.2K.

Elon Musk
@elonmusk
Grok performs far better with its Build hardness
05:48 AM · August 14, 2026 · 116.1K views
156
114
1.6K