Grok 4.6 Is a Big Jump, Just Not the One the Headlines Claim
xAI's new model leads GPT-5.6 Sol on several benchmarks and reaches 61 on the Artificial Analysis index. But the widely repeated claim that it took first place on CursorBench does not match xAI's own numbers.

xAI released Grok 4.6, and by its own benchmarks it is a substantial jump over the previous version.
On the Artificial Analysis Intelligence Index it reaches 61, matching OpenAI's GPT-5.6 Sol. It also overtakes Kimi K3, landing third in the world on that index.
One correction is needed
Over the past few days this story circulated under the headline "Grok 4.6 takes No. 1 on CursorBench, ahead of Claude Fable 5." That is not accurate.
Grok 4.6 scores 69.9% on CursorBench v3.2, up from 66.7%, a real improvement.
But in xAI's own comparison table, Fable 5 Max sits higher at 70.5%. Grok is second, not first. Why launch tables are usually arranged this way is a story in itself.
This kind of error usually appears when a story is rewritten from the launch tweet rather than the underlying table.
Where it actually leads
By xAI's figures, the model leads GPT-5.6 Sol on CursorBench, FrontierCode and AA-Briefcase.
It still trails on DeepSWE, the benchmark that measures real software-engineering issue resolution.
The fair summary: Grok 4.6 is best in several specific areas, not in others, and as with any model it depends on what you are doing.
Why this goes beyond xAI
Benchmarks are drifting into marketing across artificial intelligence. Every lab publishes a table in which its own model wins, simply by choosing which columns to show.
The fix is not complicated: before believing a headline, look at the source's own table. It is usually right there, and usually more complete than the headline. 👀




