
Explainers·3 min read
How to Read an AI Benchmark: A Complete Guide
Every week a lab announces a record. But what does a benchmark actually measure, why does the announcing company always win, and which numbers are worth anything to you?
Tags
3 stories found

Every week a lab announces a record. But what does a benchmark actually measure, why does the announcing company always win, and which numbers are worth anything to you?

For the first time a Max-tier Qwen model has shipped its weights publicly. It carries 2.4 trillion parameters but activates only 95 billion per query. The catch: almost nobody has the hardware to run it.

xAI's new model leads GPT-5.6 Sol on several benchmarks and reaches 61 on the Artificial Analysis index. But the widely repeated claim that it took first place on CursorBench does not match xAI's own numbers.