Discussion about this post

User's avatar
George Campbell's avatar

I quick addition to the piece. I'm hearing rumors about the next wave of models. Qwen, xAI/Curosr, DeepSeek, Meta, and the next wave from Anthropic (Opus 5 and Fable 5) and OpenAI (GPT 6). I suspect that the US Government will attempt to throttle the Chinese models with some kind of restrictions, especially as more US based open models come to market. The benchmarks will be very useful to understand when these all get rolled out in the upcoming weeks/months.

QUASAR's avatar

This feels like the benchmark version of a problem education has had for years.

A single score looks precise, but it often hides several different abilities underneath it.

Did the model know the answer? Could it apply that knowledge? Could it sustain the work? Was the result actually good? Or did people simply prefer how it sounded?

Those are not interchangeable.

The same is true with students. A grade can compress recall, reasoning, wording, confidence, and transfer into one number—and then tempt us to call the number “understanding.”

Maybe the first question around any benchmark should be:

What capability would this score still fail to reveal?

No posts

Ready for more?