← All posts

GuidesEge Burock

How to read AI benchmarks

What GPQA, Humanity's Last Exam, SWE-bench and overall intelligence indexes measure: a short guide to reading model comparisons.

#benchmarks#ai#models#guide

How to read AI benchmarks

Every new model launch comes with a benchmark table, and every lab calls its own model “the best”. To read these tables properly, you just need to know what each test measures.

The tests you will see most

  • GPQA Diamond: 198 PhD-level multiple-choice questions in biology, physics and chemistry. Domain experts score about 65%, skilled non-experts with web access about 34%. It measures scientific reasoning.
  • Humanity’s Last Exam (HLE): 2,500 hard questions from dozens of fields, released in January 2025 by the Center for AI Safety and Scale AI. One of the tests current models struggle with most.
  • SWE-bench Verified: 500 human-checked tasks built from real GitHub issues in 12 open-source Python projects. It measures whether a model can actually fix a real software bug.
  • Overall indexes: measures such as the Artificial Analysis Intelligence Index combine many tests into one score. Handy for quick comparisons.

3 things to watch

  1. Settings: the same model can score very differently at “low” and “max” reasoning effort. Make sure you compare numbers at the same setting.
  2. Who measured it? Scores a lab reports about itself can differ from independent measurements. Independent results are more reliable.
  3. One test is not everything: a model that leads in coding can be average at writing. Look at the tests closest to your work.

See the current ranking on our AI Models page.

A deeper look

Contamination. If a test’s questions are published online, models may have seen them during training. High scores on old, popular tests can therefore reflect memorization more than ability, which is why new and held-out test sets are more valuable.

Saturation. Once models pass 90% on a test, it can no longer tell them apart. That is why older tests like MMLU are being replaced by harder ones such as GPQA and HLE.

Cost and speed are performance too. A model that reaches the same score 10x cheaper or faster is often the better choice in practice. That is why Artificial Analysis’ “cost per task” and “speed” columns matter; the table on our AI models page includes them.

Human preference rankings. Platforms such as LMArena compute Elo scores from users voting between two models’ answers. They reflect real user taste but can favour long, polished answers.

Our take

When choosing a model, look at three things together rather than one number: results on tests close to your work, independent measurement, and cost/speed. Benchmark tables are a good start; a short trial on your own task has the final say.

Sources

  1. Wikipedia: Humanity's Last Exam
  2. Epoch AI: SWE-bench Verified
  3. Nanonets: AI benchmarks explained (GPQA, SWE-bench, Arena Elo)
  4. Artificial Analysis: model leaderboard

Related reading