Examenos

How to compare LLM benchmark scores fairly

A practical guide to comparing LLM results across benchmark versions, source types, harnesses and reasoning effort without inventing an overall ranking.

Start with the question you need answered

A benchmark measures a particular task under a particular setup. A coding result does not establish performance on every coding workflow, and a knowledge score does not replace an evaluation of your own application. Start with the task, then choose relevant evaluations.

Match the benchmark and setup

Compare the same benchmark revision, subset and scoring unit. Agent tooling, prompts, judges, model snapshots and reasoning effort can change the setup. Examenos records these caveats in the source notes when they are available. Expand a result before treating two numbers as directly comparable.

For example, vendor-reported GPQA Diamond results are a distinct page from an evaluator's measurements. A higher number from a different harness is not automatically a fair head-to-head result.

Look at evidence, not only the default number

The default is the recorded headline for the model and metric. Alternative efforts and harnesses remain in the evidence. Source links let you inspect the original report, date and configuration. A peer vendor's comparison table is identified separately from a model's own vendor.

Keep gaps visible

An em dash means no result is recorded for that cell. It is not zero and it is not a failed test. Avoid ranking models by the number of filled cells or silently averaging unrelated benchmark units.

Use the Compare tool to inspect up to five models. Use our methodology to understand headline selection and the limits of the dataset.

Sources and further reading

These explanations describe Examenos's data and display policy. Numeric examples are available in the linked model/benchmark records, each with original citations. Examenos curates results rather than conducting these evaluations.