The benchmark library
Choose what to compare.
Browse by category, then open a benchmark to see every model’s results.
Data coverage
Choose a model to see which benchmarks have recorded results and which are missing.
Explore coverage → Explore changes by dateResults over time
Choose a benchmark, then drag its timeline to see the scores we had recorded at each date.
Open benchmark timelines →Source-linked benchmark results
DeepSWE v1.1
Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).
26 / 30 models with results →Official / vendorAutomationBench v1.0.6
General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.
18 / 30 models with results →Official / vendorTerminal-Bench 2.1
Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.
17 / 30 models with results →Official / vendorGPQA Diamond
Graduate-level science QA
15 / 30 models with results →Official / vendorTerminal-Bench 4.0
Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.
15 / 30 models with results →Official / vendorAgents' Last Exam
Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.
14 / 30 models with results →Official / vendorOSWorld 2.0
Computer-use / desktop agent: operate a real GUI environment to complete open-ended tasks (partial scores must be noted).
14 / 30 models with results →Official / vendorSWE-bench Pro
Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.
12 / 30 models with results →Official / vendorHealthBench Professional
OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).
11 / 30 models with results →Official / vendorJobBench
Job/task agent benchmark
11 / 30 models with results →Official / vendorGDPval-AA v2
GDPval-AA knowledge-work Elo (v2 as reported)
10 / 30 models with results →IndependentIntelligence Index · Artificial Analysis
Recorded Intelligence Index measurements from Artificial Analysis, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
22 / 30 models with results →IndependentVals Index · Vals AI
Recorded Vals Index measurements from Vals AI, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
21 / 30 models with results →IndependentBugs fixed /105 · Bug Hunt Bench
Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
20 / 30 models with results →IndependentARC-AGI-2 · ARC Prize
Recorded ARC-AGI-2 measurements from ARC Prize, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
10 / 30 models with results →The interactive library loads every tracked metric. Browse public benchmark tables or model profiles.
| Provider | Published / checked | Evidence |
|---|