The benchmark library
Choose what to compare.
Browse by category, then open a benchmark to see every model’s results.
Data coverage
Choose a model to see which benchmarks have recorded results and which are missing.
Explore coverage → Explore changes by dateResults over time
Choose a benchmark, then drag its timeline to see the scores we had recorded at each date.
Open benchmark timelines →Source-linked benchmark results
DeepSWE v1.1
Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).
28 / 33 models with results →Official / vendorAutomationBench v1.0.6
General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.
18 / 33 models with results →Official / vendorTerminal-Bench 2.1
Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.
18 / 33 models with results →Official / vendorTerminal-Bench 4.0
Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.
18 / 33 models with results →Official / vendorGPQA Diamond
Graduate-level science QA
16 / 33 models with results →Official / vendorAgents' Last Exam
Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.
14 / 33 models with results →Official / vendorOSWorld 2.0
Computer-use / desktop agent: operate a real GUI environment to complete open-ended tasks (partial scores must be noted).
14 / 33 models with results →Official / vendorSWE-bench Pro
Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.
13 / 33 models with results →Official / vendorHealthBench Professional
OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).
12 / 33 models with results →Official / vendorFrontierCode 1.1 Main
FrontierCode main split
11 / 33 models with results →Official / vendorGray Swan IPI
Gray Swan Indirect Prompt Injection benchmark: attack success rate at K=15 attempts. Lower is better.
11 / 33 models with results →IndependentIntelligence Index · Artificial Analysis
Recorded Intelligence Index measurements from Artificial Analysis, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
24 / 33 models with results →IndependentVals Index · Vals AI
Recorded Vals Index measurements from Vals AI, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
22 / 33 models with results →IndependentBugs fixed /105 · Bug Hunt Bench
Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
21 / 33 models with results →IndependentARC-AGI-2 · ARC Prize
Recorded ARC-AGI-2 measurements from ARC Prize, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
12 / 33 models with results →The interactive library loads every tracked metric. Browse public benchmark tables or model profiles.
| Provider | Published / checked | Evidence |
|---|