Examenos

Explore benchmarks · compare every showcased model

Loading…

The benchmark library

Choose what to compare.

Browse by category, then open a benchmark to see every model’s results.

Source-linked benchmark results

Official / vendor

DeepSWE v1.1

Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).

26 / 30 models with results →
Official / vendor

AutomationBench v1.0.6

General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.

18 / 30 models with results →
Official / vendor

Terminal-Bench 2.1

Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.

17 / 30 models with results →
Official / vendor

GPQA Diamond

Graduate-level science QA

15 / 30 models with results →
Official / vendor

Terminal-Bench 4.0

Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.

15 / 30 models with results →
Official / vendor

Agents' Last Exam

Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.

14 / 30 models with results →
Official / vendor

OSWorld 2.0

Computer-use / desktop agent: operate a real GUI environment to complete open-ended tasks (partial scores must be noted).

14 / 30 models with results →
Official / vendor

SWE-bench Pro

Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.

12 / 30 models with results →
Official / vendor

HealthBench Professional

OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).

11 / 30 models with results →
Official / vendor

JobBench

Job/task agent benchmark

11 / 30 models with results →
Official / vendor

GDPval-AA v2

GDPval-AA knowledge-work Elo (v2 as reported)

10 / 30 models with results →
Independent

Intelligence Index · Artificial Analysis

Recorded Intelligence Index measurements from Artificial Analysis, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.

22 / 30 models with results →
Independent

Vals Index · Vals AI

Recorded Vals Index measurements from Vals AI, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.

21 / 30 models with results →
Independent

Bugs fixed /105 · Bug Hunt Bench

Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.

20 / 30 models with results →
Independent

ARC-AGI-2 · ARC Prize

Recorded ARC-AGI-2 measurements from ARC Prize, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.

10 / 30 models with results →

The interactive library loads every tracked metric. Browse public benchmark tables or model profiles.

Save benchmark

Collections are saved only in this browser.