The benchmark library
Choose what to compare.
Browse by category, then open a benchmark to see every model’s results.
Data coverage
Choose a model to see which benchmarks have recorded results and which are missing.
Explore coverage → Explore changes by dateResults over time
Choose a benchmark, then drag its timeline to see the scores we had recorded at each date.
Open benchmark timelines →Loading benchmark library…
Humanity's Last Exam
HLE without tools
12 with results · 22 missing · 34 models
Higher is better · %. Missing results are not zero. Expand Evidence for source dates, efforts and harness notes.
| Provider | Published / checked | Evidence | ||
|---|---|---|---|---|
| Claude Opus 5.5Anthropic | Anthropic | 64.4% | 2026-10-07 | All 5 recorded results & sources52.9% · raw 52.9 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12059% · raw 59 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12059.6% · raw 59.6 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12062.8% · raw 62.8 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12064.4% · raw 64.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=120 |
| Claude Fable 5.1Anthropic | Anthropic | 60.9% | — | All 2 recorded results & sources60.9% · raw 60.9 % Headline · unknown effort · Own vendor Source/record date: Not recorded from chart/figure Anthropic Fable page; Humanity's Last Exam no tools https://www.anthropic.com/claude/fable59.1% · raw 59.1 % Alternative · max effort · Peer vendor Source/record date: 2026-09-20 Harness: StepFun vendor-reported evaluation; No-tools HLE row; subset not restated in table. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Approved peer, max effort as labelled; not an independent run. Existing headline retained; alternate vendor configuration/evidence, no averaging. https://www.stepfun.com/step-5-preview |
| Claude Sonnet 5.5Anthropic | Anthropic | 56.9% | 2026-09-28 | All 6 recorded results & sources56.9% · raw 56.9 % Headline · max effort · Own vendor Source/record date: 2026-09-28 HLE no-tools at max https://www.anthropic.com/claude-sonnet-5-5-system-card41.4% · raw 41.4 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12043.2% · raw 43.2 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12047.9% · raw 47.9 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12053% · raw 53 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12056.9% · raw 56.9 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=120 |
| GPT-6 AstraOpenAI | OpenAI | 54.7% | 2026-09-20 | All 1 recorded result & sources54.7% · raw 54.7 % Headline · max effort · Peer vendor Source/record date: 2026-09-20 Harness: StepFun vendor-reported evaluation; No-tools HLE row; subset not restated in table. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Approved peer, max effort as labelled; not an independent run. https://www.stepfun.com/step-5-preview |
| Step 5 PreviewStepFun | StepFun | 46.5% | 2026-09-20 | All 1 recorded result & sources46.5% · raw 46.5 % Headline · high effort · Own vendor Source/record date: 2026-09-20 Harness: StepFun vendor-reported evaluation; No-tools HLE row; subset not restated in table. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort. https://www.stepfun.com/step-5-preview |
| Claude Haiku 5.5Anthropic | Anthropic | 45.9% | 2026-10-07 | All 5 recorded results & sources30.9% · raw 30.9 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12035.5% · raw 35.5 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12039.7% · raw 39.7 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12044.1% · raw 44.1 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12045.9% · raw 45.9 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=120 |
| Kimi K3Moonshot AI | Moonshot AI | 43.5% | 2026-07-16 | All 2 recorded results & sources43.5% · raw 43.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 HLE-Full without tools (with tools 56.0 stored separately) https://github.com/MoonshotAI/Kimi-K346.9% · raw 46.9 % Alternative · max effort · Peer vendor Source/record date: 2026-09-20 Harness: StepFun vendor-reported evaluation; No-tools HLE row; subset not restated in table. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Approved peer, max effort as labelled; not an independent run. Existing headline retained; alternate vendor configuration/evidence, no averaging. https://www.stepfun.com/step-5-preview |
| GLM-5.3Z.ai | Z.ai | 42% | 2026-09-10 | All 2 recorded results & sources42% · raw 42 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 from chart/figure DeepSeek V4.1 Flash table; HLE text-only subset (*) https://api-docs.deepseek.com/news/news26091042.3% · raw 42.3 % Alternative · max effort · Peer vendor Source/record date: 2026-09-20 Harness: StepFun vendor-reported evaluation; No-tools HLE row; subset not restated in table. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Approved peer, max effort as labelled; not an independent run. Existing headline retained; alternate vendor configuration/evidence, no averaging. https://www.stepfun.com/step-5-preview |
| GLM-5.3-FlashZ.ai | Z.ai | 39.9% | 2026-09-30 | All 1 recorded result & sources39.9% · raw 39.9 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Ling 3.1 FlashinclusionAI | inclusionAI | 37.4% | 2026-09-30 | All 1 recorded result & sources37.4% · raw 37.44 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| DeepSeek V4.1 FlashDeepSeek | DeepSeek | 36.8% | 2026-09-10 | All 2 recorded results & sources36.8% · raw 36.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 HLE Pass@1; dagger variant 39.1 noted in card https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash39.2% · raw 39.2 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Differs from current headline 36.8 (first-party); kept with provenance, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Qwen3.8 27BAlibaba (Qwen) | Alibaba (Qwen) | 30.8% | 2026-08-14 | All 1 recorded result & sources30.8% · raw 30.8 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Humanity's Last Exam (text; no tools called out) https://huggingface.co/Qwen/Qwen3.8-27B |
| BeamReflection AI | Reflection AI | — | — | No recorded result |
| Gemini 3.8 FlashGoogle | — | — | No recorded result | |
| Gemini 4 ArgonGoogle | — | — | No recorded result | |
| GPT-5.6 TerraOpenAI | OpenAI | — | — | No recorded result |
| GPT-6 LunaOpenAI | OpenAI | — | — | No recorded result |
| GPT-6.1 SolOpenAI | OpenAI | — | — | No recorded result |
| Grok 4.7xAI | xAI | — | — | No recorded result |
| Hy4 previewTencent | Tencent | — | — | No recorded result |
| Kolibri-1Aleph Alpha | Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | inclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | Xiaomi | — | — | No recorded result |
| MiMo-V2.6-FlashXiaomi | Xiaomi | — | — | No recorded result |
| MiMo-V2.6-ProXiaomi | Xiaomi | — | — | No recorded result |
| Mistral Large 4Mistral AI | Mistral AI | — | — | No recorded result |
| Muse Spark 1.3Meta | Meta | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | NaiveAI | — | — | No recorded result |
| Nex-N2.5-ProNex AGI | Nex AGI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | Unbiased | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Max (0902)Alibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |