The benchmark library
Choose what to compare.
Browse by category, then open a benchmark to see every model’s results.
Data coverage
Choose a model to see which benchmarks have recorded results and which are missing.
Explore coverage → Explore changes by dateResults over time
Choose a benchmark, then drag its timeline to see the scores we had recorded at each date.
Open benchmark timelines →Loading benchmark library…
FrontierCode 1.1 Main
FrontierCode main split
11 with results · 22 missing · 33 models
Higher is better · %. Missing results are not zero. Expand Evidence for source dates, efforts and harness notes.
| Provider | Published / checked | Evidence | ||
|---|---|---|---|---|
| Claude Opus 5.5Anthropic | Anthropic | 54.4% | 2026-09-22 | All 7 recorded results & sources54.4% · raw 54.4 % Headline · max effort · Own vendor Source/record date: 2026-09-22 FrontierCode v1.1 Main; adaptive max effort per table default https://www.anthropic.com/claude-opus-5-554.6% · raw 54.6 % Alternative · medium effort · Own vendor Source/record date: 2026-09-22 Default/medium effort on FrontierCode v1.1 main (cost chart text) https://www.anthropic.com/claude-opus-5-547.3% · raw 47.3 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 9061, "cost": 0.4036, "duration_min": 4.87, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode54.6% · raw 54.64 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 18519, "cost": 0.8016, "duration_min": 6.34, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode54% · raw 53.99 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 26143, "cost": 1.0896, "duration_min": 7.2, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode51.4% · raw 51.42 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 60182, "cost": 2.2545, "duration_min": 11.21, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode54.4% · raw 54.43 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 165811, "cost": 6.191, "duration_min": 25.01, "flagged_rate": 0.008}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| GPT-6 AstraOpenAI | OpenAI | 53.3% | 2026-09-03 | All 2 recorded results & sources53.3% · raw 53.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 FrontierCode 1.1 Main; footnote 8 https://openai.com/index/gpt-6-astra/53.3% · raw 53.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value. https://www.anthropic.com/claude-opus-5-5 |
| Claude Sonnet 5.5Anthropic | Anthropic | 52.1% | 2026-09-28 | All 9 recorded results & sources52.1% · raw 52.1 % Headline · xhigh effort · Own vendor Source/record date: 2026-09-28 FrontierCode v1.1 Main best at xhigh; max scores lower (46.2): Max-effort code-review skill caused timeouts and out-of-scope penalties https://www.anthropic.com/claude-sonnet-5-546.2% · raw 46.2 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 Max-effort score; below xhigh best of 52.1 https://www.anthropic.com/claude-sonnet-5-546.2% · raw 46.2 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 113. Cognition official revision 1.1; Main 100 tasks, five runs/task. Claude Code for Claude; Codex for GPT. Max-effort summary, distinct from best-effort leaderboard. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11352.1% · raw 52.1 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 113. Best published Sonnet effort is xhigh; retain max result separately. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11329.3% · raw 29.33 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 6353, "cost": 0.194, "duration_min": 3.83, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode36.5% · raw 36.48 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 7597, "cost": 0.2365, "duration_min": 4.35, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode49.4% · raw 49.38 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 13003, "cost": 0.4175, "duration_min": 5.94, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode52.1% · raw 52.09 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 47385, "cost": 1.5881, "duration_min": 12.61, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode46.2% · raw 46.21 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 595651, "cost": 20.7818, "duration_min": 62.39, "flagged_rate": 0.0465}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| Claude Fable 5.1Anthropic | Anthropic | 50.3% | 2026-09-22 | All 7 recorded results & sources50.9% · raw 50.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported in OpenAI GPT-6 Astra comparison table | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Opus 5.5 announce (anthropic.com/claude-opus-5-5); model vendor's own (Anthropic) figure preferred over peer-vendor OpenAI comparison table (rule a); Anthropic source is also later-dated (2026-09-22). Values differ (50.9 here vs 50.3 Anthropic); not reconciled. https://openai.com/index/gpt-6-astra/50.3% · raw 50.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce https://www.anthropic.com/claude-opus-5-549.8% · raw 49.82 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 19072.6, "cost": 2.3849, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode50.9% · raw 50.91 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 26109.6, "cost": 3.2845, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode50.3% · raw 50.34 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 39918.6, "cost": 5.2741, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode48.7% · raw 48.73 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 67707.6, "cost": 9.2738, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode50.3% · raw 50.28 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 91731.2, "cost": 12.8257, "flagged_rate": 0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| Grok 4.7xAI | xAI | 47.6% | 2026-09-21 | All 1 recorded result & sources47.6% · raw 47.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 Cognition FrontierCode 1.1 Main leaderboard JSON-LD ItemList: Grok 4.7 Score 47.6% (added Sep 21, 2026 changelog) https://cognition.com/frontiercode |
| Claude Haiku 5.5Anthropic | Anthropic | 46.4% | 2026-10-07 | All 7 recorded results & sources46.4% · raw 46.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 113. Cognition official revision 1.1; Main 100 tasks, five runs/task. Claude Code for Claude; Codex for GPT. Max-effort summary, distinct from best-effort leaderboard. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11345.8% · raw 45.8 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 111. Explicit xhigh score in summary table. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11134.8% · raw 34.76 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 23868, "cost": 0.0596, "duration_min": 6.71, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode41.6% · raw 41.63 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 36172, "cost": 0.13, "duration_min": 8.67, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode41.9% · raw 41.89 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 55149, "cost": 0.256, "duration_min": 11.06, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode45.8% · raw 45.83 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 100847, "cost": 0.6297, "duration_min": 16.0, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode46.4% · raw 46.36 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 181387, "cost": 1.3284, "duration_min": 23.55, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| Gemini 3.8 FlashGoogle | 43.6% | 2026-09-03 | All 1 recorded result & sources43.6% · raw 43.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported in OpenAI GPT-6 Astra table https://openai.com/index/gpt-6-astra/ | |
| GPT-6 LunaOpenAI | OpenAI | 42.4% | 2026-10-07 | All 6 recorded results & sources42.4% · raw 42.4 % Headline · max effort · Peer vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 113. Cognition official revision 1.1; Main 100 tasks, five runs/task. Claude Code for Claude; Codex for GPT. Max-effort summary, distinct from best-effort leaderboard. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11325.7% · raw 25.66 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: codex. {"tokens": 7325.0, "cost": 0.0199, "flagged_rate": 0.0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode35.5% · raw 35.53 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: codex. {"tokens": 21248.0, "cost": 0.0512, "flagged_rate": 0.0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode37.3% · raw 37.26 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: codex. {"tokens": 28999.0, "cost": 0.0646, "flagged_rate": 0.002}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode37.1% · raw 37.1 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: codex. {"tokens": 33153.0, "cost": 0.07, "flagged_rate": 0.0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode42.4% · raw 42.42 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: codex. {"tokens": 56553.0, "cost": 0.1027, "flagged_rate": 0.0}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| GPT-5.6 TerraOpenAI | OpenAI | 41.3% | 2026-08-01 | All 1 recorded result & sources41.3% · raw 41.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-01 Cognition FrontierCode 1.1 Main weighted score 41.3%; board export mirrored by evals.report (sourceUrl cognition.com) and matches llm-stats 0.413; Cognition changelog notes Terra/Luna pricing updates https://cognition.com/frontiercode |
| GLM-5.3Z.ai | Z.ai | 40.1% | 2026-09-02 | All 1 recorded result & sources40.1% · raw 40.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 Cognition FrontierCode 1.1 Main weighted score; model added Sep 2, 2026 per Cognition changelog; score 40.1% on board export mirrored by evals.report (sourceUrl cognition.com, verifiedStatus official, snapshot 2026-09-07) https://cognition.com/frontiercode |
| GLM-5.3-FlashZ.ai | Z.ai | 31.8% | 2026-09-02 | All 1 recorded result & sources31.8% · raw 31.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 Cognition FrontierCode 1.1 Main weighted score; model added Sep 2, 2026 per Cognition changelog; score 31.8% on board export mirrored by evals.report (sourceUrl cognition.com, verifiedStatus official, snapshot 2026-09-07) https://cognition.com/frontiercode |
| BeamReflection AI | Reflection AI | — | — | No recorded result |
| DeepSeek V4.1 FlashDeepSeek | DeepSeek | — | — | No recorded result |
| Gemini 4 ArgonGoogle | — | — | No recorded result | |
| GPT-6.1 SolOpenAI | OpenAI | — | — | No recorded result |
| Hy4 previewTencent | Tencent | — | — | No recorded result |
| Kimi K3Moonshot AI | Moonshot AI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | inclusionAI | — | — | No recorded result |
| Ling 3.1 FlashinclusionAI | inclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | Xiaomi | — | — | No recorded result |
| MiMo-V2.6-FlashXiaomi | Xiaomi | — | — | No recorded result |
| MiMo-V2.6-ProXiaomi | Xiaomi | — | — | No recorded result |
| Mistral Large 4Mistral AI | Mistral AI | — | — | No recorded result |
| Muse Spark 1.3Meta | Meta | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | NaiveAI | — | — | No recorded result |
| Nex-N2.5-ProNex AGI | Nex AGI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | Unbiased | — | — | No recorded result |
| Qwen3.8 27BAlibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Max (0902)Alibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | Alibaba (Qwen) | — | — | No recorded result |