Claude Opus 5.5 Benchmarks, Specifications & Availability
Explore Claude Opus 5.5 from Anthropic: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Claude Opus 5.5 with other models →Explore data coverage
Published specifications
- Provider
- Anthropic
- Access
- Proprietary
- License
- Proprietary
- Context window
- 1M
- Total parameters
- Not published
- Active parameters
- Not published
- Released
- 2026-09-22
- Modalities
- text, image, file
- Family
- Claude Opus 5
Model card · Announcement · Website · OpenRouter
Model notes
First Claude 5.5 family Opus; adaptive thinking always on (default effort medium). Not :batch. System card PDF published by Anthropic.
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | [object Object] | 2026-10-03 · OpenRouter source |
Recorded pricing notes
Matches Anthropic list $4/$20; cache read $0.20; cache write $5. Not :batch.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Terminal-Bench 4.0 | 66.4%xhigh effort | All 2 recorded results & sources66.4% · raw 66.4 % Headline · xhigh effort · Own vendor Source/record date: 2026-09-22 Terminal-Bench 4.0 at xhigh (table footnote); adaptive thinking https://www.anthropic.com/claude-opus-5-566.4% · raw 66.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| FrontierCode 1.1 Main | 54.4%max effort | All 2 recorded results & sources54.4% · raw 54.4 % Headline · max effort · Own vendor Source/record date: 2026-09-22 FrontierCode v1.1 Main; adaptive max effort per table default https://www.anthropic.com/claude-opus-5-554.6% · raw 54.6 % Alternative · medium effort · Own vendor Source/record date: 2026-09-22 Default/medium effort on FrontierCode v1.1 main (cost chart text) https://www.anthropic.com/claude-opus-5-5 |
| CursorBench 4.0 | 57.8%max effort | All 2 recorded results & sources57.8% · raw 57.8 % Headline · max effort · Own vendor Source/record date: 2026-09-22 CursorBench 4.0; adaptive max effort https://www.anthropic.com/claude-opus-5-552.5% · raw 52.5 % Alternative · medium effort · Own vendor Source/record date: 2026-09-22 Default/medium effort on CursorBench 4.0 (cost chart text) https://www.anthropic.com/claude-opus-5-5 |
| GDPval-AA 2.1 | 1846max effort | All 1 recorded result & sources1846 · raw 1846 Elo Headline · max effort · Own vendor Source/record date: 2026-09-22 GDPval-AA v2.1 Elo; adaptive max effort https://www.anthropic.com/claude-opus-5-5 |
| AutomationBench | 40%unknown effort | All 1 recorded result & sources40% · raw 40 % Headline · unknown effort · Own vendor Source/record date: 2026-09-22 AutomationBench via Zapier early-access run; safeguards-as-failures https://www.anthropic.com/claude-opus-5-5 |
| Humanity's Last Exam (w/ tools) | 67.7%max effort | All 1 recorded result & sources67.7% · raw 67.7 % Headline · max effort · Own vendor Source/record date: 2026-09-22 Humanity's Last Exam with tools https://www.anthropic.com/claude-opus-5-5 |
| Terminal-Bench Science 0.1 | 58.7%max effort | All 2 recorded results & sources58.7% · raw 58.7 % Headline · max effort · Own vendor Source/record date: 2026-09-22 Terminal-Bench-Science 0.1 https://www.anthropic.com/claude-opus-5-563.3% · raw 63.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 58.7 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| OSWorld 2.0 | 81.8%max effort | All 1 recorded result & sources81.8% · raw 81.8 % Headline · max effort · Own vendor Source/record date: 2026-09-22 OSWorld 2.0 partial score as reported https://www.anthropic.com/claude-opus-5-5 |
| Chartography (w/ tools) | 89%max effort | All 1 recorded result & sources89% · raw 89 % Headline · max effort · Own vendor Source/record date: 2026-09-22 Chartography with tools https://www.anthropic.com/claude-opus-5-5 |
| AutomationBench v1.0.6 | 40%max effort | All 2 recorded results & sources40% · raw 40 % Headline · max effort · Own vendor Source/record date: 2026-09-22 Anthropic announce table; Zapier AutomationBench without fallback; adaptive thinking max effort https://www.anthropic.com/claude-opus-5-542.5% · raw 42.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 40.0 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| DeepSWE v1.1 | 74.2%unknown effort | All 2 recorded results & sources74.2% · raw 74.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/74.2% · raw 74.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Agents' Last Exam | 34.3%unknown effort | All 2 recorded results & sources34.3% · raw 34.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/38.2% · raw 38.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 34.3 (spillover); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| SWE-bench Pro | 89.9%unknown effort | All 1 recorded result & sources89.9% · raw 89.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI; card cites the Opus 5.5 system card. https://naive.ai/en/research/ |
| FrontierSWE | 62.3%unknown effort | All 1 recorded result & sources62.3% · raw 62.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); Proximal harness https://www.anthropic.com/claude-sonnet-5-5-system-card |
| SWE-bench Multilingual | 93.9%unknown effort | All 1 recorded result & sources93.9% · raw 93.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| SWE-bench Multimodal | 61.4%unknown effort | All 1 recorded result & sources61.4% · raw 61.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Chartography | 66.3%unknown effort | All 2 recorded results & sources64.4% · raw 64.4 % Alternative · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); without tools | Demoted 2026-10-01: duplicate of headline from Google Gemini 4 Argon chart (09-30, Surge leaderboard); newer peer-comparison figure wins per protocol (c). https://www.anthropic.com/claude-sonnet-5-5-system-card66.3% · raw 66.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 64.4 (spillover); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| OfficeQA | 78.9%max effort | All 1 recorded result & sources78.9% · raw 78.9 % Headline · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| OfficeQA Pro | 67.7%max effort | All 1 recorded result & sources67.7% · raw 67.7 % Headline · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Global MMLU | 94.3%max effort | All 1 recorded result & sources94.3% · raw 94.3 % Headline · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| MILU | 93.1%max effort | All 1 recorded result & sources93.1% · raw 93.1 % Headline · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Vals Index | 67%unknown effort | All 1 recorded result & sources67% · raw 67 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Vals Finance Agent v2 | 58.6%unknown effort | All 1 recorded result & sources58.6% · raw 58.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Harvey Legal Agent Benchmark | 3.8%unknown effort | All 1 recorded result & sources3.8% · raw 3.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| FrontierSWE v2 | 62.3%unknown effort | All 1 recorded result & sources62.3% · raw 62.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Vibe Code Bench | 90.3%unknown effort | All 1 recorded result & sources90.3% · raw 90.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| PostTrainBench | 49.3%unknown effort | All 1 recorded result & sources49.3% · raw 49.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| LABBench 2 | 73.1%unknown effort | All 1 recorded result & sources73.1% · raw 73.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| RiemannBench | 69.6%unknown effort | All 1 recorded result & sources69.6% · raw 69.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GraphWalks up to 128k BFS F1 | 90.6 f1unknown effort | All 1 recorded result & sources90.6 f1 · raw 90.6 f1 Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GraphWalks 256k-1M BFS F1 | 66.8 f1unknown effort | All 1 recorded result & sources66.8 f1 · raw 66.8 f1 Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| LVBench | 83.7%unknown effort | All 1 recorded result & sources83.7% · raw 83.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| CWE-bench v1 | 67%unknown effort | All 1 recorded result & sources67% · raw 67 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Gray Swan IPI | 1%unknown effort | All 1 recorded result & sources1% · raw 1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 58max effort | All 5 recorded results & sources58 · raw 58 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-556 · raw 56 index Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-xhigh54 · raw 54 index Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-high51 · raw 51 index Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-medium42 · raw 42 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-low |
| GDPval-AA Elo · Artificial Analysis | 1846max effort | All 1 recorded result & sources1846 · raw 1846 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-23 GDPval-AA v2.1 Elo at max; AA article https://artificialanalysis.ai/articles/claude-opus-5-5 |
| AA-Briefcase Elo · Artificial Analysis | 1822max effort | All 1 recorded result & sources1822 · raw 1822 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-23 AA-Briefcase v1.1 Elo at max; AA article https://artificialanalysis.ai/articles/claude-opus-5-5 |
| Bugs fixed /105 · Bug Hunt Bench | 41.7 fixesmax effort | All 5 recorded results & sources41.7 fixes · raw 41.7 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench36 fixes · raw 36 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench31.7 fixes · raw 31.7 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench30.3 fixes · raw 30.3 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench22.3 fixes · raw 22.3 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| ARC-AGI-1 · ARC Prize | 98.5%high effort | All 5 recorded results & sources97.5% · raw 97.5 % Alternative · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-597.5% · raw 97.5 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-598.5% · raw 98.5 % Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/anthropic-claude-opus-5-597.5% · raw 97.5 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-588.5% · raw 88.5 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-5 |
| ARC-AGI-2 · ARC Prize | 93.3%high effort | All 5 recorded results & sources91.7% · raw 91.7 % Alternative · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-592.5% · raw 92.5 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-593.3% · raw 93.3 % Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/anthropic-claude-opus-5-587.5% · raw 87.5 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-570.1% · raw 70.1 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/anthropic-claude-opus-5-5 |
| Vals Index · Vals AI | 67%unknown effort | All 1 recorded result & sources67% · raw 66.97 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $32.14. https://www.vals.ai/benchmarks/vals_index |
| Vibe Code Bench v1.1 · Vals AI | 90.3%unknown effort | All 1 recorded result & sources90.3% · raw 90.29 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $57.92. https://www.vals.ai/benchmarks/vibe-code |
| AA-Omniscience Index · Artificial Analysis | 46max effort | All 1 recorded result & sources46 · raw 46 index Headline · max effort · Independent evaluator Source/record date: 2026-09-23 AA-Omniscience Index; Adaptive Reasoning Max Effort Default Fallback https://artificialanalysis.ai/evaluations/omniscience |
| Cost per Intelligence Index task · Artificial Analysis | $5.98max effort | All 5 recorded results & sources$5.98 · raw 5.98 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5$3.46 · raw 3.46 USD Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-xhigh$1.82 · raw 1.82 USD Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-high$1.34 · raw 1.34 USD Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-medium$0.55 · raw 0.55 USD Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-opus-5-5-low |
| Money gain · Andon Labs | $8735.25unknown effort | All 1 recorded result & sources$8735.25 · raw 8735.25 $ Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked. https://andonlabs.com/evals/vending-bench-2 |
| Blueprint Bench · Andon Labs | 51.2%unknown effort | All 1 recorded result & sources51.2% · raw 51.2 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100. https://andonlabs.com/evals/blueprint-bench-2 |
| Average Score · WeirdML v3 | 31.2%xhigh effort | All 1 recorded result & sources31.2% · raw 31.2 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Claude Opus 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.312003; official 80/20 aggregate (area 500k-50M tokens + final best) https://htihle.github.io/weirdml.html |
| Final Best Score · WeirdML v3 | 49.8%xhigh effort | All 1 recorded result & sources49.8% · raw 49.83 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Claude Opus 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.498310; mean final best effective score https://htihle.github.io/weirdml.html |
| Cost / Run · WeirdML v3 | $10.25xhigh effort | All 1 recorded result & sources$10.25 · raw 10.25 USD Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Claude Opus 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; mean API cost per run, same task weighting as scores https://htihle.github.io/weirdml.html |
| Output speed · Artificial Analysis | 93 tok/smax effort | All 5 recorded results & sources93 tok/s · raw 93 tok/s Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-opus-5-579 tok/s · raw 79 tok/s Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-opus-5-5-xhigh74 tok/s · raw 74 tok/s Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-opus-5-5-high72 tok/s · raw 72 tok/s Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-opus-5-5-medium76 tok/s · raw 76 tok/s Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-opus-5-5-low |
Read how we select and source scores or the comparison guide.