Examenos

Claude Opus 5.5 Benchmarks, Specifications & Availability

Explore Claude Opus 5.5 from Anthropic: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Claude Opus 5.5 with other models →Explore data coverage

Published specifications

Provider
Anthropic
Access
Proprietary
License
Proprietary
Context window
1M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-22
Modalities
text, image, file
Family
Claude Opus 5

Model card · Announcement · Website · OpenRouter

Model notes

First Claude 5.5 family Opus; adaptive thinking always on (default effort medium). Not :batch. System card PDF published by Anthropic.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

Matches Anthropic list $4/$20; cache read $0.20; cache write $5. Not :batch.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Terminal-Bench 4.066.4%xhigh effort
All 2 recorded results & sources

66.4% · raw 66.4 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-09-22

Terminal-Bench 4.0 at xhigh (table footnote); adaptive thinking

https://www.anthropic.com/claude-opus-5-5

66.4% · raw 66.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
FrontierCode 1.1 Main54.4%max effort
All 2 recorded results & sources

54.4% · raw 54.4 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

FrontierCode v1.1 Main; adaptive max effort per table default

https://www.anthropic.com/claude-opus-5-5

54.6% · raw 54.6 %

Alternative · medium effort · Own vendor

Source/record date: 2026-09-22

Default/medium effort on FrontierCode v1.1 main (cost chart text)

https://www.anthropic.com/claude-opus-5-5
CursorBench 4.057.8%max effort
All 2 recorded results & sources

57.8% · raw 57.8 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

CursorBench 4.0; adaptive max effort

https://www.anthropic.com/claude-opus-5-5

52.5% · raw 52.5 %

Alternative · medium effort · Own vendor

Source/record date: 2026-09-22

Default/medium effort on CursorBench 4.0 (cost chart text)

https://www.anthropic.com/claude-opus-5-5
GDPval-AA 2.11846max effort
All 1 recorded result & sources

1846 · raw 1846 Elo

Headline · max effort · Own vendor

Source/record date: 2026-09-22

GDPval-AA v2.1 Elo; adaptive max effort

https://www.anthropic.com/claude-opus-5-5
AutomationBench40%unknown effort
All 1 recorded result & sources

40% · raw 40 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-22

AutomationBench via Zapier early-access run; safeguards-as-failures

https://www.anthropic.com/claude-opus-5-5
Humanity's Last Exam (w/ tools)67.7%max effort
All 1 recorded result & sources

67.7% · raw 67.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

Humanity's Last Exam with tools

https://www.anthropic.com/claude-opus-5-5
Terminal-Bench Science 0.158.7%max effort
All 2 recorded results & sources

58.7% · raw 58.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

Terminal-Bench-Science 0.1

https://www.anthropic.com/claude-opus-5-5

63.3% · raw 63.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 58.7 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
OSWorld 2.081.8%max effort
All 1 recorded result & sources

81.8% · raw 81.8 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

OSWorld 2.0 partial score as reported

https://www.anthropic.com/claude-opus-5-5
Chartography (w/ tools)89%max effort
All 1 recorded result & sources

89% · raw 89 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

Chartography with tools

https://www.anthropic.com/claude-opus-5-5
AutomationBench v1.0.640%max effort
All 2 recorded results & sources

40% · raw 40 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

Anthropic announce table; Zapier AutomationBench without fallback; adaptive thinking max effort

https://www.anthropic.com/claude-opus-5-5

42.5% · raw 42.5 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 40.0 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
DeepSWE v1.174.2%unknown effort
All 2 recorded results & sources

74.2% · raw 74.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

74.2% · raw 74.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Agents' Last Exam34.3%unknown effort
All 2 recorded results & sources

34.3% · raw 34.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

38.2% · raw 38.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 34.3 (spillover); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
SWE-bench Pro89.9%unknown effort
All 1 recorded result & sources

89.9% · raw 89.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI; card cites the Opus 5.5 system card.

https://naive.ai/en/research/
FrontierSWE62.3%unknown effort
All 1 recorded result & sources

62.3% · raw 62.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); Proximal harness

https://www.anthropic.com/claude-sonnet-5-5-system-card
SWE-bench Multilingual93.9%unknown effort
All 1 recorded result & sources

93.9% · raw 93.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
SWE-bench Multimodal61.4%unknown effort
All 1 recorded result & sources

61.4% · raw 61.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
Chartography66.3%unknown effort
All 2 recorded results & sources

64.4% · raw 64.4 %

Alternative · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); without tools | Demoted 2026-10-01: duplicate of headline from Google Gemini 4 Argon chart (09-30, Surge leaderboard); newer peer-comparison figure wins per protocol (c).

https://www.anthropic.com/claude-sonnet-5-5-system-card

66.3% · raw 66.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 64.4 (spillover); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
OfficeQA78.9%max effort
All 1 recorded result & sources

78.9% · raw 78.9 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
OfficeQA Pro67.7%max effort
All 1 recorded result & sources

67.7% · raw 67.7 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
Global MMLU94.3%max effort
All 1 recorded result & sources

94.3% · raw 94.3 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
MILU93.1%max effort
All 1 recorded result & sources

93.1% · raw 93.1 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
Vals Index67%unknown effort
All 1 recorded result & sources

67% · raw 67 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Vals Finance Agent v258.6%unknown effort
All 1 recorded result & sources

58.6% · raw 58.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Harvey Legal Agent Benchmark3.8%unknown effort
All 1 recorded result & sources

3.8% · raw 3.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
FrontierSWE v262.3%unknown effort
All 1 recorded result & sources

62.3% · raw 62.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Vibe Code Bench90.3%unknown effort
All 1 recorded result & sources

90.3% · raw 90.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
PostTrainBench49.3%unknown effort
All 1 recorded result & sources

49.3% · raw 49.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
LABBench 273.1%unknown effort
All 1 recorded result & sources

73.1% · raw 73.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
RiemannBench69.6%unknown effort
All 1 recorded result & sources

69.6% · raw 69.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GraphWalks up to 128k BFS F190.6 f1unknown effort
All 1 recorded result & sources

90.6 f1 · raw 90.6 f1

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GraphWalks 256k-1M BFS F166.8 f1unknown effort
All 1 recorded result & sources

66.8 f1 · raw 66.8 f1

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
LVBench83.7%unknown effort
All 1 recorded result & sources

83.7% · raw 83.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
CWE-bench v167%unknown effort
All 1 recorded result & sources

67% · raw 67 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Gray Swan IPI1%unknown effort
All 1 recorded result & sources

1% · raw 1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis58max effort
All 5 recorded results & sources

58 · raw 58 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5

56 · raw 56 index

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-xhigh

54 · raw 54 index

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-high

51 · raw 51 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-medium

42 · raw 42 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-low
GDPval-AA Elo · Artificial Analysis1846max effort
All 1 recorded result & sources

1846 · raw 1846 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo at max; AA article

https://artificialanalysis.ai/articles/claude-opus-5-5
AA-Briefcase Elo · Artificial Analysis1822max effort
All 1 recorded result & sources

1822 · raw 1822 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

AA-Briefcase v1.1 Elo at max; AA article

https://artificialanalysis.ai/articles/claude-opus-5-5
Bugs fixed /105 · Bug Hunt Bench41.7 fixesmax effort
All 5 recorded results & sources

41.7 fixes · raw 41.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

36 fixes · raw 36 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

31.7 fixes · raw 31.7 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

30.3 fixes · raw 30.3 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

22.3 fixes · raw 22.3 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
ARC-AGI-1 · ARC Prize98.5%high effort
All 5 recorded results & sources

97.5% · raw 97.5 %

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5

97.5% · raw 97.5 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5

98.5% · raw 98.5 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/anthropic-claude-opus-5-5

97.5% · raw 97.5 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5

88.5% · raw 88.5 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5
ARC-AGI-2 · ARC Prize93.3%high effort
All 5 recorded results & sources

91.7% · raw 91.7 %

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5

92.5% · raw 92.5 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5

93.3% · raw 93.3 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/anthropic-claude-opus-5-5

87.5% · raw 87.5 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5

70.1% · raw 70.1 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-opus-5-5
Vals Index · Vals AI67%unknown effort
All 1 recorded result & sources

67% · raw 66.97 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $32.14.

https://www.vals.ai/benchmarks/vals_index
Vibe Code Bench v1.1 · Vals AI90.3%unknown effort
All 1 recorded result & sources

90.3% · raw 90.29 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $57.92.

https://www.vals.ai/benchmarks/vibe-code
AA-Omniscience Index · Artificial Analysis46max effort
All 1 recorded result & sources

46 · raw 46 index

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

AA-Omniscience Index; Adaptive Reasoning Max Effort Default Fallback

https://artificialanalysis.ai/evaluations/omniscience
Cost per Intelligence Index task · Artificial Analysis$5.98max effort
All 5 recorded results & sources

$5.98 · raw 5.98 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5

$3.46 · raw 3.46 USD

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-xhigh

$1.82 · raw 1.82 USD

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-high

$1.34 · raw 1.34 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-medium

$0.55 · raw 0.55 USD

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-opus-5-5-low
Money gain · Andon Labs$8735.25unknown effort
All 1 recorded result & sources

$8735.25 · raw 8735.25 $

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked.

https://andonlabs.com/evals/vending-bench-2
Blueprint Bench · Andon Labs51.2%unknown effort
All 1 recorded result & sources

51.2% · raw 51.2 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100.

https://andonlabs.com/evals/blueprint-bench-2
Average Score · WeirdML v331.2%xhigh effort
All 1 recorded result & sources

31.2% · raw 31.2 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Opus 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.312003; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v349.8%xhigh effort
All 1 recorded result & sources

49.8% · raw 49.83 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Opus 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.498310; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$10.25xhigh effort
All 1 recorded result & sources

$10.25 · raw 10.25 USD

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Opus 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html
Output speed · Artificial Analysis93 tok/smax effort
All 5 recorded results & sources

93 tok/s · raw 93 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-opus-5-5

79 tok/s · raw 79 tok/s

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-opus-5-5-xhigh

74 tok/s · raw 74 tok/s

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-opus-5-5-high

72 tok/s · raw 72 tok/s

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-opus-5-5-medium

76 tok/s · raw 76 tok/s

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-opus-5-5-low

Read how we select and source scores or the comparison guide.