Examenos

Claude Sonnet 5.5 Benchmarks, Specifications & Availability

Explore Claude Sonnet 5.5 from Anthropic: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Claude Sonnet 5.5 with other models →Explore data coverage

Published specifications

Provider
Anthropic
Access
Proprietary
License
Proprietary
Context window
1M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-28
Modalities
text, image, file
Family
Claude Sonnet 5

Model card · Announcement · Website · OpenRouter

Model notes

Second Claude 5.5 family model; fast Sonnet tier (30%+ faster outputs, up to 30% lower cost per task vs Sonnet 5). Adaptive thinking (default effort high). No official vendor-org weights on Hugging Face (checked 2026-09-29), so closed.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

OR prompt/completion/input_cache_read x1e6; 5m cache write $2.50 (1h $4.00 per Claude docs); ignore :batch twin. No -contribute sibling on OR API 2026-09-29.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Terminal-Bench 4.070.6%max effort
All 1 recorded result & sources

70.6% · raw 70.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

TB 4.0 at max effort; safeguards on (1.2% of requests fallback-served); 5 trials per task

https://www.anthropic.com/claude-sonnet-5-5
FrontierCode 1.1 Main52.1%xhigh effort
All 2 recorded results & sources

52.1% · raw 52.1 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-09-28

FrontierCode v1.1 Main best at xhigh; max scores lower (46.2): Max-effort code-review skill caused timeouts and out-of-scope penalties

https://www.anthropic.com/claude-sonnet-5-5

46.2% · raw 46.2 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

Max-effort score; below xhigh best of 52.1

https://www.anthropic.com/claude-sonnet-5-5
CursorBench 4.055.5%max effort
All 4 recorded results & sources

55.5% · raw 55.5 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

CursorBench 4.0 at max; run in Cursor production harness and independently measured by Cursor

https://www.anthropic.com/claude-sonnet-5-5

53.1% · raw 53.1 %

Alternative · xhigh effort · Own vendor

Source/record date: 2026-09-28

CursorBench effort curve: xhigh

https://www.anthropic.com/claude-sonnet-5-5-system-card

47.8% · raw 47.8 %

Alternative · high effort · Own vendor

Source/record date: 2026-09-28

CursorBench effort curve: high

https://www.anthropic.com/claude-sonnet-5-5-system-card

39.2% · raw 39.2 %

Alternative · medium effort · Own vendor

Source/record date: 2026-09-28

CursorBench effort curve: medium

https://www.anthropic.com/claude-sonnet-5-5-system-card
GDPval-AA 2.11844max effort
All 2 recorded results & sources

1844 · raw 1844 Elo

Headline · max effort · Own vendor

Source/record date: 2026-09-28

GDPval-AA v2.1 Elo at max; AA-ran on a pre-release deployment with a structured-output bug (since fixed; effect expected small, understating)

https://www.anthropic.com/claude-sonnet-5-5

1725 · raw 1725 Elo

Alternative · xhigh effort · Own vendor

Source/record date: 2026-09-28

GDPval-AA xhigh; uses about 67 percent fewer output tokens than max

https://www.anthropic.com/claude-sonnet-5-5-system-card
AA Briefcase v1.11811max effort
All 2 recorded results & sources

1811 · raw 1811 Elo

Headline · max effort · Own vendor

Source/record date: 2026-09-28

AA-Briefcase v1.1 Elo at max; same AA pre-release caveat as GDPval-AA

https://www.anthropic.com/claude-sonnet-5-5

1746 · raw 1746 Elo

Alternative · xhigh effort · Own vendor

Source/record date: 2026-09-28

AA-Briefcase xhigh; uses about 61 percent fewer output tokens than max

https://www.anthropic.com/claude-sonnet-5-5-system-card
Humanity's Last Exam (w/ tools)64.5%max effort
All 1 recorded result & sources

64.5% · raw 64.5 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

HLE with tools at max

https://www.anthropic.com/claude-sonnet-5-5
OSWorld 2.080.1%max effort
All 1 recorded result & sources

80.1% · raw 80.1 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

OSWorld 2.1 partial-credit at max (stored under osworld-2.0 per project convention); strict pass rate 43.5 percent

https://www.anthropic.com/claude-sonnet-5-5
Chartography61.6%max effort
All 1 recorded result & sources

61.6% · raw 61.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Chartography WITHOUT tools at max (announce table); with-tools score 90.2 stored separately

https://www.anthropic.com/claude-sonnet-5-5
SWE-bench Pro81.3%max effort
All 1 recorded result & sources

81.3% · raw 81.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

SWE-bench Pro avg over 5 trials; standard config (adaptive thinking max)

https://www.anthropic.com/claude-sonnet-5-5-system-card
SWE-bench Multilingual90.3%max effort
All 1 recorded result & sources

90.3% · raw 90.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

300 problems across 9 languages; avg over 5 trials

https://www.anthropic.com/claude-sonnet-5-5-system-card
SWE-bench Multimodal54.3%max effort
All 1 recorded result & sources

54.3% · raw 54.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Visual-context variant; avg over 5 trials

https://www.anthropic.com/claude-sonnet-5-5-system-card
DeepSWE v1.171%max effort
All 1 recorded result & sources

71% · raw 71 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

DeepSWE v1.1 avg over 5 trials

https://www.anthropic.com/claude-sonnet-5-5-system-card
FrontierCode 1.1 Extended64.4%xhigh effort
All 2 recorded results & sources

64.4% · raw 64.4 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-09-28

Extended split best at xhigh

https://www.anthropic.com/claude-sonnet-5-5-system-card

59.1% · raw 59.1 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

Extended split at max; below xhigh best of 64.4

https://www.anthropic.com/claude-sonnet-5-5-system-card
Terminal-Bench Science 0.159.9%max effort
All 1 recorded result & sources

59.9% · raw 59.9 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

TB-Science; safeguards on, no fallbacks fired

https://www.anthropic.com/claude-sonnet-5-5-system-card
FrontierSWE61.9%max effort
All 1 recorded result & sources

61.9% · raw 61.9 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

FrontierSWE v2; Proximal own harness, max reasoning, mean over 5 trials per task

https://www.anthropic.com/claude-sonnet-5-5-system-card
ArXivMath95.2%max effort
All 2 recorded results & sources

95.2% · raw 95.2 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

ArXivMath Aug-2026 set with tools (code sandbox); avg over 4 attempts

https://www.anthropic.com/claude-sonnet-5-5-system-card

86.8% · raw 86.8 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

ArXivMath without tools

https://www.anthropic.com/claude-sonnet-5-5-system-card
ProgramBench79.7%max effort
All 1 recorded result & sources

79.7% · raw 79.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

ProgramBench 166 golden tasks; mini-swe-agent harness, no time limit

https://www.anthropic.com/claude-sonnet-5-5-system-card
Humanity's Last Exam56.9%max effort
All 1 recorded result & sources

56.9% · raw 56.9 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

HLE no-tools at max

https://www.anthropic.com/claude-sonnet-5-5-system-card
DRACO87%unknown effort
All 1 recorded result & sources

87% · raw 87 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-28

DRACO best point on effort-cost curve; custom harness and Opus 4.6 judge, not comparable to paper headlines

https://www.anthropic.com/claude-sonnet-5-5-system-card
WANDR70%max effort
All 5 recorded results & sources

10% · raw 10 %

Alternative · low effort · Own vendor

Source/record date: 2026-09-28

WANDR soft F1; effort levels map low to max along the cost curve; modified harness (frozen index)

https://www.anthropic.com/claude-sonnet-5-5-system-card

29.9% · raw 29.9 %

Alternative · medium effort · Own vendor

Source/record date: 2026-09-28

WANDR soft F1; medium effort point

https://www.anthropic.com/claude-sonnet-5-5-system-card

56.3% · raw 56.3 %

Alternative · high effort · Own vendor

Source/record date: 2026-09-28

WANDR soft F1; high effort point

https://www.anthropic.com/claude-sonnet-5-5-system-card

66.6% · raw 66.6 %

Alternative · xhigh effort · Own vendor

Source/record date: 2026-09-28

WANDR soft F1; xhigh effort point

https://www.anthropic.com/claude-sonnet-5-5-system-card

70% · raw 70 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

WANDR soft F1 at max; best effort point

https://www.anthropic.com/claude-sonnet-5-5-system-card
Chartography (w/ tools)90.2%max effort
All 1 recorded result & sources

90.2% · raw 90.2 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Chartography with tools (container plus crop tool) at max

https://www.anthropic.com/claude-sonnet-5-5-system-card
OfficeQA76.9%max effort
All 1 recorded result & sources

76.9% · raw 76.9 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

OfficeQA full set at max

https://www.anthropic.com/claude-sonnet-5-5-system-card
OfficeQA Pro65.6%max effort
All 1 recorded result & sources

65.6% · raw 65.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

OfficeQA Pro 133-question subset at max

https://www.anthropic.com/claude-sonnet-5-5-system-card
Harvey Legal Agent11.7%high effort
All 1 recorded result & sources

11.7% · raw 11.7 %

Headline · high effort · Own vendor

Source/record date: 2026-09-28

LAB held-out all-pass rate at high (10.0 at max); mean criterion-pass 92.1 at high

https://www.anthropic.com/claude-sonnet-5-5-system-card
Toolathlon-Verified77.8%max effort
All 1 recorded result & sources

77.8% · raw 77.8 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Toolathlon-Verified Pass@1 over 3 trials (Pass@3 85.2)

https://www.anthropic.com/claude-sonnet-5-5-system-card
AutomationBench44.7%max effort
All 1 recorded result & sources

44.7% · raw 44.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

AutomationBench (Zapier private set rel 1.0.6) at max, fallbacks on

https://www.anthropic.com/claude-sonnet-5-5-system-card
AutomationBench v1.0.644.7%max effort
All 1 recorded result & sources

44.7% · raw 44.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Same Zapier 1.0.6 run, mirrored per project convention

https://www.anthropic.com/claude-sonnet-5-5-system-card
HealthBench65.4%max effort
All 2 recorded results & sources

65.4% · raw 65.4 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

HealthBench length-adjusted (OpenAI formula); raw 69.4

https://www.anthropic.com/claude-sonnet-5-5-system-card

69.4% · raw 69.4 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

HealthBench raw before length adjustment

https://www.anthropic.com/claude-sonnet-5-5-system-card
HealthBench Professional69.2%max effort
All 2 recorded results & sources

69.2% · raw 69.2 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

HealthBench Pro length-adjusted; raw 77.1 (level with Opus 5.5 raw)

https://www.anthropic.com/claude-sonnet-5-5-system-card

77.1% · raw 77.1 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

HealthBench Pro raw before length adjustment

https://www.anthropic.com/claude-sonnet-5-5-system-card
PhysicianBench63.2%max effort
All 1 recorded result & sources

63.2% · raw 63.2 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

PhysicianBench pass@1 over 100 EHR tasks

https://www.anthropic.com/claude-sonnet-5-5-system-card
Global MMLU92.1%max effort
All 1 recorded result & sources

92.1% · raw 92.1 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Global MMLU 42-language avg; single trial, no tools

https://www.anthropic.com/claude-sonnet-5-5-system-card
MILU91.6%max effort
All 1 recorded result & sources

91.6% · raw 91.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

MILU 11-language avg over 5 trials

https://www.anthropic.com/claude-sonnet-5-5-system-card
BenchCAD96.3%max effort
All 2 recorded results & sources

96.3% · raw 96.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

BenchCAD Vision2Code voxel IoU with tools; published 0.963, converted to percent per project scale rule

https://www.anthropic.com/claude-sonnet-5-5-system-card

74.7% · raw 74.7 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

BenchCAD Vision2Code voxel IoU without tools; published 0.747, converted to percent

https://www.anthropic.com/claude-sonnet-5-5-system-card

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis56max effort
All 5 recorded results & sources

56 · raw 56 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5

52 · raw 52 index

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh

47 · raw 47 index

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-high

41 · raw 41 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-medium

36 · raw 36 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-low
Cost per Intelligence Index task · Artificial Analysis$7.67max effort
All 5 recorded results & sources

$7.67 · raw 7.67 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5

$2.75 · raw 2.75 USD

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh

$1.12 · raw 1.12 USD

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-high

$0.59 · raw 0.59 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-medium

$0.42 · raw 0.42 USD

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-sonnet-5-5-low
Output speed · Artificial Analysis105 tok/sxhigh effort
All 5 recorded results & sources

105 tok/s · raw 105 tok/s

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh

139 tok/s · raw 139 tok/s

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-sonnet-5-5

102 tok/s · raw 102 tok/s

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-sonnet-5-5-high

99 tok/s · raw 99 tok/s

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-sonnet-5-5-medium

91 tok/s · raw 91 tok/s

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-sonnet-5-5-low
GDPval-AA Elo · Artificial Analysis1844max effort
All 1 recorded result & sources

1844 · raw 1844 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-28

GDPval-AA v2.1 Elo at max; run independently by AA on a pre-release deployment (structured-output bug caveat, since fixed)

https://www.anthropic.com/claude-sonnet-5-5-system-card
AA-Briefcase Elo · Artificial Analysis1811max effort
All 1 recorded result & sources

1811 · raw 1811 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-28

AA-Briefcase v1.1 Elo at max; same AA pre-release caveat

https://www.anthropic.com/claude-sonnet-5-5-system-card
Vals Index · Vals AI67%max effort
All 1 recorded result & sources

67% · raw 67.04 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $21.34.

https://www.vals.ai/benchmarks/vals_index
Vibe Code Bench v1.1 · Vals AI92.4%max effort
All 1 recorded result & sources

92.4% · raw 92.39 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $31.25.

https://www.vals.ai/benchmarks/vibe-code
Bugs fixed /105 · Bug Hunt Bench51.3 fixesmax effort
All 5 recorded results & sources

51.3 fixes · raw 51.3 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

36 fixes · raw 36 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

34.3 fixes · raw 34.3 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

23.3 fixes · raw 23.3 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

19 fixes · raw 19 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Average Score · WeirdML v320%xhigh effort
All 1 recorded result & sources

20% · raw 20 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Sonnet 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.199968; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v337%xhigh effort
All 1 recorded result & sources

37% · raw 37.02 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Sonnet 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.370190; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$9.17xhigh effort
All 1 recorded result & sources

$9.17 · raw 9.17 USD

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Sonnet 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html

Read how we select and source scores or the comparison guide.