Claude Sonnet 5.5 Benchmarks, Specifications & Availability
Explore Claude Sonnet 5.5 from Anthropic: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Claude Sonnet 5.5 with other models →Explore data coverage
Published specifications
- Provider
- Anthropic
- Access
- Proprietary
- License
- Proprietary
- Context window
- 1M
- Total parameters
- Not published
- Active parameters
- Not published
- Released
- 2026-09-28
- Modalities
- text, image, file
- Family
- Claude Sonnet 5
Model card · Announcement · Website · OpenRouter
Model notes
Second Claude 5.5 family model; fast Sonnet tier (30%+ faster outputs, up to 30% lower cost per task vs Sonnet 5). Adaptive thinking (default effort high). No official vendor-org weights on Hugging Face (checked 2026-09-29), so closed.
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | [object Object] | 2026-10-03 · OpenRouter source |
Recorded pricing notes
OR prompt/completion/input_cache_read x1e6; 5m cache write $2.50 (1h $4.00 per Claude docs); ignore :batch twin. No -contribute sibling on OR API 2026-09-29.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Terminal-Bench 4.0 | 70.6%max effort | All 1 recorded result & sources70.6% · raw 70.6 % Headline · max effort · Own vendor Source/record date: 2026-09-28 TB 4.0 at max effort; safeguards on (1.2% of requests fallback-served); 5 trials per task https://www.anthropic.com/claude-sonnet-5-5 |
| FrontierCode 1.1 Main | 52.1%xhigh effort | All 2 recorded results & sources52.1% · raw 52.1 % Headline · xhigh effort · Own vendor Source/record date: 2026-09-28 FrontierCode v1.1 Main best at xhigh; max scores lower (46.2): Max-effort code-review skill caused timeouts and out-of-scope penalties https://www.anthropic.com/claude-sonnet-5-546.2% · raw 46.2 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 Max-effort score; below xhigh best of 52.1 https://www.anthropic.com/claude-sonnet-5-5 |
| CursorBench 4.0 | 55.5%max effort | All 4 recorded results & sources55.5% · raw 55.5 % Headline · max effort · Own vendor Source/record date: 2026-09-28 CursorBench 4.0 at max; run in Cursor production harness and independently measured by Cursor https://www.anthropic.com/claude-sonnet-5-553.1% · raw 53.1 % Alternative · xhigh effort · Own vendor Source/record date: 2026-09-28 CursorBench effort curve: xhigh https://www.anthropic.com/claude-sonnet-5-5-system-card47.8% · raw 47.8 % Alternative · high effort · Own vendor Source/record date: 2026-09-28 CursorBench effort curve: high https://www.anthropic.com/claude-sonnet-5-5-system-card39.2% · raw 39.2 % Alternative · medium effort · Own vendor Source/record date: 2026-09-28 CursorBench effort curve: medium https://www.anthropic.com/claude-sonnet-5-5-system-card |
| GDPval-AA 2.1 | 1844max effort | All 2 recorded results & sources1844 · raw 1844 Elo Headline · max effort · Own vendor Source/record date: 2026-09-28 GDPval-AA v2.1 Elo at max; AA-ran on a pre-release deployment with a structured-output bug (since fixed; effect expected small, understating) https://www.anthropic.com/claude-sonnet-5-51725 · raw 1725 Elo Alternative · xhigh effort · Own vendor Source/record date: 2026-09-28 GDPval-AA xhigh; uses about 67 percent fewer output tokens than max https://www.anthropic.com/claude-sonnet-5-5-system-card |
| AA Briefcase v1.1 | 1811max effort | All 2 recorded results & sources1811 · raw 1811 Elo Headline · max effort · Own vendor Source/record date: 2026-09-28 AA-Briefcase v1.1 Elo at max; same AA pre-release caveat as GDPval-AA https://www.anthropic.com/claude-sonnet-5-51746 · raw 1746 Elo Alternative · xhigh effort · Own vendor Source/record date: 2026-09-28 AA-Briefcase xhigh; uses about 61 percent fewer output tokens than max https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Humanity's Last Exam (w/ tools) | 64.5%max effort | All 1 recorded result & sources64.5% · raw 64.5 % Headline · max effort · Own vendor Source/record date: 2026-09-28 HLE with tools at max https://www.anthropic.com/claude-sonnet-5-5 |
| OSWorld 2.0 | 80.1%max effort | All 1 recorded result & sources80.1% · raw 80.1 % Headline · max effort · Own vendor Source/record date: 2026-09-28 OSWorld 2.1 partial-credit at max (stored under osworld-2.0 per project convention); strict pass rate 43.5 percent https://www.anthropic.com/claude-sonnet-5-5 |
| Chartography | 61.6%max effort | All 1 recorded result & sources61.6% · raw 61.6 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Chartography WITHOUT tools at max (announce table); with-tools score 90.2 stored separately https://www.anthropic.com/claude-sonnet-5-5 |
| SWE-bench Pro | 81.3%max effort | All 1 recorded result & sources81.3% · raw 81.3 % Headline · max effort · Own vendor Source/record date: 2026-09-28 SWE-bench Pro avg over 5 trials; standard config (adaptive thinking max) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| SWE-bench Multilingual | 90.3%max effort | All 1 recorded result & sources90.3% · raw 90.3 % Headline · max effort · Own vendor Source/record date: 2026-09-28 300 problems across 9 languages; avg over 5 trials https://www.anthropic.com/claude-sonnet-5-5-system-card |
| SWE-bench Multimodal | 54.3%max effort | All 1 recorded result & sources54.3% · raw 54.3 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Visual-context variant; avg over 5 trials https://www.anthropic.com/claude-sonnet-5-5-system-card |
| DeepSWE v1.1 | 71%max effort | All 1 recorded result & sources71% · raw 71 % Headline · max effort · Own vendor Source/record date: 2026-09-28 DeepSWE v1.1 avg over 5 trials https://www.anthropic.com/claude-sonnet-5-5-system-card |
| FrontierCode 1.1 Extended | 64.4%xhigh effort | All 2 recorded results & sources64.4% · raw 64.4 % Headline · xhigh effort · Own vendor Source/record date: 2026-09-28 Extended split best at xhigh https://www.anthropic.com/claude-sonnet-5-5-system-card59.1% · raw 59.1 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 Extended split at max; below xhigh best of 64.4 https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Terminal-Bench Science 0.1 | 59.9%max effort | All 1 recorded result & sources59.9% · raw 59.9 % Headline · max effort · Own vendor Source/record date: 2026-09-28 TB-Science; safeguards on, no fallbacks fired https://www.anthropic.com/claude-sonnet-5-5-system-card |
| FrontierSWE | 61.9%max effort | All 1 recorded result & sources61.9% · raw 61.9 % Headline · max effort · Own vendor Source/record date: 2026-09-28 FrontierSWE v2; Proximal own harness, max reasoning, mean over 5 trials per task https://www.anthropic.com/claude-sonnet-5-5-system-card |
| ArXivMath | 95.2%max effort | All 2 recorded results & sources95.2% · raw 95.2 % Headline · max effort · Own vendor Source/record date: 2026-09-28 ArXivMath Aug-2026 set with tools (code sandbox); avg over 4 attempts https://www.anthropic.com/claude-sonnet-5-5-system-card86.8% · raw 86.8 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 ArXivMath without tools https://www.anthropic.com/claude-sonnet-5-5-system-card |
| ProgramBench | 79.7%max effort | All 1 recorded result & sources79.7% · raw 79.7 % Headline · max effort · Own vendor Source/record date: 2026-09-28 ProgramBench 166 golden tasks; mini-swe-agent harness, no time limit https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Humanity's Last Exam | 56.9%max effort | All 1 recorded result & sources56.9% · raw 56.9 % Headline · max effort · Own vendor Source/record date: 2026-09-28 HLE no-tools at max https://www.anthropic.com/claude-sonnet-5-5-system-card |
| DRACO | 87%unknown effort | All 1 recorded result & sources87% · raw 87 % Headline · unknown effort · Own vendor Source/record date: 2026-09-28 DRACO best point on effort-cost curve; custom harness and Opus 4.6 judge, not comparable to paper headlines https://www.anthropic.com/claude-sonnet-5-5-system-card |
| WANDR | 70%max effort | All 5 recorded results & sources10% · raw 10 % Alternative · low effort · Own vendor Source/record date: 2026-09-28 WANDR soft F1; effort levels map low to max along the cost curve; modified harness (frozen index) https://www.anthropic.com/claude-sonnet-5-5-system-card29.9% · raw 29.9 % Alternative · medium effort · Own vendor Source/record date: 2026-09-28 WANDR soft F1; medium effort point https://www.anthropic.com/claude-sonnet-5-5-system-card56.3% · raw 56.3 % Alternative · high effort · Own vendor Source/record date: 2026-09-28 WANDR soft F1; high effort point https://www.anthropic.com/claude-sonnet-5-5-system-card66.6% · raw 66.6 % Alternative · xhigh effort · Own vendor Source/record date: 2026-09-28 WANDR soft F1; xhigh effort point https://www.anthropic.com/claude-sonnet-5-5-system-card70% · raw 70 % Headline · max effort · Own vendor Source/record date: 2026-09-28 WANDR soft F1 at max; best effort point https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Chartography (w/ tools) | 90.2%max effort | All 1 recorded result & sources90.2% · raw 90.2 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Chartography with tools (container plus crop tool) at max https://www.anthropic.com/claude-sonnet-5-5-system-card |
| OfficeQA | 76.9%max effort | All 1 recorded result & sources76.9% · raw 76.9 % Headline · max effort · Own vendor Source/record date: 2026-09-28 OfficeQA full set at max https://www.anthropic.com/claude-sonnet-5-5-system-card |
| OfficeQA Pro | 65.6%max effort | All 1 recorded result & sources65.6% · raw 65.6 % Headline · max effort · Own vendor Source/record date: 2026-09-28 OfficeQA Pro 133-question subset at max https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Harvey Legal Agent | 11.7%high effort | All 1 recorded result & sources11.7% · raw 11.7 % Headline · high effort · Own vendor Source/record date: 2026-09-28 LAB held-out all-pass rate at high (10.0 at max); mean criterion-pass 92.1 at high https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Toolathlon-Verified | 77.8%max effort | All 1 recorded result & sources77.8% · raw 77.8 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Toolathlon-Verified Pass@1 over 3 trials (Pass@3 85.2) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| AutomationBench | 44.7%max effort | All 1 recorded result & sources44.7% · raw 44.7 % Headline · max effort · Own vendor Source/record date: 2026-09-28 AutomationBench (Zapier private set rel 1.0.6) at max, fallbacks on https://www.anthropic.com/claude-sonnet-5-5-system-card |
| AutomationBench v1.0.6 | 44.7%max effort | All 1 recorded result & sources44.7% · raw 44.7 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Same Zapier 1.0.6 run, mirrored per project convention https://www.anthropic.com/claude-sonnet-5-5-system-card |
| HealthBench | 65.4%max effort | All 2 recorded results & sources65.4% · raw 65.4 % Headline · max effort · Own vendor Source/record date: 2026-09-28 HealthBench length-adjusted (OpenAI formula); raw 69.4 https://www.anthropic.com/claude-sonnet-5-5-system-card69.4% · raw 69.4 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 HealthBench raw before length adjustment https://www.anthropic.com/claude-sonnet-5-5-system-card |
| HealthBench Professional | 69.2%max effort | All 2 recorded results & sources69.2% · raw 69.2 % Headline · max effort · Own vendor Source/record date: 2026-09-28 HealthBench Pro length-adjusted; raw 77.1 (level with Opus 5.5 raw) https://www.anthropic.com/claude-sonnet-5-5-system-card77.1% · raw 77.1 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 HealthBench Pro raw before length adjustment https://www.anthropic.com/claude-sonnet-5-5-system-card |
| PhysicianBench | 63.2%max effort | All 1 recorded result & sources63.2% · raw 63.2 % Headline · max effort · Own vendor Source/record date: 2026-09-28 PhysicianBench pass@1 over 100 EHR tasks https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Global MMLU | 92.1%max effort | All 1 recorded result & sources92.1% · raw 92.1 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Global MMLU 42-language avg; single trial, no tools https://www.anthropic.com/claude-sonnet-5-5-system-card |
| MILU | 91.6%max effort | All 1 recorded result & sources91.6% · raw 91.6 % Headline · max effort · Own vendor Source/record date: 2026-09-28 MILU 11-language avg over 5 trials https://www.anthropic.com/claude-sonnet-5-5-system-card |
| BenchCAD | 96.3%max effort | All 2 recorded results & sources96.3% · raw 96.3 % Headline · max effort · Own vendor Source/record date: 2026-09-28 BenchCAD Vision2Code voxel IoU with tools; published 0.963, converted to percent per project scale rule https://www.anthropic.com/claude-sonnet-5-5-system-card74.7% · raw 74.7 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 BenchCAD Vision2Code voxel IoU without tools; published 0.747, converted to percent https://www.anthropic.com/claude-sonnet-5-5-system-card |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 56max effort | All 5 recorded results & sources56 · raw 56 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-552 · raw 52 index Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh47 · raw 47 index Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-high41 · raw 41 index Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-medium36 · raw 36 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-low |
| Cost per Intelligence Index task · Artificial Analysis | $7.67max effort | All 5 recorded results & sources$7.67 · raw 7.67 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5$2.75 · raw 2.75 USD Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh$1.12 · raw 1.12 USD Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-high$0.59 · raw 0.59 USD Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-medium$0.42 · raw 0.42 USD Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. https://artificialanalysis.ai/models/claude-sonnet-5-5-low |
| Output speed · Artificial Analysis | 105 tok/sxhigh effort | All 5 recorded results & sources105 tok/s · raw 105 tok/s Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-sonnet-5-5-xhigh139 tok/s · raw 139 tok/s Alternative · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-sonnet-5-5102 tok/s · raw 102 tok/s Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-sonnet-5-5-high99 tok/s · raw 99 tok/s Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-sonnet-5-5-medium91 tok/s · raw 91 tok/s Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/claude-sonnet-5-5-low |
| GDPval-AA Elo · Artificial Analysis | 1844max effort | All 1 recorded result & sources1844 · raw 1844 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-28 GDPval-AA v2.1 Elo at max; run independently by AA on a pre-release deployment (structured-output bug caveat, since fixed) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| AA-Briefcase Elo · Artificial Analysis | 1811max effort | All 1 recorded result & sources1811 · raw 1811 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-28 AA-Briefcase v1.1 Elo at max; same AA pre-release caveat https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Vals Index · Vals AI | 67%max effort | All 1 recorded result & sources67% · raw 67.04 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $21.34. https://www.vals.ai/benchmarks/vals_index |
| Vibe Code Bench v1.1 · Vals AI | 92.4%max effort | All 1 recorded result & sources92.4% · raw 92.39 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $31.25. https://www.vals.ai/benchmarks/vibe-code |
| Bugs fixed /105 · Bug Hunt Bench | 51.3 fixesmax effort | All 5 recorded results & sources51.3 fixes · raw 51.3 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench36 fixes · raw 36 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench34.3 fixes · raw 34.3 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench23.3 fixes · raw 23.3 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench19 fixes · raw 19 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Average Score · WeirdML v3 | 20%xhigh effort | All 1 recorded result & sources20% · raw 20 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Claude Sonnet 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.199968; official 80/20 aggregate (area 500k-50M tokens + final best) https://htihle.github.io/weirdml.html |
| Final Best Score · WeirdML v3 | 37%xhigh effort | All 1 recorded result & sources37% · raw 37.02 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Claude Sonnet 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; raw 0.370190; mean final best effective score https://htihle.github.io/weirdml.html |
| Cost / Run · WeirdML v3 | $9.17xhigh effort | All 1 recorded result & sources$9.17 · raw 9.17 USD Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Claude Sonnet 5.5 (xhigh); harness claude_code 2.1.280; values from prepared data JSON; mean API cost per run, same task weighting as scores https://htihle.github.io/weirdml.html |
Read how we select and source scores or the comparison guide.