Claude Haiku 5.5
Explore Claude Haiku 5.5 from Anthropic: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Claude Haiku 5.5 with other models →Calculate costExplore data coverage
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | $0.1 | $0.5 | $0.01 | $0.125 | 2026-10-08 · OpenRouter source |
Recorded pricing notes
Base rates through 100,000 prompt tokens; long-context pricing changes above 100,000: $0.50 input / $2.50 output / $0.05 cache read per million. Cache write 5m $0.125/$0.625, 1h $0.20/$1.00 (short/long). Dashboard endpoint minima are not an exact request quote. No contributor sibling in current catalog.; min-healthy endpoint minima 2026-10-08. Long-context override applies to request pricing, not the displayed base minimum.
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| SWE-bench Pro | 64.8%max effort | All 1 recorded result & sources64.8% · raw 64.8 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 111. Adaptive max, mean of five trials; Anthropic coding harness. See pp. 111–112 for evaluation-specific configuration. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=111 |
| SWE-bench Multilingual | 83.7%max effort | All 1 recorded result & sources83.7% · raw 83.7 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 111. Adaptive max, mean of five trials; Anthropic coding harness. See pp. 111–112 for evaluation-specific configuration. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=111 |
| SWE-bench Multimodal | 30.7%max effort | All 1 recorded result & sources30.7% · raw 30.7 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 111. Adaptive max, mean of five trials; Anthropic coding harness. See pp. 111–112 for evaluation-specific configuration. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=111 |
| FrontierCode 1.1 Main | 46.4%max effort | All 7 recorded results & sources46.4% · raw 46.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 113. Cognition official revision 1.1; Main 100 tasks, five runs/task. Claude Code for Claude; Codex for GPT. Max-effort summary, distinct from best-effort leaderboard. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11345.8% · raw 45.8 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 111. Explicit xhigh score in summary table. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11134.8% · raw 34.76 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 23868, "cost": 0.0596, "duration_min": 6.71, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode41.6% · raw 41.63 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 36172, "cost": 0.13, "duration_min": 8.67, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode41.9% · raw 41.89 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 55149, "cost": 0.256, "duration_min": 11.06, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode45.8% · raw 45.83 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 100847, "cost": 0.6297, "duration_min": 16.0, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode46.4% · raw 46.36 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 181387, "cost": 1.3284, "duration_min": 23.55, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| FrontierCode 1.1 Extended | 58.4%max effort | All 6 recorded results & sources58.4% · raw 58.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 114. Cognition revision 1.1; Extended 150 tasks, five runs/task; Claude Code. Max-effort result. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11448.2% · raw 48.19 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 20240, "cost": 0.0486, "duration_min": 6.08, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode55% · raw 55.04 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 30398, "cost": 0.1021, "duration_min": 7.68, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode55.9% · raw 55.88 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 46347, "cost": 0.1998, "duration_min": 9.89, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode58.1% · raw 58.12 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 84943, "cost": 0.4876, "duration_min": 13.92, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode58.4% · raw 58.43 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Cognition primary public data.json, v1_1 weighted new_score ×100, not unweighted correct. Harness: claude-code. {"tokens": 157710, "cost": 1.0803, "duration_min": 20.95, "flagged_rate": null}. Cost is source-observed USD/task; not an API rate. Data checked 8 October; evaluation date unpublished. Anthropic launch headline retained; exact board precision/effort alternative, no averaging. https://cognition.com/frontiercode |
| Humanity's Last Exam | 45.9%max effort | All 5 recorded results & sources30.9% · raw 30.9 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12035.5% · raw 35.5 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12039.7% · raw 39.7 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12044.1% · raw 44.1 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12045.9% · raw 45.9 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 120. No tools; adaptive effort matrix, five trials. Published labeled chart values. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=120 |
| Humanity's Last Exam (w/ tools) | 57.4%max effort | All 5 recorded results & sources38.6% · raw 38.6 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 119. Web search/fetch and code execution; contamination-source restrictions; adaptive effort matrix, five trials. Web-search fees excluded from chart costs. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11945% · raw 45 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 119. Web search/fetch and code execution; contamination-source restrictions; adaptive effort matrix, five trials. Web-search fees excluded from chart costs. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11950.1% · raw 50.1 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 119. Web search/fetch and code execution; contamination-source restrictions; adaptive effort matrix, five trials. Web-search fees excluded from chart costs. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11954.8% · raw 54.8 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 119. Web search/fetch and code execution; contamination-source restrictions; adaptive effort matrix, five trials. Web-search fees excluded from chart costs. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=11957.4% · raw 57.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 119. Web search/fetch and code execution; contamination-source restrictions; adaptive effort matrix, five trials. Web-search fees excluded from chart costs. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=119 |
| Terminal-Bench 4.0 | 39.2%max effort | All 5 recorded results & sources12.7% · raw 12.7 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch interactive chart labels; Haiku ten runs/task, no fallback; offline-egress Anthropic harness. Sonnet five runs/task with safeguard fallback. Do not merge independent AA fallback runs. https://www.anthropic.com/claude-haiku-5-520.3% · raw 20.3 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch interactive chart labels; Haiku ten runs/task, no fallback; offline-egress Anthropic harness. Sonnet five runs/task with safeguard fallback. Do not merge independent AA fallback runs. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-524.8% · raw 24.8 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch interactive chart labels; Haiku ten runs/task, no fallback; offline-egress Anthropic harness. Sonnet five runs/task with safeguard fallback. Do not merge independent AA fallback runs. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-531.5% · raw 31.5 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch interactive chart labels; Haiku ten runs/task, no fallback; offline-egress Anthropic harness. Sonnet five runs/task with safeguard fallback. Do not merge independent AA fallback runs. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-539.2% · raw 39.2 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch interactive chart labels; Haiku ten runs/task, no fallback; offline-egress Anthropic harness. Sonnet five runs/task with safeguard fallback. Do not merge independent AA fallback runs. https://www.anthropic.com/claude-haiku-5-5 |
| Terminal-Bench Science 0.1 | 20.6%max effort | All 1 recorded result & sources20.6% · raw 20.6 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 116. Haiku ten runs/task (700 trials), no fallback; 2.3% ended by safeguards. Offline-egress Anthropic harness; peers five runs/task. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=116 |
| FrontierSWE v2 | 43.8%max effort | All 1 recorded result & sources43.8% · raw 43.8 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 117. 34 tasks, five runs/task, 20-hour wall-clock budget, Proximal agent harness. No unapproved Fable 5 peer attached to Fable 5.1. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=117 |
| ProgramBench (166-task reference-filtered) | 82%max effort | All 1 recorded result & sources82% · raw 82 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 117. Hidden-test pass rate, mini-swe-agent, no six-hour timeout. 166 of 200 tasks; remove 34 with reference pass rate below 0.9; only reference-passing tests. Distinct from full-set Almost@1. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=117 |
| DRACO (Anthropic, Opus 4.6 grader) | 81.5%max effort | All 5 recorded results & sources64.3% · raw 64.3 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 121. 100 tasks; 980k-token budget, Anthropic web/code agent; final report file only. Opus 4.6 grader, five grading runs/response; not comparable to original Gemini 3 Pro grading. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12172.4% · raw 72.4 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 121. 100 tasks; 980k-token budget, Anthropic web/code agent; final report file only. Opus 4.6 grader, five grading runs/response; not comparable to original Gemini 3 Pro grading. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12177.8% · raw 77.8 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 121. 100 tasks; 980k-token budget, Anthropic web/code agent; final report file only. Opus 4.6 grader, five grading runs/response; not comparable to original Gemini 3 Pro grading. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12180.5% · raw 80.5 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 121. 100 tasks; 980k-token budget, Anthropic web/code agent; final report file only. Opus 4.6 grader, five grading runs/response; not comparable to original Gemini 3 Pro grading. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12181.5% · raw 81.5 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 121. 100 tasks; 980k-token budget, Anthropic web/code agent; final report file only. Opus 4.6 grader, five grading runs/response; not comparable to original Gemini 3 Pro grading. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=121 |
| WANDR (Anthropic, Opus 4.8 grader) | 49.9%max effort | All 5 recorded results & sources3.5% · raw 3.5 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 123. Soft F1, 980k-token task budget; Anthropic harness and Opus 4.8 grader. Distinct from prior Opus 4.6/frozen-index or Perplexity GPT-5.4 grading. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12312.9% · raw 12.9 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 123. Soft F1, 980k-token task budget; Anthropic harness and Opus 4.8 grader. Distinct from prior Opus 4.6/frozen-index or Perplexity GPT-5.4 grading. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12337.3% · raw 37.3 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 123. Soft F1, 980k-token task budget; Anthropic harness and Opus 4.8 grader. Distinct from prior Opus 4.6/frozen-index or Perplexity GPT-5.4 grading. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12345.9% · raw 45.9 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 123. Soft F1, 980k-token task budget; Anthropic harness and Opus 4.8 grader. Distinct from prior Opus 4.6/frozen-index or Perplexity GPT-5.4 grading. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=12349.9% · raw 49.9 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 123. Soft F1, 980k-token task budget; Anthropic harness and Opus 4.8 grader. Distinct from prior Opus 4.6/frozen-index or Perplexity GPT-5.4 grading. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=123 |
| Chartography | 46.4%max effort | All 1 recorded result & sources46.4% · raw 46.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 124. 100 tasks, no tools. Claude adaptive max, mean five runs; Gemini 3.5 Flash grader. Non-Claude values reproduced from Surge board; their effort is not specified in this figure. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=124 |
| Chartography (w/ tools) | 86.2%max effort | All 1 recorded result & sources86.2% · raw 86.2 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 124. 100 tasks, five runs at max; image files, Python libraries and image cropping tool; Gemini 3.5 Flash grader. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=124 |
| BenchCAD Vision2Code (1,000-file subset) | 67%max effort | All 1 recorded result & sources67% · raw 67 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 126. Random 1,000-file subset of 17,900 Vision2Code files, five runs at max. Raw voxel IoU 0–1 multiplied by 100; tools variant has image files/Python/cropping. Distinct from full-set BenchCAD. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=126 |
| BenchCAD Vision2Code (1,000-file subset, tools) | 87%max effort | All 1 recorded result & sources87% · raw 87 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 126. Random 1,000-file subset of 17,900 Vision2Code files, five runs at max. Raw voxel IoU 0–1 multiplied by 100; tools variant has image files/Python/cropping. Distinct from full-set BenchCAD. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=126 |
| OSWorld 2.1 (offline subset, partial credit) | 72.4%max effort | All 5 recorded results & sources42% · raw 42 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch chart exact labels; 82/108 offline tasks, no VM internet, 1080p, 500 actions, five attempts/task; Opus 4.8 grader. Partial checkpoint credit; distinct from strict completion and full OSWorld. https://www.anthropic.com/claude-haiku-5-553.3% · raw 53.3 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch chart exact labels; 82/108 offline tasks, no VM internet, 1080p, 500 actions, five attempts/task; Opus 4.8 grader. Partial checkpoint credit; distinct from strict completion and full OSWorld. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-561.3% · raw 61.3 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch chart exact labels; 82/108 offline tasks, no VM internet, 1080p, 500 actions, five attempts/task; Opus 4.8 grader. Partial checkpoint credit; distinct from strict completion and full OSWorld. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-567.6% · raw 67.6 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch chart exact labels; 82/108 offline tasks, no VM internet, 1080p, 500 actions, five attempts/task; Opus 4.8 grader. Partial checkpoint credit; distinct from strict completion and full OSWorld. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-572.4% · raw 72.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Official launch chart exact labels; 82/108 offline tasks, no VM internet, 1080p, 500 actions, five attempts/task; Opus 4.8 grader. Partial checkpoint credit; distinct from strict completion and full OSWorld. https://www.anthropic.com/claude-haiku-5-5 |
| OSWorld 2.1 (offline subset, strict completion) | 37.1%max effort | All 1 recorded result & sources37.1% · raw 37.1 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 128. 82/108 offline tasks, five attempts/task, max; all checkpoints must pass. No VM internet, 1080p, 500 actions, Opus 4.8 grader. Not partial-credit score. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=128 |
| OfficeQA | 73.5%max effort | All 1 recorded result & sources73.5% · raw 73.5 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 130. Relevant documents preselected and supplied as extracted text; max, mean five runs. Pro is the 133-question subset; not end-to-end document retrieval. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=130 |
| OfficeQA Pro | 60.3%max effort | All 1 recorded result & sources60.3% · raw 60.3 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 130. Relevant documents preselected and supplied as extracted text; max, mean five runs. Pro is the 133-question subset; not end-to-end document retrieval. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=130 |
| GDPval-AA 2.1 | 1620max effort | All 5 recorded results & sources1125 · raw 1125 Elo Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Vendor-reported GDPval-AA v2.1 native Elo; exact launch chart labels, rounded vendor evidence kept apart from independently sourced AA decimal values. https://www.anthropic.com/claude-haiku-5-51277 · raw 1277 Elo Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Vendor-reported GDPval-AA v2.1 native Elo; exact launch chart labels, rounded vendor evidence kept apart from independently sourced AA decimal values. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-51420 · raw 1420 Elo Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Vendor-reported GDPval-AA v2.1 native Elo; exact launch chart labels, rounded vendor evidence kept apart from independently sourced AA decimal values. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-51513 · raw 1513 Elo Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Vendor-reported GDPval-AA v2.1 native Elo; exact launch chart labels, rounded vendor evidence kept apart from independently sourced AA decimal values. Alternative evidence; existing headline/configuration retained. No averaging. https://www.anthropic.com/claude-haiku-5-51620 · raw 1620 Elo Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card / launch. Vendor-reported GDPval-AA v2.1 native Elo; exact launch chart labels, rounded vendor evidence kept apart from independently sourced AA decimal values. https://www.anthropic.com/claude-haiku-5-5 |
| AA Briefcase v1.1 | 1578max effort | All 2 recorded results & sources1578 · raw 1578 Elo Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 131. Vendor-reported AA-Briefcase v1.1 native Elo, max; independently sourced AA decimal values remain separate. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=1311372 · raw 1372 Elo Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 131. Medium-effort result, native Elo. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=131 |
| HealthBench | 61.6%max effort | All 5 recorded results & sources59.9% · raw 59.9 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled. Low–xhigh one trial/level (5 Oct); max mean five trials (2 Oct). Raw score is separate. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13460.2% · raw 60.2 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled. Low–xhigh one trial/level (5 Oct); max mean five trials (2 Oct). Raw score is separate. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13460.3% · raw 60.3 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled. Low–xhigh one trial/level (5 Oct); max mean five trials (2 Oct). Raw score is separate. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13461.1% · raw 61.1 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled. Low–xhigh one trial/level (5 Oct); max mean five trials (2 Oct). Raw score is separate. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13461.6% · raw 61.6 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled. Low–xhigh one trial/level (5 Oct); max mean five trials (2 Oct). Raw score is separate. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=134 |
| HealthBench Professional | 64.8%max effort | All 6 recorded results & sources57.9% · raw 57.9 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled; five runs/level. Low–xhigh 4 Oct; max 1–2 Oct. Raw score is separate. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13459.9% · raw 59.9 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled; five runs/level. Low–xhigh 4 Oct; max 1–2 Oct. Raw score is separate. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13461.3% · raw 61.3 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled; five runs/level. Low–xhigh 4 Oct; max 1–2 Oct. Raw score is separate. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13461% · raw 61 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled; five runs/level. Low–xhigh 4 Oct; max 1–2 Oct. Raw score is separate. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13464.8% · raw 64.8 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. Length-adjusted; Opus 4.8 grader, safeguards enabled; five runs/level. Low–xhigh 4 Oct; max 1–2 Oct. Raw score is separate. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13464.2% · raw 64.2 % Alternative · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 133. Length-adjusted repeat on 4 October, mean thirteen runs; alternative to headline five-run 1–2 October result. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=133 |
| HealthBench (raw, Anthropic grader) | 64.2%max effort | All 1 recorded result & sources64.2% · raw 64.2 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 132. Raw score before response-length adjustment; max, Opus 4.8 grader, safeguards enabled. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=132 |
| HealthBench Professional (raw, Anthropic grader) | 71%max effort | All 1 recorded result & sources71% · raw 71 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 132. Raw score before response-length adjustment; max, Opus 4.8 grader, safeguards enabled. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=132 |
| PhysicianBench | 43%max effort | All 5 recorded results & sources17.8% · raw 17.8 % Alternative · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. 100 EHR tasks; pass@1 mean five attempts/task, Opus 5 grader; safeguards enabled. Low–xhigh 5 October, max 1–2 October. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13425.2% · raw 25.2 % Alternative · medium effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. 100 EHR tasks; pass@1 mean five attempts/task, Opus 5 grader; safeguards enabled. Low–xhigh 5 October, max 1–2 October. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13431.6% · raw 31.6 % Alternative · high effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. 100 EHR tasks; pass@1 mean five attempts/task, Opus 5 grader; safeguards enabled. Low–xhigh 5 October, max 1–2 October. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13435.8% · raw 35.8 % Alternative · xhigh effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. 100 EHR tasks; pass@1 mean five attempts/task, Opus 5 grader; safeguards enabled. Low–xhigh 5 October, max 1–2 October. Alternative evidence; existing headline/configuration retained. No averaging. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13443% · raw 43 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 134. 100 EHR tasks; pass@1 mean five attempts/task, Opus 5 grader; safeguards enabled. Low–xhigh 5 October, max 1–2 October. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=134 |
| Global MMLU | 87.8%max effort | All 1 recorded result & sources87.8% · raw 87.8 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 135. GMMLU: 42 languages, one trial; MILU: 11 languages, five trials. Max; rare blocked examples excluded (<0.1% GMMLU / <0.2% MILU). https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=135 |
| MILU | 87.6%max effort | All 1 recorded result & sources87.6% · raw 87.6 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 136. GMMLU: 42 languages, one trial; MILU: 11 languages, five trials. Max; rare blocked examples excluded (<0.1% GMMLU / <0.2% MILU). https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=136 |
| SpatialBench Verified | 67.7%max effort | All 1 recorded result & sources67.7% · raw 67.7 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 137. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=137 |
| SingleCellBench | 56.2%max effort | All 1 recorded result & sources56.2% · raw 56.2 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 137. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=137 |
| Morphology-to-molecule matching | 25.4%max effort | All 1 recorded result & sources25.4% · raw 25.4 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 138. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=138 |
| Medicinal chemistry (ADME) | 58%max effort | All 1 recorded result & sources58% · raw 58 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 138. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=138 |
| Protein design: sequence generation | 33.1%max effort | All 1 recorded result & sources33.1% · raw 33.1 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 138. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=138 |
| Protein design: library ranking | 45.3%max effort | All 1 recorded result & sources45.3% · raw 45.3 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 138. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=138 |
| Biomedical image analysis | 43.5%max effort | All 1 recorded result & sources43.5% · raw 43.5 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 138. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=138 |
| Protocol troubleshooting | 58%max effort | All 1 recorded result & sources58% · raw 58 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 139. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=139 |
| Protocol understanding v2 | 64.1%max effort | All 1 recorded result & sources64.1% · raw 64.1 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 139. Life sciences, adaptive max; Haiku biology safeguards disabled to measure underlying capability, not deployed behavior. Five runs unless specified: sequence generation and biomedical image analysis one; library ranking three. Figure 8.13.A peer values are rounded raw fractions ×100; text values take precedence where supplied. ADME 504 questions; biomedical scores normalized to competition-best results. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=139 |
| VCT multimodal virology (350 questions) | 48%unknown effort | All 1 recorded result & sources48% · raw 48 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 13. 350-question multimodal virology troubleshooting; raw accuracy 0–1 multiplied by 100. Biology safeguards off; capability assessment. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=13 |
| Anthropic ECI (internal benchmark fit) | 167.1 scoreunknown effort | All 1 recorded result & sources167.1 score · raw 167.11 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 20. Anthropic internal benchmark IRT fit; native index. Not Epoch’s public ECI. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=20 |
| ExploitBench V8 (AutoNudge mean flags) | 6.6 flagsunknown effort | All 1 recorded result & sources6.6 flags · raw 6.56 flags Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 25. 41 V8 vulnerabilities from 2024 onward; 300-turn uniform author harness, five trials/arm. Cyber safeguards off, escape classifiers on. Mean flags native 0–16; Cap rate randomly sampled trials; ACE counts both plain and AutoNudge, 410 trials total. Distinct from Jun–Aug 2026 challenge set. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=25 |
| ExploitBench V8 (AutoNudge cap rate) | 49%unknown effort | All 1 recorded result & sources49% · raw 49 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 25. 41 V8 vulnerabilities from 2024 onward; 300-turn uniform author harness, five trials/arm. Cyber safeguards off, escape classifiers on. Mean flags native 0–16; Cap rate randomly sampled trials; ACE counts both plain and AutoNudge, 410 trials total. Distinct from Jun–Aug 2026 challenge set. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=25 |
| ExploitBench V8 (full ACE count) | 4 exploitsunknown effort | All 1 recorded result & sources4 exploits · raw 4 exploits Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 25. 41 V8 vulnerabilities from 2024 onward; 300-turn uniform author harness, five trials/arm. Cyber safeguards off, escape classifiers on. Mean flags native 0–16; Cap rate randomly sampled trials; ACE counts both plain and AutoNudge, 410 trials total. Distinct from Jun–Aug 2026 challenge set. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=25 |
| ExploitBench V8 (plain mean flags) | 5.2 flagsunknown effort | All 1 recorded result & sources5.2 flags · raw 5.15 flags Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 25. 41 V8 vulnerabilities from 2024 onward; 300-turn uniform author harness, five trials/arm. Cyber safeguards off, escape classifiers on. Mean flags native 0–16; Cap rate randomly sampled trials; ACE counts both plain and AutoNudge, 410 trials total. Distinct from Jun–Aug 2026 challenge set. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=25 |
| ExploitBench V8 (plain cap rate) | 38%unknown effort | All 1 recorded result & sources38% · raw 38 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 25. 41 V8 vulnerabilities from 2024 onward; 300-turn uniform author harness, five trials/arm. Cyber safeguards off, escape classifiers on. Mean flags native 0–16; Cap rate randomly sampled trials; ACE counts both plain and AutoNudge, 410 trials total. Distinct from Jun–Aug 2026 challenge set. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=25 |
| ExploitBench V8 (full ACE rate) | 1%unknown effort | All 1 recorded result & sources1% · raw 1 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 24. Published 1.0% (4/410), both plain and AutoNudge arms; 41 vulnerabilities, cyber safeguards disabled. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=24 |
| Binary Exploitation (Aug 2026 harness, identification) | 58.2%unknown effort | All 1 recorded result & sources58.2% · raw 58.2 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 28. 831 entrypoints, 228 OSS-Fuzz projects; August 2026 revised harness, not older OSS-Fuzz harness. Pass@1 vulnerability identification. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=28 |
| Binary Exploitation (Aug 2026 harness, control-flow hijacks) | 3 exploitsunknown effort | All 1 recorded result & sources3 exploits · raw 3 exploits Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 28. 831 entrypoints, 228 projects; revised August 2026 harness. Published successful control-flow hijack count. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=28 |
| ExploitGym (2h, exploit count) | 79 exploitsunknown effort | All 1 recorded result & sources79 exploits · raw 79 exploits Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 29. 869 known vulnerabilities; successful exploits must capture dynamic secret flag and use intended vulnerability. Wall-clock budget per task; mitigations disabled as specified on p. 29. Count, not percent. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=29 |
| ExploitGym (6h, exploit count) | 82 exploitsunknown effort | All 1 recorded result & sources82 exploits · raw 82 exploits Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 29. 869 known vulnerabilities; successful exploits must capture dynamic secret flag and use intended vulnerability. Wall-clock budget per task; mitigations disabled as specified on p. 29. Count, not percent. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=29 |
| Gray Swan IPI (k=1) | 0.7%unknown effort | All 1 recorded result & sources0.7% · raw 0.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 50. Q1+Q2 2026 IPI suite; probability attacker succeeds in k attempts; extended thinking, no product-specific prompt-injection protections. Do not mix with older challenge sets. Gemini/Kimi effort unspecified. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=50 |
| Gray Swan IPI (k=10) | 5.5%unknown effort | All 1 recorded result & sources5.5% · raw 5.5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 50. Q1+Q2 2026 IPI suite; probability attacker succeeds in k attempts; extended thinking, no product-specific prompt-injection protections. Do not mix with older challenge sets. Gemini/Kimi effort unspecified. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=50 |
| Gray Swan IPI | 7.1%unknown effort | All 1 recorded result & sources7.1% · raw 7.1 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 50. Q1+Q2 2026 IPI suite; probability attacker succeeds in k attempts; extended thinking, no product-specific prompt-injection protections. Do not mix with older challenge sets. Gemini/Kimi effort unspecified. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=50 |
| Gray Swan IPI k=15 (coding) | 0.2%unknown effort | All 1 recorded result & sources0.2% · raw 0.2 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 50. Domain-specific attack success rate, Q1+Q2 2026 suite; no product prompt-injection protections. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=50 |
| Gray Swan IPI k=15 (tool-use) | 4%unknown effort | All 1 recorded result & sources4% · raw 4 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 50. Domain-specific attack success rate, Q1+Q2 2026 suite; no product prompt-injection protections. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=50 |
| Gray Swan IPI k=15 (gui) | 24.4%unknown effort | All 1 recorded result & sources24.4% · raw 24.4 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 50. Domain-specific attack success rate, Q1+Q2 2026 suite; no product prompt-injection protections. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=50 |
| Single-turn harmless response rate (API) | 98.4%unknown effort | All 1 recorded result & sources98.4% · raw 98.39 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 31. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=31 |
| Single-turn harmless response rate (Claude.ai) | 99.7%unknown effort | All 1 recorded result & sources99.7% · raw 99.71 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 31. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=31 |
| Single-turn benign over-refusal (API) | 0.2%unknown effort | All 1 recorded result & sources0.2% · raw 0.17 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 32. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=32 |
| Single-turn benign over-refusal (Claude.ai) | 0.8%unknown effort | All 1 recorded result & sources0.8% · raw 0.82 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 32. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=32 |
| Child safety harmless (API) | 99.3%unknown effort | All 1 recorded result & sources99.3% · raw 99.29 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 34. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=34 |
| Child safety over-refusal (API) | 0%unknown effort | All 1 recorded result & sources0% · raw 0.04 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 34. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=34 |
| Child safety harmless (Claude.ai) | 99.9%unknown effort | All 1 recorded result & sources99.9% · raw 99.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 34. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=34 |
| Child safety over-refusal (Claude.ai) | 0%unknown effort | All 1 recorded result & sources0% · raw 0 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 34. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=34 |
| Suicide/self-harm safety harmless (API) | 99.6%unknown effort | All 1 recorded result & sources99.6% · raw 99.61 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 36. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=36 |
| Suicide/self-harm safety over-refusal (API) | 0%unknown effort | All 1 recorded result & sources0% · raw 0 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 36. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=36 |
| Suicide/self-harm safety harmless (Claude.ai) | 100%unknown effort | All 1 recorded result & sources100% · raw 100 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 36. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=36 |
| Suicide/self-harm safety over-refusal (Claude.ai) | 0.4%unknown effort | All 1 recorded result & sources0.4% · raw 0.41 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 36. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=36 |
| Disordered eating safety harmless (API) | 97.4%unknown effort | All 1 recorded result & sources97.4% · raw 97.43 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 39. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=39 |
| Disordered eating safety over-refusal (API) | 0%unknown effort | All 1 recorded result & sources0% · raw 0.03 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 39. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=39 |
| Disordered eating safety harmless (Claude.ai) | 99.8%unknown effort | All 1 recorded result & sources99.8% · raw 99.83 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 39. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=39 |
| Disordered eating safety over-refusal (Claude.ai) | 0.1%unknown effort | All 1 recorded result & sources0.1% · raw 0.11 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 39. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=39 |
| Child safety multi-turn appropriate responses (API) | 96%unknown effort | All 1 recorded result & sources96% · raw 96 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 35. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=35 |
| Child safety multi-turn appropriate responses (Claude.ai) | 99%unknown effort | All 1 recorded result & sources99% · raw 99 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 35. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=35 |
| Suicide/self-harm multi-turn appropriate responses (API) | 70%unknown effort | All 1 recorded result & sources70% · raw 70 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 37. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=37 |
| Suicide/self-harm multi-turn appropriate responses (Claude.ai) | 90%unknown effort | All 1 recorded result & sources90% · raw 90 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 37. Anthropic internal evaluation; API without system prompt and Claude.ai default system prompt are distinct. Updated evaluation set as of Haiku 5.5 card; do not infer real-world incidence. Reported percentage, confidence intervals in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=37 |
| BBQ disambiguated accuracy | 55.6%none effort | All 1 recorded result & sources55.6% · raw 55.56 % Headline · none effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 43. No system prompt; thinking disabled where supported. Accuracy, not signed bias. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=43 |
| BBQ ambiguous accuracy | 99.7%none effort | All 1 recorded result & sources99.7% · raw 99.65 % Headline · none effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 43. No system prompt; thinking disabled where supported. Accuracy, not signed bias. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=43 |
| Claude Code malicious refusal | 84.3%unknown effort | All 1 recorded result & sources84.3% · raw 84.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 48. Product agent harness; malicious computer-use is average of with/without thinking for Haiku/Sonnet, thinking-only for Opus. Specific evaluated task set, not general safety. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=48 |
| Claude Code dual-use/benign success | 98.9%unknown effort | All 1 recorded result & sources98.9% · raw 98.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 48. Product agent harness; malicious computer-use is average of with/without thinking for Haiku/Sonnet, thinking-only for Opus. Specific evaluated task set, not general safety. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=48 |
| Malicious computer-use refusal | 82.6%unknown effort | All 1 recorded result & sources82.6% · raw 82.59 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 49. Product agent harness; malicious computer-use is average of with/without thinking for Haiku/Sonnet, thinking-only for Opus. Specific evaluated task set, not general safety. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=49 |
| Shade IPI coding (without probes) | 0.1%unknown effort | All 1 recorded result & sources0.1% · raw 0.08 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 51. Adaptive red-team attempt-level attack success rate; 200 attempts/scenario. Coding 40 scenarios, GUI 14. Haiku has no model fallback; larger Claude models may fall back under protections. Scenario counts preserved in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=51 |
| Shade IPI coding (with probes) | 0%unknown effort | All 1 recorded result & sources0% · raw 0 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 51. Adaptive red-team attempt-level attack success rate; 200 attempts/scenario. Coding 40 scenarios, GUI 14. Haiku has no model fallback; larger Claude models may fall back under protections. Scenario counts preserved in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=51 |
| Shade IPI gui (without probes) | 0.1%unknown effort | All 1 recorded result & sources0.1% · raw 0.07 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 53. Adaptive red-team attempt-level attack success rate; 200 attempts/scenario. Coding 40 scenarios, GUI 14. Haiku has no model fallback; larger Claude models may fall back under protections. Scenario counts preserved in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=53 |
| Shade IPI gui (with probes) | 0.1%unknown effort | All 1 recorded result & sources0.1% · raw 0.07 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 53. Adaptive red-team attempt-level attack success rate; 200 attempts/scenario. Coding 40 scenarios, GUI 14. Haiku has no model fallback; larger Claude models may fall back under protections. Scenario counts preserved in source. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=53 |
| Cowork browser IPI (auto mode) | 0%unknown effort | All 1 recorded result & sources0% · raw 0 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 54. 110 scenarios, ten attacks/scenario; attempt-level ASR, 0/110 successful scenarios; auto mode on. Saturated evaluation scheduled for deprecation. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=54 |
| Anthropic audit: Misaligned behavior | 3.2 scoreunknown effort | All 1 recorded result & sources3.2 score · raw 3.19 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Misaligned with constitution | 3 scoreunknown effort | All 1 recorded result & sources3 score · raw 2.96 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Misaligned behavior in Claude Code sandboxes | 3 scoreunknown effort | All 1 recorded result & sources3 score · raw 3.02 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Misaligned behavior in GUI | 4 scoreunknown effort | All 1 recorded result & sources4 score · raw 4 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Cooperation with human misuse | 1.8 scoreunknown effort | All 1 recorded result & sources1.8 score · raw 1.83 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Compliance with deception toward user | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.134 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Harmful system prompt compliance | 1.2 scoreunknown effort | All 1 recorded result & sources1.2 score · raw 1.2 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Exfiltration or safeguard tampering | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.075 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Full-turn prefill susceptibility | 1.3 scoreunknown effort | All 1 recorded result & sources1.3 score · raw 1.276 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 61. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=61 |
| Anthropic audit: Accepting unverifiable authorization | 2.7 scoreunknown effort | All 1 recorded result & sources2.7 score · raw 2.69 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Complying with prompt injections | 1.6 scoreunknown effort | All 1 recorded result & sources1.6 score · raw 1.56 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Fraud | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.125 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Military-grade weapons | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.061 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Biological weapons uplift | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.017 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Chemical weapons uplift | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.028 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Radiological or nuclear weapons uplift | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.004 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Explosive weapons uplift | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.017 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Controlled substance synthesis uplift | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.047 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Harmful sexual content | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.054 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Disallowed cyberoffense | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.099 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Undermining liberal democracy | 1.2 scoreunknown effort | All 1 recorded result & sources1.2 score · raw 1.22 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 62. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=62 |
| Anthropic audit: Overrefusal | 2.2 scoreunknown effort | All 1 recorded result & sources2.2 score · raw 2.203 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 64. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=64 |
| Anthropic audit: Ignoring explicit constraints | 2.1 scoreunknown effort | All 1 recorded result & sources2.1 score · raw 2.06 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 64. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=64 |
| Anthropic audit: Reckless tool use | 2.4 scoreunknown effort | All 1 recorded result & sources2.4 score · raw 2.36 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 64. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=64 |
| Anthropic audit: User deception | 1.4 scoreunknown effort | All 1 recorded result & sources1.4 score · raw 1.36 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Sycophancy | 1.8 scoreunknown effort | All 1 recorded result & sources1.8 score · raw 1.77 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Encouragement of user delusion | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.038 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Evasiveness on controversial topics | 1.3 scoreunknown effort | All 1 recorded result & sources1.3 score · raw 1.34 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Input hallucination | 1.9 scoreunknown effort | All 1 recorded result & sources1.9 score · raw 1.88 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Important omissions | 2 scoreunknown effort | All 1 recorded result & sources2 score · raw 2.04 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Failure to disclose bad or lazy behavior | 1.7 scoreunknown effort | All 1 recorded result & sources1.7 score · raw 1.73 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: False completion claims | 1.9 scoreunknown effort | All 1 recorded result & sources1.9 score · raw 1.89 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 66. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=66 |
| Anthropic audit: Motivated reasoning | 2.3 scoreunknown effort | All 1 recorded result & sources2.3 score · raw 2.32 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 67. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=67 |
| Anthropic audit: Unsanctioned sandbox escape attempt | 1.3 scoreunknown effort | All 1 recorded result & sources1.3 score · raw 1.34 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 67. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=67 |
| Anthropic audit: Dangerous actions in possibly real setting | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.076 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 67. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=67 |
| Anthropic audit: Whistleblowing | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.034 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Self preservation | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.03 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Self serving bias | 1.2 scoreunknown effort | All 1 recorded result & sources1.2 score · raw 1.16 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Evidence of misaligned goals | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.03 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Indirect evidence of misalignment | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.086 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Unsanctioned third-party contact | 1.2 scoreunknown effort | All 1 recorded result & sources1.2 score · raw 1.2 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Unprompted boundary probing | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.114 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Internalized propaganda | 1 scoreunknown effort | All 1 recorded result & sources1 score · raw 1.045 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Approval-gate bypass | 1.7 scoreunknown effort | All 1 recorded result & sources1.7 score · raw 1.71 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 69. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=69 |
| Anthropic audit: Circumventing constraints or controls | 1.7 scoreunknown effort | All 1 recorded result & sources1.7 score · raw 1.65 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 70. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=70 |
| Anthropic audit: Security shortcut for expedience | 2.1 scoreunknown effort | All 1 recorded result & sources2.1 score · raw 2.09 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 70. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=70 |
| Anthropic audit: Behavior consistency | 6.2 scoreunknown effort | All 1 recorded result & sources6.2 score · raw 6.22 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 71. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=71 |
| Anthropic audit: Coherence between actions and views | 6.1 scoreunknown effort | All 1 recorded result & sources6.1 score · raw 6.12 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 71. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=71 |
| Anthropic audit: Unfaithful thinking | 1.9 scoreunknown effort | All 1 recorded result & sources1.9 score · raw 1.85 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 71. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=71 |
| Anthropic audit: Illegible thinking | 2.6 scoreunknown effort | All 1 recorded result & sources2.6 score · raw 2.58 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 71. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=71 |
| Anthropic audit: Good for user | 6.7 scoreunknown effort | All 1 recorded result & sources6.7 score · raw 6.74 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Supporting user autonomy | 7 scoreunknown effort | All 1 recorded result & sources7 score · raw 7.01 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Creative mastery | 6.5 scoreunknown effort | All 1 recorded result & sources6.5 score · raw 6.45 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Admirable behavior | 6.5 scoreunknown effort | All 1 recorded result & sources6.5 score · raw 6.45 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Fun or funny behavior | 2.5 scoreunknown effort | All 1 recorded result & sources2.5 score · raw 2.49 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Intellectual depth | 7.4 scoreunknown effort | All 1 recorded result & sources7.4 score · raw 7.39 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Warmth | 5.9 scoreunknown effort | All 1 recorded result & sources5.9 score · raw 5.94 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Character drift | 1.1 scoreunknown effort | All 1 recorded result & sources1.1 score · raw 1.082 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Wet blanket | 2.1 scoreunknown effort | All 1 recorded result & sources2.1 score · raw 2.089 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Anthropic audit: Condescension toward user | 1.6 scoreunknown effort | All 1 recorded result & sources1.6 score · raw 1.58 score Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 73. Automated internal behavioral audit; model-judged native 1–10 scale, not a rate. Up to ~4,100 investigations / ~1,850 seeds; some metrics apply only to scenario subsets. Simulated settings may not predict deployment; see pp. 58–76. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=73 |
| Sandbox boundaries (28 scenarios): all attempted crossings | 8.7%unknown effort | All 1 recorded result & sources8.7% · raw 8.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 75. 28 handmade scenarios ×14 trials (392 runs). Attempt severity is judged; freshly re-evaluated on current version, distinct from 1,753-scenario opportunistic sandbox test. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=75 |
| Sandbox boundaries (28 scenarios): low-severity attempts | 6.6%unknown effort | All 1 recorded result & sources6.6% · raw 6.6 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 75. 28 handmade scenarios ×14 trials (392 runs). Attempt severity is judged; freshly re-evaluated on current version, distinct from 1,753-scenario opportunistic sandbox test. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=75 |
| Sandbox boundaries (28 scenarios): moderate-or-higher attempts | 2%unknown effort | All 1 recorded result & sources2% · raw 2 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 75. 28 handmade scenarios ×14 trials (392 runs). Attempt severity is judged; freshly re-evaluated on current version, distinct from 1,753-scenario opportunistic sandbox test. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=75 |
| Sandbox boundaries (1,753 scenarios): attempts | 30%unknown effort | All 1 recorded result & sources30% · raw 30 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 78. Same 1,753 deliberately weak sandboxes, one attempt/scenario. Haiku 524 attempted crossings; incomplete instrumentation may underestimate. Distinct from 28-scenario severity evaluation. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=78 |
| MASK public split honesty | 96.8%unknown effort | All 1 recorded result & sources96.8% · raw 96.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 81. Public MASK split; fraction of scored pressured answers consistent with elicited belief. Evasions/inconsistent beliefs excluded. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=81 |
| MASK honesty: continuations | 99%unknown effort | All 1 recorded result & sources99% · raw 99 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 82. Public MASK scored subset by pressure category; figure rounded to whole percent except explicit text precision for Haiku disinformation. Not separate global honesty score. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=82 |
| MASK honesty: disinformation | 88.5%unknown effort | All 1 recorded result & sources88.5% · raw 88.5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 82. Public MASK scored subset by pressure category; figure rounded to whole percent except explicit text precision for Haiku disinformation. Not separate global honesty score. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=82 |
| MASK honesty: doubling down | 100%unknown effort | All 1 recorded result & sources100% · raw 100 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 82. Public MASK scored subset by pressure category; figure rounded to whole percent except explicit text precision for Haiku disinformation. Not separate global honesty score. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=82 |
| MASK honesty: known facts | 98%unknown effort | All 1 recorded result & sources98% · raw 98 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 82. Public MASK scored subset by pressure category; figure rounded to whole percent except explicit text precision for Haiku disinformation. Not separate global honesty score. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=82 |
| MASK honesty: provided facts | 97%unknown effort | All 1 recorded result & sources97% · raw 97 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 82. Public MASK scored subset by pressure category; figure rounded to whole percent except explicit text precision for Haiku disinformation. Not separate global honesty score. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=82 |
| Silent copying of leaked coding answers | 17.3%unknown effort | All 1 recorded result & sources17.3% · raw 17.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 83. Held-out coding tasks with reachable leaked solutions; substantial use without disclosure, as judged by Anthropic. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=83 |
| Spontaneous disclosure of concealed changes | 95.9%unknown effort | All 1 recorded result & sources95.9% · raw 95.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 84. Summary request on transcripts containing concealed repository changes; spontaneous disclosure rate. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=84 |
| Secret retention: replies | 57%unknown effort | All 1 recorded result & sources57% · raw 57 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 92. 39 simulated confidential jobs; Opus 5.5 counterpart presses for up to 20 turns. End-of-conversation fraction without leak. Reply-only and reasoning-visible conditions use different environments. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=92 |
| Secret retention: reasoning and replies | 3%unknown effort | All 1 recorded result & sources3% · raw 3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card, p. 92. 39 simulated confidential jobs; Opus 5.5 counterpart presses for up to 20 turns. End-of-conversation fraction without leak. Reply-only and reasoning-visible conditions use different environments. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=92 |
| Multi-turn biological weapons (api) | 94%unknown effort | All 1 recorded result & sources94% · raw 94 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn biological weapons (claude-ai) | 89%unknown effort | All 1 recorded result & sources89% · raw 89 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn cyberattacks (api) | 98%unknown effort | All 1 recorded result & sources98% · raw 98 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn cyberattacks (claude-ai) | 100%unknown effort | All 1 recorded result & sources100% · raw 100 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn deadly weapons (api) | 78%unknown effort | All 1 recorded result & sources78% · raw 78 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn deadly weapons (claude-ai) | 78%unknown effort | All 1 recorded result & sources78% · raw 78 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn hate and discrimination (api) | 99%unknown effort | All 1 recorded result & sources99% · raw 99 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn hate and discrimination (claude-ai) | 98%unknown effort | All 1 recorded result & sources98% · raw 98 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn influence operations (api) | 87%unknown effort | All 1 recorded result & sources87% · raw 87 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn influence operations (claude-ai) | 88%unknown effort | All 1 recorded result & sources88% · raw 88 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn romance scams (api) | 96%unknown effort | All 1 recorded result & sources96% · raw 96 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn romance scams (claude-ai) | 94%unknown effort | All 1 recorded result & sources94% · raw 94 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn tracking and surveillance (api) | 95%unknown effort | All 1 recorded result & sources95% · raw 95 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn tracking and surveillance (claude-ai) | 98%unknown effort | All 1 recorded result & sources98% · raw 98 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn violent extremism (api) | 90%unknown effort | All 1 recorded result & sources90% · raw 90 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Multi-turn violent extremism (claude-ai) | 90%unknown effort | All 1 recorded result & sources90% · raw 90 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 33. Appropriate throughout conversation; API without system prompt versus Claude.ai default prompt; synthetic user Opus 4.6, category-specific rubrics. Updated set; scores are not comparable across categories. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=33 |
| Political even handedness (api) | 97.7%unknown effort | All 1 recorded result & sources97.7% · raw 97.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 41. Pairwise political evaluation, 1,350 paired prompts across 150 topics and nine tasks; surfaces have distinct system prompts. Combined Fable/Mythos bar not assigned to either model. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=41 |
| Political even handedness (claude-ai) | 99.1%unknown effort | All 1 recorded result & sources99.1% · raw 99.1 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 41. Pairwise political evaluation, 1,350 paired prompts across 150 topics and nine tasks; surfaces have distinct system prompts. Combined Fable/Mythos bar not assigned to either model. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=41 |
| Political opposing perspectives (api) | 35.9%unknown effort | All 1 recorded result & sources35.9% · raw 35.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 41. Pairwise political evaluation, 1,350 paired prompts across 150 topics and nine tasks; surfaces have distinct system prompts. Combined Fable/Mythos bar not assigned to either model. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=41 |
| Political opposing perspectives (claude-ai) | 67.8%unknown effort | All 1 recorded result & sources67.8% · raw 67.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 41. Pairwise political evaluation, 1,350 paired prompts across 150 topics and nine tasks; surfaces have distinct system prompts. Combined Fable/Mythos bar not assigned to either model. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=41 |
| Political refusal (api) | 3.3%unknown effort | All 1 recorded result & sources3.3% · raw 3.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 42. Pairwise political evaluation, 1,350 paired prompts across 150 topics and nine tasks; surfaces have distinct system prompts. Combined Fable/Mythos bar not assigned to either model. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=42 |
| Political refusal (claude-ai) | 5%unknown effort | All 1 recorded result & sources5% · raw 5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 42. Pairwise political evaluation, 1,350 paired prompts across 150 topics and nine tasks; surfaces have distinct system prompts. Combined Fable/Mythos bar not assigned to either model. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=42 |
| Election integrity harmless api | 100%unknown effort | All 1 recorded result & sources100% · raw 100 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 45. Single-turn election evaluation; 300 harmful and 300 benign prompts; API without prompt versus Claude.ai default prompt. Specific evaluation, not general election capability. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=45 |
| Election integrity overrefusal api | 0.3%unknown effort | All 1 recorded result & sources0.3% · raw 0.33 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 45. Single-turn election evaluation; 300 harmful and 300 benign prompts; API without prompt versus Claude.ai default prompt. Specific evaluation, not general election capability. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=45 |
| Election integrity harmless claude ai | 100%unknown effort | All 1 recorded result & sources100% · raw 100 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 45. Single-turn election evaluation; 300 harmful and 300 benign prompts; API without prompt versus Claude.ai default prompt. Specific evaluation, not general election capability. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=45 |
| Election integrity overrefusal claude ai | 0.3%unknown effort | All 1 recorded result & sources0.3% · raw 0.33 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 45. Single-turn election evaluation; 300 harmful and 300 benign prompts; API without prompt versus Claude.ai default prompt. Specific evaluation, not general election capability. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=45 |
| Election integrity multi-turn (api) | 99%unknown effort | All 1 recorded result & sources99% · raw 99 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 45. Appropriate throughout multi-turn election conversations; synthetic users/rubric. API without system prompt versus Claude.ai default prompt. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=45 |
| Election integrity multi-turn (claude-ai) | 98%unknown effort | All 1 recorded result & sources98% · raw 98 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 45. Appropriate throughout multi-turn election conversations; synthetic users/rubric. API without system prompt versus Claude.ai default prompt. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=45 |
| LinuxArena stealth (hidden) | 5.9%max effort | All 1 recorded result & sources5.9% · raw 5.9 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 88. Successful harmful side task without Opus 4.8 monitor flag, best optimized prompt per model/setting; Haiku one run/task pair, peers five. Hidden reasoning withheld from monitor; other conditions expose full trajectory. Revised prompt optimization may change earlier scores. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=88 |
| LinuxArena stealth (none) | 0.4%none effort | All 1 recorded result & sources0.4% · raw 0.4 % Headline · none effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 88. Successful harmful side task without Opus 4.8 monitor flag, best optimized prompt per model/setting; Haiku one run/task pair, peers five. Hidden reasoning withheld from monitor; other conditions expose full trajectory. Revised prompt optimization may change earlier scores. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=88 |
| LinuxArena stealth (visible) | 0.1%max effort | All 1 recorded result & sources0.1% · raw 0.1 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 88. Successful harmful side task without Opus 4.8 monitor flag, best optimized prompt per model/setting; Haiku one run/task pair, peers five. Hidden reasoning withheld from monitor; other conditions expose full trajectory. Revised prompt optimization may change earlier scores. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=88 |
| SHADE-Arena stealth (hidden) | 1.2%max effort | All 1 recorded result & sources1.2% · raw 1.2 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 90. 24 task pairs, >=25 attempts/pair, best optimized prompt per setting, Opus 4.8 monitor. Haiku API safety interventions on; ended episodes count unsuccessful. Different from Shade IPI attacks. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=90 |
| SHADE-Arena stealth (none) | 5.2%none effort | All 1 recorded result & sources5.2% · raw 5.2 % Headline · none effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 90. 24 task pairs, >=25 attempts/pair, best optimized prompt per setting, Opus 4.8 monitor. Haiku API safety interventions on; ended episodes count unsuccessful. Different from Shade IPI attacks. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=90 |
| SHADE-Arena stealth (visible-low) | 2.3%low effort | All 1 recorded result & sources2.3% · raw 2.3 % Headline · low effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 90. 24 task pairs, >=25 attempts/pair, best optimized prompt per setting, Opus 4.8 monitor. Haiku API safety interventions on; ended episodes count unsuccessful. Different from Shade IPI attacks. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=90 |
| SHADE-Arena stealth (visible-max) | 0%max effort | All 1 recorded result & sources0% · raw 0 % Headline · max effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 90. 24 task pairs, >=25 attempts/pair, best optimized prompt per setting, Opus 4.8 monitor. Haiku API safety interventions on; ended episodes count unsuccessful. Different from Shade IPI attacks. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=90 |
| AA-Omniscience public split, Anthropic: net | 12%unknown effort | All 1 recorded result & sources12% · raw 12 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 80. Anthropic evaluation on public split, no tools; not independent AA private evaluation. Published fractions ×100; net = correct minus incorrect in percentage points. Chart precision retained. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=80 |
| AA-Omniscience public split, Anthropic: correct | 44%unknown effort | All 1 recorded result & sources44% · raw 44 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 80. Anthropic evaluation on public split, no tools; not independent AA private evaluation. Published fractions ×100; net = correct minus incorrect in percentage points. Chart precision retained. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=80 |
| AA-Omniscience public split, Anthropic: incorrect | 32%unknown effort | All 1 recorded result & sources32% · raw 32 % Headline · unknown effort · Own vendor Source/record date: 2026-10-07 Anthropic Haiku 5.5 system card p. 80. Anthropic evaluation on public split, no tools; not independent AA private evaluation. Published fractions ×100; net = correct minus incorrect in percentage points. Chart precision retained. https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf#page=80 |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 43.4max effort | All 5 recorded results & sources43.4 · raw 43.3950199670746 index Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-541.2 · raw 41.2489304677024 index Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh37.8 · raw 37.8240211671094 index Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high34.5 · raw 34.4646519474657 index Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium29.4 · raw 29.4494981304947 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Briefcase Elo · Artificial Analysis | 1577.7max effort | All 5 recorded results & sources1577.7 · raw 1577.72 Elo Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-51532.7 · raw 1532.67 Elo Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh1442.6 · raw 1442.55 Elo Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-high1371.7 · raw 1371.7 Elo Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-medium1111.6 · raw 1111.58 Elo Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| GDPval-AA Elo · Artificial Analysis | 1620.1max effort | All 5 recorded results & sources1620.1 · raw 1620.05 Elo Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-51510.8 · raw 1510.76 Elo Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh1420 · raw 1420.03 Elo Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-high1276.8 · raw 1276.84 Elo Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-medium1124.9 · raw 1124.94 Elo Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AutomationBench (AA) · Artificial Analysis | 35.4%max effort | All 5 recorded results & sources35.4% · raw 35.411015132784 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-536% · raw 35.953889773428 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh33.7% · raw 33.726853117672 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high28.6% · raw 28.589614814293 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium22.9% · raw 22.910812181111 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| Terminal-Bench 4.0 (AA) · Artificial Analysis | 32.8%max effort | All 5 recorded results & sources32.8% · raw 32.828282828283 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-529.3% · raw 29.292929292929 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh21.7% · raw 21.717171717172 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high15.2% · raw 15.151515151515 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium12.6% · raw 12.626262626263 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| SciCode (AA) · Artificial Analysis | 55%max effort | All 5 recorded results & sources55% · raw 54.976851851852 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-551.7% · raw 51.736111111111 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh48.7% · raw 48.726851851852 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high49% · raw 48.958333333333 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium49.2% · raw 49.189814814815 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| Humanity’s Last Exam (AA) · Artificial Analysis | 44.4%max effort | All 5 recorded results & sources44.4% · raw 44.392956441149 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-542.7% · raw 42.678405931418 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh37.3% · raw 37.25671918443 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high33.8% · raw 33.781278962002 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium27% · raw 27.015755329008 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| GDP.PDF (AA) · Artificial Analysis | 20.8%max effort | All 5 recorded results & sources20.8% · raw 20.8 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-518.2% · raw 18.2 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh17.2% · raw 17.2 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high15.2% · raw 15.2 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium11.2% · raw 11.2 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| CritPt (AA) · Artificial Analysis | 18.9%max effort | All 5 recorded results & sources18.9% · raw 18.857142857143 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-522.6% · raw 22.571428571429 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh18.6% · raw 18.571428571429 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high12.9% · raw 12.857142857143 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium9.1% · raw 9.142857142857 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Omniscience Index · Artificial Analysis | 10.7 scoremax effort | All 5 recorded results & sources10.7 score · raw 10.666666666666666 score Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-56.1 score · raw 6.05 score Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh5.8 score · raw 5.75 score Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-high4.5 score · raw 4.466666666666667 score Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-medium3.3 score · raw 3.25 score Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Native index/Elo. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-LCR · Artificial Analysis | 82.7%max effort | All 5 recorded results & sources82.7% · raw 82.666666666667 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-578.3% · raw 78.333333333333 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh77.3% · raw 77.333333333333 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high77.3% · raw 77.333333333333 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium72% · raw 72 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Omniscience accuracy · Artificial Analysis | 36.4%max effort | All 5 recorded results & sources36.4% · raw 36.383333333333 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-535% · raw 34.95 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh34.8% · raw 34.783333333333 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high34% · raw 34.033333333333 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium33.3% · raw 33.25 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Omniscience hallucination rate · Artificial Analysis | 40.4%max effort | All 5 recorded results & sources40.4% · raw 40.424417081478 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-544.4% · raw 44.427363566487 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh44.5% · raw 44.518272425249 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high44.8% · raw 44.820616472966 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium44.9% · raw 44.943820224719 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| Terminal-Bench Science 0.1 (AA) · Artificial Analysis | 20%max effort | All 5 recorded results & sources20% · raw 20 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-518.1% · raw 18.095238095238 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh10% · raw 10 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high7.1% · raw 7.142857142857 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium1.9% · raw 1.904761904762 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| HLAB (AA) · Artificial Analysis | 89.9%max effort | All 5 recorded results & sources89.9% · raw 89.868287740628 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-589.9% · raw 89.897235489941 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh89.5% · raw 89.520914748878 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high89.2% · raw 89.231437255753 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium87.9% · raw 87.85641916341 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA industry index: financeAndAccounting · Artificial Analysis | 43.9 scoremax effort | All 5 recorded results & sources43.9 score · raw 43.899544721027 score Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-540.8 score · raw 40.7857632782172 score Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh37.7 score · raw 37.6782515170274 score Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high35.4 score · raw 35.4471752244645 score Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium28.8 score · raw 28.7822173180696 score Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA industry index: strategyAndOps · Artificial Analysis | 40.4 scoremax effort | All 5 recorded results & sources40.4 score · raw 40.3835291326876 score Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-539.3 score · raw 39.281621532151 score Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh37.5 score · raw 37.4933078688108 score Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high33.5 score · raw 33.5360066238334 score Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium28 score · raw 27.9750207106805 score Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA industry index: legal · Artificial Analysis | 42.3 scoremax effort | All 5 recorded results & sources42.3 score · raw 42.2752066606861 score Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-539.9 score · raw 39.9006717721577 score Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh37.6 score · raw 37.6058438212847 score Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high35.5 score · raw 35.5422528285353 score Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium30.6 score · raw 30.647798971037 score Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA industry index: engineering · Artificial Analysis | 44.9 scoremax effort | All 5 recorded results & sources44.9 score · raw 44.8823073189862 score Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-543.1 score · raw 43.1016145693664 score Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh39.2 score · raw 39.1642979209545 score Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high35.2 score · raw 35.231640545599 score Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium31.2 score · raw 31.2103312647192 score Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA industry index: economics · Artificial Analysis | 51.2 scoremax effort | All 5 recorded results & sources51.2 score · raw 51.2271805877356 score Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-548.7 score · raw 48.704358742663 score Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh45.8 score · raw 45.8402683812172 score Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high43.3 score · raw 43.2746768033673 score Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium37.5 score · raw 37.4612643651529 score Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Briefcase rubric pass rate · Artificial Analysis | 52.3%max effort | All 5 recorded results & sources52.3% · raw 52.323232323232 % Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-551.4% · raw 51.414141414141 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh47% · raw 46.969696969697 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-high44.9% · raw 44.949494949495 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-medium33.8% · raw 33.838383838384 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. Raw fraction ×100. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Briefcase analytical quality Elo · Artificial Analysis | 1907.6max effort | All 5 recorded results & sources1907.6 · raw 1907.63 Elo Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-51812.7 · raw 1812.7 Elo Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh1702.4 · raw 1702.38 Elo Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high1599 · raw 1598.98 Elo Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium1225.1 · raw 1225.1 Elo Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| AA-Briefcase presentation Elo · Artificial Analysis | 1429.1max effort | All 5 recorded results & sources1429.1 · raw 1429.1 Elo Headline · max effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 max configuration; Claude Haiku 5.5 (Max, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-51393.4 · raw 1393.42 Elo Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 xhigh configuration; Claude Haiku 5.5 (Xhigh, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-xhigh1303.5 · raw 1303.45 Elo Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 high configuration; Claude Haiku 5.5 (High, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-high1220.8 · raw 1220.84 Elo Alternative · medium effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 medium configuration; Claude Haiku 5.5 (Medium, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-medium990.3 · raw 990.32 Elo Alternative · low effort · Independent evaluator Source/record date: 2026-10-08 AA primary exact Haiku 5.5 low configuration; Claude Haiku 5.5 (Low, Default Fallback). Non-estimated published values. Default safeguard fallback harness differs from Anthropic no-fallback runs. Evaluation date unpublished; checked 8 October. https://artificialanalysis.ai/models/claude-haiku-5-5-low |
| Bugs fixed /105 · Bug Hunt Bench | 22 fixesmax effort | All 3 recorded results & sources22 fixes · raw 22 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-08 Claude Code harness; 1 runs, evaluated 7 October 2026; published raw fixes out of 105, not percentage. https://github.com/phuryn/bug-hunt-bench18.3 fixes · raw 18.3 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-08 Claude Code harness; 3 runs, evaluated 7 October 2026; published raw fixes out of 105, not percentage. https://github.com/phuryn/bug-hunt-bench15 fixes · raw 15 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-08 Claude Code harness; 1 runs, evaluated 7 October 2026; published raw fixes out of 105, not percentage. https://github.com/phuryn/bug-hunt-bench |
Read how we select and source scores or the comparison guide.