Examenos

Claude Fable 5.1 Benchmarks, Specifications & Availability

Explore Claude Fable 5.1 from Anthropic: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Claude Fable 5.1 with other models →Explore data coverage

Published specifications

Provider
Anthropic
Access
Proprietary
License
Proprietary
Context window
1M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-01
Modalities
text, image, file
Family
Claude Fable 5

Model card · Announcement · Website · OpenRouter

Model notes

Generally available Fable 5.1; Mythos 5.1 is same base under stricter safeguards.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
ARC-AGI-197.5%unknown effort
All 1 recorded result & sources

97.5% · raw 97.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra comparison table

https://openai.com/index/gpt-6-astra/
ARC-AGI-290%unknown effort
All 1 recorded result & sources

90% · raw 90 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra comparison table

https://openai.com/index/gpt-6-astra/
AutomationBench v1.0.631.4%unknown effort
All 2 recorded results & sources

31.4% · raw 31.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/

31.4% · raw 31.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
DeepSWE v1.167.4%unknown effort
All 2 recorded results & sources

67.4% · raw 67.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/

67.4% · raw 67.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
ExploitGym30.4%unknown effort
All 1 recorded result & sources

30.4% · raw 30.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
FrontierCode 1.1 Extended63.6%unknown effort
All 1 recorded result & sources

63.6% · raw 63.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra comparison table

https://openai.com/index/gpt-6-astra/
FrontierCode 1.1 Main50.3%unknown effort
All 2 recorded results & sources

50.9% · raw 50.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra comparison table | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Opus 5.5 announce (anthropic.com/claude-opus-5-5); model vendor's own (Anthropic) figure preferred over peer-vendor OpenAI comparison table (rule a); Anthropic source is also later-dated (2026-09-22). Values differ (50.9 here vs 50.3 Anthropic); not reconciled.

https://openai.com/index/gpt-6-astra/

50.3% · raw 50.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5
FrontierMath Tier 487.8%unknown effort
All 1 recorded result & sources

87.8% · raw 87.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
GPQA Diamond93.7%unknown effort
All 1 recorded result & sources

93.7% · raw 93.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
HealthBench Professional58.1%unknown effort
All 1 recorded result & sources

58.1% · raw 58.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
Humanity's Last Exam (w/ tools)65.6%unknown effort
All 2 recorded results & sources

65% · raw 65 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Opus 5.5 announce (anthropic.com/claude-opus-5-5); model vendor's own (Anthropic) figure preferred over peer-vendor OpenAI comparison table (rule a); Anthropic source is also later-dated (2026-09-22). Values differ (65.0 here vs 65.6 Anthropic); not reconciled. source_type relabeled official→vendor-comparison (OpenAI reporting an Anthropic model).

https://openai.com/index/gpt-6-astra/

65.6% · raw 65.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5
Terminal-Bench 4.055.8%unknown effort
All 3 recorded results & sources

55.8% · raw 55.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI GPT-6 Astra announcement table | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Opus 5.5 announce (anthropic.com/claude-opus-5-5); model vendor's own (Anthropic) figure preferred over peer-vendor OpenAI comparison table (rule a); Anthropic source is also later-dated (2026-09-22); same value. source_type relabeled official→vendor-comparison (OpenAI reporting an Anthropic model).

https://openai.com/index/gpt-6-astra/

55.8% · raw 55.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5

57.9% · raw 57.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 55.8 (spillover); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Terminal-Bench Science 0.152.6%unknown effort
All 3 recorded results & sources

52.6% · raw 52.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-01

Anthropic announcement; also Simon Willison summary

https://www.anthropic.com/claude/fable

52.6% · raw 52.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Fable launch page (anthropic.com/claude/fable); model's own launch announcement preferred over later spillover in Opus 5.5 comparison table (rule a); same value.

https://www.anthropic.com/claude-opus-5-5

52.6% · raw 52.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
CursorBench 4.051.8%unknown effort
All 1 recorded result & sources

51.8% · raw 51.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5
GDPval-AA 2.11735unknown effort
All 1 recorded result & sources

1735 · raw 1735 Elo

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5
AutomationBench31.4%unknown effort
All 1 recorded result & sources

31.4% · raw 31.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce (Zapier)

https://www.anthropic.com/claude-opus-5-5
OSWorld 2.077.9%unknown effort
All 2 recorded results & sources

80.7% · raw 80.7 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

OSWorld 2.0 partial; as reported by Anthropic Opus 5.5 announce; demoted as headline after Fable own-page figure harvest

https://www.anthropic.com/claude-opus-5-5

77.9% · raw 77.9 %

Headline · unknown effort · Own vendor

Source/record date: Not recorded

from chart/figure Anthropic Fable page; OSWorld 2.0 partial (strict 41.7% also shown)

https://www.anthropic.com/claude/fable
Chartography (w/ tools)88.4%unknown effort
All 1 recorded result & sources

88.4% · raw 88.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5
GDPval-AA v21853unknown effort
All 1 recorded result & sources

1853 · raw 1853 Elo

Headline · unknown effort · Own vendor

Source/record date: Not recorded

from chart/figure Anthropic Fable page benchmark table; GDPval-AA v2 Elo

https://www.anthropic.com/claude/fable
Humanity's Last Exam60.9%unknown effort
All 1 recorded result & sources

60.9% · raw 60.9 %

Headline · unknown effort · Own vendor

Source/record date: Not recorded

from chart/figure Anthropic Fable page; Humanity's Last Exam no tools

https://www.anthropic.com/claude/fable
CursorBench 3.2.073.4%unknown effort
All 1 recorded result & sources

73.4% · raw 73.4 %

Headline · unknown effort · Own vendor

Source/record date: Not recorded

from chart/figure Anthropic Fable page; CursorBench 3.2.0 (distinct from 4.0)

https://www.anthropic.com/claude/fable
AA Intelligence Index (vendor-cited)66unknown effort
All 1 recorded result & sources

66 · raw 66 index

Headline · unknown effort · Peer vendor

Source/record date: Not recorded

from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); Claude Fable 5.1 (max with fallback)

https://huggingface.co/inclusionAI/Ling-3.0-flash-VL
FrontierSWE56.3%unknown effort
All 1 recorded result & sources

56.3% · raw 56.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); Proximal harness

https://www.anthropic.com/claude-sonnet-5-5-system-card
Global MMLU94%max effort
All 1 recorded result & sources

94% · raw 94 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
MILU93%max effort
All 1 recorded result & sources

93% · raw 93 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card)

https://www.anthropic.com/claude-sonnet-5-5-system-card
Vals Index65.8%unknown effort
All 1 recorded result & sources

65.8% · raw 65.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Vals Finance Agent v258.9%unknown effort
All 1 recorded result & sources

58.9% · raw 58.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Harvey Legal Agent Benchmark6.7%unknown effort
All 1 recorded result & sources

6.7% · raw 6.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
FrontierSWE v256.3%unknown effort
All 1 recorded result & sources

56.3% · raw 56.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Vibe Code Bench90.3%unknown effort
All 1 recorded result & sources

90.3% · raw 90.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
PostTrainBench40.2%unknown effort
All 1 recorded result & sources

40.2% · raw 40.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
LABBench 268.6%unknown effort
All 1 recorded result & sources

68.6% · raw 68.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
RiemannBench65.6%unknown effort
All 1 recorded result & sources

65.6% · raw 65.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GraphWalks up to 128k BFS F191.4 f1unknown effort
All 1 recorded result & sources

91.4 f1 · raw 91.4 f1

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GraphWalks 256k-1M BFS F165 f1unknown effort
All 1 recorded result & sources

65 f1 · raw 65 f1

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Chartography46.2%unknown effort
All 1 recorded result & sources

46.2% · raw 46.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
LVBench79.7%unknown effort
All 1 recorded result & sources

79.7% · raw 79.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
CWE-bench v158%unknown effort
All 1 recorded result & sources

58% · raw 58 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Gray Swan IPI1%unknown effort
All 1 recorded result & sources

1% · raw 1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis53max effort
All 5 recorded results & sources

53 · raw 53 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1

53 · raw 53 index

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-xhigh

51 · raw 51 index

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-high

49 · raw 49 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-medium

47 · raw 47 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-low
Cost per Intelligence Index task · Artificial Analysis$7.63max effort
All 5 recorded results & sources

$7.63 · raw 7.63 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1

$5.98 · raw 5.98 USD

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-xhigh

$3.91 · raw 3.91 USD

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-high

$2.98 · raw 2.98 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-medium

$2.37 · raw 2.37 USD

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration.

https://artificialanalysis.ai/models/claude-fable-5-1-low
Output speed · Artificial Analysis68 tok/smax effort
All 5 recorded results & sources

68 tok/s · raw 68 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-fable-5-1

66 tok/s · raw 66 tok/s

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-fable-5-1-xhigh

54 tok/s · raw 54 tok/s

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-fable-5-1-high

54 tok/s · raw 54 tok/s

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-fable-5-1-medium

53 tok/s · raw 53 tok/s

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; with fallback reasoning configuration. Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/claude-fable-5-1-low
Vals Index · Vals AI65.8%unknown effort
All 1 recorded result & sources

65.8% · raw 65.83 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $28.71.

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench43 fixesmax effort
All 5 recorded results & sources

43 fixes · raw 43 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort max; 1 runs; evaluation 2026-09-01. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

33 fixes · raw 33 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort high; 1 runs; evaluation 2026-09-01. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

29 fixes · raw 29 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort low; 1 runs; evaluation 2026-09-02. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

29 fixes · raw 29 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort xhigh; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

21 fixes · raw 21 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort medium; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Pass Rate · MCP Atlas (Scale Labs)87.2%unknown effort
All 1 recorded result & sources

87.2% · raw 87.2 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

Scale Labs Performance Comparison chart (Fable 5.1); ±2.05 CI

https://labs.scale.com/leaderboard/mcp_atlas
Vibe Code Bench v1.1 · Vals AI90.3%unknown effort
All 1 recorded result & sources

90.3% · raw 90.26 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $33.37.

https://www.vals.ai/benchmarks/vibe-code
CUDA board % of roofline · KernelBench (community board)48.4%max effort
All 1 recorded result & sources

48.4% · raw 48.4 %

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

CUDA deck 4/4 pass; Mega 22.95×; or-fable max harness

https://kernelbench.com/models/claude-fable-5-1
ARC-AGI-1 · ARC Prize97.5%max effort
All 5 recorded results & sources

97.5% · raw 97.5 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/anthropic-claude-fable-5-1

96.5% · raw 96.5 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1

96% · raw 96 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1

94.5% · raw 94.5 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1

90% · raw 90 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1
ARC-AGI-2 · ARC Prize90%max effort
All 5 recorded results & sources

90% · raw 90 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/anthropic-claude-fable-5-1

90% · raw 90 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1

88.8% · raw 88.8 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1

86.3% · raw 86.3 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1

78.3% · raw 78.3 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/anthropic-claude-fable-5-1
AA-Omniscience Index · Artificial Analysis43max effort
All 1 recorded result & sources

43 · raw 43 index

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

AA-Omniscience Index; Adaptive Reasoning Max Effort Default Fallback

https://artificialanalysis.ai/evaluations/omniscience
GDPval-AA Elo · Artificial Analysis1735max effort
All 1 recorded result & sources

1735 · raw 1735 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; Claude Fable 5.1 Adaptive Reasoning Max Effort Default Fallback

https://artificialanalysis.ai/evaluations/gdpval-aa
Money gain · Andon Labs$4921.56unknown effort
All 1 recorded result & sources

$4921.56 · raw 4921.56 $

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked.

https://andonlabs.com/evals/vending-bench-2
Blueprint Bench · Andon Labs41.9%unknown effort
All 1 recorded result & sources

41.9% · raw 41.9 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100.

https://andonlabs.com/evals/blueprint-bench-2
Average Score · WeirdML v326%xhigh effort
All 1 recorded result & sources

26% · raw 25.97 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Fable 5.1 (xhigh); harness claude_code 2.1.270; values from prepared data JSON; raw 0.259702; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v343.5%xhigh effort
All 1 recorded result & sources

43.5% · raw 43.53 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Fable 5.1 (xhigh); harness claude_code 2.1.270; values from prepared data JSON; raw 0.435272; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$31.93xhigh effort
All 1 recorded result & sources

$31.93 · raw 31.93 USD

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Claude Fable 5.1 (xhigh); harness claude_code 2.1.270; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html

Read how we select and source scores or the comparison guide.