Examenos

GPT-6 Astra Benchmarks, Specifications & Availability

Explore GPT-6 Astra from OpenAI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare GPT-6 Astra with other models →Explore data coverage

Published specifications

Provider
OpenAI
Access
Proprietary
License
Proprietary
Context window
1.05M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-03
Modalities
text, image, file
Family
GPT-6

Model card · Announcement · Website · OpenRouter

Model notes

Flagship GPT-6 generation; computer use / agents / science / cyber focus.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

Higher rates above ~272k prompt tokens on OpenRouter overrides.; min-healthy endpoint minima 2026-10-01; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-01; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-02; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-03; time-window overrides apply on a winning provider

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
AA Coding Agent Index v1.467unknown effort
All 1 recorded result & sources

67 · raw 67 index

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

Artificial Analysis Coding Agent Index v1.4 as listed by OpenAI

https://openai.com/index/gpt-6-astra/
Agents' Last Exam59.3%unknown effort
All 3 recorded results & sources

59.3% · raw 59.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

33.3% · raw 33.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

34.2% · raw 34.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 59.3 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
ARC-AGI-198.5%unknown effort
All 1 recorded result & sources

98.5% · raw 98.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

ARC-AGI-1

https://openai.com/index/gpt-6-astra/
ARC-AGI-295%unknown effort
All 1 recorded result & sources

95% · raw 95 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

ARC-AGI-2

https://openai.com/index/gpt-6-astra/
ARC-AGI-399.9%unknown effort
All 1 recorded result & sources

99.9% · raw 99.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/
AutomationBench v1.0.641.4%unknown effort
All 2 recorded results & sources

41.4% · raw 41.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

41.4% · raw 41.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
DeepSWE v1.174.1%unknown effort
All 2 recorded results & sources

74.1% · raw 74.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

74.1% · raw 74.1 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
ExploitBench100%unknown effort
All 1 recorded result & sources

100% · raw 100 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/
ExploitBench (Jun–Aug 2026)39%unknown effort
All 1 recorded result & sources

39% · raw 39 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

Internal contamination-aware ExploitBench Jun–Aug 2026

https://openai.com/index/gpt-6-astra/
ExploitGym42.4%unknown effort
All 1 recorded result & sources

42.4% · raw 42.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/
FrontierCode 1.1 Extended64.5%unknown effort
All 1 recorded result & sources

64.5% · raw 64.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

FrontierCode 1.1 Extended; footnote 8 on OpenAI page

https://openai.com/index/gpt-6-astra/
FrontierCode 1.1 Main53.3%unknown effort
All 2 recorded results & sources

53.3% · raw 53.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

FrontierCode 1.1 Main; footnote 8

https://openai.com/index/gpt-6-astra/

53.3% · raw 53.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value.

https://www.anthropic.com/claude-opus-5-5
FrontierMath Tier 497.6%unknown effort
All 1 recorded result & sources

97.6% · raw 97.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

FrontierMath Tier 4 (v2); prose also cites 98% saturation

https://openai.com/index/gpt-6-astra/
GeneBench Pro37.1%unknown effort
All 1 recorded result & sources

37.1% · raw 37.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

GeneBench Pro

https://openai.com/index/gpt-6-astra/
GPQA Diamond96%unknown effort
All 1 recorded result & sources

96% · raw 96 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/
HealthBench Professional64.7%unknown effort
All 2 recorded results & sources

64.7% · raw 64.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-22

length-adjusted (raw 68.2, mean length 3185 chars); Sep 22 system-card correction supersedes earlier launch-post 63.4

https://deploymentsafety.openai.com/gpt-6-astra

70.3% · raw 70.3 %

Alternative · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); length-adjusted with Anthropic grader. Kept non-headline under Astra own-card 64.7 (different harness)

https://www.anthropic.com/claude-sonnet-5-5-system-card
Humanity's Last Exam (w/ tools)57.2%unknown effort
All 2 recorded results & sources

57.2% · raw 57.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

57.2% · raw 57.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value.

https://www.anthropic.com/claude-opus-5-5
OSWorld 2.072.6%unknown effort
All 2 recorded results & sources

72.6% · raw 72.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

v2026.08.08 offline set, partial score; ~40 min/task

https://openai.com/index/gpt-6-astra/

72.6% · raw 72.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
ScreenSpot-Pro92.7%unknown effort
All 1 recorded result & sources

92.7% · raw 92.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

ScreenSpot-Pro no tools

https://openai.com/index/gpt-6-astra/
SEC-Bench Pro85.4%unknown effort
All 1 recorded result & sources

85.4% · raw 85.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/
SRE-Bench88%unknown effort
All 1 recorded result & sources

88% · raw 88 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

SRE-Bench reverse engineering

https://openai.com/index/gpt-6-astra/
Terminal-Bench 4.057.9%unknown effort
All 3 recorded results & sources

57.9% · raw 57.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

57.9% · raw 57.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce (OpenAI figure) | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value.

https://www.anthropic.com/claude-opus-5-5

58.2% · raw 58.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 57.9 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Terminal-Bench Science 0.164.6%unknown effort
All 3 recorded results & sources

64.6% · raw 64.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

64.6% · raw 64.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce (OpenAI figure) | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value.

https://www.anthropic.com/claude-opus-5-5

68.1% · raw 68.1 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 64.6 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
HealthBench Hard36.6%unknown effort
All 1 recorded result & sources

36.6% · raw 36.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-22

length-adjusted (raw 34.2, mean length 1697 chars); Astra system card HealthBench table (Sep 22 corrected values)

https://deploymentsafety.openai.com/gpt-6-astra
GDPval-AA 2.11542unknown effort
All 1 recorded result & sources

1542 · raw 1542 Elo

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5
AutomationBench41.4%unknown effort
All 1 recorded result & sources

41.4% · raw 41.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce (Zapier public)

https://www.anthropic.com/claude-opus-5-5
AA Intelligence Index (vendor-cited)61unknown effort
All 1 recorded result & sources

61 · raw 61 index

Headline · unknown effort · Peer vendor

Source/record date: Not recorded

from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); GPT-6 Astra (max)

https://huggingface.co/inclusionAI/Ling-3.0-flash-VL
FrontierSWE65.5%unknown effort
All 1 recorded result & sources

65.5% · raw 65.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); Proximal harness

https://www.anthropic.com/claude-sonnet-5-5-system-card
Chartography71%unknown effort
All 2 recorded results & sources

71% · raw 71 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); without tools, per Surge AI public reporting

https://www.anthropic.com/claude-sonnet-5-5-system-card

71% · raw 71 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
HealthBench60%max effort
All 1 recorded result & sources

60% · raw 60 %

Headline · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); length-adjusted Sep-2026 figure with Anthropic grader and harness, differs from Astra own-card setup

https://www.anthropic.com/claude-sonnet-5-5-system-card
Vals Index63.1%unknown effort
All 1 recorded result & sources

63.1% · raw 63.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Vals Finance Agent v253.5%unknown effort
All 1 recorded result & sources

53.5% · raw 53.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Harvey Legal Agent Benchmark5.4%unknown effort
All 1 recorded result & sources

5.4% · raw 5.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
FrontierSWE v265.5%unknown effort
All 1 recorded result & sources

65.5% · raw 65.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Vibe Code Bench89.6%unknown effort
All 1 recorded result & sources

89.6% · raw 89.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
PostTrainBench44.3%unknown effort
All 1 recorded result & sources

44.3% · raw 44.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
LABBench 285.4%unknown effort
All 1 recorded result & sources

85.4% · raw 85.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
RiemannBench72%unknown effort
All 1 recorded result & sources

72% · raw 72 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GraphWalks up to 128k BFS F198.7 f1unknown effort
All 1 recorded result & sources

98.7 f1 · raw 98.7 f1

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GraphWalks 256k-1M BFS F171.8 f1unknown effort
All 1 recorded result & sources

71.8 f1 · raw 71.8 f1

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
LVBench87.5%unknown effort
All 1 recorded result & sources

87.5% · raw 87.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
CWE-bench v168%unknown effort
All 1 recorded result & sources

68% · raw 68 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Gray Swan IPI8.5%unknown effort
All 1 recorded result & sources

8.5% · raw 8.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology

https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis53max effort
All 5 recorded results & sources

53 · raw 53 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra

52 · raw 52 index

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-xhigh

51 · raw 51 index

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-high

50 · raw 50 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-medium

46 · raw 46 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-low
Cost per Intelligence Index task · Artificial Analysis$3.26max effort
All 5 recorded results & sources

$3.26 · raw 3.26 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra

$2.31 · raw 2.31 USD

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-xhigh

$1.73 · raw 1.73 USD

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-high

$1.54 · raw 1.54 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-medium

$0.82 · raw 0.82 USD

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-astra-low
Output speed · Artificial Analysis54 tok/smax effort
All 5 recorded results & sources

54 tok/s · raw 54 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-astra

49 tok/s · raw 49 tok/s

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-astra-xhigh

45 tok/s · raw 45 tok/s

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-astra-high

45 tok/s · raw 45 tok/s

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-astra-medium

46 tok/s · raw 46 tok/s

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-astra-low
Coding Agent Index · Artificial Analysis62unknown effort
All 1 recorded result & sources

62 · raw 62 index

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-22

Codex harness; ties Fable 5.1

https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
Vals Index · Vals AI63.1%unknown effort
All 1 recorded result & sources

63.1% · raw 63.13 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $18.46.

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench45 fixesmax effort
All 5 recorded results & sources

45 fixes · raw 45 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

43 fixes · raw 43 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

35 fixes · raw 35 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort high; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

34 fixes · raw 34 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

27 fixes · raw 27 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort low; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Vibe Code Bench v1.1 · Vals AI89.6%unknown effort
All 1 recorded result & sources

89.6% · raw 89.59 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $38.51.

https://www.vals.ai/benchmarks/vibe-code
ARC-AGI-1 · ARC Prize98.5%xhigh effort
All 6 recorded results & sources

97.5% · raw 97.5 %

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

98.5% · raw 98.5 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-astra

98.5% · raw 98.5 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

97.5% · raw 97.5 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

96.5% · raw 96.5 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

86% · raw 86 %

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra
ARC-AGI-2 · ARC Prize95%max effort
All 6 recorded results & sources

95% · raw 95 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-astra

93.3% · raw 93.3 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

92.1% · raw 92.1 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

92.1% · raw 92.1 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

85.4% · raw 85.4 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra

59.6% · raw 59.6 %

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-astra
ARC-AGI-3 (Standard) · ARC Prize62.7%max effort
All 6 recorded results & sources

62.7% · raw 62.71 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry). Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-astra

59.3% · raw 59.34 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-astra

54.8% · raw 54.82 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-astra

38.6% · raw 38.59 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-astra

17.5% · raw 17.45 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-astra

35.2% · raw 35.18 %

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-astra
ARC-AGI-3 · ARC Prize100%high effort
All 6 recorded results & sources

98.6% · raw 98.55 %

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-astra

98.4% · raw 98.44 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-astra

100% · raw 99.95 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-astra

98.4% · raw 98.44 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-astra

98% · raw 98.03 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-astra

96.7% · raw 96.72 %

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-astra
AA-Omniscience Index · Artificial Analysis44high effort
All 1 recorded result & sources

44 · raw 44 index

Headline · high effort · Independent evaluator

Source/record date: 2026-09-23

AA-Omniscience Index; GPT-6 Astra (high) as cited on evaluation page

https://artificialanalysis.ai/evaluations/omniscience
GDPval-AA Elo · Artificial Analysis1542max effort
All 1 recorded result & sources

1542 · raw 1542 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; GPT-6 Astra (max)

https://artificialanalysis.ai/evaluations/gdpval-aa
Money gain · Andon Labs$15014.70unknown effort
All 1 recorded result & sources

$15014.70 · raw 15014.7 $

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked.

https://andonlabs.com/evals/vending-bench-2
Blueprint Bench · Andon Labs49.7%unknown effort
All 1 recorded result & sources

49.7% · raw 49.7 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100.

https://andonlabs.com/evals/blueprint-bench-2
Average Score · WeirdML v342.2%xhigh effort
All 1 recorded result & sources

42.2% · raw 42.22 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant GPT-6 Astra (xhigh); harness codex_cli 0.154.0; values from prepared data JSON; raw 0.422202; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v354.1%xhigh effort
All 1 recorded result & sources

54.1% · raw 54.05 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant GPT-6 Astra (xhigh); harness codex_cli 0.154.0; values from prepared data JSON; raw 0.540453; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$27.85xhigh effort
All 1 recorded result & sources

$27.85 · raw 27.85 USD

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant GPT-6 Astra (xhigh); harness codex_cli 0.154.0; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html

Read how we select and source scores or the comparison guide.