GPT-6 Astra Benchmarks, Specifications & Availability
Explore GPT-6 Astra from OpenAI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare GPT-6 Astra with other models →Explore data coverage
Published specifications
- Provider
- OpenAI
- Access
- Proprietary
- License
- Proprietary
- Context window
- 1.05M
- Total parameters
- Not published
- Active parameters
- Not published
- Released
- 2026-09-03
- Modalities
- text, image, file
- Family
- GPT-6
Model card · Announcement · Website · OpenRouter
Model notes
Flagship GPT-6 generation; computer use / agents / science / cyber focus.
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | [object Object] | 2026-10-03 · OpenRouter source |
Recorded pricing notes
Higher rates above ~272k prompt tokens on OpenRouter overrides.; min-healthy endpoint minima 2026-10-01; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-01; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-02; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-03; time-window overrides apply on a winning provider
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| AA Coding Agent Index v1.4 | 67unknown effort | All 1 recorded result & sources67 · raw 67 index Headline · unknown effort · Own vendor Source/record date: 2026-09-03 Artificial Analysis Coding Agent Index v1.4 as listed by OpenAI https://openai.com/index/gpt-6-astra/ |
| Agents' Last Exam | 59.3%unknown effort | All 3 recorded results & sources59.3% · raw 59.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/33.3% · raw 33.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/34.2% · raw 34.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 59.3 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| ARC-AGI-1 | 98.5%unknown effort | All 1 recorded result & sources98.5% · raw 98.5 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 ARC-AGI-1 https://openai.com/index/gpt-6-astra/ |
| ARC-AGI-2 | 95%unknown effort | All 1 recorded result & sources95% · raw 95 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 ARC-AGI-2 https://openai.com/index/gpt-6-astra/ |
| ARC-AGI-3 | 99.9%unknown effort | All 1 recorded result & sources99.9% · raw 99.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/ |
| AutomationBench v1.0.6 | 41.4%unknown effort | All 2 recorded results & sources41.4% · raw 41.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/41.4% · raw 41.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| DeepSWE v1.1 | 74.1%unknown effort | All 2 recorded results & sources74.1% · raw 74.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/74.1% · raw 74.1 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| ExploitBench | 100%unknown effort | All 1 recorded result & sources100% · raw 100 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/ |
| ExploitBench (Jun–Aug 2026) | 39%unknown effort | All 1 recorded result & sources39% · raw 39 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 Internal contamination-aware ExploitBench Jun–Aug 2026 https://openai.com/index/gpt-6-astra/ |
| ExploitGym | 42.4%unknown effort | All 1 recorded result & sources42.4% · raw 42.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/ |
| FrontierCode 1.1 Extended | 64.5%unknown effort | All 1 recorded result & sources64.5% · raw 64.5 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 FrontierCode 1.1 Extended; footnote 8 on OpenAI page https://openai.com/index/gpt-6-astra/ |
| FrontierCode 1.1 Main | 53.3%unknown effort | All 2 recorded results & sources53.3% · raw 53.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 FrontierCode 1.1 Main; footnote 8 https://openai.com/index/gpt-6-astra/53.3% · raw 53.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value. https://www.anthropic.com/claude-opus-5-5 |
| FrontierMath Tier 4 | 97.6%unknown effort | All 1 recorded result & sources97.6% · raw 97.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 FrontierMath Tier 4 (v2); prose also cites 98% saturation https://openai.com/index/gpt-6-astra/ |
| GeneBench Pro | 37.1%unknown effort | All 1 recorded result & sources37.1% · raw 37.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 GeneBench Pro https://openai.com/index/gpt-6-astra/ |
| GPQA Diamond | 96%unknown effort | All 1 recorded result & sources96% · raw 96 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/ |
| HealthBench Professional | 64.7%unknown effort | All 2 recorded results & sources64.7% · raw 64.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-22 length-adjusted (raw 68.2, mean length 3185 chars); Sep 22 system-card correction supersedes earlier launch-post 63.4 https://deploymentsafety.openai.com/gpt-6-astra70.3% · raw 70.3 % Alternative · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); length-adjusted with Anthropic grader. Kept non-headline under Astra own-card 64.7 (different harness) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Humanity's Last Exam (w/ tools) | 57.2%unknown effort | All 2 recorded results & sources57.2% · raw 57.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/57.2% · raw 57.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value. https://www.anthropic.com/claude-opus-5-5 |
| OSWorld 2.0 | 72.6%unknown effort | All 2 recorded results & sources72.6% · raw 72.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 v2026.08.08 offline set, partial score; ~40 min/task https://openai.com/index/gpt-6-astra/72.6% · raw 72.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| ScreenSpot-Pro | 92.7%unknown effort | All 1 recorded result & sources92.7% · raw 92.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 ScreenSpot-Pro no tools https://openai.com/index/gpt-6-astra/ |
| SEC-Bench Pro | 85.4%unknown effort | All 1 recorded result & sources85.4% · raw 85.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/ |
| SRE-Bench | 88%unknown effort | All 1 recorded result & sources88% · raw 88 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 SRE-Bench reverse engineering https://openai.com/index/gpt-6-astra/ |
| Terminal-Bench 4.0 | 57.9%unknown effort | All 3 recorded results & sources57.9% · raw 57.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/57.9% · raw 57.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce (OpenAI figure) | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value. https://www.anthropic.com/claude-opus-5-558.2% · raw 58.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 57.9 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Terminal-Bench Science 0.1 | 64.6%unknown effort | All 3 recorded results & sources64.6% · raw 64.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/64.6% · raw 64.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce (OpenAI figure) | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value. https://www.anthropic.com/claude-opus-5-568.1% · raw 68.1 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 64.6 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| HealthBench Hard | 36.6%unknown effort | All 1 recorded result & sources36.6% · raw 36.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-22 length-adjusted (raw 34.2, mean length 1697 chars); Astra system card HealthBench table (Sep 22 corrected values) https://deploymentsafety.openai.com/gpt-6-astra |
| GDPval-AA 2.1 | 1542unknown effort | All 1 recorded result & sources1542 · raw 1542 Elo Headline · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce https://www.anthropic.com/claude-opus-5-5 |
| AutomationBench | 41.4%unknown effort | All 1 recorded result & sources41.4% · raw 41.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce (Zapier public) https://www.anthropic.com/claude-opus-5-5 |
| AA Intelligence Index (vendor-cited) | 61unknown effort | All 1 recorded result & sources61 · raw 61 index Headline · unknown effort · Peer vendor Source/record date: Not recorded from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); GPT-6 Astra (max) https://huggingface.co/inclusionAI/Ling-3.0-flash-VL |
| FrontierSWE | 65.5%unknown effort | All 1 recorded result & sources65.5% · raw 65.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); Proximal harness https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Chartography | 71%unknown effort | All 2 recorded results & sources71% · raw 71 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); without tools, per Surge AI public reporting https://www.anthropic.com/claude-sonnet-5-5-system-card71% · raw 71 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| HealthBench | 60%max effort | All 1 recorded result & sources60% · raw 60 % Headline · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); length-adjusted Sep-2026 figure with Anthropic grader and harness, differs from Astra own-card setup https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Vals Index | 63.1%unknown effort | All 1 recorded result & sources63.1% · raw 63.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Vals Finance Agent v2 | 53.5%unknown effort | All 1 recorded result & sources53.5% · raw 53.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Harvey Legal Agent Benchmark | 5.4%unknown effort | All 1 recorded result & sources5.4% · raw 5.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| FrontierSWE v2 | 65.5%unknown effort | All 1 recorded result & sources65.5% · raw 65.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Vibe Code Bench | 89.6%unknown effort | All 1 recorded result & sources89.6% · raw 89.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| PostTrainBench | 44.3%unknown effort | All 1 recorded result & sources44.3% · raw 44.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| LABBench 2 | 85.4%unknown effort | All 1 recorded result & sources85.4% · raw 85.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| RiemannBench | 72%unknown effort | All 1 recorded result & sources72% · raw 72 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GraphWalks up to 128k BFS F1 | 98.7 f1unknown effort | All 1 recorded result & sources98.7 f1 · raw 98.7 f1 Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GraphWalks 256k-1M BFS F1 | 71.8 f1unknown effort | All 1 recorded result & sources71.8 f1 · raw 71.8 f1 Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| LVBench | 87.5%unknown effort | All 1 recorded result & sources87.5% · raw 87.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| CWE-bench v1 | 68%unknown effort | All 1 recorded result & sources68% · raw 68 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Gray Swan IPI | 8.5%unknown effort | All 1 recorded result & sources8.5% · raw 8.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 53max effort | All 5 recorded results & sources53 · raw 53 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra52 · raw 52 index Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-xhigh51 · raw 51 index Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-high50 · raw 50 index Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-medium46 · raw 46 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-low |
| Cost per Intelligence Index task · Artificial Analysis | $3.26max effort | All 5 recorded results & sources$3.26 · raw 3.26 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra$2.31 · raw 2.31 USD Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-xhigh$1.73 · raw 1.73 USD Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-high$1.54 · raw 1.54 USD Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-medium$0.82 · raw 0.82 USD Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gpt-6-astra-low |
| Output speed · Artificial Analysis | 54 tok/smax effort | All 5 recorded results & sources54 tok/s · raw 54 tok/s Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/gpt-6-astra49 tok/s · raw 49 tok/s Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/gpt-6-astra-xhigh45 tok/s · raw 45 tok/s Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/gpt-6-astra-high45 tok/s · raw 45 tok/s Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/gpt-6-astra-medium46 tok/s · raw 46 tok/s Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/gpt-6-astra-low |
| Coding Agent Index · Artificial Analysis | 62unknown effort | All 1 recorded result & sources62 · raw 62 index Headline · unknown effort · Independent evaluator Source/record date: 2026-09-22 Codex harness; ties Fable 5.1 https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra |
| Vals Index · Vals AI | 63.1%unknown effort | All 1 recorded result & sources63.1% · raw 63.13 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $18.46. https://www.vals.ai/benchmarks/vals_index |
| Bugs fixed /105 · Bug Hunt Bench | 45 fixesmax effort | All 5 recorded results & sources45 fixes · raw 45 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench43 fixes · raw 43 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench35 fixes · raw 35 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort high; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench34 fixes · raw 34 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench27 fixes · raw 27 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort low; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Vibe Code Bench v1.1 · Vals AI | 89.6%unknown effort | All 1 recorded result & sources89.6% · raw 89.59 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $38.51. https://www.vals.ai/benchmarks/vibe-code |
| ARC-AGI-1 · ARC Prize | 98.5%xhigh effort | All 6 recorded results & sources97.5% · raw 97.5 % Alternative · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra98.5% · raw 98.5 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/openai-gpt-6-astra98.5% · raw 98.5 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra97.5% · raw 97.5 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra96.5% · raw 96.5 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra86% · raw 86 % Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra |
| ARC-AGI-2 · ARC Prize | 95%max effort | All 6 recorded results & sources95% · raw 95 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/openai-gpt-6-astra93.3% · raw 93.3 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra92.1% · raw 92.1 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra92.1% · raw 92.1 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra85.4% · raw 85.4 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra59.6% · raw 59.6 % Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/openai-gpt-6-astra |
| ARC-AGI-3 (Standard) · ARC Prize | 62.7%max effort | All 6 recorded results & sources62.7% · raw 62.71 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Standard (notes carry). Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/openai-gpt-6-astra59.3% · raw 59.34 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Standard (notes carry). https://arcprize.org/results/openai-gpt-6-astra54.8% · raw 54.82 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Standard (notes carry). https://arcprize.org/results/openai-gpt-6-astra38.6% · raw 38.59 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Standard (notes carry). https://arcprize.org/results/openai-gpt-6-astra17.5% · raw 17.45 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Standard (notes carry). https://arcprize.org/results/openai-gpt-6-astra35.2% · raw 35.18 % Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Standard (notes carry). https://arcprize.org/results/openai-gpt-6-astra |
| ARC-AGI-3 · ARC Prize | 100%high effort | All 6 recorded results & sources98.6% · raw 98.55 % Alternative · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Provider Adapter. https://arcprize.org/results/openai-gpt-6-astra98.4% · raw 98.44 % Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Provider Adapter. https://arcprize.org/results/openai-gpt-6-astra100% · raw 99.95 % Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Provider Adapter. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/openai-gpt-6-astra98.4% · raw 98.44 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Provider Adapter. https://arcprize.org/results/openai-gpt-6-astra98% · raw 98.03 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Provider Adapter. https://arcprize.org/results/openai-gpt-6-astra96.7% · raw 96.72 % Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Harness: Provider Adapter. https://arcprize.org/results/openai-gpt-6-astra |
| AA-Omniscience Index · Artificial Analysis | 44high effort | All 1 recorded result & sources44 · raw 44 index Headline · high effort · Independent evaluator Source/record date: 2026-09-23 AA-Omniscience Index; GPT-6 Astra (high) as cited on evaluation page https://artificialanalysis.ai/evaluations/omniscience |
| GDPval-AA Elo · Artificial Analysis | 1542max effort | All 1 recorded result & sources1542 · raw 1542 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-23 GDPval-AA v2.1 Elo; GPT-6 Astra (max) https://artificialanalysis.ai/evaluations/gdpval-aa |
| Money gain · Andon Labs | $15014.70unknown effort | All 1 recorded result & sources$15014.70 · raw 15014.7 $ Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked. https://andonlabs.com/evals/vending-bench-2 |
| Blueprint Bench · Andon Labs | 49.7%unknown effort | All 1 recorded result & sources49.7% · raw 49.7 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100. https://andonlabs.com/evals/blueprint-bench-2 |
| Average Score · WeirdML v3 | 42.2%xhigh effort | All 1 recorded result & sources42.2% · raw 42.22 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant GPT-6 Astra (xhigh); harness codex_cli 0.154.0; values from prepared data JSON; raw 0.422202; official 80/20 aggregate (area 500k-50M tokens + final best) https://htihle.github.io/weirdml.html |
| Final Best Score · WeirdML v3 | 54.1%xhigh effort | All 1 recorded result & sources54.1% · raw 54.05 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant GPT-6 Astra (xhigh); harness codex_cli 0.154.0; values from prepared data JSON; raw 0.540453; mean final best effective score https://htihle.github.io/weirdml.html |
| Cost / Run · WeirdML v3 | $27.85xhigh effort | All 1 recorded result & sources$27.85 · raw 27.85 USD Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant GPT-6 Astra (xhigh); harness codex_cli 0.154.0; values from prepared data JSON; mean API cost per run, same task weighting as scores https://htihle.github.io/weirdml.html |
Read how we select and source scores or the comparison guide.