Examenos

GPT-6.1 Sol Benchmarks, Specifications & Availability

Explore GPT-6.1 Sol from OpenAI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare GPT-6.1 Sol with other models →Explore data coverage

Published specifications

Provider
OpenAI
Access
Proprietary
License
Proprietary
Context window
1.05M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-29
Modalities
file, image, text
Family
GPT-6

Model card · Announcement · Website · OpenRouter

Model notes

Mid GPT-6 tier refresh replacing GPT-6 Sol; near-Astra coding/computer-use/professional scores at ~1/5 Astra token prices. Knowledge cutoff Apr 30, 2026. Closed: HF API search for official vendor weights empty 2026-09-30; AA lists it proprietary.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

Announcement rate card $2/$10, cached input $0.10 (95% off input; half of GPT-6 Sol $0.20). Prompts >272K input billed 2x input/cache and 1.5x output for full request. No -contribute sibling on OpenRouter.; min-healthy endpoint minima 2026-10-01; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-01; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-02; time-window overrides apply on a winning provider; min-healthy endpoint minima 2026-10-03; time-window overrides apply on a winning provider

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
DeepSWE v1.175.2%high effort
All 2 recorded results & sources

75.2% · raw 75.2 %

Headline · high effort · Own vendor

Source/record date: 2026-09-29

Announcement chart: matches Astra at ~1/5 cost; +6.4 over GPT-6 Sol best at lower effort and cost

https://openai.com/index/introducing-gpt-6-1-sol/

71.9% · raw 71.9 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-29

Max-effort point on same chart; falls back from high-effort 75.2

https://openai.com/index/introducing-gpt-6-1-sol/
OSWorld 2.071.4%max effort
All 1 recorded result & sources

71.4% · raw 71.4 %

Headline · max effort · Own vendor

Source/record date: 2026-09-29

Offline set partial reward v2026.08.08; +7.0 over GPT-6 Sol at less than half cost; -2.1 vs Astra at ~1/7 cost/task

https://openai.com/index/introducing-gpt-6-1-sol/
GDP.PDF32%high effort
All 1 recorded result & sources

32% · raw 32 %

Headline · high effort · Own vendor

Source/record date: 2026-09-29

Higher reasoning setting; +3.2 over Opus 5.5 with fallbacks; -0.2 vs Astra 32.2

https://openai.com/index/introducing-gpt-6-1-sol/
AutomationBench v1.0.636.1%max effort
All 2 recorded results & sources

36.1% · raw 36.1 %

Headline · max effort · Own vendor

Source/record date: 2026-09-29

Max effort; +2.9 over GPT-6 Sol; -5.3 vs Astra 41.4; -6.4 vs Opus 5.5 with fallbacks 42.5

https://openai.com/index/introducing-gpt-6-1-sol/

31.7% · raw 31.7 %

Alternative · medium effort · Own vendor

Source/record date: 2026-09-29

Medium effort: +2.2 over Opus 5.5 (29.5), +4.8 over GPT-6 Sol same setting

https://openai.com/index/introducing-gpt-6-1-sol/
Terminal-Bench Science 0.157%max effort
All 1 recorded result & sources

57% · raw 57 %

Headline · max effort · Own vendor

Source/record date: 2026-09-29

Max effort $5.47/task vs $23.80 Astra and $23.21 Opus 5.5; Astra best 68.1

https://openai.com/index/introducing-gpt-6-1-sol/
ExploitBench99.7%max effort
All 1 recorded result & sources

99.7% · raw 99.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-29

Max effort; vs 81.7 GPT-6 Sol, 100 Astra; vendor cautions possible contamination from historical vulns

https://deploymentsafety.openai.com/gpt-6-1-sol
ExploitBench (Jun–Aug 2026)21.5%unknown effort
All 1 recorded result & sources

21.5% · raw 21.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Internal Port code execution on Jun-Aug 2026 vulns (contamination-aware); vs 5.5 Sol, 31.5 Astra, 3.5 GPT-5.6 Sol

https://deploymentsafety.openai.com/gpt-6-1-sol
SEC-Bench Pro78.8%unknown effort
All 1 recorded result & sources

78.8% · raw 78.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

pass@1; vs 66.3 Sol, 85.4 Astra, 79.1 GPT-5.6 Sol

https://deploymentsafety.openai.com/gpt-6-1-sol
ExploitGym35.1%unknown effort
All 1 recorded result & sources

35.1% · raw 35.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Intended-vulnerability rate per attempt; vs 22.1 Sol, 42.4 Astra, 30.3 GPT-5.6 Sol

https://deploymentsafety.openai.com/gpt-6-1-sol
HealthBench Professional64.2%unknown effort
All 1 recorded result & sources

64.2% · raw 64.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Length-adjusted (raw 67.2, mean length 3038 chars); +3.4 vs Sol; within 0.5 of Astra 64.7

https://deploymentsafety.openai.com/gpt-6-1-sol
HealthBench58.5%unknown effort
All 1 recorded result & sources

58.5% · raw 58.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Length-adjusted (raw 56.7, mean length 1701); +5.3 vs Sol

https://deploymentsafety.openai.com/gpt-6-1-sol
HealthBench Hard36.2%unknown effort
All 1 recorded result & sources

36.2% · raw 36.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Length-adjusted (raw 33.4, mean length 1646); +6.1 vs Sol

https://deploymentsafety.openai.com/gpt-6-1-sol
HealthBench Consensus96%unknown effort
All 1 recorded result & sources

96% · raw 96 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Length-adjusted (raw 95.9, mean length 1686); -0.2 vs Sol

https://deploymentsafety.openai.com/gpt-6-1-sol
MentalHealthBench57.9%max effort
All 1 recorded result & sources

57.9% · raw 57.9 %

Headline · max effort · Own vendor

Source/record date: 2026-09-29

Overall mean +-1.0 SE over 1215 tasks at max effort; vs 54.2 Sol, 58.7 Astra

https://deploymentsafety.openai.com/gpt-6-1-sol

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis52max effort
All 5 recorded results & sources

52 · raw 52 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol

51 · raw 51 index

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-xhigh

50 · raw 50 index

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-high

48 · raw 48 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-medium

42 · raw 42 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-low
Cost per Intelligence Index task · Artificial Analysis$0.72max effort
All 5 recorded results & sources

$0.72 · raw 0.72 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol

$0.39 · raw 0.39 USD

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-xhigh

$0.32 · raw 0.32 USD

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-high

$0.21 · raw 0.21 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-medium

$0.13 · raw 0.13 USD

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gpt-6-1-sol-low
Output speed · Artificial Analysis63 tok/smax effort
All 5 recorded results & sources

63 tok/s · raw 63 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-1-sol

53 tok/s · raw 53 tok/s

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-1-sol-xhigh

58 tok/s · raw 58 tok/s

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-1-sol-high

53 tok/s · raw 53 tok/s

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-1-sol-medium

60 tok/s · raw 60 tok/s

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gpt-6-1-sol-low
Average Score · WeirdML v335.5%xhigh effort
All 1 recorded result & sources

35.5% · raw 35.47 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant GPT-6.1 Sol (xhigh); harness codex_cli 0.156.0; values from prepared data JSON; raw 0.354727; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v347.5%xhigh effort
All 1 recorded result & sources

47.5% · raw 47.51 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant GPT-6.1 Sol (xhigh); harness codex_cli 0.156.0; values from prepared data JSON; raw 0.475139; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$4.99xhigh effort
All 1 recorded result & sources

$4.99 · raw 4.99 USD

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant GPT-6.1 Sol (xhigh); harness codex_cli 0.156.0; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html
Vals Index · Vals AI61.2%unknown effort
All 1 recorded result & sources

61.2% · raw 61.15 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $3.24.

https://www.vals.ai/benchmarks/vals_index
Vibe Code Bench v1.1 · Vals AI88.9%unknown effort
All 1 recorded result & sources

88.9% · raw 88.93 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $6.23.

https://www.vals.ai/benchmarks/vibe-code
Bugs fixed /105 · Bug Hunt Bench44.3 fixesmax effort
All 5 recorded results & sources

44.3 fixes · raw 44.3 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

42.7 fixes · raw 42.7 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort xhigh; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

36.5 fixes · raw 36.5 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort high; 2 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

29 fixes · raw 29 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort medium; 2 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

22.5 fixes · raw 22.5 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort low; 2 runs; evaluation 2026-10-01. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
ARC-AGI-1 · ARC Prize98.5%xhigh effort
All 5 recorded results & sources

96.5% · raw 96.5 %

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol

98.5% · raw 98.5 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-1-sol

98.5% · raw 98.5 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol

95.5% · raw 95.5 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol

93.5% · raw 93.5 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol
ARC-AGI-2 · ARC Prize94.2%max effort
All 5 recorded results & sources

94.2% · raw 94.2 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-1-sol

91.7% · raw 91.7 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol

91.7% · raw 91.7 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol

86.7% · raw 86.7 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol

76.7% · raw 76.7 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/openai-gpt-6-1-sol
ARC-AGI-3 (Standard) · ARC Prize52.7%max effort
All 5 recorded results & sources

52.7% · raw 52.73 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry). Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-1-sol

39.9% · raw 39.93 %

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-1-sol

26.7% · raw 26.72 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-1-sol

10.6% · raw 10.57 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-1-sol

3.9% · raw 3.92 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/openai-gpt-6-1-sol
ARC-AGI-3 · ARC Prize96.4%xhigh effort
All 5 recorded results & sources

96.2% · raw 96.18 %

Alternative · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-1-sol

96.4% · raw 96.37 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/openai-gpt-6-1-sol

95% · raw 94.99 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-1-sol

91% · raw 91.02 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-1-sol

82.8% · raw 82.79 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/openai-gpt-6-1-sol

Read how we select and source scores or the comparison guide.