Examenos

GPQA Diamond Results Across LLMs

Vendor-reported GPQA Diamond results across showcased LLMs. Graduate-level science QA

Official / vendor Higher is better · unit: %

What this table contains

Graduate-level science QA

15 models with results · 15 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

GPQA Diamond: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
GPT-6 AstraOpenAI96%unknown effort2026-09-03
All 1 recorded result & sources

96% · raw 96 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/
Gemini 3.8 FlashGoogle95.3%unknown effort2026-09-03
All 1 recorded result & sources

95.3% · raw 95.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
Claude Fable 5.1Anthropic93.7%unknown effort2026-09-03
All 1 recorded result & sources

93.7% · raw 93.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
Kimi K3Moonshot AI93.5%max effort2026-07-16
All 1 recorded result & sources

93.5% · raw 93.5 %

Headline · max effort · Own vendor

Source/record date: 2026-07-16

GPQA Diamond; reasoning_effort=max, temp=1.0, top_p=0.95

https://github.com/MoonshotAI/Kimi-K3
GPT-5.6 TerraOpenAI92.9%unknown effort2026-07-09
All 1 recorded result & sources

92.9% · raw 92.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-07-09

https://openai.com/index/gpt-5-6
Qwen3.8 Max (0902)Alibaba (Qwen)92.6%unknown effort2026-09-02
All 1 recorded result & sources

92.6% · raw 92.6 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

GPQA Diamond

https://www.datacamp.com/blog/qwen3-8-max
Pareto 26.10 PreviewUnbiased92.4%unknown effort2026-10-01
All 1 recorded result & sources

92.4% · raw 92.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-01

Unbiased preliminary vendor result; mean task cost $0.004. Unbiased says results may change and denominator and cost methods need confirmation.

https://unbiased.ai/blog/pareto-26-10-preview/
Hy4 previewTencent92.3%unknown effort2026-08-28
All 1 recorded result & sources

92.3% · raw 92.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

Vendor chart transcription

https://hy4.site/benchmarks/hy4-benchmarks
Qwen3.8 FlashAlibaba (Qwen)91.7%unknown effort2026-08-26
All 1 recorded result & sources

91.7% · raw 91.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

Qwen3.8-Flash-Next blog table (GPQA Diamond)

https://qwen.ai/blog?id=qwen3.8-flash-next
Qwen3.8 Omni FlashAlibaba (Qwen)91%xhigh effort2026-09-18
All 1 recorded result & sources

91% · raw 91 %

Headline · xhigh effort · Third-party

Source/record date: 2026-09-18

GPQA Diamond — launch secondary transcription

https://www.cometapi.com/what-is-qwen3-8-omni-flash-and-how-to-access/
DeepSeek V4.1 FlashDeepSeek90.9%unknown effort2026-09-10
All 1 recorded result & sources

90.9% · raw 90.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

GPQA Diamond Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Qwen3.8 27BAlibaba (Qwen)89.2%unknown effort2026-08-14
All 2 recorded results & sources

89.2% · raw 89.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

GPQA Diamond

https://huggingface.co/Qwen/Qwen3.8-27B

89.2% · raw 89.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Chain-of-thought, avg@8. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
GLM-5.3Z.ai88.1%unknown effort2026-09-10
All 1 recorded result & sources

88.1% · raw 88.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-10

from chart/figure as reported on DeepSeek V4.1 Flash announce table (peer spillover)

https://api-docs.deepseek.com/news/news260910
Kolibri-1Aleph Alpha84.3%high effort2026-10-03
All 1 recorded result & sources

84.3% · raw 84.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Chain-of-thought, avg@8.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Mercury 2.5Inception79%unknown effort2026-09-08
All 1 recorded result & sources

79% · raw 79 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-08

from chart/figure on Inception Mercury 2.5 release post (alt: Mercury 2.5 vs Mercury 2.0)

https://www.inceptionlabs.ai/blog/introducing-mercury-2-5
Claude Opus 5.5Anthropic——No recorded result
Claude Sonnet 5.5Anthropic——No recorded result
Gemini 4 ArgonGoogle——No recorded result
GLM-5.3-FlashZ.ai——No recorded result
GPT-6 LunaOpenAI——No recorded result
GPT-6.1 SolOpenAI——No recorded result
Grok 4.7xAI——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Ling 3.1 FlashinclusionAI——No recorded result
MiMo-V2.6-Distill-Qwen-9BXiaomi——No recorded result
MiMo-V2.6-FlashXiaomi——No recorded result
MiMo-V2.6-ProXiaomi——No recorded result
Muse Spark 1.3Meta——No recorded result
Naive-N0.5-FlashNaiveAI——No recorded result
Nex-N2.5-ProNex AGI——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.