Examenos

HealthBench Professional Results Across LLMs

Vendor-reported HealthBench Professional results across showcased LLMs. OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).

Official / vendor Higher is better · unit: %

What this table contains

OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).

11 models with results · 19 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

HealthBench Professional: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
Claude Sonnet 5.5Anthropic69.2%max effort2026-09-28
All 2 recorded results & sources

69.2% · raw 69.2 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

HealthBench Pro length-adjusted; raw 77.1 (level with Opus 5.5 raw)

https://www.anthropic.com/claude-sonnet-5-5-system-card

77.1% · raw 77.1 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-28

HealthBench Pro raw before length adjustment

https://www.anthropic.com/claude-sonnet-5-5-system-card
Ling 3.1 FlashinclusionAI65.4%unknown effort2026-09-30
All 2 recorded results & sources

65.4% · raw 65.35 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-30

AQ environment; evaluation result, not clinical certification | Effort undisclosed.

https://x.com/AntLingAGI/status/2105335238725697819

65.4% · raw 65.35 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as first-party X post; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
GPT-6 AstraOpenAI64.7%unknown effort2026-09-22
All 2 recorded results & sources

64.7% · raw 64.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-22

length-adjusted (raw 68.2, mean length 3185 chars); Sep 22 system-card correction supersedes earlier launch-post 63.4

https://deploymentsafety.openai.com/gpt-6-astra

70.3% · raw 70.3 %

Alternative · max effort · Peer vendor

Source/record date: 2026-09-28

As reported by Anthropic (Sonnet 5.5 card); length-adjusted with Anthropic grader. Kept non-headline under Astra own-card 64.7 (different harness)

https://www.anthropic.com/claude-sonnet-5-5-system-card
GPT-6.1 SolOpenAI64.2%unknown effort2026-09-29
All 1 recorded result & sources

64.2% · raw 64.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-29

Length-adjusted (raw 67.2, mean length 3038 chars); +3.4 vs Sol; within 0.5 of Astra 64.7

https://deploymentsafety.openai.com/gpt-6-1-sol
GPT-6 LunaOpenAI60.8%unknown effort2026-09-22
All 1 recorded result & sources

60.8% · raw 60.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-22

length-adjusted (raw 61.2, mean length 2119 chars); Astra card Sol/Luna appendix

https://deploymentsafety.openai.com/gpt-6-astra
Claude Fable 5.1Anthropic58.1%unknown effort2026-09-03
All 1 recorded result & sources

58.1% · raw 58.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
GPT-5.6 TerraOpenAI57.7%unknown effort2026-07-09
All 1 recorded result & sources

57.7% · raw 57.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-07-09

https://openai.com/index/gpt-5-6
Grok 4.7xAI56.7%unknown effort2026-09-21
All 1 recorded result & sources

56.7% · raw 56.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://x.ai/news/grok-4-7
Gemini 3.8 FlashGoogle52.1%unknown effort2026-09-03
All 1 recorded result & sources

52.1% · raw 52.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
DeepSeek V4.1 FlashDeepSeek50.4%unknown effort2026-09-30
All 1 recorded result & sources

50.4% · raw 50.37 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
GLM-5.3-FlashZ.ai49.1%unknown effort2026-09-30
All 1 recorded result & sources

49.1% · raw 49.07 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Claude Opus 5.5Anthropic——No recorded result
Gemini 4 ArgonGoogle——No recorded result
GLM-5.3Z.ai——No recorded result
Hy4 previewTencent——No recorded result
Kimi K3Moonshot AI——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
MiMo-V2.6-Distill-Qwen-9BXiaomi——No recorded result
MiMo-V2.6-FlashXiaomi——No recorded result
MiMo-V2.6-ProXiaomi——No recorded result
Muse Spark 1.3Meta——No recorded result
Naive-N0.5-FlashNaiveAI——No recorded result
Nex-N2.5-ProNex AGI——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result
Qwen3.8 27BAlibaba (Qwen)——No recorded result
Qwen3.8 FlashAlibaba (Qwen)——No recorded result
Qwen3.8 Max (0902)Alibaba (Qwen)——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.