HealthBench Professional Results Across LLMs
Vendor-reported HealthBench Professional results across showcased LLMs. OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).
What this table contains
OpenAI HealthBench Professional: evaluates model capability and safety on real clinician-oriented chats (copilot-style medical professional use), graded with physician rubrics. Distinct from HealthBench Hard (hardest consumer/health conversation subset).
11 models with results · 19 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| Claude Sonnet 5.5Anthropic | 69.2%max effort | 2026-09-28 | All 2 recorded results & sources69.2% · raw 69.2 % Headline · max effort · Own vendor Source/record date: 2026-09-28 HealthBench Pro length-adjusted; raw 77.1 (level with Opus 5.5 raw) https://www.anthropic.com/claude-sonnet-5-5-system-card77.1% · raw 77.1 % Alternative · max effort · Own vendor Source/record date: 2026-09-28 HealthBench Pro raw before length adjustment https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Ling 3.1 FlashinclusionAI | 65.4%unknown effort | 2026-09-30 | All 2 recorded results & sources65.4% · raw 65.35 % Headline · unknown effort · Own vendor Source/record date: 2026-09-30 AQ environment; evaluation result, not clinical certification | Effort undisclosed. https://x.com/AntLingAGI/status/210533523872569781965.4% · raw 65.35 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as first-party X post; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| GPT-6 AstraOpenAI | 64.7%unknown effort | 2026-09-22 | All 2 recorded results & sources64.7% · raw 64.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-22 length-adjusted (raw 68.2, mean length 3185 chars); Sep 22 system-card correction supersedes earlier launch-post 63.4 https://deploymentsafety.openai.com/gpt-6-astra70.3% · raw 70.3 % Alternative · max effort · Peer vendor Source/record date: 2026-09-28 As reported by Anthropic (Sonnet 5.5 card); length-adjusted with Anthropic grader. Kept non-headline under Astra own-card 64.7 (different harness) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| GPT-6.1 SolOpenAI | 64.2%unknown effort | 2026-09-29 | All 1 recorded result & sources64.2% · raw 64.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-29 Length-adjusted (raw 67.2, mean length 3038 chars); +3.4 vs Sol; within 0.5 of Astra 64.7 https://deploymentsafety.openai.com/gpt-6-1-sol |
| GPT-6 LunaOpenAI | 60.8%unknown effort | 2026-09-22 | All 1 recorded result & sources60.8% · raw 60.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-22 length-adjusted (raw 61.2, mean length 2119 chars); Astra card Sol/Luna appendix https://deploymentsafety.openai.com/gpt-6-astra |
| Claude Fable 5.1Anthropic | 58.1%unknown effort | 2026-09-03 | All 1 recorded result & sources58.1% · raw 58.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported by OpenAI https://openai.com/index/gpt-6-astra/ |
| GPT-5.6 TerraOpenAI | 57.7%unknown effort | 2026-07-09 | All 1 recorded result & sources57.7% · raw 57.7 % Headline · unknown effort · Own vendor Source/record date: 2026-07-09 https://openai.com/index/gpt-5-6 |
| Grok 4.7xAI | 56.7%unknown effort | 2026-09-21 | All 1 recorded result & sources56.7% · raw 56.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://x.ai/news/grok-4-7 |
| Gemini 3.8 FlashGoogle | 52.1%unknown effort | 2026-09-03 | All 1 recorded result & sources52.1% · raw 52.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported by OpenAI https://openai.com/index/gpt-6-astra/ |
| DeepSeek V4.1 FlashDeepSeek | 50.4%unknown effort | 2026-09-30 | All 1 recorded result & sources50.4% · raw 50.37 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| GLM-5.3-FlashZ.ai | 49.1%unknown effort | 2026-09-30 | All 1 recorded result & sources49.1% · raw 49.07 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Claude Opus 5.5Anthropic | — | — | No recorded result |
| Gemini 4 ArgonGoogle | — | — | No recorded result |
| GLM-5.3Z.ai | — | — | No recorded result |
| Hy4 previewTencent | — | — | No recorded result |
| Kimi K3Moonshot AI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | — | — | No recorded result |
| MiMo-V2.6-FlashXiaomi | — | — | No recorded result |
| MiMo-V2.6-ProXiaomi | — | — | No recorded result |
| Muse Spark 1.3Meta | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | — | — | No recorded result |
| Nex-N2.5-ProNex AGI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
| Qwen3.8 27BAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Max (0902)Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Benchmark source references
- HealthBench Professional PDF · first_party
- OpenAI HealthBench · first_party
- OpenAI GPT-6 Astra system card · vendor_system_card
Use the fair-comparison guide before interpreting results from different configurations.