Terminal-Bench 4.0 Results Across LLMs
Vendor-reported Terminal-Bench 4.0 results across showcased LLMs. Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.
What this table contains
Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.
15 models with results · 15 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| Claude Sonnet 5.5Anthropic | 70.6%max effort | 2026-09-28 | All 1 recorded result & sources70.6% · raw 70.6 % Headline · max effort · Own vendor Source/record date: 2026-09-28 TB 4.0 at max effort; safeguards on (1.2% of requests fallback-served); 5 trials per task https://www.anthropic.com/claude-sonnet-5-5 |
| Claude Opus 5.5Anthropic | 66.4%xhigh effort | 2026-09-22 | All 2 recorded results & sources66.4% · raw 66.4 % Headline · xhigh effort · Own vendor Source/record date: 2026-09-22 Terminal-Bench 4.0 at xhigh (table footnote); adaptive thinking https://www.anthropic.com/claude-opus-5-566.4% · raw 66.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GPT-6 AstraOpenAI | 57.9%unknown effort | 2026-09-03 | All 3 recorded results & sources57.9% · raw 57.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/57.9% · raw 57.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce (OpenAI figure) | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value. https://www.anthropic.com/claude-opus-5-558.2% · raw 58.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 57.9 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Gemini 4 ArgonGoogle | 57.4%max effort | 2026-09-30 | All 1 recorded result & sources57.4% · raw 57.4 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed; peers from official public leaderboard (highest thinking level) | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Claude Fable 5.1Anthropic | 55.8%unknown effort | 2026-09-22 | All 3 recorded results & sources55.8% · raw 55.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported by OpenAI GPT-6 Astra announcement table | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Opus 5.5 announce (anthropic.com/claude-opus-5-5); model vendor's own (Anthropic) figure preferred over peer-vendor OpenAI comparison table (rule a); Anthropic source is also later-dated (2026-09-22); same value. source_type relabeled official→vendor-comparison (OpenAI reporting an Anthropic model). https://openai.com/index/gpt-6-astra/55.8% · raw 55.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-22 As reported by Anthropic Opus 5.5 announce https://www.anthropic.com/claude-opus-5-557.9% · raw 57.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 55.8 (spillover); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Pareto 26.10 PreviewUnbiased | 50.8%unknown effort | 2026-10-01 | All 1 recorded result & sources50.8% · raw 50.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-01 Unbiased preliminary vendor result; mean task cost $0.48. Unbiased says results may change and denominator and cost methods need confirmation. https://unbiased.ai/blog/pareto-26-10-preview/ |
| Ling 3.1 FlashinclusionAI | 40.4%unknown effort | 2026-09-30 | All 1 recorded result & sources40.4% · raw 40.4 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Grok 4.7xAI | 38%unknown effort | 2026-09-21 | All 1 recorded result & sources38% · raw 38 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://x.ai/news/grok-4-7 |
| GLM-5.3Z.ai | 37.9%unknown effort | 2026-09-10 | All 1 recorded result & sources37.9% · raw 37.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| MiMo-V2.6-ProXiaomi | 34.9%unknown effort | 2026-09-21 | All 1 recorded result & sources34.9% · raw 34.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL |
| GLM-5.3-FlashZ.ai | 32.8%unknown effort | 2026-09-30 | All 1 recorded result & sources32.8% · raw 32.8 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| DeepSeek V4.1 FlashDeepSeek | 31.2%unknown effort | 2026-09-10 | All 1 recorded result & sources31.2% · raw 31.2 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 Secondary guide citing DeepSeek official agent snapshot https://deepseekagent.io/deepseek-v4-1-flash |
| MiMo-V2.6-FlashXiaomi | 28.8%unknown effort | 2026-09-21 | All 1 recorded result & sources28.8% · raw 28.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL |
| GPT-5.6 TerraOpenAI | 21.5%max effort | 2026-09-23 | All 1 recorded result & sources21.5% · raw 21.5 % Headline · max effort · Own vendor Source/record date: 2026-09-23 Official Terminal-Bench 4.0 board: GPT-5.6 Terra (max) + Codex; resolution rate 21.5% ±3.3 https://www.tbench.ai/leaderboard/terminal-bench/4.0 |
| Gemini 3.8 FlashGoogle | 19.1%unknown effort | 2026-09-03 | All 1 recorded result & sources19.1% · raw 19.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://deepmind.google/models/model-cards/gemini-3-8-flash |
| GPT-6 LunaOpenAI | — | — | No recorded result |
| GPT-6.1 SolOpenAI | — | — | No recorded result |
| Hy4 previewTencent | — | — | No recorded result |
| Kimi K3Moonshot AI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | — | — | No recorded result |
| Muse Spark 1.3Meta | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | — | — | No recorded result |
| Nex-N2.5-ProNex AGI | — | — | No recorded result |
| Qwen3.8 27BAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Max (0902)Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Benchmark source references
- Terminal-Bench · first_party
- Terminal-Bench leaderboard · leaderboard
- OpenAI GPT-6 Astra system card · vendor_system_card
Use the fair-comparison guide before interpreting results from different configurations.