Examenos

Terminal-Bench 4.0 Results Across LLMs

Vendor-reported Terminal-Bench 4.0 results across showcased LLMs. Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.

Official / vendor Higher is better · unit: %

What this table contains

Multi-hour agentic terminal/coding tasks: shell, tools, and long-running project work under a hard time budget.

15 models with results · 15 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

Terminal-Bench 4.0: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
Claude Sonnet 5.5Anthropic70.6%max effort2026-09-28
All 1 recorded result & sources

70.6% · raw 70.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

TB 4.0 at max effort; safeguards on (1.2% of requests fallback-served); 5 trials per task

https://www.anthropic.com/claude-sonnet-5-5
Claude Opus 5.5Anthropic66.4%xhigh effort2026-09-22
All 2 recorded results & sources

66.4% · raw 66.4 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-09-22

Terminal-Bench 4.0 at xhigh (table footnote); adaptive thinking

https://www.anthropic.com/claude-opus-5-5

66.4% · raw 66.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GPT-6 AstraOpenAI57.9%unknown effort2026-09-03
All 3 recorded results & sources

57.9% · raw 57.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

57.9% · raw 57.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce (OpenAI figure) | Demoted 2026-09-25: duplicate of headline from OpenAI GPT-6 Astra announce (openai.com/index/gpt-6-astra); vendor official preferred over Anthropic peer comparison table (rule a); same value.

https://www.anthropic.com/claude-opus-5-5

58.2% · raw 58.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 57.9 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Gemini 4 ArgonGoogle57.4%max effort2026-09-30
All 1 recorded result & sources

57.4% · raw 57.4 %

Headline · max effort · Own vendor

Source/record date: 2026-09-30

Self-computed; peers from official public leaderboard (highest thinking level) | Highest thinking settings per Google eval methodology.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Claude Fable 5.1Anthropic55.8%unknown effort2026-09-22
All 3 recorded results & sources

55.8% · raw 55.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI GPT-6 Astra announcement table | Demoted 2026-09-25: duplicate of headline from Anthropic Claude Opus 5.5 announce (anthropic.com/claude-opus-5-5); model vendor's own (Anthropic) figure preferred over peer-vendor OpenAI comparison table (rule a); Anthropic source is also later-dated (2026-09-22); same value. source_type relabeled official→vendor-comparison (OpenAI reporting an Anthropic model).

https://openai.com/index/gpt-6-astra/

55.8% · raw 55.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-22

As reported by Anthropic Opus 5.5 announce

https://www.anthropic.com/claude-opus-5-5

57.9% · raw 57.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 55.8 (spillover); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Pareto 26.10 PreviewUnbiased50.8%unknown effort2026-10-01
All 1 recorded result & sources

50.8% · raw 50.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-01

Unbiased preliminary vendor result; mean task cost $0.48. Unbiased says results may change and denominator and cost methods need confirmation.

https://unbiased.ai/blog/pareto-26-10-preview/
Ling 3.1 FlashinclusionAI40.4%unknown effort2026-09-30
All 1 recorded result & sources

40.4% · raw 40.4 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Grok 4.7xAI38%unknown effort2026-09-21
All 1 recorded result & sources

38% · raw 38 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://x.ai/news/grok-4-7
GLM-5.3Z.ai37.9%unknown effort2026-09-10
All 1 recorded result & sources

37.9% · raw 37.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-10

As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column)

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
MiMo-V2.6-ProXiaomi34.9%unknown effort2026-09-21
All 1 recorded result & sources

34.9% · raw 34.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
GLM-5.3-FlashZ.ai32.8%unknown effort2026-09-30
All 1 recorded result & sources

32.8% · raw 32.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
DeepSeek V4.1 FlashDeepSeek31.2%unknown effort2026-09-10
All 1 recorded result & sources

31.2% · raw 31.2 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

Secondary guide citing DeepSeek official agent snapshot

https://deepseekagent.io/deepseek-v4-1-flash
MiMo-V2.6-FlashXiaomi28.8%unknown effort2026-09-21
All 1 recorded result & sources

28.8% · raw 28.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
GPT-5.6 TerraOpenAI21.5%max effort2026-09-23
All 1 recorded result & sources

21.5% · raw 21.5 %

Headline · max effort · Own vendor

Source/record date: 2026-09-23

Official Terminal-Bench 4.0 board: GPT-5.6 Terra (max) + Codex; resolution rate 21.5% ±3.3

https://www.tbench.ai/leaderboard/terminal-bench/4.0
Gemini 3.8 FlashGoogle19.1%unknown effort2026-09-03
All 1 recorded result & sources

19.1% · raw 19.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://deepmind.google/models/model-cards/gemini-3-8-flash
GPT-6 LunaOpenAI——No recorded result
GPT-6.1 SolOpenAI——No recorded result
Hy4 previewTencent——No recorded result
Kimi K3Moonshot AI——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
MiMo-V2.6-Distill-Qwen-9BXiaomi——No recorded result
Muse Spark 1.3Meta——No recorded result
Naive-N0.5-FlashNaiveAI——No recorded result
Nex-N2.5-ProNex AGI——No recorded result
Qwen3.8 27BAlibaba (Qwen)——No recorded result
Qwen3.8 FlashAlibaba (Qwen)——No recorded result
Qwen3.8 Max (0902)Alibaba (Qwen)——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.