Examenos

Terminal-Bench 2.1 Results Across LLMs

Vendor-reported Terminal-Bench 2.1 results across showcased LLMs. Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.

Official / vendor Higher is better · unit: %

What this table contains

Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.

17 models with results · 13 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

Terminal-Bench 2.1: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
DeepSeek V4.1 FlashDeepSeek90.6%max effort2026-09-10
All 3 recorded results & sources

90.6% · raw 90.6 %

Headline · max effort · Third-party

Source/record date: 2026-09-10

Secondary guide citing DeepSeek official agent snapshot; max effort

https://deepseekagent.io/deepseek-v4-1-flash

90.6% · raw 90.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

90.6% · raw 90.6 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
MiMo-V2.6-ProXiaomi89.9%unknown effort2026-09-21
All 1 recorded result & sources

89.9% · raw 89.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Gemini 3.8 FlashGoogle89.4%unknown effort2026-09-03
All 1 recorded result & sources

89.4% · raw 89.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://deepmind.google/models/model-cards/gemini-3-8-flash
Muse Spark 1.3Meta88.8%unknown effort2026-09-02
All 2 recorded results & sources

88.8% · raw 88.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

Meta published table; ties GPT-5.6 Sol; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3

88.8% · raw 88.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI; matches the Meta figure value.

https://naive.ai/en/research/
Kimi K3Moonshot AI88.3%max effort2026-07-16
All 2 recorded results & sources

88.3% · raw 88.3 %

Headline · max effort · Own vendor

Source/record date: 2026-07-16

Terminal-Bench 2.1; Kimi Code harness

https://github.com/MoonshotAI/Kimi-K3

88.3% · raw 88.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GLM-5.3Z.ai88.2%unknown effort2026-09-10
All 3 recorded results & sources

88.2% · raw 88.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-10

As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column)

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

88.2% · raw 88.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from DeepSeek-V4.1-Flash HF comparison table; both rows are peer-vendor spillover with the same value (no Z.ai official TB2.1 row), kept the later-dated source (2026-09-10 vs 2026-09-08) (rule b).

https://huggingface.co/nex-agi/Nex-N2.5-Pro

88.2% · raw 88.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
MiMo-V2.6-FlashXiaomi87.6%unknown effort2026-09-21
All 1 recorded result & sources

87.6% · raw 87.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
GPT-5.6 TerraOpenAI87.4%unknown effort2026-07-09
All 1 recorded result & sources

87.4% · raw 87.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-07-09

https://openai.com/index/gpt-5-6
Naive-N0.5-FlashNaiveAI86.7%unknown effort2026-09-27
All 1 recorded result & sources

86.7% · raw 86.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-27

Vendor chart figure; harness Claude Code 2.1.207.

https://naive.ai/en/research/
Qwen3.8 Max (0902)Alibaba (Qwen)86.6%max effort2026-09-02
All 3 recorded results & sources

86.6% · raw 86.6 %

Headline · max effort · Third-party

Source/record date: 2026-09-02

Terminal-Bench 2.1 — Qwen3.8-Max family table

https://www.datacamp.com/blog/qwen3-8-max

86.6% · raw 86.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value (Nex row effort unknown, vendor row max).

https://huggingface.co/nex-agi/Nex-N2.5-Pro

88.8% · raw 88.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
Hy4 previewTencent85.4%unknown effort2026-08-28
All 2 recorded results & sources

85.4% · raw 85.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

Transcribed from Tencent vendor launch charts; official page is chart-heavy

https://hy4.site/benchmarks/hy4-benchmarks

85.4% · raw 85.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GLM-5.3-FlashZ.ai84.3%unknown effort2026-08-26
All 3 recorded results & sources

84.3% · raw 84.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash

84.3% · raw 84.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

84.3% · raw 84.3 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Nex-N2.5-ProNex AGI82.7%unknown effort2026-09-08
All 1 recorded result & sources

82.7% · raw 82.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-08

NexAU harness; temp=0.7

https://huggingface.co/nex-agi/Nex-N2.5-Pro
Ling 3.1 FlashinclusionAI81.2%unknown effort2026-09-30
All 1 recorded result & sources

81.2% · raw 81.18 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Qwen3.8 27BAlibaba (Qwen)73%unknown effort2026-08-14
All 2 recorded results & sources

73% · raw 73 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Terminal Bench 2.1 (Terminus); HF card HTML table

https://huggingface.co/Qwen/Qwen3.8-27B

76.8% · raw 76.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
MiMo-V2.6-Distill-Qwen-9BXiaomi37.1%unknown effort2026-09-21
All 1 recorded result & sources

37.1% · raw 37.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

Released SFT checkpoint, avg@1; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Kolibri-1Aleph Alpha27.7%high effort2026-10-03
All 1 recorded result & sources

27.7% · raw 27.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Claude Fable 5.1Anthropic——No recorded result
Claude Opus 5.5Anthropic——No recorded result
Claude Sonnet 5.5Anthropic——No recorded result
Gemini 4 ArgonGoogle——No recorded result
GPT-6 AstraOpenAI——No recorded result
GPT-6 LunaOpenAI——No recorded result
GPT-6.1 SolOpenAI——No recorded result
Grok 4.7xAI——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result
Qwen3.8 FlashAlibaba (Qwen)——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.