Terminal-Bench 2.1 Results Across LLMs
Vendor-reported Terminal-Bench 2.1 results across showcased LLMs. Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.
What this table contains
Earlier Terminal-Bench agent suite (shorter/earlier protocol than 3.0/4.0); keep separate from later versions.
17 models with results · 13 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| DeepSeek V4.1 FlashDeepSeek | 90.6%max effort | 2026-09-10 | All 3 recorded results & sources90.6% · raw 90.6 % Headline · max effort · Third-party Source/record date: 2026-09-10 Secondary guide citing DeepSeek official agent snapshot; max effort https://deepseekagent.io/deepseek-v4-1-flash90.6% · raw 90.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/90.6% · raw 90.6 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| MiMo-V2.6-ProXiaomi | 89.9%unknown effort | 2026-09-21 | All 1 recorded result & sources89.9% · raw 89.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL |
| Gemini 3.8 FlashGoogle | 89.4%unknown effort | 2026-09-03 | All 1 recorded result & sources89.4% · raw 89.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://deepmind.google/models/model-cards/gemini-3-8-flash |
| Muse Spark 1.3Meta | 88.8%unknown effort | 2026-09-02 | All 2 recorded results & sources88.8% · raw 88.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 Meta published table; ties GPT-5.6 Sol; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-388.8% · raw 88.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI; matches the Meta figure value. https://naive.ai/en/research/ |
| Kimi K3Moonshot AI | 88.3%max effort | 2026-07-16 | All 2 recorded results & sources88.3% · raw 88.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Terminal-Bench 2.1; Kimi Code harness https://github.com/MoonshotAI/Kimi-K388.3% · raw 88.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GLM-5.3Z.ai | 88.2%unknown effort | 2026-09-10 | All 3 recorded results & sources88.2% · raw 88.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash88.2% · raw 88.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from DeepSeek-V4.1-Flash HF comparison table; both rows are peer-vendor spillover with the same value (no Z.ai official TB2.1 row), kept the later-dated source (2026-09-10 vs 2026-09-08) (rule b). https://huggingface.co/nex-agi/Nex-N2.5-Pro88.2% · raw 88.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| MiMo-V2.6-FlashXiaomi | 87.6%unknown effort | 2026-09-21 | All 1 recorded result & sources87.6% · raw 87.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL |
| GPT-5.6 TerraOpenAI | 87.4%unknown effort | 2026-07-09 | All 1 recorded result & sources87.4% · raw 87.4 % Headline · unknown effort · Own vendor Source/record date: 2026-07-09 https://openai.com/index/gpt-5-6 |
| Naive-N0.5-FlashNaiveAI | 86.7%unknown effort | 2026-09-27 | All 1 recorded result & sources86.7% · raw 86.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| Qwen3.8 Max (0902)Alibaba (Qwen) | 86.6%max effort | 2026-09-02 | All 3 recorded results & sources86.6% · raw 86.6 % Headline · max effort · Third-party Source/record date: 2026-09-02 Terminal-Bench 2.1 — Qwen3.8-Max family table https://www.datacamp.com/blog/qwen3-8-max86.6% · raw 86.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value (Nex row effort unknown, vendor row max). https://huggingface.co/nex-agi/Nex-N2.5-Pro88.8% · raw 88.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted. https://naive.ai/en/research/ |
| Hy4 previewTencent | 85.4%unknown effort | 2026-08-28 | All 2 recorded results & sources85.4% · raw 85.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-28 Transcribed from Tencent vendor launch charts; official page is chart-heavy https://hy4.site/benchmarks/hy4-benchmarks85.4% · raw 85.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GLM-5.3-FlashZ.ai | 84.3%unknown effort | 2026-08-26 | All 3 recorded results & sources84.3% · raw 84.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-26 https://z.ai/blog/glm-5.3-flash84.3% · raw 84.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/84.3% · raw 84.3 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Nex-N2.5-ProNex AGI | 82.7%unknown effort | 2026-09-08 | All 1 recorded result & sources82.7% · raw 82.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-08 NexAU harness; temp=0.7 https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| Ling 3.1 FlashinclusionAI | 81.2%unknown effort | 2026-09-30 | All 1 recorded result & sources81.2% · raw 81.18 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Qwen3.8 27BAlibaba (Qwen) | 73%unknown effort | 2026-08-14 | All 2 recorded results & sources73% · raw 73 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Terminal Bench 2.1 (Terminus); HF card HTML table https://huggingface.co/Qwen/Qwen3.8-27B76.8% · raw 76.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | 37.1%unknown effort | 2026-09-21 | All 1 recorded result & sources37.1% · raw 37.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 Released SFT checkpoint, avg@1; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
| Kolibri-1Aleph Alpha | 27.7%high effort | 2026-10-03 | All 1 recorded result & sources27.7% · raw 27.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Claude Fable 5.1Anthropic | — | — | No recorded result |
| Claude Opus 5.5Anthropic | — | — | No recorded result |
| Claude Sonnet 5.5Anthropic | — | — | No recorded result |
| Gemini 4 ArgonGoogle | — | — | No recorded result |
| GPT-6 AstraOpenAI | — | — | No recorded result |
| GPT-6 LunaOpenAI | — | — | No recorded result |
| GPT-6.1 SolOpenAI | — | — | No recorded result |
| Grok 4.7xAI | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Benchmark source references
- Terminal-Bench · first_party
- Terminal-Bench leaderboard · leaderboard
- Vendor announce tables · vendor_release
- Kolibri technical report · vendor_system_card
Use the fair-comparison guide before interpreting results from different configurations.