JobBench Results Across LLMs
Vendor-reported JobBench results across showcased LLMs. Job/task agent benchmark
What this table contains
Job/task agent benchmark
11 models with results · 19 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| Muse Spark 1.3Meta | 64.9%unknown effort | 2026-09-02 | All 1 recorded result & sources64.9% · raw 64.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 JobBench; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| Qwen3.8 Max (0902)Alibaba (Qwen) | 64%unknown effort | 2026-09-02 | All 2 recorded results & sources64% · raw 64 % Headline · unknown effort · Third-party Source/record date: 2026-09-02 Alibaba published comparison as reported by DataCamp https://www.datacamp.com/blog/qwen3-8-max53.4% · raw 53.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp (64.0); vendor official 0902 figure preferred (rules a+b). 53.4 equals the pre-0902 Qwen3.8-Max (2026-08-03) value in Alibaba's table, so Nex's "Qwen3.8-Max" column appears to be the original snapshot, not 0902. https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| MiMo-V2.6-ProXiaomi | 62%unknown effort | 2026-09-21 | All 1 recorded result & sources62% · raw 62 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL |
| Hy4 previewTencent | 61.7%unknown effort | 2026-08-28 | All 1 recorded result & sources61.7% · raw 61.7 % Headline · unknown effort · Own vendor Source/record date: 2026-08-28 Vendor chart transcription https://hy4.site/benchmarks/hy4-benchmarks |
| MiMo-V2.6-FlashXiaomi | 61.2%unknown effort | 2026-09-21 | All 1 recorded result & sources61.2% · raw 61.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL |
| GLM-5.3Z.ai | 58.2%unknown effort | 2026-09-08 | All 1 recorded result & sources58.2% · raw 58.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| Qwen3.8 FlashAlibaba (Qwen) | 55.7%xhigh effort | 2026-08-26 | All 1 recorded result & sources55.7% · raw 55.7 % Headline · xhigh effort · Own vendor Source/record date: 2026-08-26 JobBench — QwenCloud latest-model page https://docs.qwencloud.com/developer-guides/getting-started/latest-model |
| Kimi K3Moonshot AI | 54.3%max effort | 2026-07-16 | All 1 recorded result & sources54.3% · raw 54.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 JobBench https://github.com/MoonshotAI/Kimi-K3 |
| Nex-N2.5-ProNex AGI | 41.4%unknown effort | 2026-09-08 | All 1 recorded result & sources41.4% · raw 41.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-08 Job Bench https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| Qwen3.8 27BAlibaba (Qwen) | 33.4%unknown effort | 2026-08-14 | All 1 recorded result & sources33.4% · raw 33.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 JobBench https://huggingface.co/Qwen/Qwen3.8-27B |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | 18.3%unknown effort | 2026-09-21 | All 1 recorded result & sources18.3% · raw 18.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 Released SFT checkpoint, avg@1; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
| Claude Fable 5.1Anthropic | — | — | No recorded result |
| Claude Opus 5.5Anthropic | — | — | No recorded result |
| Claude Sonnet 5.5Anthropic | — | — | No recorded result |
| DeepSeek V4.1 FlashDeepSeek | — | — | No recorded result |
| Gemini 3.8 FlashGoogle | — | — | No recorded result |
| Gemini 4 ArgonGoogle | — | — | No recorded result |
| GLM-5.3-FlashZ.ai | — | — | No recorded result |
| GPT-5.6 TerraOpenAI | — | — | No recorded result |
| GPT-6 AstraOpenAI | — | — | No recorded result |
| GPT-6 LunaOpenAI | — | — | No recorded result |
| GPT-6.1 SolOpenAI | — | — | No recorded result |
| Grok 4.7xAI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Ling 3.1 FlashinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Use the fair-comparison guide before interpreting results from different configurations.