Examenos

JobBench Results Across LLMs

Vendor-reported JobBench results across showcased LLMs. Job/task agent benchmark

Official / vendor Higher is better · unit: %

What this table contains

Job/task agent benchmark

11 models with results · 19 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

JobBench: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
Muse Spark 1.3Meta64.9%unknown effort2026-09-02
All 1 recorded result & sources

64.9% · raw 64.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

JobBench; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3
Qwen3.8 Max (0902)Alibaba (Qwen)64%unknown effort2026-09-02
All 2 recorded results & sources

64% · raw 64 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max

53.4% · raw 53.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp (64.0); vendor official 0902 figure preferred (rules a+b). 53.4 equals the pre-0902 Qwen3.8-Max (2026-08-03) value in Alibaba's table, so Nex's "Qwen3.8-Max" column appears to be the original snapshot, not 0902.

https://huggingface.co/nex-agi/Nex-N2.5-Pro
MiMo-V2.6-ProXiaomi62%unknown effort2026-09-21
All 1 recorded result & sources

62% · raw 62 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Hy4 previewTencent61.7%unknown effort2026-08-28
All 1 recorded result & sources

61.7% · raw 61.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

Vendor chart transcription

https://hy4.site/benchmarks/hy4-benchmarks
MiMo-V2.6-FlashXiaomi61.2%unknown effort2026-09-21
All 1 recorded result & sources

61.2% · raw 61.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
GLM-5.3Z.ai58.2%unknown effort2026-09-08
All 1 recorded result & sources

58.2% · raw 58.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
Qwen3.8 FlashAlibaba (Qwen)55.7%xhigh effort2026-08-26
All 1 recorded result & sources

55.7% · raw 55.7 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

JobBench — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
Kimi K3Moonshot AI54.3%max effort2026-07-16
All 1 recorded result & sources

54.3% · raw 54.3 %

Headline · max effort · Own vendor

Source/record date: 2026-07-16

JobBench

https://github.com/MoonshotAI/Kimi-K3
Nex-N2.5-ProNex AGI41.4%unknown effort2026-09-08
All 1 recorded result & sources

41.4% · raw 41.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-08

Job Bench

https://huggingface.co/nex-agi/Nex-N2.5-Pro
Qwen3.8 27BAlibaba (Qwen)33.4%unknown effort2026-08-14
All 1 recorded result & sources

33.4% · raw 33.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

JobBench

https://huggingface.co/Qwen/Qwen3.8-27B
MiMo-V2.6-Distill-Qwen-9BXiaomi18.3%unknown effort2026-09-21
All 1 recorded result & sources

18.3% · raw 18.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

Released SFT checkpoint, avg@1; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Claude Fable 5.1Anthropic——No recorded result
Claude Opus 5.5Anthropic——No recorded result
Claude Sonnet 5.5Anthropic——No recorded result
DeepSeek V4.1 FlashDeepSeek——No recorded result
Gemini 3.8 FlashGoogle——No recorded result
Gemini 4 ArgonGoogle——No recorded result
GLM-5.3-FlashZ.ai——No recorded result
GPT-5.6 TerraOpenAI——No recorded result
GPT-6 AstraOpenAI——No recorded result
GPT-6 LunaOpenAI——No recorded result
GPT-6.1 SolOpenAI——No recorded result
Grok 4.7xAI——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Ling 3.1 FlashinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
Naive-N0.5-FlashNaiveAI——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Use the fair-comparison guide before interpreting results from different configurations.