Examenos

Qwen3.8 Max (0902) Benchmarks, Specifications & Availability

Explore Qwen3.8 Max (0902) from Alibaba (Qwen): published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Qwen3.8 Max (0902) with other models →Explore data coverage

Published specifications

Provider
Alibaba (Qwen)
Access
Open weights
License
Qwen3.8-Max License
Context window
1M
Total parameters
2.4T
Active parameters
95B
Released
2026-09-02
Modalities
text, image, video
Family
Qwen3.8

Model card · Announcement · Website · OpenRouter

Model notes

Post-training upgrade of Qwen3.8-Max focused on coding/agents (0902 snapshot). Open weights: Qwen/Qwen3.8-2.4T-A95B on HF, ungated, Qwen3.8-Max License (checked 2026-09-28); README confirms 2.4T total / 95B activated.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
CoWorkBench76.1%unknown effort
All 1 recorded result & sources

76.1% · raw 76.1 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max
DeepSWE v1.169.3%unknown effort
All 3 recorded results & sources

69.3% · raw 69.3 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max

69.3% · raw 69.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value.

https://huggingface.co/nex-agi/Nex-N2.5-Pro

56.6% · raw 56.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
GPQA Diamond92.6%unknown effort
All 1 recorded result & sources

92.6% · raw 92.6 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

GPQA Diamond

https://www.datacamp.com/blog/qwen3-8-max
HealthBench60.2%unknown effort
All 1 recorded result & sources

60.2% · raw 60.2 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

HealthBench

https://www.datacamp.com/blog/qwen3-8-max
Humanity's Last Exam43.6%unknown effort
All 1 recorded result & sources

43.6% · raw 43.6 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Humanity's Last Exam

https://www.datacamp.com/blog/qwen3-8-max
IFBench82.8%unknown effort
All 1 recorded result & sources

82.8% · raw 82.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

IFBench instruction following

https://www.datacamp.com/blog/qwen3-8-max
JobBench64%unknown effort
All 2 recorded results & sources

64% · raw 64 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max

53.4% · raw 53.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp (64.0); vendor official 0902 figure preferred (rules a+b). 53.4 equals the pre-0902 Qwen3.8-Max (2026-08-03) value in Alibaba's table, so Nex's "Qwen3.8-Max" column appears to be the original snapshot, not 0902.

https://huggingface.co/nex-agi/Nex-N2.5-Pro
LVBench81.8%max effort
All 1 recorded result & sources

81.8% · raw 81.8 %

Headline · max effort · Third-party

Source/record date: 2026-09-02

LVBench — Qwen3.8-Max family table

https://www.datacamp.com/blog/qwen3-8-max
MLS-Bench-Lite50.1%unknown effort
All 1 recorded result & sources

50.1% · raw 50.1 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

MLS-Bench-Lite

https://www.datacamp.com/blog/qwen3-8-max
MobileWorld77.8%unknown effort
All 1 recorded result & sources

77.8% · raw 77.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

MobileWorld

https://www.datacamp.com/blog/qwen3-8-max
NL2Repo-Bench64.9%unknown effort
All 2 recorded results & sources

64.9% · raw 64.9 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max

55.9% · raw 55.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
OSWorld-Verified86.1%unknown effort
All 2 recorded results & sources

86.1% · raw 86.1 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

OSWorld-Verified

https://www.datacamp.com/blog/qwen3-8-max

86.1% · raw 86.1 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value.

https://huggingface.co/nex-agi/Nex-N2.5-Pro
PaperBench93%max effort
All 1 recorded result & sources

93% · raw 93 %

Headline · max effort · Third-party

Source/record date: 2026-09-02

PaperBench — DataCamp table for Qwen3.8-Max (family); confirm if 0902-identical

https://www.datacamp.com/blog/qwen3-8-max
PerceptionBench63.5%max effort
All 1 recorded result & sources

63.5% · raw 63.5 %

Headline · max effort · Third-party

Source/record date: 2026-09-02

PerceptionBench — Qwen3.8-Max family table

https://www.datacamp.com/blog/qwen3-8-max
ProgramBench (Almost Solved)28%unknown effort
All 1 recorded result & sources

28% · raw 28 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

ProgramBench Almost Solved; Alibaba 0902 table via DataCamp

https://www.datacamp.com/blog/qwen3-8-max
QwenSWEBench V270%unknown effort
All 1 recorded result & sources

70% · raw 70 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max
SWE-Atlas QnA66.3%unknown effort
All 1 recorded result & sources

66.3% · raw 66.3 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

SWE-Atlas QnA

https://www.datacamp.com/blog/qwen3-8-max
SWE-bench Pro67.7%unknown effort
All 3 recorded results & sources

67.7% · raw 67.7 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

SWE-Pro / SWE-bench Pro

https://www.datacamp.com/blog/qwen3-8-max

67.7% · raw 67.7 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value.

https://huggingface.co/nex-agi/Nex-N2.5-Pro

67.7% · raw 67.7 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
SWE-Marathon44.8%unknown effort
All 1 recorded result & sources

44.8% · raw 44.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

SWE-Marathon

https://www.datacamp.com/blog/qwen3-8-max
Terminal-Bench 2.186.6%max effort
All 3 recorded results & sources

86.6% · raw 86.6 %

Headline · max effort · Third-party

Source/record date: 2026-09-02

Terminal-Bench 2.1 — Qwen3.8-Max family table

https://www.datacamp.com/blog/qwen3-8-max

86.6% · raw 86.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value (Nex row effort unknown, vendor row max).

https://huggingface.co/nex-agi/Nex-N2.5-Pro

88.8% · raw 88.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
Terminal-Bench 3.029%unknown effort
All 1 recorded result & sources

29% · raw 29 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max
Toolathlon-Verified73.3%unknown effort
All 2 recorded results & sources

73.3% · raw 73.3 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max

72.5% · raw 72.5 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp (73.3); vendor official 0902 figure preferred (rules a+b). 72.5 equals the pre-0902 Qwen3.8-Max (2026-08-03) value in Alibaba's table, so Nex's "Qwen3.8-Max" column appears to be the original snapshot, not 0902.

https://huggingface.co/nex-agi/Nex-N2.5-Pro
Vision2Web69%unknown effort
All 2 recorded results & sources

69% · raw 69 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Vision2Web

https://www.datacamp.com/blog/qwen3-8-max

75.1% · raw 75.1 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Harness: Nex-AGI Vision2Web eval (avg of Frontend/Webpage/Website; Gemini-3.5-Flash VLM judge, GLM-5V-Turbo (Claude Code) GUI agent; Nex card footnote 7) — different harness from Alibaba's figure. Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max table via DataCamp (69.0); vendor official preferred (rule a).

https://huggingface.co/nex-agi/Nex-N2.5-Pro
WorkArena1468unknown effort
All 1 recorded result & sources

1468 · raw 1468 Elo

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Elo; Alibaba via DataCamp

https://www.datacamp.com/blog/qwen3-8-max
AutomationBench v1.0.639.8%unknown effort
All 1 recorded result & sources

39.8% · raw 39.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
GDPval-AA v21717unknown effort
All 1 recorded result & sources

1717 · raw 1717 Elo

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
OSWorld 2.046.7%unknown effort
All 1 recorded result & sources

46.7% · raw 46.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
WebTest52.3%unknown effort
All 1 recorded result & sources

52.3% · raw 52.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
WebArena-Verified66.8%unknown effort
All 1 recorded result & sources

66.8% · raw 66.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
OSWorld-G84.9%unknown effort
All 1 recorded result & sources

84.9% · raw 84.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
SWE-MM39.2%unknown effort
All 1 recorded result & sources

39.2% · raw 39.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
OmniDoc92.1%unknown effort
All 1 recorded result & sources

92.1% · raw 92.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
MMMU Pro (no tools)82.7%unknown effort
All 1 recorded result & sources

82.7% · raw 82.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-23

As reported from Alibaba Qwen3.8-Max-0902 published comparison table (DataCamp transcription); MMMU-Pro

https://www.datacamp.com/blog/qwen3-8-max
FrontierSWE73.5%unknown effort
All 1 recorded result & sources

73.5% · raw 73.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis45max effort
All 1 recorded result & sources

45 · raw 45 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-max
Cost per Intelligence Index task · Artificial Analysis$5.41max effort
All 1 recorded result & sources

$5.41 · raw 5.41 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-max
Output speed · Artificial Analysis39 tok/smax effort
All 1 recorded result & sources

39 tok/s · raw 39 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/qwen3-8-max
Vals Index · Vals AI51.8%max effort
All 1 recorded result & sources

51.8% · raw 51.84 %

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

Vals lists alibaba/qwen3.8-max (not an explicit 0902 slug); mapped to site model qwen3.8-max-0902

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench25.7 fixesmax effort
All 1 recorded result & sources

25.7 fixes · raw 25.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

Claude Code / Alibaba API max; 3-run mean — listed as Qwen3.8-Max; board data/benchmark.json updated 2026-09-23

https://github.com/phuryn/bug-hunt-bench
Hard board % of roofline · KernelBench (community board)24.1%max effort
All 1 recorded result & sources

24.1% · raw 24.1 %

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

Hard 5/6 on qwen3.8-max board; CUDA 19.3% 2/4; some cells audit-flagged

https://kernelbench.com/models/qwen3.8-max
GDPval-AA Elo · Artificial Analysis1668unknown effort
All 1 recorded result & sources

1668 · raw 1668 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; Qwen3.8 Max (0902)

https://artificialanalysis.ai/evaluations/gdpval-aa
CUDA board % of roofline · KernelBench (community board)19.3%max effort
All 1 recorded result & sources

19.3% · raw 19.3 %

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

CUDA 2/4; Hard 24.1% already ingested; some cells audit-flagged

https://kernelbench.com/models/qwen3.8-max

Read how we select and source scores or the comparison guide.