Examenos

Qwen3.8 Flash Benchmarks, Specifications & Availability

Explore Qwen3.8 Flash from Alibaba (Qwen): published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Qwen3.8 Flash with other models →Explore data coverage

Published specifications

Provider
Alibaba (Qwen)
Access
Open weights
License
Qwen Community License 1.0
Context window
1M
Total parameters
125B
Active parameters
6B
Released
2026-08-26
Modalities
text, image, video
Family
Qwen3.8

Model card · Announcement · Website · OpenRouter

Model notes

Production QwenCloud API for Qwen3.8-Flash-Next architecture (+51B n-gram embeddings off-accelerator). Open-weight Flash-Next sibling on HF/ModelScope. OR slug verified 2026-09-23: only qwen/qwen3.8-flash (no dated Flash suffix; Max uses qwen/qwen3.8-max-0902). Open weights verified on Hugging Face https://huggingface.co/Qwen/Qwen3.8-Flash-Next (2026-09-25). Previously mislabeled closed. Weights published as Qwen3.8-Flash-Next (Qwen Community License 1.0); Qwen card and release blog say the hosted qwen3.8-flash is the official version of Flash-Next served on QwenCloud with production extras (1M ctx default, built-in tools).

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
SWE-bench Pro62.5%xhigh effort
All 1 recorded result & sources

62.5% · raw 62.5 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

SWE-bench Pro — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
DeepSWE v1.158.7%xhigh effort
All 1 recorded result & sources

58.7% · raw 58.7 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

DeepSWE 1.1 — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
SWE-bench Multilingual81%xhigh effort
All 1 recorded result & sources

81% · raw 81 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

SWE-bench Multilingual — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
CoWorkBench73.9%xhigh effort
All 1 recorded result & sources

73.9% · raw 73.9 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

CoWorkBench — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
JobBench55.7%xhigh effort
All 1 recorded result & sources

55.7% · raw 55.7 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

JobBench — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
Toolathlon73.5%xhigh effort
All 1 recorded result & sources

73.5% · raw 73.5 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

Toolathlon — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
AndroidWorld84.5%xhigh effort
All 1 recorded result & sources

84.5% · raw 84.5 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

AndroidWorld — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
MathVision95.7%xhigh effort
All 1 recorded result & sources

95.7% · raw 95.7 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

MathVision — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
LVBench76.6%xhigh effort
All 1 recorded result & sources

76.6% · raw 76.6 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

LVBench — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
GPQA Diamond91.7%unknown effort
All 1 recorded result & sources

91.7% · raw 91.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

Qwen3.8-Flash-Next blog table (GPQA Diamond)

https://qwen.ai/blog?id=qwen3.8-flash-next
OSWorld 2.052.3%unknown effort
All 1 recorded result & sources

52.3% · raw 52.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

OSWorld 2.0 partial reward from Qwen3.8-Flash-Next blog (binary 19.4 also listed)

https://qwen.ai/blog?id=qwen3.8-flash-next

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Bugs fixed /105 · Bug Hunt Bench26 fixesmax effort
All 2 recorded results & sources

26 fixes · raw 26 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Alibaba API; effort max; 1 runs; evaluation 2026-09-11. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

23 fixes · raw 23 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Alibaba API; effort low; 1 runs; evaluation 2026-09-11. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
GDPval-AA Elo · Artificial Analysis1612unknown effort
All 1 recorded result & sources

1612 · raw 1612 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; Qwen3.8-Flash-Next

https://artificialanalysis.ai/evaluations/gdpval-aa
Intelligence Index · Artificial Analysis40unknown effort
All 1 recorded result & sources

40 · raw 40 index

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

Qwen3.8-Flash-Next AA model page

https://artificialanalysis.ai/models/qwen3-8-flash-next
Output speed · Artificial Analysis53.2 tok/sunknown effort
All 1 recorded result & sources

53.2 tok/s · raw 53.2 tok/s

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

AA model page output tokens/sec

https://artificialanalysis.ai/models/qwen3-8-flash-next
Cost per Intelligence Index task · Artificial Analysis$0.37unknown effort
All 1 recorded result & sources

$0.37 · raw 0.37 $

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

AA model page

https://artificialanalysis.ai/models/qwen3-8-flash-next

Read how we select and source scores or the comparison guide.