Examenos

GLM-5.3-Flash Benchmarks, Specifications & Availability

Explore GLM-5.3-Flash from Z.ai: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare GLM-5.3-Flash with other models →Explore data coverage

Published specifications

Provider
Z.ai
Access
Open weights
License
MIT
Context window
1M
Total parameters
320B
Active parameters
18B
Released
2026-08-26
Modalities
text, image, video
Family
GLM-5.3

Model card · Announcement · Website · OpenRouter

Model notes

First natively multimodal GLM-5 series model; efficiency tier. Open weights verified on Hugging Face https://huggingface.co/zai-org/GLM-5.3-Flash (2026-09-25).

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object]—2026-10-03 · OpenRouter source
Recorded pricing notes

Re-pulled 2026-09-25 from OR API; was in/out/cache 0.15/0.5/0.05 on 2026-09-23.; min-healthy endpoint minima 2026-10-01; discount flag 0.5 on DeepInfra (undocumented, not applied); min-healthy endpoint minima 2026-10-01; discount flag 0.5 on DeepInfra (undocumented, not applied); min-healthy endpoint minima 2026-10-02; discount flag 0.5 on DeepInfra (undocumented, not applied); min-healthy endpoint minima 2026-10-03; discount flag 0.5 on DeepInfra (undocumented, not applied); discount flag 0.5 on DeepInfra (undocumented, not applied)

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
AA Intelligence Index (vendor-cited)57unknown effort
All 1 recorded result & sources

57 · raw 57 index

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

AA Intelligence Index v4.1.1 as cited by z.ai ($0.045/task discounted)

https://z.ai/blog/glm-5.3-flash
Agents' Last Exam26.3%unknown effort
All 2 recorded results & sources

26.3% · raw 26.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash

26.3% · raw 26.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
AutomationBench v1.0.648.8%unknown effort
All 1 recorded result & sources

48.8% · raw 48.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash
DeepSWE v1.163.4%unknown effort
All 3 recorded results & sources

63.4% · raw 63.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash

63.4% · raw 63.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

63.4% · raw 63.4 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
ExtractBench Short96.3%unknown effort
All 1 recorded result & sources

96.3% · raw 96.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

ExtractBench Short from HF eval badge

https://huggingface.co/zai-org/GLM-5.3-Flash
GDPval-AA v21773unknown effort
All 1 recorded result & sources

1773 · raw 1773 Elo

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

Elo GDPval-AA v2

https://z.ai/blog/glm-5.3-flash
Humanity's Last Exam (w/ tools)55.3%unknown effort
All 1 recorded result & sources

55.3% · raw 55.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash
MVBench77.8%unknown effort
All 1 recorded result & sources

77.8% · raw 77.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

MVBench

https://z.ai/blog/glm-5.3-flash
NL2Repo56.3%unknown effort
All 1 recorded result & sources

56.3% · raw 56.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

NL2Repo on z.ai blog table

https://z.ai/blog/glm-5.3-flash
NL2Repo-Bench56.3%unknown effort
All 1 recorded result & sources

56.3% · raw 56.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash
OSWorld 2.059.1%unknown effort
All 1 recorded result & sources

59.1% · raw 59.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

OSWorld 2.0

https://z.ai/blog/glm-5.3-flash
Terminal-Bench 2.184.3%unknown effort
All 3 recorded results & sources

84.3% · raw 84.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash

84.3% · raw 84.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

84.3% · raw 84.3 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Toolathlon-Verified78.4%unknown effort
All 1 recorded result & sources

78.4% · raw 78.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash
Z.ai Code Bench v1.029%max effort
All 1 recorded result & sources

29% · raw 29 %

Headline · max effort · Own vendor

Source/record date: 2026-08-26

max effort; vs Opus 4.8 29.5

https://z.ai/blog/glm-5.3-flash
OSWorld-Verified62.3%unknown effort
All 1 recorded result & sources

62.3% · raw 62.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
OSWorld-G83.3%unknown effort
All 1 recorded result & sources

83.3% · raw 83.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
SWE-MM20.6%unknown effort
All 1 recorded result & sources

20.6% · raw 20.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
FrontierCode 1.1 Main31.8%unknown effort
All 1 recorded result & sources

31.8% · raw 31.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

Cognition FrontierCode 1.1 Main weighted score; model added Sep 2, 2026 per Cognition changelog; score 31.8% on board export mirrored by evals.report (sourceUrl cognition.com, verifiedStatus official, snapshot 2026-09-07)

https://cognition.com/frontiercode
SWE-Pro63.1%unknown effort
All 1 recorded result & sources

63.1% · raw 63.06 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Terminal-Bench 4.032.8%unknown effort
All 1 recorded result & sources

32.8% · raw 32.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
HealthBench Professional49.1%unknown effort
All 1 recorded result & sources

49.1% · raw 49.07 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
DRACO78.6%unknown effort
All 1 recorded result & sources

78.6% · raw 78.55 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
WideSearch80.2%unknown effort
All 1 recorded result & sources

80.2% · raw 80.24 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Humanity's Last Exam39.9%unknown effort
All 1 recorded result & sources

39.9% · raw 39.9 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
MultiChallenge62.6%unknown effort
All 1 recorded result & sources

62.6% · raw 62.6 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis42unknown effort
All 1 recorded result & sources

42 · raw 42 index

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/glm-5-3-flash
Cost per Intelligence Index task · Artificial Analysis$0.25unknown effort
All 1 recorded result & sources

$0.25 · raw 0.25 USD

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/glm-5-3-flash
Output speed · Artificial Analysis54 tok/sunknown effort
All 1 recorded result & sources

54 tok/s · raw 54 tok/s

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/glm-5-3-flash
Vals Index · Vals AI47.2%unknown effort
All 1 recorded result & sources

47.2% · raw 47.22 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench17.7 fixesmax effort
All 4 recorded results & sources

17.7 fixes · raw 17.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

16 fixes · raw 16 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort high; 1 runs; evaluation 2026-09-13. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

13 fixes · raw 13 fixes

Alternative · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

9 fixes · raw 9 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort low; 1 runs; evaluation 2026-09-13. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
GDPval-AA Elo · Artificial Analysis1641unknown effort
All 1 recorded result & sources

1641 · raw 1641 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; GLM 5.3 Flash

https://artificialanalysis.ai/evaluations/gdpval-aa
Hard board % of roofline · KernelBench (community board)6.5%unknown effort
All 1 recorded result & sources

6.5% · raw 6.5 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

Community board lists model as GLM-5.3 Flash (slug ox-alpha). Hard deck % of roofline (3/6); Mega 13.64×; CUDA 4.2% 2/4. Not KernelGen 1P.

https://kernelbench.com/models/ox-alpha
Vibe Code Bench v1.1 · Vals AI30.8%unknown effort
All 1 recorded result & sources

30.8% · raw 30.76 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $2.35.

https://www.vals.ai/benchmarks/vibe-code
ARC-AGI-1 · ARC Prize91%max effort
All 3 recorded results & sources

91% · raw 91 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/zai-glm-5-3-flash

71.8% · raw 71.8 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/zai-glm-5-3-flash

47% · raw 47 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/zai-glm-5-3-flash
ARC-AGI-2 · ARC Prize65.8%max effort
All 3 recorded results & sources

65.8% · raw 65.8 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/zai-glm-5-3-flash

50.1% · raw 50.1 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/zai-glm-5-3-flash

27.9% · raw 27.9 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/zai-glm-5-3-flash

Read how we select and source scores or the comparison guide.