GLM-5.3 Benchmarks, Specifications & Availability
Explore GLM-5.3 from Z.ai: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare GLM-5.3 with other models →Explore data coverage
Published specifications
- Provider
- Z.ai
- Access
- Open weights
- License
- GLM-5.3 License
- Context window
- 1M
- Total parameters
- 753B
- Active parameters
- Not published
- Released
- 2026-08-14
- Modalities
- text
- Family
- GLM-5.3
Model card · Announcement · Website · OpenRouter
Model notes
Flagship post-training upgrade on GLM-5.2 base; strong coding/cyber claims. Open weights verified on Hugging Face https://huggingface.co/zai-org/GLM-5.3 (2026-09-25). Previously mislabeled closed/"Proprietary (API)"; official zai-org repo has downloadable weights (custom GLM-5.3 License, MIT-style permissive). Size 753B total from the Hugging Face safetensors parameter count (zai-org/GLM-5.3, 2026-09-25); active params not published on the card.
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | — | 2026-10-03 · OpenRouter source |
Recorded pricing notes
Re-pulled 2026-09-25 from OR API; was in/out/cache 0.84/2.64/0.156 on 2026-09-23.; min-healthy endpoint minima 2026-10-01; discount flag 0.889 on Baidu (undocumented, not applied); discount flag 0.889 on Baidu (undocumented, not applied); min-healthy endpoint minima 2026-10-01; discount flag 0.889 on Baidu (undocumented, not applied); discount flag 0.889 on Baidu (undocumented, not applied); min-healthy endpoint minima 2026-10-02; discount flag 0.889 on Baidu (undocumented, not applied); discount flag 0.889 on Baidu (undocumented, not applied); min-healthy endpoint minima 2026-10-03; discount flag 0.7 on Novita (undocumented, not applied); discount flag 0.7 on Novita (undocumented, not applied)
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Agents' Last Exam | 28.5%unknown effort | All 2 recorded results & sources28.5% · raw 28.5 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 https://docs.z.ai/guides/llm/glm-5.328.5% · raw 28.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| AutomationBench v1.0.6 | 48.8%unknown effort | All 2 recorded results & sources48.8% · raw 48.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash48.2% · raw 48.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| CyberGym | 84.5%unknown effort | All 1 recorded result & sources84.5% · raw 84.5 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 https://docs.z.ai/guides/llm/glm-5.3 |
| DeepSWE v1.1 | 66.9%unknown effort | All 3 recorded results & sources66.9% · raw 66.9 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 https://docs.z.ai/guides/llm/glm-5.366.9% · raw 66.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Z.ai GLM-5.3 docs (docs.z.ai); vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value. https://huggingface.co/nex-agi/Nex-N2.5-Pro66.9% · raw 66.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| ExploitBench | 54.4%unknown effort | All 1 recorded result & sources54.4% · raw 54.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 https://docs.z.ai/guides/llm/glm-5.3 |
| ExploitGym | 15%unknown effort | All 1 recorded result & sources15% · raw 15 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| Humanity's Last Exam (w/ tools) | 62.5%unknown effort | All 1 recorded result & sources62.5% · raw 62.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| NL2Repo-Bench | 58%unknown effort | All 2 recorded results & sources58% · raw 58 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash58% · raw 58 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| ProgramBench | 19%unknown effort | All 2 recorded results & sources19% · raw 19 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash19% · raw 19 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Terminal-Bench 2.1 | 88.2%unknown effort | All 3 recorded results & sources88.2% · raw 88.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash88.2% · raw 88.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from DeepSeek-V4.1-Flash HF comparison table; both rows are peer-vendor spillover with the same value (no Z.ai official TB2.1 row), kept the later-dated source (2026-09-10 vs 2026-09-08) (rule b). https://huggingface.co/nex-agi/Nex-N2.5-Pro88.2% · raw 88.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Terminal-Bench 3.0 | 28.3%unknown effort | All 1 recorded result & sources28.3% · raw 28.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 docs cite improvement 4.6→28.3 https://docs.z.ai/guides/llm/glm-5.3 |
| Terminal-Bench 4.0 | 37.9%unknown effort | All 1 recorded result & sources37.9% · raw 37.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| SWE-bench Pro | 64.6%unknown effort | All 1 recorded result & sources64.6% · raw 64.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| Toolathlon-Verified | 73%unknown effort | All 1 recorded result & sources73% · raw 73 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| GDPval-AA v2 | 1769unknown effort | All 2 recorded results & sources1763 · raw 1763 Elo Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro1769 · raw 1769 Elo Headline · unknown effort · Own vendor Source/record date: Not recorded from chart/figure Z.ai GLM-5.3 docs LLM Performance Evaluation chart; Elo https://docs.z.ai/guides/llm/glm-5.3 |
| JobBench | 58.2%unknown effort | All 1 recorded result & sources58.2% · raw 58.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| GPQA Diamond | 88.1%unknown effort | All 1 recorded result & sources88.1% · raw 88.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 from chart/figure as reported on DeepSeek V4.1 Flash announce table (peer spillover) https://api-docs.deepseek.com/news/news260910 |
| Humanity's Last Exam | 42%unknown effort | All 1 recorded result & sources42% · raw 42 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 from chart/figure DeepSeek V4.1 Flash table; HLE text-only subset (*) https://api-docs.deepseek.com/news/news260910 |
| AA Intelligence Index (vendor-cited) | 60unknown effort | All 1 recorded result & sources60 · raw 60 index Headline · unknown effort · Peer vendor Source/record date: Not recorded from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); GLM-5.3 (max) https://huggingface.co/inclusionAI/Ling-3.0-flash-VL |
| FrontierCode 1.1 Main | 40.1%unknown effort | All 1 recorded result & sources40.1% · raw 40.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 Cognition FrontierCode 1.1 Main weighted score; model added Sep 2, 2026 per Cognition changelog; score 40.1% on board export mirrored by evals.report (sourceUrl cognition.com, verifiedStatus official, snapshot 2026-09-07) https://cognition.com/frontiercode |
| FrontierSWE | 78.1%unknown effort | All 1 recorded result & sources78.1% · raw 78.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| PostTrainBench | 39.8%unknown effort | All 1 recorded result & sources39.8% · raw 39.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Gray Swan IPI | 31.5%unknown effort | All 1 recorded result & sources31.5% · raw 31.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon Gray Swan chart (K=15 attack success rate) https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 45max effort | All 2 recorded results & sources45 · raw 45 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/glm-5-334 · raw 34 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/glm-5-3-low |
| Cost per Intelligence Index task · Artificial Analysis | $2.01max effort | All 2 recorded results & sources$2.01 · raw 2.01 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/glm-5-3$0.85 · raw 0.85 USD Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/glm-5-3-low |
| Output speed · Artificial Analysis | 71 tok/smax effort | All 2 recorded results & sources71 tok/s · raw 71 tok/s Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/glm-5-367 tok/s · raw 67 tok/s Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/glm-5-3-low |
| Vals Index · Vals AI | 53.5%unknown effort | All 1 recorded result & sources53.5% · raw 53.51 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $7.25. https://www.vals.ai/benchmarks/vals_index |
| Bugs fixed /105 · Bug Hunt Bench | 19 fixesmax effort | All 1 recorded result & sources19 fixes · raw 19 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Z.ai API; effort max; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| Pass Rate · MCP Atlas (Scale Labs) | 84.2%unknown effort | All 1 recorded result & sources84.2% · raw 84.2 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 Scale Labs Performance Comparison chart (GLM 5.3); ±2.15 CI https://labs.scale.com/leaderboard/mcp_atlas |
| Hard board % of roofline · KernelBench (community board) | 24.6%unknown effort | All 1 recorded result & sources24.6% · raw 24.6 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 Hard 5/6; CUDA 10.0% 1/4; Mega 19.43× https://kernelbench.com/models/glm-5.3 |
| GDPval-AA Elo · Artificial Analysis | 1646max effort | All 1 recorded result & sources1646 · raw 1646 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-23 GDPval-AA v2.1 Elo; GLM-5.3 (max) https://artificialanalysis.ai/evaluations/gdpval-aa |
| CUDA board % of roofline · KernelBench (community board) | 10%unknown effort | All 1 recorded result & sources10% · raw 10 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 CUDA 1/4 publishable; Hard 24.6% already ingested https://kernelbench.com/models/glm-5.3 |
| Money gain · Andon Labs | $7663.61unknown effort | All 1 recorded result & sources$7663.61 · raw 7663.61 $ Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked. https://andonlabs.com/evals/vending-bench-2 |
| Vibe Code Bench v1.1 · Vals AI | 78.1%unknown effort | All 1 recorded result & sources78.1% · raw 78.13 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $12.45. https://www.vals.ai/benchmarks/vibe-code |
Read how we select and source scores or the comparison guide.