DeepSeek V4.1 Flash Benchmarks, Specifications & Availability
Explore DeepSeek V4.1 Flash from DeepSeek: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare DeepSeek V4.1 Flash with other models →Explore data coverage
Published specifications
- Provider
- DeepSeek
- Access
- Open weights
- License
- MIT
- Context window
- 1M
- Total parameters
- 552B
- Active parameters
- 8B in / 16B out
- Released
- 2026-09-10
- Modalities
- text, image
- Family
- DeepSeek V4.1
Model card · Announcement · Website · OpenRouter
Model notes
Asymmetric MoE Causal Encoder–Decoder; native multimodal; KV cache compression focus. Open weights verified on Hugging Face https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash (2026-09-25).
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | — | 2026-10-03 · OpenRouter source |
Recorded pricing notes
OpenRouter base = peak weekday rates; off-peak / weekend overrides drop to ~$0.15/$0.60 (see OR overrides).; min-healthy endpoint minima 2026-10-01; discount flag 0.53 on StreamLake (undocumented, not applied); min-healthy endpoint minima 2026-10-01; discount flag 0.53 on StreamLake (undocumented, not applied); min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03; discount flag 0.53 on StreamLake (undocumented, not applied)
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Agents' Last Exam | 31.8%unknown effort | All 1 recorded result & sources31.8% · raw 31.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 Agent's Last Exam Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| AutomationBench v1.0.6 | 54.8%unknown effort | All 1 recorded result & sources54.8% · raw 54.8 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 VentureBeat quoting DeepSeek https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5 |
| BabyVision (w/ tools) | 89.6%unknown effort | All 1 recorded result & sources89.6% · raw 89.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 BabyVision w/ tools https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| Chartography (w/ tools) | 78.9%unknown effort | All 1 recorded result & sources78.9% · raw 78.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 Chartography w/ tools https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| Codeforces Rating | 3471unknown effort | All 1 recorded result & sources3471 · raw 3471 rating Headline · unknown effort · Own vendor Source/record date: 2026-09-10 Codeforces rating https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| CyberGym | 88.1%unknown effort | All 2 recorded results & sources88.1% · raw 88.1 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 VentureBeat quoting DeepSeek https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-588.1% · raw 88.1 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| DeepSWE v1.1 | 74.2%unknown effort | All 3 recorded results & sources74.2% · raw 74.2 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 VentureBeat quoting DeepSeek reports; also deepseekagent.io 74.2 https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-574.2% · raw 74.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/74.2% · raw 74.2 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| ExploitGym | 15.3%unknown effort | All 1 recorded result & sources15.3% · raw 15.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 ExploitGym Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| GPQA Diamond | 90.9%unknown effort | All 1 recorded result & sources90.9% · raw 90.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 GPQA Diamond Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| Humanity's Last Exam | 36.8%unknown effort | All 2 recorded results & sources36.8% · raw 36.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 HLE Pass@1; dagger variant 39.1 noted in card https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash39.2% · raw 39.2 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Differs from current headline 36.8 (first-party); kept with provenance, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Humanity's Last Exam (w/ tools) | 63.9%unknown effort | All 1 recorded result & sources63.9% · raw 63.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 HLE w/ tools Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| MathArena Apex | 65.6%unknown effort | All 1 recorded result & sources65.6% · raw 65.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 MathArena Apex Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| NL2Repo-Bench | 65.4%unknown effort | All 2 recorded results & sources65.4% · raw 65.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 from chart/figure on DeepSeek V4.1 Flash announce (benchmark table PNG); was 64.0 from secondary guide https://api-docs.deepseek.com/news/news26091064% · raw 64 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| ProgramBench | 20.3%max effort | All 2 recorded results & sources20.3% · raw 20.3 % Headline · max effort · Own vendor Source/record date: 2026-09-10 ProgramBench Almost@1; DSH Minimal, max effort https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash20.3% · raw 20.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| SEC-Bench Pro | 62.8%unknown effort | All 1 recorded result & sources62.8% · raw 62.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 SEC-Bench Pro Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| Terminal-Bench 2.1 | 90.6%max effort | All 3 recorded results & sources90.6% · raw 90.6 % Headline · max effort · Third-party Source/record date: 2026-09-10 Secondary guide citing DeepSeek official agent snapshot; max effort https://deepseekagent.io/deepseek-v4-1-flash90.6% · raw 90.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/90.6% · raw 90.6 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Terminal-Bench 3.0 | 30%unknown effort | All 1 recorded result & sources30% · raw 30 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 Secondary guide citing DeepSeek official agent snapshot https://deepseekagent.io/deepseek-v4-1-flash |
| Terminal-Bench 4.0 | 31.2%unknown effort | All 1 recorded result & sources31.2% · raw 31.2 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 Secondary guide citing DeepSeek official agent snapshot https://deepseekagent.io/deepseek-v4-1-flash |
| ZeroBench-main (w/ tools) | 49%unknown effort | All 1 recorded result & sources49% · raw 49 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 ZeroBench-main w/ tools Pass@5 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| SWE-Pro | 56.8%unknown effort | All 1 recorded result & sources56.8% · raw 56.77 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| HealthBench Professional | 50.4%unknown effort | All 1 recorded result & sources50.4% · raw 50.37 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| DRACO | 79.9%unknown effort | All 1 recorded result & sources79.9% · raw 79.85 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| WideSearch | 80.8%unknown effort | All 1 recorded result & sources80.8% · raw 80.81 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| MultiChallenge | 72.1%unknown effort | All 1 recorded result & sources72.1% · raw 72.12 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 39max effort | All 2 recorded results & sources39 · raw 39 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/deepseek-v4-1-flash25 · raw 25 index Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/deepseek-v4-1-flash-non-reasoning |
| Cost per Intelligence Index task · Artificial Analysis | $0.27max effort | All 2 recorded results & sources$0.27 · raw 0.27 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/deepseek-v4-1-flash$0.15 · raw 0.15 USD Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/deepseek-v4-1-flash-non-reasoning |
| Output speed · Artificial Analysis | 209 tok/smax effort | All 2 recorded results & sources209 tok/s · raw 209 tok/s Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/deepseek-v4-1-flash215 tok/s · raw 215 tok/s Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/deepseek-v4-1-flash-non-reasoning |
| Vals Index · Vals AI | 51.3%unknown effort | All 1 recorded result & sources51.3% · raw 51.32 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $0.33. https://www.vals.ai/benchmarks/vals_index |
| Bugs fixed /105 · Bug Hunt Bench | 21.7 fixesmax effort | All 2 recorded results & sources21.7 fixes · raw 21.7 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / DeepSeek API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench19 fixes · raw 19 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / DeepSeek API; effort high; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Hard board % of roofline · KernelBench (community board) | 18.8%unknown effort | All 1 recorded result & sources18.8% · raw 18.8 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 Hard 6/6; also CUDA 23.4% 4/4, Mega 17.10× https://kernelbench.com/models/deepseek-flash |
| Vibe Code Bench v1.1 · Vals AI | 84.7%unknown effort | All 1 recorded result & sources84.7% · raw 84.74 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $0.41. https://www.vals.ai/benchmarks/vibe-code |
| GDPval-AA Elo · Artificial Analysis | 1600max effort | All 1 recorded result & sources1600 · raw 1600 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-23 GDPval-AA v2.1 Elo; DeepSeek V4.1 Flash (Reasoning, Max Effort); board anchor https://artificialanalysis.ai/evaluations/gdpval-aa |
| CUDA board % of roofline · KernelBench (community board) | 23.4%unknown effort | All 1 recorded result & sources23.4% · raw 23.4 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 CUDA 4/4; also Hard 18.8% already ingested https://kernelbench.com/models/deepseek-flash |
| Average Score · WeirdML v3 | 6.2%high effort | All 2 recorded results & sources6.2% · raw 6.23 % Headline · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant DeepSeek V4.1 Flash (high, Novita) via Codex; harness codex_cli 0.156.0; values from prepared data JSON; raw 0.062291; official 80/20 aggregate (area 500k-50M tokens + final best) https://htihle.github.io/weirdml.html6% · raw 5.95 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant DeepSeek V4.1 Flash (high, Novita) via OpenCode; harness opencode 1.18.30; values from prepared data JSON; raw 0.059463; official 80/20 aggregate (area 500k-50M tokens + final best) | Harness variant: non-headline, Codex run is the headline. https://htihle.github.io/weirdml.html |
| Final Best Score · WeirdML v3 | 13.3%high effort | All 1 recorded result & sources13.3% · raw 13.28 % Headline · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant DeepSeek V4.1 Flash (high, Novita) via Codex; harness codex_cli 0.156.0; values from prepared data JSON; raw 0.132778; mean final best effective score https://htihle.github.io/weirdml.html |
| Cost / Run · WeirdML v3 | $0.48high effort | All 1 recorded result & sources$0.48 · raw 0.48 USD Headline · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant DeepSeek V4.1 Flash (high, Novita) via Codex; harness codex_cli 0.156.0; values from prepared data JSON; mean API cost per run, same task weighting as scores https://htihle.github.io/weirdml.html |
Read how we select and source scores or the comparison guide.