Examenos

DeepSeek V4.1 Flash Benchmarks, Specifications & Availability

Explore DeepSeek V4.1 Flash from DeepSeek: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare DeepSeek V4.1 Flash with other models →Explore data coverage

Published specifications

Provider
DeepSeek
Access
Open weights
License
MIT
Context window
1M
Total parameters
552B
Active parameters
8B in / 16B out
Released
2026-09-10
Modalities
text, image
Family
DeepSeek V4.1

Model card · Announcement · Website · OpenRouter

Model notes

Asymmetric MoE Causal Encoder–Decoder; native multimodal; KV cache compression focus. Open weights verified on Hugging Face https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash (2026-09-25).

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object]—2026-10-03 · OpenRouter source
Recorded pricing notes

OpenRouter base = peak weekday rates; off-peak / weekend overrides drop to ~$0.15/$0.60 (see OR overrides).; min-healthy endpoint minima 2026-10-01; discount flag 0.53 on StreamLake (undocumented, not applied); min-healthy endpoint minima 2026-10-01; discount flag 0.53 on StreamLake (undocumented, not applied); min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03; discount flag 0.53 on StreamLake (undocumented, not applied)

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Agents' Last Exam31.8%unknown effort
All 1 recorded result & sources

31.8% · raw 31.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

Agent's Last Exam Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
AutomationBench v1.0.654.8%unknown effort
All 1 recorded result & sources

54.8% · raw 54.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

VentureBeat quoting DeepSeek

https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5
BabyVision (w/ tools)89.6%unknown effort
All 1 recorded result & sources

89.6% · raw 89.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

BabyVision w/ tools

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Chartography (w/ tools)78.9%unknown effort
All 1 recorded result & sources

78.9% · raw 78.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

Chartography w/ tools

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Codeforces Rating3471unknown effort
All 1 recorded result & sources

3471 · raw 3471 rating

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

Codeforces rating

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
CyberGym88.1%unknown effort
All 2 recorded results & sources

88.1% · raw 88.1 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

VentureBeat quoting DeepSeek

https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5

88.1% · raw 88.1 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
DeepSWE v1.174.2%unknown effort
All 3 recorded results & sources

74.2% · raw 74.2 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

VentureBeat quoting DeepSeek reports; also deepseekagent.io 74.2

https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5

74.2% · raw 74.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

74.2% · raw 74.2 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
ExploitGym15.3%unknown effort
All 1 recorded result & sources

15.3% · raw 15.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

ExploitGym Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
GPQA Diamond90.9%unknown effort
All 1 recorded result & sources

90.9% · raw 90.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

GPQA Diamond Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Humanity's Last Exam36.8%unknown effort
All 2 recorded results & sources

36.8% · raw 36.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

HLE Pass@1; dagger variant 39.1 noted in card

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

39.2% · raw 39.2 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Differs from current headline 36.8 (first-party); kept with provenance, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Humanity's Last Exam (w/ tools)63.9%unknown effort
All 1 recorded result & sources

63.9% · raw 63.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

HLE w/ tools Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
MathArena Apex65.6%unknown effort
All 1 recorded result & sources

65.6% · raw 65.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

MathArena Apex Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
NL2Repo-Bench65.4%unknown effort
All 2 recorded results & sources

65.4% · raw 65.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

from chart/figure on DeepSeek V4.1 Flash announce (benchmark table PNG); was 64.0 from secondary guide

https://api-docs.deepseek.com/news/news260910

64% · raw 64 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
ProgramBench20.3%max effort
All 2 recorded results & sources

20.3% · raw 20.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-10

ProgramBench Almost@1; DSH Minimal, max effort

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

20.3% · raw 20.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
SEC-Bench Pro62.8%unknown effort
All 1 recorded result & sources

62.8% · raw 62.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

SEC-Bench Pro Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Terminal-Bench 2.190.6%max effort
All 3 recorded results & sources

90.6% · raw 90.6 %

Headline · max effort · Third-party

Source/record date: 2026-09-10

Secondary guide citing DeepSeek official agent snapshot; max effort

https://deepseekagent.io/deepseek-v4-1-flash

90.6% · raw 90.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

90.6% · raw 90.6 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Terminal-Bench 3.030%unknown effort
All 1 recorded result & sources

30% · raw 30 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

Secondary guide citing DeepSeek official agent snapshot

https://deepseekagent.io/deepseek-v4-1-flash
Terminal-Bench 4.031.2%unknown effort
All 1 recorded result & sources

31.2% · raw 31.2 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

Secondary guide citing DeepSeek official agent snapshot

https://deepseekagent.io/deepseek-v4-1-flash
ZeroBench-main (w/ tools)49%unknown effort
All 1 recorded result & sources

49% · raw 49 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

ZeroBench-main w/ tools Pass@5

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
SWE-Pro56.8%unknown effort
All 1 recorded result & sources

56.8% · raw 56.77 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
HealthBench Professional50.4%unknown effort
All 1 recorded result & sources

50.4% · raw 50.37 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
DRACO79.9%unknown effort
All 1 recorded result & sources

79.9% · raw 79.85 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
WideSearch80.8%unknown effort
All 1 recorded result & sources

80.8% · raw 80.81 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
MultiChallenge72.1%unknown effort
All 1 recorded result & sources

72.1% · raw 72.12 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis39max effort
All 2 recorded results & sources

39 · raw 39 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/deepseek-v4-1-flash

25 · raw 25 index

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/deepseek-v4-1-flash-non-reasoning
Cost per Intelligence Index task · Artificial Analysis$0.27max effort
All 2 recorded results & sources

$0.27 · raw 0.27 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/deepseek-v4-1-flash

$0.15 · raw 0.15 USD

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/deepseek-v4-1-flash-non-reasoning
Output speed · Artificial Analysis209 tok/smax effort
All 2 recorded results & sources

209 tok/s · raw 209 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/deepseek-v4-1-flash

215 tok/s · raw 215 tok/s

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/deepseek-v4-1-flash-non-reasoning
Vals Index · Vals AI51.3%unknown effort
All 1 recorded result & sources

51.3% · raw 51.32 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $0.33.

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench21.7 fixesmax effort
All 2 recorded results & sources

21.7 fixes · raw 21.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / DeepSeek API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

19 fixes · raw 19 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / DeepSeek API; effort high; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Hard board % of roofline · KernelBench (community board)18.8%unknown effort
All 1 recorded result & sources

18.8% · raw 18.8 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

Hard 6/6; also CUDA 23.4% 4/4, Mega 17.10×

https://kernelbench.com/models/deepseek-flash
Vibe Code Bench v1.1 · Vals AI84.7%unknown effort
All 1 recorded result & sources

84.7% · raw 84.74 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $0.41.

https://www.vals.ai/benchmarks/vibe-code
GDPval-AA Elo · Artificial Analysis1600max effort
All 1 recorded result & sources

1600 · raw 1600 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; DeepSeek V4.1 Flash (Reasoning, Max Effort); board anchor

https://artificialanalysis.ai/evaluations/gdpval-aa
CUDA board % of roofline · KernelBench (community board)23.4%unknown effort
All 1 recorded result & sources

23.4% · raw 23.4 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

CUDA 4/4; also Hard 18.8% already ingested

https://kernelbench.com/models/deepseek-flash
Average Score · WeirdML v36.2%high effort
All 2 recorded results & sources

6.2% · raw 6.23 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant DeepSeek V4.1 Flash (high, Novita) via Codex; harness codex_cli 0.156.0; values from prepared data JSON; raw 0.062291; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html

6% · raw 5.95 %

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant DeepSeek V4.1 Flash (high, Novita) via OpenCode; harness opencode 1.18.30; values from prepared data JSON; raw 0.059463; official 80/20 aggregate (area 500k-50M tokens + final best) | Harness variant: non-headline, Codex run is the headline.

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v313.3%high effort
All 1 recorded result & sources

13.3% · raw 13.28 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant DeepSeek V4.1 Flash (high, Novita) via Codex; harness codex_cli 0.156.0; values from prepared data JSON; raw 0.132778; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$0.48high effort
All 1 recorded result & sources

$0.48 · raw 0.48 USD

Headline · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant DeepSeek V4.1 Flash (high, Novita) via Codex; harness codex_cli 0.156.0; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html

Read how we select and source scores or the comparison guide.