Naive-N0.5-Flash Benchmarks, Specifications & Availability
Explore Naive-N0.5-Flash from NaiveAI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Naive-N0.5-Flash with other models →Explore data coverage
Published specifications
- Provider
- NaiveAI
- Access
- Open weights
- License
- MIT
- Context window
- 1M
- Total parameters
- 309B
- Active parameters
- 15.5B
- Released
- 2026-09-27
- Modalities
- text
- Family
- Naive N
Model card · Announcement · Website
Model notes
309B MoE / 15.5B active; coding + AI R&D; SWA-DSA hybrid with no full-attention layers; native 1M context. Built on the MiMo-V2.5 base. Not on OpenRouter as of 2026-09-27; vendor API $0.10/$0.40/$0.01 per 1M input/output/cache. Open weights verified on Hugging Face https://huggingface.co/NaiveAI/Naive-N0.5-Flash (2026-09-27, MIT, ungated).
Pricing · OpenRouter
No OpenRouter pricing is recorded for this model. Missing rates are not free.
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| DeepSWE v1.1 | 67.8%unknown effort | All 1 recorded result & sources67.8% · raw 67.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure; harness Claude Code 2.1.207, 1M context, temp 1.0, top-p 0.95. https://naive.ai/en/research/ |
| Agents' Last Exam | 32.4%unknown effort | All 1 recorded result & sources32.4% · raw 32.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 ALE-CLI per card sources; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| Terminal-Bench 2.1 | 86.7%unknown effort | All 1 recorded result & sources86.7% · raw 86.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| SWE-bench Pro | 73.6%unknown effort | All 1 recorded result & sources73.6% · raw 73.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 PDF chart text and figure pixels read 73.6; the release page alt text says 68.8 and looks stale. Chart value used; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| ProgramBench | 17.5%unknown effort | All 1 recorded result & sources17.5% · raw 17.5 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Almost@1 per card; vendor chart figure. https://naive.ai/en/research/ |
| NL2Repo-Bench | 71.9%unknown effort | All 1 recorded result & sources71.9% · raw 71.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| FrontierSWE | 78.2%unknown effort | All 1 recorded result & sources78.2% · raw 78.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Dominance score per card; vendor chart figure. https://naive.ai/en/research/ |
| PostTrainBench | 37.5%unknown effort | All 1 recorded result & sources37.5% · raw 37.5 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure. https://naive.ai/en/research/ |
| MLE-bench-30 | 73.7%unknown effort | All 1 recorded result & sources73.7% · raw 73.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Average position score per Gemini 3.6 Flash card protocol; vendor chart figure. https://naive.ai/en/research/ |
| PaperBench | 63.2%unknown effort | All 1 recorded result & sources63.2% · raw 63.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure. https://naive.ai/en/research/ |
| SOL-ExecBench | 72.8 ptsunknown effort | All 1 recorded result & sources72.8 pts · raw 72.81 pts Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Mean SOL score; in-house AutoResearch harness; peer Recursive Superintelligence 67.56 is not a table model. https://naive.ai/en/research/ |
| NanoChat AutoResearch | 0.9 BPBunknown effort | All 1 recorded result & sources0.9 BPB · raw 0.9051 BPB Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Validation BPB, lower is better; in-house AutoResearch harness. https://naive.ai/en/research/ |
| NanoGPT SpeedRun | 73.8 sunknown effort | All 1 recorded result & sources73.8 s · raw 73.8 s Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Training time seconds, lower is better; in-house harness. https://naive.ai/en/research/ |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
No independent evaluator results are recorded for this model.
Read how we select and source scores or the comparison guide.