Kimi K3 Benchmarks, Specifications & Availability
Explore Kimi K3 from Moonshot AI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Kimi K3 with other models →Explore data coverage
Published specifications
- Provider
- Moonshot AI
- Access
- Open weights
- License
- Kimi K3 License
- Context window
- 1M
- Total parameters
- 2.8T
- Active parameters
- 104B
- Released
- 2026-07-16
- Modalities
- text, image, video
- Family
- Kimi K3
Model card · Announcement · Website · OpenRouter
Model notes
2.8T MoE / 104B active; native multimodal (text/image/video per OR; GH summary also lists text+image). Ignore OR :batch twin for main pricing. Tech report arXiv:2607.24653. Open weights verified on Hugging Face https://huggingface.co/moonshotai/Kimi-K3 (2026-09-25).
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | — | 2026-10-03 · OpenRouter source |
Recorded pricing notes
OR prompt/completion/input_cache_read ×1e6; ignore :batch twin. No -contribute sibling.; min-healthy endpoint minima 2026-10-01; discount flag 0.15 on Phala (undocumented, not applied); min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; discount flag 0.25 on Phala (undocumented, not applied); min-healthy endpoint minima 2026-10-03; discount flag 0.25 on Phala (undocumented, not applied)
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| GPQA Diamond | 93.5%max effort | All 1 recorded result & sources93.5% · raw 93.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 GPQA Diamond; reasoning_effort=max, temp=1.0, top_p=0.95 https://github.com/MoonshotAI/Kimi-K3 |
| CritPt | 23.4%max effort | All 1 recorded result & sources23.4% · raw 23.4 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 CritPt; Moonshot cites Artificial Analysis as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| AA-LCR | 74.7%max effort | All 1 recorded result & sources74.7% · raw 74.7 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 AA-LCR; Moonshot cites Artificial Analysis as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| Humanity's Last Exam | 43.5%max effort | All 1 recorded result & sources43.5% · raw 43.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 HLE-Full without tools (with tools 56.0 stored separately) https://github.com/MoonshotAI/Kimi-K3 |
| Humanity's Last Exam (w/ tools) | 56%max effort | All 1 recorded result & sources56% · raw 56 % Headline · max effort · Own vendor Source/record date: 2026-07-16 HLE-Full with general tools https://github.com/MoonshotAI/Kimi-K3 |
| DeepSWE v1.1 | 67.5%max effort | All 2 recorded results & sources67.5% · raw 67.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 DeepSWE v1.1 tasks; Kimi Code harness (mini-SWE-agent board reports 67.3) https://github.com/MoonshotAI/Kimi-K367.5% · raw 67.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| ProgramBench | 77.8%max effort | All 1 recorded result & sources77.8% · raw 77.8 % Headline · max effort · Own vendor Source/record date: 2026-07-16 ProgramBench; Kimi Code harness https://github.com/MoonshotAI/Kimi-K3 |
| Terminal-Bench 2.1 | 88.3%max effort | All 2 recorded results & sources88.3% · raw 88.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Terminal-Bench 2.1; Kimi Code harness https://github.com/MoonshotAI/Kimi-K388.3% · raw 88.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| FrontierSWE | 81.2%max effort | All 2 recorded results & sources81.2% · raw 81.2 % Headline · max effort · Own vendor Source/record date: 2026-07-16 FrontierSWE dominance; Kimi Code harness; as of 2026-07-16 https://github.com/MoonshotAI/Kimi-K381.2% · raw 81.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| SWE-Marathon | 42%max effort | All 1 recorded result & sources42% · raw 42 % Headline · max effort · Own vendor Source/record date: 2026-07-16 SWE-Marathon; Claude Code harness; H20-calibrated branch pre-v1.1 https://github.com/MoonshotAI/Kimi-K3 |
| PostTrainBench | 36.6%max effort | All 2 recorded results & sources36.6% · raw 36.6 % Headline · max effort · Own vendor Source/record date: 2026-07-16 PostTrainBench; Harbor + Claude Code harness; avg of 3 on H20 https://github.com/MoonshotAI/Kimi-K336.6% · raw 36.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| MLS-Bench-Lite | 48.3%max effort | All 1 recorded result & sources48.3% · raw 48.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MLS-Bench-Lite; Kimi Code harness https://github.com/MoonshotAI/Kimi-K3 |
| SciCode | 58.7%max effort | All 1 recorded result & sources58.7% · raw 58.7 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 SciCode; Moonshot cites Artificial Analysis as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| Kimi Code Bench 2.0 | 72.9%max effort | All 1 recorded result & sources72.9% · raw 72.9 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Kimi Code Bench 2.0 (in-house); Kimi Code harness (73.7 with Claude Code noted) https://github.com/MoonshotAI/Kimi-K3 |
| BrowseComp | 91.2%max effort | All 1 recorded result & sources91.2% · raw 91.2 % Headline · max effort · Own vendor Source/record date: 2026-07-16 BrowseComp with context-compaction at 300K (90.4 with full 1M / no compaction) https://github.com/MoonshotAI/Kimi-K3 |
| DeepSearchQA | 95%max effort | All 1 recorded result & sources95% · raw 95 % Headline · max effort · Own vendor Source/record date: 2026-07-16 DeepSearchQA F1 https://github.com/MoonshotAI/Kimi-K3 |
| ResearchRubrics | 76.2%max effort | All 1 recorded result & sources76.2% · raw 76.2 % Headline · max effort · Own vendor Source/record date: 2026-07-16 ResearchRubrics https://github.com/MoonshotAI/Kimi-K3 |
| GDPval-AA v2 | 1686max effort | All 1 recorded result & sources1686 · raw 1686 Elo Headline · max effort · Peer vendor Source/record date: 2026-07-16 GDPval-AA v2 Elo; Moonshot cites Artificial Analysis as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| Toolathlon-Verified | 76.5%max effort | All 1 recorded result & sources76.5% · raw 76.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Toolathlon-Verified https://github.com/MoonshotAI/Kimi-K3 |
| MCPMark-Verified | 94.5%max effort | All 1 recorded result & sources94.5% · raw 94.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MCPMark-Verified https://github.com/MoonshotAI/Kimi-K3 |
| MCP Atlas | 84.2%max effort | All 1 recorded result & sources84.2% · raw 84.2 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MCP-Atlas 500-task public subset, 100-turn limit, Gemini 3.1 Pro judge https://github.com/MoonshotAI/Kimi-K3 |
| AutomationBench v1.0.6 | 30.8%max effort | All 1 recorded result & sources30.8% · raw 30.8 % Headline · max effort · Own vendor Source/record date: 2026-07-16 AutomationBench 600-task public subset (mapped to automationbench-v1.0.6) https://github.com/MoonshotAI/Kimi-K3 |
| JobBench | 54.3%max effort | All 1 recorded result & sources54.3% · raw 54.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 JobBench https://github.com/MoonshotAI/Kimi-K3 |
| AA Briefcase v1.1 | 1548max effort | All 1 recorded result & sources1548 · raw 1548 Elo Headline · max effort · Peer vendor Source/record date: 2026-07-16 AA-Briefcase Elo; Moonshot cites Artificial Analysis as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| Agents' Last Exam | 28.3%max effort | All 2 recorded results & sources28.3% · raw 28.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Agents' Last Exam primary pass-rate; Kimi Code harness; from official leaderboard as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K328.3% · raw 28.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| APEX-Agents | 41%max effort | All 1 recorded result & sources41% · raw 41 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 APEX-Agents; Moonshot cites APEX leaderboard as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| OfficeQA Pro | 63.3%max effort | All 1 recorded result & sources63.3% · raw 63.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 OfficeQA Pro; entire PDF corpus as images; Claude Code harness https://github.com/MoonshotAI/Kimi-K3 |
| SpreadsheetBench 2 | 34.8%max effort | All 1 recorded result & sources34.8% · raw 34.8 % Headline · max effort · Own vendor Source/record date: 2026-07-16 SpreadsheetBench 2; Claude Code harness https://github.com/MoonshotAI/Kimi-K3 |
| OSWorld-Verified | 84.8%max effort | All 1 recorded result & sources84.8% · raw 84.8 % Headline · max effort · Own vendor Source/record date: 2026-07-16 OSWorld-Verified https://github.com/MoonshotAI/Kimi-K3 |
| OSWorld 2.0 | 58.3%max effort | All 1 recorded result & sources58.3% · raw 58.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 OSWorld 2.0 https://github.com/MoonshotAI/Kimi-K3 |
| SaaS-Bench | 60.1%max effort | All 1 recorded result & sources60.1% · raw 60.1 % Headline · max effort · Own vendor Source/record date: 2026-07-16 SaaS-Bench https://github.com/MoonshotAI/Kimi-K3 |
| τ³-Bench Banking | 33.4%max effort | All 1 recorded result & sources33.4% · raw 33.4 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 τ³-Banking; Moonshot cites Artificial Analysis as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| Harvey Legal Agent | 94.6%max effort | All 1 recorded result & sources94.6% · raw 94.6 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 Harvey Lab-AA criterion pass rate; Moonshot cites AA as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K3 |
| CorpFin v2 | 71.6%max effort | All 1 recorded result & sources71.6% · raw 71.6 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 CorpFin v2; Moonshot cites Vals AI https://github.com/MoonshotAI/Kimi-K3 |
| Vals Finance Agent v2 | 54.4%max effort | All 1 recorded result & sources54.4% · raw 54.4 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 Finance Agent v2; Moonshot cites Vals AI https://github.com/MoonshotAI/Kimi-K3 |
| Legal Research Bench | 44.2%max effort | All 1 recorded result & sources44.2% · raw 44.2 % Headline · max effort · Peer vendor Source/record date: 2026-07-16 Legal Research Bench; Moonshot cites Vals AI https://github.com/MoonshotAI/Kimi-K3 |
| WorldVQA | 51%max effort | All 1 recorded result & sources51% · raw 51 % Headline · max effort · Own vendor Source/record date: 2026-07-16 WorldVQA ForceAnswer https://github.com/MoonshotAI/Kimi-K3 |
| OmniDocBench 1.5 | 91.1%max effort | All 1 recorded result & sources91.1% · raw 91.1 % Headline · max effort · Own vendor Source/record date: 2026-07-16 OmniDocBench (mapped to omnidocbench-1.5; version not further specified on GH) https://github.com/MoonshotAI/Kimi-K3 |
| PerceptionBench | 58.5%max effort | All 1 recorded result & sources58.5% · raw 58.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 PerceptionBench (in-house atomic visual perception) https://github.com/MoonshotAI/Kimi-K3 |
| Video-MME | 90%max effort | All 1 recorded result & sources90% · raw 90 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Video-MME (w. subtitles) https://github.com/MoonshotAI/Kimi-K3 |
| MMVU | 82.1%max effort | All 1 recorded result & sources82.1% · raw 82.1 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MMVU https://github.com/MoonshotAI/Kimi-K3 |
| BabyVision (w/ tools) | 85.7%max effort | All 1 recorded result & sources85.7% · raw 85.7 % Headline · max effort · Own vendor Source/record date: 2026-07-16 BabyVision w/ python https://github.com/MoonshotAI/Kimi-K3 |
| MMMU Pro (no tools) | 81.6%max effort | All 1 recorded result & sources81.6% · raw 81.6 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MMMU-Pro without tools (with tools 83.4 stored separately) https://github.com/MoonshotAI/Kimi-K3 |
| MMMU Pro (with tools) | 83.4%max effort | All 1 recorded result & sources83.4% · raw 83.4 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MMMU-Pro with Python tools https://github.com/MoonshotAI/Kimi-K3 |
| CharXiv RQ | 91.3%max effort | All 2 recorded results & sources91.3% · raw 91.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 CharXiv (RQ) with tools (without tools 84.8 noted) https://github.com/MoonshotAI/Kimi-K384.8% · raw 84.8 % Alternative · max effort · Own vendor Source/record date: 2026-07-16 CharXiv (RQ) without tools https://github.com/MoonshotAI/Kimi-K3 |
| MathVision | 97.8%max effort | All 2 recorded results & sources97.8% · raw 97.8 % Headline · max effort · Own vendor Source/record date: 2026-07-16 MathVision with tools (without tools 94.3 noted) https://github.com/MoonshotAI/Kimi-K394.3% · raw 94.3 % Alternative · max effort · Own vendor Source/record date: 2026-07-16 MathVision without tools https://github.com/MoonshotAI/Kimi-K3 |
| ZeroBench-main (w/ tools) | 41%max effort | All 1 recorded result & sources41% · raw 41 % Headline · max effort · Own vendor Source/record date: 2026-07-16 ZeroBench pass@5 with tools (without tools 23.0 noted) https://github.com/MoonshotAI/Kimi-K3 |
| Gray Swan IPI | 52.7%unknown effort | All 1 recorded result & sources52.7% · raw 52.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon Gray Swan chart (K=15 attack success rate) https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp |
| CyberGym | 80%unknown effort | All 1 recorded result & sources80% · raw 80 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| SWE-Pro | 64.8%unknown effort | All 1 recorded result & sources64.8% · raw 64.84 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 44max effort | All 2 recorded results & sources44 · raw 44 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/kimi-k330 · raw 30 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/kimi-k3-low |
| Output speed · Artificial Analysis | 34 tok/smax effort | All 2 recorded results & sources34 tok/s · raw 34 tok/s Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/kimi-k336 tok/s · raw 36 tok/s Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/kimi-k3-low |
| Vals Index · Vals AI | 50.3%max effort | All 1 recorded result & sources50.3% · raw 50.3 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $6.38. https://www.vals.ai/benchmarks/vals_index |
| Bugs fixed /105 · Bug Hunt Bench | 21 fixesunknown effort | All 1 recorded result & sources21 fixes · raw 21 fixes Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-07-26. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| ARC-AGI-1 · ARC Prize | 94.5%max effort | All 3 recorded results & sources94.5% · raw 94.5 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/moonshot-kimi-k386.7% · raw 86.7 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/moonshot-kimi-k365.7% · raw 65.7 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/moonshot-kimi-k3 |
| ARC-AGI-2 · ARC Prize | 60.4%max effort | All 3 recorded results & sources60.4% · raw 60.4 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/moonshot-kimi-k355% · raw 55 % Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/moonshot-kimi-k312.4% · raw 12.4 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/moonshot-kimi-k3 |
| Hard board % of roofline · KernelBench (community board) | 18.9%unknown effort | All 1 recorded result & sources18.9% · raw 18.9 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 Kimi K3 (1M) kinetic-claude; Hard 6/6; Mega 9.79× 1/1; CUDA 10.6% 4/4 https://kernelbench.com/models/kinetic-0715-1m |
| CUDA board % of roofline · KernelBench (community board) | 10.6%unknown effort | All 1 recorded result & sources10.6% · raw 10.6 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 Kimi K3 (1M); CUDA 4/4; Hard 18.9% already ingested https://kernelbench.com/models/kinetic-0715-1m |
| Money gain · Andon Labs | $4665.04unknown effort | All 2 recorded results & sources$4665.04 · raw 4665.04 $ Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Moonshot provider. Direct model-vendor deployment is headline among provider variants. VB2 final balance minus $500. Corrects prior plain Kimi 2693.75 gain, which came from the vb1 block rather than vb2. https://andonlabs.com/evals/vending-bench-2$4406.96 · raw 4406.96 $ Alternative · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Fireworks provider. VB2 final balance minus $500; alternate deployment. https://andonlabs.com/evals/vending-bench-2 |
| Blueprint Bench · Andon Labs | 29.5%unknown effort | All 1 recorded result & sources29.5% · raw 29.5 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100. https://andonlabs.com/evals/blueprint-bench-2 |
| Average Score · WeirdML v3 | 7.4%high effort | All 1 recorded result & sources7.4% · raw 7.42 % Headline · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Kimi K3 (high, Fireworks); harness kimi_code 2.0.0; values from prepared data JSON; raw 0.074209; official 80/20 aggregate (area 500k-50M tokens + final best) https://htihle.github.io/weirdml.html |
| Final Best Score · WeirdML v3 | 15.6%high effort | All 1 recorded result & sources15.6% · raw 15.62 % Headline · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Kimi K3 (high, Fireworks); harness kimi_code 2.0.0; values from prepared data JSON; raw 0.156217; mean final best effective score https://htihle.github.io/weirdml.html |
| Cost / Run · WeirdML v3 | $19.53high effort | All 1 recorded result & sources$19.53 · raw 19.53 USD Headline · high effort · Independent evaluator Source/record date: 2026-10-02 WeirdML variant Kimi K3 (high, Fireworks); harness kimi_code 2.0.0; values from prepared data JSON; mean API cost per run, same task weighting as scores https://htihle.github.io/weirdml.html |
| Vibe Code Bench v1.1 · Vals AI | 85%unknown effort | All 1 recorded result & sources85% · raw 84.97 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $10.01. https://www.vals.ai/benchmarks/vibe-code |
| Cost per Intelligence Index task · Artificial Analysis | $2.00max effort | All 2 recorded results & sources$2.00 · raw 2 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/kimi-k3$1.15 · raw 1.15 USD Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/kimi-k3-low |
Read how we select and source scores or the comparison guide.