Mistral Large 4
Explore Mistral Large 4 from Mistral AI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Mistral Large 4 with other models →Calculate costExplore data coverage
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | $0.68 | $2.09 | $0.07 | — | 2026-10-06 · OpenRouter source |
Recorded pricing notes
Min-healthy endpoint minima: Mistral, status 0, 30m uptime 99.8473%. Published endpoint prices include the current reduction. Discount flag 0.5 is undocumented and not multiplied again. 524288-token context; 262144 max output. No contributor sibling on 2026-10-06.
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| DeepSWE v1.1 | 61.7%unknown effort | All 2 recorded results & sources62% · raw 62 % Alternative · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 0; as reported by Mistral. Harness: Artificial Analysis private prelaunch evaluation; peer harnesses: Claude Code (Qwen), Cloak (DeepSeek), OpenCode (GLM), Kimi Code CLI (Kimi). Demoted 2026-10-06: precise own launch text from https://mistral.ai/news/mistral-large-4 takes precedence. https://mistral.ai/_astro/artificial-analysis---deepswe-1.1%201_Z1nS7pb.webp?dpl=6ac4fc2e64b64f0a0d45c1ff61.7% · raw 61.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 AA private prelaunch evaluation; article is more precise than chart 62%. https://mistral.ai/news/mistral-large-4 |
| Terminal-Bench 4.0 | 28.3%unknown effort | All 2 recorded results & sources28.3% · raw 28.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 1; as reported by Mistral. Harness: Artificial Analysis private prelaunch evaluation. Summary figure; conflicts with detailed chart for some peers; do not average. https://mistral.ai/_astro/artificial-analysis---terminal-bench-4.0%201_B8JMR.webp?dpl=6ac4fc2e64b64f0a0d45c1ff28% · raw 28 % Alternative · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 10; as reported by Mistral. Harness: AA private prelaunch evaluation; detailed chart differs from summary chart (Kimi 21 vs 12.9, GLM 40 vs 41.9). Configuration is unspecified; retain discrepancy. Alternative evidence; existing headline retained under source/configuration precedence. No averaging. https://mistral.ai/_astro/code-benchmarks---terminal-bench-4%201_Z1s1c31.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| AA Cyber Index (vendor-reported successes) | 50%unknown effort | All 1 recorded result & sources50% · raw 50 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 2; as reported by Mistral. Successes, not safety blocks; third-party evaluation reported by Mistral. https://mistral.ai/_astro/cybersecurity-benchmarks---aa-cyber-index%201_Z8sIQc.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| AutomationBench | 59.9%unknown effort | All 1 recorded result & sources59.9% · raw 59.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 3; as reported by Mistral. 657 business workflows; AA harness; version not stated. Distinct from public split and v1.0.6. https://mistral.ai/_astro/artificial-analysis---automationbench%201_Z2vplmz.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| Vals Finance Agent v2 | 54.7%unknown effort | All 1 recorded result & sources54.7% · raw 54.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 4; as reported by Mistral. Vendor-reported Vals AI result; independently sourced precise results remain in evaluator columns. https://mistral.ai/_astro/vals.ai---finance-agent-v2%201_18mXuR.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| Harvey Legal Agent Benchmark | 15.8%unknown effort | All 1 recorded result & sources15.8% · raw 15.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 5; as reported by Mistral. Vendor-reported Vals AI HLAB; independently sourced precise results remain in evaluator columns. https://mistral.ai/_astro/vals.ai---harvey's-legal-agent-benchmark%201_ZQeTGa.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| Dense200 (bounding boxes) | 42%unknown effort | All 1 recorded result & sources42% · raw 42 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 6; as reported by Mistral. Bounding-box visual grounding. Article rounds GPT-6 Astra to 41%; use figure's 41.5%. https://mistral.ai/_astro/dense200-(bbox)%201_2flnuY.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| CyberGym-E2E (AA harness, vendor-reported) | 82%unknown effort | All 1 recorded result & sources82% · raw 82 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 7; as reported by Mistral. Vulnerability reproduction and patching; distinct from ordinary CyberGym. https://mistral.ai/_astro/cybersecurity-benchmarks---cybergym-e2e-(aa)%201_Z2dNL7P.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| Cybench | 93%unknown effort | All 1 recorded result & sources93% · raw 93 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 9; as reported by Mistral. 40 security-competition challenges; rounded vendor figure. https://mistral.ai/_astro/cybersecurity-benchmarks---cybench%201_1dsNEh.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| SWE-Atlas QnA | 59.4%unknown effort | All 2 recorded results & sources59% · raw 59 % Alternative · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 11; as reported by Mistral. Harness: AA private prelaunch evaluation; source does not establish equivalence to Codebase QnA launch split. Demoted 2026-10-06: precise own launch text from https://mistral.ai/news/mistral-large-4 takes precedence. https://mistral.ai/_astro/code-benchmarks---swe-atlas-qna%201_Z26BdIm.webp?dpl=6ac4fc2e64b64f0a0d45c1ff59.4% · raw 59.4 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 AA private prelaunch evaluation; article is more precise than chart 59%. https://mistral.ai/news/mistral-large-4 |
| ChartQA Pro | 63.1%unknown effort | All 1 recorded result & sources63.1% · raw 63.1 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 13; as reported by Mistral. Chart question answering; exact Pro revision. https://mistral.ai/_astro/multimodal-benchmarks---chartqa-pro%201_1sTH6.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| GDP.PDF | 18.6%unknown effort | All 1 recorded result & sources18.6% · raw 18.6 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 14; as reported by Mistral. AA all-pass document benchmark reported in Mistral launch. https://mistral.ai/_astro/multimodal-benchmarks---gdp.pdf---aa%201_1TQNQM.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| SciCode-Verified (pass@1, n=6) | 91.8%unknown effort | All 1 recorded result & sources91.8% · raw 91.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 15; as reported by Mistral. Verified scientific-code workflows, pass@1, n=6; distinct from original SciCode. https://mistral.ai/_astro/scicode-verified-pass@1-(n_6)-alt_ZYLgs8.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| FinWorkBench (Finch) | 67.4%unknown effort | All 1 recorded result & sources67.4% · raw 67.4 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 17; as reported by Mistral. Spreadsheet creation/editing for finance/accounting; vendor-reported Finch. https://mistral.ai/_astro/finch-(finworkbench)%201_1fqgOj.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| B3 Agent Security (attack resistance) | 93.3%unknown effort | All 1 recorded result & sources93.3% · raw 93.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 18; as reported by Mistral. Lakera B3 indirect prompt-injection attack resistance; linked weak dataset. https://mistral.ai/_astro/b3-agent-security-benchmark%201_Z8xhR4.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| Harmful cyber requests (Mistral refusal average) | 95.3%unknown effort | All 1 recorded result & sources95.3% · raw 95.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 19; as reported by Mistral. Average refusal across JailbreakBench, StrongREJECT and AgentHarm; descriptive refusal rate, not cyber capability. https://mistral.ai/_astro/refusal-of-harmful-cyber-requests%201_ZTfDmf.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| KORA aggregate | 1.7 scoreunknown effort | All 2 recorded results & sources1.7 score · raw 1.7 score Alternative · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 20; as reported by Mistral. 0–2 aggregate, NOT 0–1 or percent. Chart rounds to one decimal; article states Mistral 1.691. Demoted 2026-10-06: precise own launch text from https://mistral.ai/news/mistral-large-4 takes precedence. https://mistral.ai/_astro/kora-benchmark%201_1iS6Sj.webp?dpl=6ac4fc2e64b64f0a0d45c1ff1.7 score · raw 1.691 score Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Native 0–2 scale; article is more precise than rounded chart 1.7. https://mistral.ai/news/mistral-large-4 |
| ML4 vs GLM-5.3: STEM preference | 68%unknown effort | All 2 recorded results & sources68% · raw 68 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation, weighted preference against GLM-5.3 only. Detailed STEM figure reports 69%, retained as alternative. https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff69% · raw 69 % Alternative · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 16; as reported by Mistral. Detailed STEM weighted win rate; overview reports 68%. Category shares: much better 7%, better 43%, slightly better 15%, tie 22%, slightly worse 14%; rounded shares, do not recompute. Alternative evidence; existing headline retained under source/configuration precedence. No averaging. https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-stem-win-rate-breakdown%201_Z21ewlj.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| ML4 vs GLM-5.3: CAD preference | 62%unknown effort | All 1 recorded result & sources62% · raw 62 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation; weighted preference against GLM-5.3 only. https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| ML4 vs GLM-5.3: finance preference | 50%unknown effort | All 1 recorded result & sources50% · raw 50 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation; weighted preference against GLM-5.3 only. https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| ML4 vs GLM-5.3: coding preference | 48%unknown effort | All 1 recorded result & sources48% · raw 48 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation; weighted preference against GLM-5.3 only. https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff |
| AA Coding Agent Index (vendor-reported) | 49.8%unknown effort | All 1 recorded result & sources49.8% · raw 49.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-06 As reported by Mistral. AA private prelaunch harness aggregate, as reported by Mistral; separate from evaluator source. https://mistral.ai/news/mistral-large-4 |
| AA-Briefcase (vendor-reported) | 1393unknown effort | All 1 recorded result & sources1393 · raw 1393 Elo Headline · unknown effort · Own vendor Source/record date: 2026-10-06 As reported by Mistral. Long-horizon knowledge work, vendor-rounded Elo; separate from independently sourced AA value. https://mistral.ai/news/mistral-large-4 |
| Surge coding quality (Mistral blind study) | 3.7 scoreunknown effort | All 1 recorded result & sources3.7 score · raw 3.74 score Headline · unknown effort · Own vendor Source/record date: 2026-10-06 As reported by Mistral. Blind professional annotator rating on 1–5 scale, five models; vendor-published Surge study, not a universal ranking. https://mistral.ai/news/mistral-large-4 |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 38.4unknown effort | All 1 recorded result & sources38.4 · raw 38.3770408917494 index Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published non-estimated AA index; exact preview. Reasoning enabled, effort not specified. https://artificialanalysis.ai/models/mistral-large-4 |
| Cost per Intelligence Index task · Artificial Analysis | $1.13unknown effort | All 1 recorded result & sources$1.13 · raw 1.1317590533008346 USD Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Evaluator observed cost/task; not API list price. https://artificialanalysis.ai/models/mistral-large-4 |
| Output speed · Artificial Analysis | 116.1 tok/sunknown effort | All 1 recorded result & sources116.1 tok/s · raw 116.080085205733 tok/s Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA median output speed via Mistral; independent from OpenRouter speed snapshot. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Briefcase Elo · Artificial Analysis | 1392.5unknown effort | All 1 recorded result & sources1392.5 · raw 1392.53 Elo Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Native published scale. https://artificialanalysis.ai/models/mistral-large-4 |
| GDPval-AA Elo · Artificial Analysis | 1423.9unknown effort | All 1 recorded result & sources1423.9 · raw 1423.91 Elo Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Native published scale. https://artificialanalysis.ai/models/mistral-large-4 |
| AutomationBench (AA) · Artificial Analysis | 59.9%unknown effort | All 1 recorded result & sources59.9% · raw 59.901362468087 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| Terminal-Bench 4.0 (AA) · Artificial Analysis | 26.8%unknown effort | All 1 recorded result & sources26.8% · raw 26.767676767677 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| SciCode (AA) · Artificial Analysis | 54.2%unknown effort | All 1 recorded result & sources54.2% · raw 54.166666666667 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| Humanity’s Last Exam (AA) · Artificial Analysis | 35%unknown effort | All 1 recorded result & sources35% · raw 35.032437442076 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| GDP.PDF (AA) · Artificial Analysis | 18.6%unknown effort | All 1 recorded result & sources18.6% · raw 18.6 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| CritPt (AA) · Artificial Analysis | 10.6%unknown effort | All 1 recorded result & sources10.6% · raw 10.571428571429 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Omniscience Index · Artificial Analysis | -5.3 scoreunknown effort | All 1 recorded result & sources-5.3 score · raw -5.3 score Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Native published scale. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-LCR · Artificial Analysis | 81.3%unknown effort | All 1 recorded result & sources81.3% · raw 81.333333333333 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA exact preview; Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| MMMU-Pro (AA) · Artificial Analysis | 76.4%unknown effort | All 1 recorded result & sources76.4% · raw 76.416184971098 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published raw 0–1 fraction converted to percent; accuracy and hallucination are separate metrics. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Omniscience accuracy · Artificial Analysis | 25.8%unknown effort | All 1 recorded result & sources25.8% · raw 25.816666666667 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published raw 0–1 fraction converted to percent; accuracy and hallucination are separate metrics. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Omniscience hallucination rate · Artificial Analysis | 41.9%unknown effort | All 1 recorded result & sources41.9% · raw 41.945630195462 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published raw 0–1 fraction converted to percent; accuracy and hallucination are separate metrics. https://artificialanalysis.ai/models/mistral-large-4 |
| AA industry index: Finance and accounting · Artificial Analysis | 38.3 scoreunknown effort | All 1 recorded result & sources38.3 score · raw 38.262600132251 score Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published AA industry index; native index scale, not percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA industry index: Strategy and operations · Artificial Analysis | 42.2 scoreunknown effort | All 1 recorded result & sources42.2 score · raw 42.2092826635119 score Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published AA industry index; native index scale, not percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA industry index: Legal · Artificial Analysis | 37 scoreunknown effort | All 1 recorded result & sources37 score · raw 36.9687229783312 score Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published AA industry index; native index scale, not percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA industry index: Engineering · Artificial Analysis | 37.6 scoreunknown effort | All 1 recorded result & sources37.6 score · raw 37.5808314171772 score Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published AA industry index; native index scale, not percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA industry index: Economics · Artificial Analysis | 43.6 scoreunknown effort | All 1 recorded result & sources43.6 score · raw 43.5568947713933 score Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published AA industry index; native index scale, not percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Briefcase rubric pass rate · Artificial Analysis | 45.2%unknown effort | All 1 recorded result & sources45.2% · raw 45.151515151515 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 Published raw 0–1 fraction converted to percent. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Briefcase analytical quality Elo · Artificial Analysis | 1515.4unknown effort | All 1 recorded result & sources1515.4 · raw 1515.38 Elo Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA-Briefcase component; native Elo. https://artificialanalysis.ai/models/mistral-large-4 |
| AA-Briefcase presentation Elo · Artificial Analysis | 1354.5unknown effort | All 1 recorded result & sources1354.5 · raw 1354.54 Elo Headline · unknown effort · Independent evaluator Source/record date: 2026-10-06 AA-Briefcase component; native Elo. https://artificialanalysis.ai/models/mistral-large-4 |
| Vals Index · Vals AI | 48%high effort | All 1 recorded result & sources48% · raw 48.045 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals Index v2.1; GDP-weighted finance/coding/legal/tax. Published accuracy already percent; no conversion. https://www.vals.ai/benchmarks/vals_index |
| Vibe Code Bench v1.1 · Vals AI | 78.4%high effort | All 1 recorded result & sources78.4% · raw 78.403 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. Harness: OpenHands. https://www.vals.ai/benchmarks/vibe-code |
| Finance Agent v2 · Vals AI | 54.7%high effort | All 1 recorded result & sources54.7% · raw 54.678 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. https://www.vals.ai/benchmarks/fabv2 |
| Excel Modeling Benchmark · Vals AI | 56.2%high effort | All 1 recorded result & sources56.2% · raw 56.213 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. https://www.vals.ai/benchmarks/emb |
| Terminal-Bench 4.0 · Vals AI | 22.7%high effort | All 1 recorded result & sources22.7% · raw 22.727 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. https://www.vals.ai/benchmarks/terminal-bench-4 |
| Code Migration · Vals AI | 30.6%high effort | All 1 recorded result & sources30.6% · raw 30.561 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. Full standalone benchmark, not Index subset (22.799%). https://www.vals.ai/benchmarks/code-migration |
| Legal Research Bench · Vals AI | 31.7%high effort | All 1 recorded result & sources31.7% · raw 31.731 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. https://www.vals.ai/benchmarks/legal_research |
| HLAB · Vals AI | 15.8%high effort | All 1 recorded result & sources15.8% · raw 15.833 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. https://www.vals.ai/benchmarks/hlab |
| Tax Agent Bench · Vals AI | 63.3%high effort | All 1 recorded result & sources63.3% · raw 63.294 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. https://www.vals.ai/benchmarks/tax_agent_bench |
| Code Migration (Vals Index subset) · Vals AI | 22.8%high effort | All 1 recorded result & sources22.8% · raw 22.799 % Headline · high effort · Independent evaluator Source/record date: 2026-10-06 Index subset: 50/120 CLI tasks plus 10 COBOL tasks, weighted 75%/25%; distinct from full standalone benchmark. Published percent. https://www.vals.ai/benchmarks/vals_index |
Read how we select and source scores or the comparison guide.