Examenos

Official vendor benches · independent evaluators · OpenRouter pricing

Model profilesv 06df80d

Mistral Large 4

Explore Mistral Large 4 from Mistral AI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Mistral Large 4 with other models →Calculate costExplore data coverage

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default$0.68$2.09$0.07—2026-10-06 · OpenRouter source
Recorded pricing notes

Min-healthy endpoint minima: Mistral, status 0, 30m uptime 99.8473%. Published endpoint prices include the current reduction. Discount flag 0.5 is undocumented and not multiplied again. 524288-token context; 262144 max output. No contributor sibling on 2026-10-06.

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
DeepSWE v1.161.7%unknown effort
All 2 recorded results & sources

62% · raw 62 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 0; as reported by Mistral. Harness: Artificial Analysis private prelaunch evaluation; peer harnesses: Claude Code (Qwen), Cloak (DeepSeek), OpenCode (GLM), Kimi Code CLI (Kimi). Demoted 2026-10-06: precise own launch text from https://mistral.ai/news/mistral-large-4 takes precedence.

https://mistral.ai/_astro/artificial-analysis---deepswe-1.1%201_Z1nS7pb.webp?dpl=6ac4fc2e64b64f0a0d45c1ff

61.7% · raw 61.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

AA private prelaunch evaluation; article is more precise than chart 62%.

https://mistral.ai/news/mistral-large-4
Terminal-Bench 4.028.3%unknown effort
All 2 recorded results & sources

28.3% · raw 28.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 1; as reported by Mistral. Harness: Artificial Analysis private prelaunch evaluation. Summary figure; conflicts with detailed chart for some peers; do not average.

https://mistral.ai/_astro/artificial-analysis---terminal-bench-4.0%201_B8JMR.webp?dpl=6ac4fc2e64b64f0a0d45c1ff

28% · raw 28 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 10; as reported by Mistral. Harness: AA private prelaunch evaluation; detailed chart differs from summary chart (Kimi 21 vs 12.9, GLM 40 vs 41.9). Configuration is unspecified; retain discrepancy. Alternative evidence; existing headline retained under source/configuration precedence. No averaging.

https://mistral.ai/_astro/code-benchmarks---terminal-bench-4%201_Z1s1c31.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
AA Cyber Index (vendor-reported successes)50%unknown effort
All 1 recorded result & sources

50% · raw 50 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 2; as reported by Mistral. Successes, not safety blocks; third-party evaluation reported by Mistral.

https://mistral.ai/_astro/cybersecurity-benchmarks---aa-cyber-index%201_Z8sIQc.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
AutomationBench59.9%unknown effort
All 1 recorded result & sources

59.9% · raw 59.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 3; as reported by Mistral. 657 business workflows; AA harness; version not stated. Distinct from public split and v1.0.6.

https://mistral.ai/_astro/artificial-analysis---automationbench%201_Z2vplmz.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
Vals Finance Agent v254.7%unknown effort
All 1 recorded result & sources

54.7% · raw 54.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 4; as reported by Mistral. Vendor-reported Vals AI result; independently sourced precise results remain in evaluator columns.

https://mistral.ai/_astro/vals.ai---finance-agent-v2%201_18mXuR.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
Harvey Legal Agent Benchmark15.8%unknown effort
All 1 recorded result & sources

15.8% · raw 15.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 5; as reported by Mistral. Vendor-reported Vals AI HLAB; independently sourced precise results remain in evaluator columns.

https://mistral.ai/_astro/vals.ai---harvey's-legal-agent-benchmark%201_ZQeTGa.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
Dense200 (bounding boxes)42%unknown effort
All 1 recorded result & sources

42% · raw 42 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 6; as reported by Mistral. Bounding-box visual grounding. Article rounds GPT-6 Astra to 41%; use figure's 41.5%.

https://mistral.ai/_astro/dense200-(bbox)%201_2flnuY.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
CyberGym-E2E (AA harness, vendor-reported)82%unknown effort
All 1 recorded result & sources

82% · raw 82 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 7; as reported by Mistral. Vulnerability reproduction and patching; distinct from ordinary CyberGym.

https://mistral.ai/_astro/cybersecurity-benchmarks---cybergym-e2e-(aa)%201_Z2dNL7P.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
Cybench93%unknown effort
All 1 recorded result & sources

93% · raw 93 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 9; as reported by Mistral. 40 security-competition challenges; rounded vendor figure.

https://mistral.ai/_astro/cybersecurity-benchmarks---cybench%201_1dsNEh.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
SWE-Atlas QnA59.4%unknown effort
All 2 recorded results & sources

59% · raw 59 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 11; as reported by Mistral. Harness: AA private prelaunch evaluation; source does not establish equivalence to Codebase QnA launch split. Demoted 2026-10-06: precise own launch text from https://mistral.ai/news/mistral-large-4 takes precedence.

https://mistral.ai/_astro/code-benchmarks---swe-atlas-qna%201_Z26BdIm.webp?dpl=6ac4fc2e64b64f0a0d45c1ff

59.4% · raw 59.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

AA private prelaunch evaluation; article is more precise than chart 59%.

https://mistral.ai/news/mistral-large-4
ChartQA Pro63.1%unknown effort
All 1 recorded result & sources

63.1% · raw 63.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 13; as reported by Mistral. Chart question answering; exact Pro revision.

https://mistral.ai/_astro/multimodal-benchmarks---chartqa-pro%201_1sTH6.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
GDP.PDF18.6%unknown effort
All 1 recorded result & sources

18.6% · raw 18.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 14; as reported by Mistral. AA all-pass document benchmark reported in Mistral launch.

https://mistral.ai/_astro/multimodal-benchmarks---gdp.pdf---aa%201_1TQNQM.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
SciCode-Verified (pass@1, n=6)91.8%unknown effort
All 1 recorded result & sources

91.8% · raw 91.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 15; as reported by Mistral. Verified scientific-code workflows, pass@1, n=6; distinct from original SciCode.

https://mistral.ai/_astro/scicode-verified-pass@1-(n_6)-alt_ZYLgs8.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
FinWorkBench (Finch)67.4%unknown effort
All 1 recorded result & sources

67.4% · raw 67.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 17; as reported by Mistral. Spreadsheet creation/editing for finance/accounting; vendor-reported Finch.

https://mistral.ai/_astro/finch-(finworkbench)%201_1fqgOj.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
B3 Agent Security (attack resistance)93.3%unknown effort
All 1 recorded result & sources

93.3% · raw 93.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 18; as reported by Mistral. Lakera B3 indirect prompt-injection attack resistance; linked weak dataset.

https://mistral.ai/_astro/b3-agent-security-benchmark%201_Z8xhR4.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
Harmful cyber requests (Mistral refusal average)95.3%unknown effort
All 1 recorded result & sources

95.3% · raw 95.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 19; as reported by Mistral. Average refusal across JailbreakBench, StrongREJECT and AgentHarm; descriptive refusal rate, not cyber capability.

https://mistral.ai/_astro/refusal-of-harmful-cyber-requests%201_ZTfDmf.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
KORA aggregate1.7 scoreunknown effort
All 2 recorded results & sources

1.7 score · raw 1.7 score

Alternative · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 20; as reported by Mistral. 0–2 aggregate, NOT 0–1 or percent. Chart rounds to one decimal; article states Mistral 1.691. Demoted 2026-10-06: precise own launch text from https://mistral.ai/news/mistral-large-4 takes precedence.

https://mistral.ai/_astro/kora-benchmark%201_1iS6Sj.webp?dpl=6ac4fc2e64b64f0a0d45c1ff

1.7 score · raw 1.691 score

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Native 0–2 scale; article is more precise than rounded chart 1.7.

https://mistral.ai/news/mistral-large-4
ML4 vs GLM-5.3: STEM preference68%unknown effort
All 2 recorded results & sources

68% · raw 68 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation, weighted preference against GLM-5.3 only. Detailed STEM figure reports 69%, retained as alternative.

https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff

69% · raw 69 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 16; as reported by Mistral. Detailed STEM weighted win rate; overview reports 68%. Category shares: much better 7%, better 43%, slightly better 15%, tie 22%, slightly worse 14%; rounded shares, do not recompute. Alternative evidence; existing headline retained under source/configuration precedence. No averaging.

https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-stem-win-rate-breakdown%201_Z21ewlj.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
ML4 vs GLM-5.3: CAD preference62%unknown effort
All 1 recorded result & sources

62% · raw 62 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation; weighted preference against GLM-5.3 only.

https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
ML4 vs GLM-5.3: finance preference50%unknown effort
All 1 recorded result & sources

50% · raw 50 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation; weighted preference against GLM-5.3 only.

https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
ML4 vs GLM-5.3: coding preference48%unknown effort
All 1 recorded result & sources

48% · raw 48 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

Mistral launch figure 21; as reported by Mistral. Internal expert human evaluation; weighted preference against GLM-5.3 only.

https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-weighted-win-rate-by-domain%201_2urKvC.webp?dpl=6ac4fc2e64b64f0a0d45c1ff
AA Coding Agent Index (vendor-reported)49.8%unknown effort
All 1 recorded result & sources

49.8% · raw 49.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

As reported by Mistral. AA private prelaunch harness aggregate, as reported by Mistral; separate from evaluator source.

https://mistral.ai/news/mistral-large-4
AA-Briefcase (vendor-reported)1393unknown effort
All 1 recorded result & sources

1393 · raw 1393 Elo

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

As reported by Mistral. Long-horizon knowledge work, vendor-rounded Elo; separate from independently sourced AA value.

https://mistral.ai/news/mistral-large-4
Surge coding quality (Mistral blind study)3.7 scoreunknown effort
All 1 recorded result & sources

3.7 score · raw 3.74 score

Headline · unknown effort · Own vendor

Source/record date: 2026-10-06

As reported by Mistral. Blind professional annotator rating on 1–5 scale, five models; vendor-published Surge study, not a universal ranking.

https://mistral.ai/news/mistral-large-4

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis38.4unknown effort
All 1 recorded result & sources

38.4 · raw 38.3770408917494 index

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published non-estimated AA index; exact preview. Reasoning enabled, effort not specified.

https://artificialanalysis.ai/models/mistral-large-4
Cost per Intelligence Index task · Artificial Analysis$1.13unknown effort
All 1 recorded result & sources

$1.13 · raw 1.1317590533008346 USD

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Evaluator observed cost/task; not API list price.

https://artificialanalysis.ai/models/mistral-large-4
Output speed · Artificial Analysis116.1 tok/sunknown effort
All 1 recorded result & sources

116.1 tok/s · raw 116.080085205733 tok/s

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA median output speed via Mistral; independent from OpenRouter speed snapshot.

https://artificialanalysis.ai/models/mistral-large-4
AA-Briefcase Elo · Artificial Analysis1392.5unknown effort
All 1 recorded result & sources

1392.5 · raw 1392.53 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Native published scale.

https://artificialanalysis.ai/models/mistral-large-4
GDPval-AA Elo · Artificial Analysis1423.9unknown effort
All 1 recorded result & sources

1423.9 · raw 1423.91 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Native published scale.

https://artificialanalysis.ai/models/mistral-large-4
AutomationBench (AA) · Artificial Analysis59.9%unknown effort
All 1 recorded result & sources

59.9% · raw 59.901362468087 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
Terminal-Bench 4.0 (AA) · Artificial Analysis26.8%unknown effort
All 1 recorded result & sources

26.8% · raw 26.767676767677 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
SciCode (AA) · Artificial Analysis54.2%unknown effort
All 1 recorded result & sources

54.2% · raw 54.166666666667 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
Humanity’s Last Exam (AA) · Artificial Analysis35%unknown effort
All 1 recorded result & sources

35% · raw 35.032437442076 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
GDP.PDF (AA) · Artificial Analysis18.6%unknown effort
All 1 recorded result & sources

18.6% · raw 18.6 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
CritPt (AA) · Artificial Analysis10.6%unknown effort
All 1 recorded result & sources

10.6% · raw 10.571428571429 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
AA-Omniscience Index · Artificial Analysis-5.3 scoreunknown effort
All 1 recorded result & sources

-5.3 score · raw -5.3 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Native published scale.

https://artificialanalysis.ai/models/mistral-large-4
AA-LCR · Artificial Analysis81.3%unknown effort
All 1 recorded result & sources

81.3% · raw 81.333333333333 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA exact preview; Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
MMMU-Pro (AA) · Artificial Analysis76.4%unknown effort
All 1 recorded result & sources

76.4% · raw 76.416184971098 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published raw 0–1 fraction converted to percent; accuracy and hallucination are separate metrics.

https://artificialanalysis.ai/models/mistral-large-4
AA-Omniscience accuracy · Artificial Analysis25.8%unknown effort
All 1 recorded result & sources

25.8% · raw 25.816666666667 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published raw 0–1 fraction converted to percent; accuracy and hallucination are separate metrics.

https://artificialanalysis.ai/models/mistral-large-4
AA-Omniscience hallucination rate · Artificial Analysis41.9%unknown effort
All 1 recorded result & sources

41.9% · raw 41.945630195462 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published raw 0–1 fraction converted to percent; accuracy and hallucination are separate metrics.

https://artificialanalysis.ai/models/mistral-large-4
AA industry index: Finance and accounting · Artificial Analysis38.3 scoreunknown effort
All 1 recorded result & sources

38.3 score · raw 38.262600132251 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published AA industry index; native index scale, not percent.

https://artificialanalysis.ai/models/mistral-large-4
AA industry index: Strategy and operations · Artificial Analysis42.2 scoreunknown effort
All 1 recorded result & sources

42.2 score · raw 42.2092826635119 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published AA industry index; native index scale, not percent.

https://artificialanalysis.ai/models/mistral-large-4
AA industry index: Legal · Artificial Analysis37 scoreunknown effort
All 1 recorded result & sources

37 score · raw 36.9687229783312 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published AA industry index; native index scale, not percent.

https://artificialanalysis.ai/models/mistral-large-4
AA industry index: Engineering · Artificial Analysis37.6 scoreunknown effort
All 1 recorded result & sources

37.6 score · raw 37.5808314171772 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published AA industry index; native index scale, not percent.

https://artificialanalysis.ai/models/mistral-large-4
AA industry index: Economics · Artificial Analysis43.6 scoreunknown effort
All 1 recorded result & sources

43.6 score · raw 43.5568947713933 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published AA industry index; native index scale, not percent.

https://artificialanalysis.ai/models/mistral-large-4
AA-Briefcase rubric pass rate · Artificial Analysis45.2%unknown effort
All 1 recorded result & sources

45.2% · raw 45.151515151515 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

Published raw 0–1 fraction converted to percent.

https://artificialanalysis.ai/models/mistral-large-4
AA-Briefcase analytical quality Elo · Artificial Analysis1515.4unknown effort
All 1 recorded result & sources

1515.4 · raw 1515.38 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA-Briefcase component; native Elo.

https://artificialanalysis.ai/models/mistral-large-4
AA-Briefcase presentation Elo · Artificial Analysis1354.5unknown effort
All 1 recorded result & sources

1354.5 · raw 1354.54 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-06

AA-Briefcase component; native Elo.

https://artificialanalysis.ai/models/mistral-large-4
Vals Index · Vals AI48%high effort
All 1 recorded result & sources

48% · raw 48.045 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals Index v2.1; GDP-weighted finance/coding/legal/tax. Published accuracy already percent; no conversion.

https://www.vals.ai/benchmarks/vals_index
Vibe Code Bench v1.1 · Vals AI78.4%high effort
All 1 recorded result & sources

78.4% · raw 78.403 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. Harness: OpenHands.

https://www.vals.ai/benchmarks/vibe-code
Finance Agent v2 · Vals AI54.7%high effort
All 1 recorded result & sources

54.7% · raw 54.678 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000.

https://www.vals.ai/benchmarks/fabv2
Excel Modeling Benchmark · Vals AI56.2%high effort
All 1 recorded result & sources

56.2% · raw 56.213 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000.

https://www.vals.ai/benchmarks/emb
Terminal-Bench 4.0 · Vals AI22.7%high effort
All 1 recorded result & sources

22.7% · raw 22.727 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000.

https://www.vals.ai/benchmarks/terminal-bench-4
Code Migration · Vals AI30.6%high effort
All 1 recorded result & sources

30.6% · raw 30.561 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000. Full standalone benchmark, not Index subset (22.799%).

https://www.vals.ai/benchmarks/code-migration
Legal Research Bench · Vals AI31.7%high effort
All 1 recorded result & sources

31.7% · raw 31.731 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000.

https://www.vals.ai/benchmarks/legal_research
HLAB · Vals AI15.8%high effort
All 1 recorded result & sources

15.8% · raw 15.833 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000.

https://www.vals.ai/benchmarks/hlab
Tax Agent Bench · Vals AI63.3%high effort
All 1 recorded result & sources

63.3% · raw 63.294 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Vals primary standalone benchmark; accuracy already percent. Provider: Mistral AI; high effort, temperature 1, top_p 0.95, max output 256000.

https://www.vals.ai/benchmarks/tax_agent_bench
Code Migration (Vals Index subset) · Vals AI22.8%high effort
All 1 recorded result & sources

22.8% · raw 22.799 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-06

Index subset: 50/120 CLI tasks plus 10 COBOL tasks, weighted 75%/25%; distinct from full standalone benchmark. Published percent.

https://www.vals.ai/benchmarks/vals_index

Read how we select and source scores or the comparison guide.