Examenos

Official vendor benches · independent evaluators · OpenRouter pricing

Model profilesv 61747bc

Step 5 Preview

Explore Step 5 Preview from StepFun: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Step 5 Preview with other models →Compare subscription plansExplore data coverage

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default$1$2.7$0.05—2026-10-08 · OpenRouter source
Recorded pricing notes

Healthy endpoint minima (status 0, uptime >=98%). No contributor sibling found on 8 October 2026. Cache-write price unpublished. OpenRouter fp8 provider; 1,000,000 context / 64,000 maximum output tokens. Input/output/cache rates are per million tokens.

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
GPQA Diamond93.5%high effort
All 1 recorded result & sources

93.5% · raw 93.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
Humanity's Last Exam46.5%high effort
All 1 recorded result & sources

46.5% · raw 46.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; No-tools HLE row; subset not restated in table. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
AA-LCR v1.188.3%high effort
All 1 recorded result & sources

88.3% · raw 88.3 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Explicit revision 1.1; kept separate from unspecified AA-LCR. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
CritPt20.9%high effort
All 1 recorded result & sources

20.9% · raw 20.9 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
DeepSWE v1.167.7%high effort
All 1 recorded result & sources

67.7% · raw 67.7 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; SWE-agent harness, temperature=1.0, top_p=0.95; not Datacurve mini-swe-agent leaderboard. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
Terminal-Bench 2.185%high effort
All 1 recorded result & sources

85% · raw 85 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Explicit v2.1. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
Terminal-Bench 4.033.3%high effort
All 1 recorded result & sources

33.3% · raw 33.3 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Explicit v4; not Vals, Hermes or Anthropic harness results. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
CyberGym84.7%high effort
All 1 recorded result & sources

84.7% · raw 84.7 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
SciCode58.9%high effort
All 1 recorded result & sources

58.9% · raw 58.9 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
RoadmapBench54.3%high effort
All 1 recorded result & sources

54.3% · raw 54.3 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Version, harness and task count unpublished. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
ProgramBench (StepFun pass rate)80.5%high effort
All 1 recorded result & sources

80.5% · raw 80.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Pass Rate as labelled, not Almost@1 and not Anthropic 166-task reference-filtered subset; denominator/filter unspecified. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
SWE-Marathon v1.1 (partial score)72.7%high effort
All 1 recorded result & sources

72.7% · raw 72.7 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Explicit v1.1 partial credit; distinct from unspecified SWE-Marathon. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
MLS-Bench-Lite40.5%high effort
All 1 recorded result & sources

40.5% · raw 40.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Lite subset. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
SWE-Atlas QnA63.6%high effort
All 1 recorded result & sources

63.6% · raw 63.6 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; QnA track. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
SWE Atlas Test Writing50.8%high effort
All 1 recorded result & sources

50.8% · raw 50.8 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Test-writing track. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
StepCodeBench49%high effort
All 1 recorded result & sources

49% · raw 49 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Internal suite: 553 repositories, 9 task categories, 20 domains, 33 languages; avg@4. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
StepCode-Bench-Daily64.9%high effort
All 1 recorded result & sources

64.9% · raw 64.9 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Internal suite, separate daily result; task count/scoring detail unpublished. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
StepCode-Bench-General65%high effort
All 1 recorded result & sources

65% · raw 65 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Internal suite, separate general result; task count/scoring detail unpublished. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
GDPval-AA 2.11566high effort
All 1 recorded result & sources

1566 · raw 1566 Elo

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Native Elo; vendor quotes AA as of 20 September 2026. Not current independent AA decimal result. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
τ³-Bench Banking42.5%high effort
All 1 recorded result & sources

42.5% · raw 42.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Banking domain, not tau2 or telecom. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
AutomationBench-AA (vendor-reported)51%high effort
All 1 recorded result & sources

51% · raw 51 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; AA-labelled suite as quoted by StepFun; independent observations remain under AA evaluator. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
AutomationBench (public split)44%high effort
All 1 recorded result & sources

44% · raw 44 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Explicit public split; not AA suite or versioned full benchmark. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
AA Briefcase v1.11433high effort
All 1 recorded result & sources

1433 · raw 1433 Elo

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Native vendor Elo; not independent AA live value. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
Toolathlon-Verified74.1%high effort
All 1 recorded result & sources

74.1% · raw 74.1 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Verified variant. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
MCP Atlas85.6%high effort
All 1 recorded result & sources

85.6% · raw 85.6 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Vendor table; separate from Scale evaluator. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
PresentBench76.8%high effort
All 1 recorded result & sources

76.8% · raw 76.8 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Task count, revision and rubric unpublished. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
OfficeQA Pro60.3%high effort
All 1 recorded result & sources

60.3% · raw 60.3 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Pro subset; retrieval/document preprocessing unspecified. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
SpeadSheet v2 (as reported by StepFun)29.4%high effort
All 1 recorded result & sources

29.4% · raw 29.4 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Exact source spelling SpeadSheet v2; do not silently equate to SpreadsheetBench 2 without primary clarification. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
JobBench59%high effort
All 1 recorded result & sources

59% · raw 59 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
APEX-Agents37.8%high effort
All 1 recorded result & sources

37.8% · raw 37.8 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
DRACO83.3%high effort
All 1 recorded result & sources

83.3% · raw 83.3 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Grader unspecified; do not assume Anthropic Opus 4.6 grading. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
BrowseComp88.7%high effort
All 1 recorded result & sources

88.7% · raw 88.7 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
HLE with tools (text-only)59.4%high effort
All 1 recorded result & sources

59.4% · raw 59.4 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Text-only subset, explicitly distinguished by source footnote; not full HLE with tools. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
FinStepBench LiveSearch74.5%high effort
All 1 recorded result & sources

74.5% · raw 74.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Internal financial information retrieval/verification suite. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
FinStepBench CorporateValuation60.6%high effort
All 1 recorded result & sources

60.6% · raw 60.6 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Internal forecast/valuation suite. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
FinStepBench DeepResearch55.8%high effort
All 1 recorded result & sources

55.8% · raw 55.8 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Table label FinanceDR, corresponding DeepResearch chart uses identical values. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
FrontierFinance66.4%high effort
All 1 recorded result & sources

66.4% · raw 66.4 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; External benchmark; vendor run, 220 questions and 11,543 criteria across six investment use cases. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
Agents' Last Exam (ALE-CLI)29.5%high effort
All 1 recorded result & sources

29.5% · raw 29.5 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; CLI subset, distinct from full ALE. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
MMMU Pro (no tools)76%high effort
All 1 recorded result & sources

76% · raw 76 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Vendor score 76.0, not independent AA 76.36. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
GDP.PDF14.8%high effort
All 1 recorded result & sources

14.8% · raw 14.8 %

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; All-pass professional document reasoning. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/step-5-preview
GDPval-AA v21571high effort
All 1 recorded result & sources

1571 · raw 1571 Elo

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Dated launch image; GDPval-AA v2 as of 19 September. Current product table is v2.1. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/assets/step-5-preview-img03-BxUbgaBo.webp
MLA-512 H100 kernel optimization (best of 4)508 TFLOPShigh effort
All 1 recorded result & sources

508 TFLOPS · raw 508 TFLOPS

Headline · high effort · Own vendor

Source/record date: 2026-09-20

Harness: StepFun vendor-reported evaluation; Labelled peak achieved forward+backward TFLOPS, best of four 24-hour attempts; H100 MLA-512 case study. Full agent/grader/task-subset settings unpublished unless stated. Source inspected 8 October 2026; evaluation date unknown unless stated. Step 5 Preview high effort.

https://www.stepfun.com/assets/inference-combined-7fkNnlSY.png

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis43.7unknown effort
All 1 recorded result & sources

43.7 · raw 43.7343049141614 index

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Intelligence Index v4.3.2, ten evaluations.

https://artificialanalysis.ai/models/step-5
GDPval-AA Elo · Artificial Analysis1584.2unknown effort
All 1 recorded result & sources

1584.2 · raw 1584.24 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Native Elo/index.

https://artificialanalysis.ai/models/step-5
AutomationBench (AA) · Artificial Analysis51%unknown effort
All 1 recorded result & sources

51% · raw 50.956594190963 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
Terminal-Bench 4.0 (AA) · Artificial Analysis33.3%unknown effort
All 1 recorded result & sources

33.3% · raw 33.333333333333 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
SciCode (AA) · Artificial Analysis58.9%unknown effort
All 1 recorded result & sources

58.9% · raw 58.912037037037 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100. Source marks this evaluation under review.

https://artificialanalysis.ai/models/step-5
Humanity’s Last Exam (AA) · Artificial Analysis46.5%unknown effort
All 1 recorded result & sources

46.5% · raw 46.478220574606 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
GDP.PDF (AA) · Artificial Analysis14.8%unknown effort
All 1 recorded result & sources

14.8% · raw 14.8 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
CritPt (AA) · Artificial Analysis20.9%unknown effort
All 1 recorded result & sources

20.9% · raw 20.857142857143 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100. Source marks this evaluation under review.

https://artificialanalysis.ai/models/step-5
AA-Omniscience Index · Artificial Analysis16.4 scoreunknown effort
All 1 recorded result & sources

16.4 score · raw 16.383333333333333 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Native Elo/index.

https://artificialanalysis.ai/models/step-5
AA-LCR · Artificial Analysis88.3%unknown effort
All 1 recorded result & sources

88.3% · raw 88.333333333333 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
AA-Omniscience accuracy · Artificial Analysis41.5%unknown effort
All 1 recorded result & sources

41.5% · raw 41.533333333333 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
AA-Omniscience hallucination rate · Artificial Analysis43%unknown effort
All 1 recorded result & sources

43% · raw 43.015963511973 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
Terminal-Bench Science 0.1 (AA) · Artificial Analysis2.4%unknown effort
All 1 recorded result & sources

2.4% · raw 2.380952380952 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
HLAB (AA) · Artificial Analysis93.4%unknown effort
All 1 recorded result & sources

93.4% · raw 93.443334780721 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100. Harvey criterion pass rate; not Vals HLAB task-level pass.

https://artificialanalysis.ai/models/step-5
MMMU-Pro (AA) · Artificial Analysis76.4%unknown effort
All 1 recorded result & sources

76.4% · raw 76.35838150289 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
MLCR-AA · Artificial Analysis16.7%unknown effort
All 1 recorded result & sources

16.7% · raw 16.666666666667 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
ITBench-AA · Artificial Analysis55.6%unknown effort
All 1 recorded result & sources

55.6% · raw 55.555555555556 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
AA-AnalystAgent · Artificial Analysis35%unknown effort
All 1 recorded result & sources

35% · raw 35 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100. pass^5; not pass@1.

https://artificialanalysis.ai/models/step-5
EnterpriseOps-Gym-AA · Artificial Analysis47.2%unknown effort
All 1 recorded result & sources

47.2% · raw 47.179946284691 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
APEX-Agents (AA) · Artificial Analysis38%unknown effort
All 1 recorded result & sources

38% · raw 37.979351032448 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published raw 0–1 fraction ×100.

https://artificialanalysis.ai/models/step-5
AA-Briefcase Elo · Artificial Analysis1425.2unknown effort
All 1 recorded result & sources

1425.2 · raw 1425.23 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. AA-Briefcase v1.1; rubric raw fraction ×100, native Elo components otherwise.

https://artificialanalysis.ai/models/step-5
AA-Briefcase rubric pass rate · Artificial Analysis47.8%unknown effort
All 1 recorded result & sources

47.8% · raw 47.777777777778 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. AA-Briefcase v1.1; rubric raw fraction ×100, native Elo components otherwise.

https://artificialanalysis.ai/models/step-5
AA-Briefcase analytical quality Elo · Artificial Analysis1543.8unknown effort
All 1 recorded result & sources

1543.8 · raw 1543.83 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. AA-Briefcase v1.1; rubric raw fraction ×100, native Elo components otherwise.

https://artificialanalysis.ai/models/step-5
AA-Briefcase presentation Elo · Artificial Analysis1394.1unknown effort
All 1 recorded result & sources

1394.1 · raw 1394.11 Elo

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. AA-Briefcase v1.1; rubric raw fraction ×100, native Elo components otherwise.

https://artificialanalysis.ai/models/step-5
AA industry index: Finance and accounting · Artificial Analysis45.4 scoreunknown effort
All 1 recorded result & sources

45.4 score · raw 45.3728047646677 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published capability composite, not percent.

https://artificialanalysis.ai/models/step-5
AA industry index: Strategy and operations · Artificial Analysis47.1 scoreunknown effort
All 1 recorded result & sources

47.1 score · raw 47.1241657884068 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published capability composite, not percent.

https://artificialanalysis.ai/models/step-5
AA industry index: Legal · Artificial Analysis46.2 scoreunknown effort
All 1 recorded result & sources

46.2 score · raw 46.2280066508375 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published capability composite, not percent.

https://artificialanalysis.ai/models/step-5
AA industry index: Healthcare and medical · Artificial Analysis40.8 scoreunknown effort
All 1 recorded result & sources

40.8 score · raw 40.8171769208551 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published capability composite, not percent.

https://artificialanalysis.ai/models/step-5
AA industry index: Engineering · Artificial Analysis45 scoreunknown effort
All 1 recorded result & sources

45 score · raw 44.9763545147623 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published capability composite, not percent.

https://artificialanalysis.ai/models/step-5
AA industry index: Economics · Artificial Analysis52.8 scoreunknown effort
All 1 recorded result & sources

52.8 score · raw 52.8491063677788 score

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Published capability composite, not percent.

https://artificialanalysis.ai/models/step-5
Output speed · Artificial Analysis87.6 tok/sunknown effort
All 1 recorded result & sources

87.6 tok/s · raw 87.5504243985988 tok/s

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Median output speed on StepFun provider; distinct from live OpenRouter throughput.

https://artificialanalysis.ai/models/step-5
Cost per Intelligence Index task · Artificial Analysis$1.026/taskunknown effort
All 1 recorded result & sources

$1.026/task · raw 1.0264068637684054 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Observed weighted suite cost/task at AA measured price/cache settings, not a token price or guaranteed request cost.

https://artificialanalysis.ai/models/step-5
AA-Briefcase (AA) cost per task · Artificial Analysis$2.373/taskunknown effort
All 1 recorded result & sources

$2.373/task · raw 2.372851216942416 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
GDPval-AA (AA) cost per task · Artificial Analysis$0.979/taskunknown effort
All 1 recorded result & sources

$0.979/task · raw 0.979128029896363 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
AutomationBench-AA (AA) cost per task · Artificial Analysis$0.225/taskunknown effort
All 1 recorded result & sources

$0.225/task · raw 0.2247917635599391 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
Terminal-Bench 4.0 (AA) cost per task · Artificial Analysis$5.175/taskunknown effort
All 1 recorded result & sources

$5.175/task · raw 5.175400077426664 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
SciCode (AA) cost per task · Artificial Analysis$0.02/taskunknown effort
All 1 recorded result & sources

$0.02/task · raw 0.02031152013888889 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
Humanity's Last Exam (AA) cost per task · Artificial Analysis$0.065/taskunknown effort
All 1 recorded result & sources

$0.065/task · raw 0.06486026482854496 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
GDP.PDF (AA) cost per task · Artificial Analysis$0.122/taskunknown effort
All 1 recorded result & sources

$0.122/task · raw 0.122188105 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
CritPt (AA) cost per task · Artificial Analysis$0.158/taskunknown effort
All 1 recorded result & sources

$0.158/task · raw 0.15762519 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
AA-Omniscience (AA) cost per task · Artificial Analysis$0.015/taskunknown effort
All 1 recorded result & sources

$0.015/task · raw 0.015089091799999997 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
AA-LCR (AA) cost per task · Artificial Analysis$0.1/taskunknown effort
All 1 recorded result & sources

$0.1/task · raw 0.10049821099999999 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Artificial Analysis exact Step 5 Preview public model-page data, non-estimated. AA effort not labelled; do not infer high from vendor or medium from API default. Evaluation date unpublished; checked 8 October 2026. Source per-evaluation unweighted observed costPerTask; not weighted suite contribution, token rate or calculator assumption.

https://artificialanalysis.ai/models/step-5
Vals Index · Vals AI49.4%unknown effort
All 1 recorded result & sources

49.4% · raw 49.35 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±1.22 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Code Migration · Vals AI39%unknown effort
All 1 recorded result & sources

39% · raw 38.98 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±4.53 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Excel Modeling Benchmark · Vals AI57.1%unknown effort
All 1 recorded result & sources

57.1% · raw 57.12 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±3.17 percentage points. Source label EMB.

https://www.vals.ai/models/stepfun_step-5-preview
Finance Agent v2 · Vals AI50.7%unknown effort
All 1 recorded result & sources

50.7% · raw 50.67 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±0.32 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Legal Research Bench · Vals AI40.4%unknown effort
All 1 recorded result & sources

40.4% · raw 40.38 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±3.41 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
MysteryMechanism · Vals AI18.5%unknown effort
All 1 recorded result & sources

18.5% · raw 18.47 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±2.61 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
ProofBench v1.1 · Vals AI42%unknown effort
All 1 recorded result & sources

42% · raw 42 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±4.96 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Public Benefits Bench v1.1 · Vals AI58.1%unknown effort
All 1 recorded result & sources

58.1% · raw 58.12 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±1.28 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Tax Agent Bench · Vals AI59.5%unknown effort
All 1 recorded result & sources

59.5% · raw 59.54 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±3.09 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Vibe Code Bench v1.1 · Vals AI70.4%unknown effort
All 1 recorded result & sources

70.4% · raw 70.4 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±4.27 percentage points. Harness: OpenHands, verified on Vibe Code board.

https://www.vals.ai/models/stepfun_step-5-preview
BioMysteryBench · Vals AI67.4%unknown effort
All 1 recorded result & sources

67.4% · raw 67.41 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±0.37 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
HLAB · Vals AI9.2%unknown effort
All 1 recorded result & sources

9.2% · raw 9.17 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±2.44 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
IOI · Vals AI38.1%unknown effort
All 1 recorded result & sources

38.1% · raw 38.06 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±6.69 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Terminal-Bench 4.0 · Vals AI31.8%unknown effort
All 1 recorded result & sources

31.8% · raw 31.82 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±2.62 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Terminal-Bench Science · Vals AI2.9%unknown effort
All 1 recorded result & sources

2.9% · raw 2.86 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Published uncertainty ±2.01 percentage points.

https://www.vals.ai/models/stepfun_step-5-preview
Vals Index cost per task · Vals AI$2.626/taskunknown effort
All 1 recorded result & sources

$2.626/task · raw 2.626 USD/task

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-08

Vals exact Step 5 Preview model page; evaluator date unpublished, checked 8 October. Default provider StepFun, temperature 1, top_p 0.95, source max-output setting 1,024,000 (harness setting, not vendor product limit); suite overrides possible. Effort not published. Source Cost / Test; observed Vals Index task cost, no token-price derivation.

https://www.vals.ai/models/stepfun_step-5-preview

Read how we select and source scores or the comparison guide.