Step 5 Preview

Step 5 Preview is a reasoning model from StepFun in the Step 5 family. 24 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.72.5 ±5.3
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.83/s
Input / 1MUS dollars per 1M input tokens.$1
Output / 1MUS dollars per 1M output tokens.$2.7
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
74.6
CodingCode writing and repair.
73.9
ReasoningLogic problems and puzzles.
75.8
MultimodalTasks with images and text.
61.4
KnowledgeFacts and expert knowledge.
70.0
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
GPQA DiamondKnowledge93.5%64.6David Rein et al.
BrowseCompAgentic88.7%74.2OpenAI
Artificial Analysis Long Context ReasoningReasoning88.3%69.3Artificial Analysis
MCP AtlasAgentic85.6%72.6OpenAI
Terminal-Bench 2.1 (provider run)Agentic85.0%74.2DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic85.0%74.2DeepSeek-AI
CyberGymAgentic84.7%74.1Zhun Wang et al.
Data Research and Analysis with Complex OperationsAgentic83.3%Anthropic
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding80.5%77.1John Yang et al.
Artificial Analysis MMMU-ProMultimodal76.4%60.5Artificial Analysis
Massive Multi-discipline Multimodal Understanding ProMultimodal76.0%51.0MMMU-Pro authors
Toolathlon-VerifiedAgentic74.1%72.3Moonshot AI
SWE-MarathonCoding72.7%Abundant AI and BenchFlow
DeepSWEAgentic67.7%73.0Datacurve AI
OfficeQA ProMultimodal60.3%72.6OfficeQA Pro authors
JobBenchAgentic59.0%74.7Yuetai Li et al.
Scientific Code BenchmarkCoding58.9%70.0Benchmark authors
Artificial Analysis SciCodeCoding58.9%74.6Artificial Analysis
GDPval-AA normalizedAgentic53.3%76.5Artificial Analysis
Humanity's Last ExamKnowledge46.5%68.2Center for AI Safety et al.
Artificial Analysis Humanity's Last ExamKnowledge46.5%75.0Artificial Analysis
AutomationBenchAgentic44.0%89.2Moonshot AI
Artificial Analysis Intelligence IndexKnowledge43.7%77.1Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge41.5%65.2Artificial Analysis
MLS-Bench LiteCoding40.5%MLS-Bench
APEX-AgentsAgentic37.8%70.8Moonshot AI / APEX-Agents benchmark authors
Agents' Last ExamAgentic29.5%69.0DeepSeek-AI
Critical Physics TasksReasoning20.9%82.2Artificial Analysis

24 benchmarks count, from 25 of 28 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from StepFun

Step 3.7 Flash56.4Step 3.5 FlashUnrankedStep-AudioUnrankedStep-Audio-Chat 130BUnranked