StepFun
availableShows if the model has enough results for an index.Step 3.7 Flash
Step 3.7 Flash is a reasoning model from StepFun. 22 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.56.4 ±4.9
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.93/s
Input / 1MUS dollars per 1M input tokens.$0.2
Output / 1MUS dollars per 1M output tokens.$1.15
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.N/A
50 is the middle of the board. The range shows the doubt in the index.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
22 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 98.5% | 69.5 | — | — | Victor Barres et al. |
| V* | Multimodal | 95.3% | 55.6 | — | — | Z.AI |
| DeepSearchQA | Agentic | 92.8% | 71.9 | — | — | Meta AI |
| Artificial Analysis GPQA Diamond | Knowledge | 80.9% | 51.5 | — | — | Artificial Analysis |
| SimpleVQA | Multimodal | 79.2% | 73.2 | — | — | Z.AI |
| BrowseComp | Agentic | 75.8% | 63.5 | — | — | OpenAI |
| Artificial Analysis MMMU-Pro | Multimodal | 75.3% | 59.1 | — | — | Artificial Analysis |
| Artificial Analysis Long Context Reasoning | Reasoning | 73.7% | 59.2 | — | — | Artificial Analysis |
| Artificial Analysis IFBench | Instruction | 67.3% | 57.4 | — | — | Artificial Analysis |
| Claw-Eval | Agentic | 67.1% | 63.9 | — | — | Bowen Ye et al. |
| SWE-bench Pro | Coding | 56.3% | 58.4 | — | — | Xiang Deng et al. |
| Gert Labs Composite Game Benchmark | Agentic | 51.6% | 58.3 | — | — | Gert Labs |
| Toolathlon | Agentic | 49.5% | 64.1 | — | — | OpenAI |
| Humanity's Last Exam with tools | Agentic | 47.2% | 60.3 | — | — | DeepSeek-AI |
| Artificial Analysis SciCode | Coding | 43.9% | 53.8 | — | — | Artificial Analysis |
| Artificial Analysis Coding Index | Coding | 39.6% | 46.8 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 25.8% | 55.4 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 25.8% | 45.8 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 21.4% | 47.8 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 19.5% | 46.8 | — | — | Artificial Analysis |
| APEX-Agents-AA | Agentic | 14.8% | 52.8 | — | — | Artificial Analysis / Mercor |
| Critical Physics Tasks | Reasoning | 2.3% | 43.2 | — | — | Artificial Analysis |
22 benchmarks count, from 22 of 22 results. A grey row does not count. Too few models took that benchmark.