OpenAI
availableShows if the model has enough results for an index.GPT-5.4 mini
GPT-5.4 mini is a reasoning model from OpenAI in the GPT-5.4 family. 47 benchmarks count toward its score, in 8 categories.
IndexOverall score. 50 is the middle.56.8 ±2.4
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.62/s
Input / 1MUS dollars per 1M input tokens.$0.75 batch $0.375
Output / 1MUS dollars per 1M output tokens.$4.5 batch $2.25 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.400K
EloLMArena rating and rank.1412 (#127)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
59,387 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
47 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| AIME | Math | 95.6% | 58.3 | — | 16 Apr 2026 | Vals AI |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 93.4% | 65.8 | — | — | Victor Barres et al. |
| OTIS Mock AIME 2024-2025 | Math | 88.9% | 61.2 | xhigh effort | — | Epoch AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 88.0% | 59.5 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 87.5% | 58.2 | — | — | Artificial Analysis |
| GPQA diamond | Knowledge | 86.9% | 58.4 | xhigh effort | — | Epoch AI |
| MMLU Pro | Knowledge | 84.6% | 53.8 | — | 1 Sept 2026 | Vals AI |
| GPQA Diamond | Knowledge | 83.1% | 54.9 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 81.5% | 58.4 | — | 1 Sept 2026 | Vals AI |
| MMMU Pro | Multimodal | 79.2% | 56.3 | — | 1 Sept 2026 | Vals AI |
| LiveBench Mathematics | Math | 78.5% | 52.7 | xhigh effort | 25 Jun 2026 | LiveBench |
| MMMU-Pro with Python | Multimodal | 78.0% | — | — | — | OpenAI |
| Artificial Analysis Long Context Reasoning | Reasoning | 77.0% | 61.5 | — | — | Artificial Analysis |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 76.6% | 52.0 | — | — | MMMU-Pro authors |
| Artificial Analysis MMMU-Pro | Multimodal | 73.3% | 56.7 | — | — | Artificial Analysis |
| Artificial Analysis IFBench | Instruction | 73.3% | 63.5 | — | — | Artificial Analysis |
| SWE-bench | Coding | 73.0% | 56.1 | — | 1 Sept 2026 | Vals AI |
| OSWorld-Verified | Agentic | 72.1% | 60.8 | — | — | Tianbao Xie et al. |
| LiveBench Coding | Coding | 71.6% | 56.6 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Reasoning | Reasoning | 71.3% | 55.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Language | Knowledge | 71.0% | 56.2 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Data Analysis | Reasoning | 70.8% | 54.3 | xhigh effort | 25 Jun 2026 | LiveBench |
| EuroEval Swedish | Multilingual | 68.9% | 88.3 | high effort | — | EuroEval |
| EuroEval French | Multilingual | 67.4% | 86.4 | high effort | — | EuroEval |
| EuroEval French | Multilingual | 67.3% | 86.3 | medium effort | — | EuroEval |
| EuroEval Swedish | Multilingual | 67.3% | 86.3 | medium effort | — | EuroEval |
| EuroEval French | Multilingual | 67.1% | 86.1 | low effort | — | EuroEval |
| EuroEval Italian | Multilingual | 66.2% | 84.9 | high effort | — | EuroEval |
| EuroEval Italian | Multilingual | 65.8% | 84.5 | medium effort | — | EuroEval |
| EuroEval Swedish | Multilingual | 65.0% | 83.5 | low effort | — | EuroEval |
| EuroEval Portuguese | Multilingual | 64.7% | 83.1 | high effort | — | EuroEval |
| EuroEval Portuguese | Multilingual | 63.7% | 81.8 | medium effort | — | EuroEval |
| ARC-AGI-1 (semi-private) | Reasoning | 63.7% | 57.3 | xhigh effort | — | ARC Prize Foundation |
| EuroEval Dutch | Multilingual | 63.1% | 81.1 | high effort | — | EuroEval |
| EuroEval Dutch | Multilingual | 63.0% | 81.0 | medium effort | — | EuroEval |
| EuroEval Spanish | Multilingual | 62.8% | 80.7 | high effort | — | EuroEval |
| EuroEval Polish | Multilingual | 62.5% | 80.3 | high effort | — | EuroEval |
| EuroEval Italian | Multilingual | 62.5% | 80.3 | low effort | — | EuroEval |
| EuroEval Dutch | Multilingual | 62.1% | 79.8 | low effort | — | EuroEval |
| EuroEval Spanish | Multilingual | 60.8% | 78.2 | medium effort | — | EuroEval |
| EuroEval Portuguese | Multilingual | 60.6% | 77.9 | low effort | — | EuroEval |
| EuroEval Polish | Multilingual | 60.2% | 77.5 | medium effort | — | EuroEval |
| LiveBench Instruction Following | Instruction | 59.8% | 54.6 | xhigh effort | 25 Jun 2026 | LiveBench |
| EuroEval Polish | Multilingual | 58.8% | 75.7 | low effort | — | EuroEval |
| EuroEval Spanish | Multilingual | 57.9% | 74.6 | low effort | — | EuroEval |
| EuroEval German | Multilingual | 57.8% | 74.4 | high effort | — | EuroEval |
| MCP Atlas | Agentic | 57.7% | 51.7 | — | — | OpenAI |
| EuroEval German | Multilingual | 56.3% | 72.6 | medium effort | — | EuroEval |
| Artificial Analysis Coding Index | Coding | 56.1% | 58.5 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 54.7% | 56.3 | — | 21 Sept 2026 | Vals AI |
| EuroEval German | Multilingual | 54.6% | 70.5 | low effort | — | EuroEval |
| Artificial Analysis SciCode | Coding | 52.1% | 65.2 | — | — | Artificial Analysis |
| FrontierMath-Tiers-1-3-v2-Private | Math | 51.2% | 60.3 | xhigh effort | — | Epoch AI |
| Vibe Code Bench v1.1 | Coding | 48.0% | 62.1 | OpenHands | 21 Sept 2026 | Vals AI |
| Terminal-Bench 2.0 | Agentic | 44.9% | 54.8 | — | 4 Jun 2026 | Vals AI |
| Toolathlon | Agentic | 42.9% | 58.0 | — | — | OpenAI |
| LiveBench Agentic Coding | Agentic | 41.7% | 55.9 | xhigh effort | 25 Jun 2026 | LiveBench |
| Humanity's Last Exam | Knowledge | 41.5% | 63.9 | — | — | Center for AI Safety et al. |
| Artificial Analysis Omniscience Accuracy | Knowledge | 37.5% | 60.3 | — | — | Artificial Analysis |
| SimpleQA Verified | Knowledge | 29.4% | 48.8 | high effort | — | Epoch AI |
| FrontierMath-2025-02-28-Private | Math | 28.3% | 57.0 | high effort | — | Epoch AI |
| APEX-Agents-AA | Agentic | 28.2% | 63.3 | — | — | Artificial Analysis / Mercor |
| Humanity's Last Exam without tools | Knowledge | 28.2% | 52.7 | — | — | OpenAI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 28.1% | 55.0 | — | — | Artificial Analysis |
| FrontierCode 1.1 Main | Coding | 27.0% | 57.4 | — | — | Cognition |
| GDPval-AA normalized | Agentic | 25.0% | 54.8 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 24.1% | 52.5 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 24.0% | 54.6 | xhigh effort | — | Epoch AI |
| Artificial Analysis Agentic Index | Agentic | 19.6% | 53.2 | — | — | Artificial Analysis |
| ARC-AGI-2 (semi-private) | Reasoning | 18.9% | 49.2 | xhigh effort | — | ARC Prize Foundation |
| Code Migration | Coding | 12.9% | 53.2 | — | 21 Sept 2026 | Vals AI |
| Critical Physics Tasks | Reasoning | 10.0% | 59.3 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-v2-Private | Math | 9.8% | 52.6 | xhigh effort | — | Epoch AI |
| Mystery Game Puzzles | Reasoning | 7.0% | 41.4 | medium effort | — | Epoch AI |
| IOI v1 | Coding | 6.4% | 43.6 | — | 9 Aug 2026 | Vals AI |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 2.1% | 45.7 | high effort | — | Epoch AI |
| ProgramBench | Coding | 0.0% | — | — | 21 Sept 2026 | Vals AI |
47 benchmarks count, from 75 of 77 results. A grey row does not count. Too few models took that benchmark.