OpenAI
availableShows if the model has enough results for an index.GPT-5.2-Codex
GPT-5.2-Codex is a reasoning model from OpenAI. 20 benchmarks count toward its score, in 7 categories.
IndexOverall score. 50 is the middle.62.1 ±4.1
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.15/s
Input / 1MUS dollars per 1M input tokens.$1.75
Output / 1MUS dollars per 1M output tokens.$14
ContextMaximum tokens in one request.400K
EloLMArena rating and rank.N/A
50 is the middle of the board. The range shows the doubt in the index.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
20 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 92.1% | 64.9 | — | — | Victor Barres et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 89.9% | 60.7 | — | — | Artificial Analysis |
| LiveBench Mathematics | Math | 88.8% | 66.5 | — | 25 Jun 2026 | LiveBench |
| LiveCodeBench | Coding | 88.0% | 64.3 | — | 1 Sept 2026 | Vals AI |
| LiveBench Coding | Coding | 83.6% | 76.4 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis Long Context Reasoning | Reasoning | 82.3% | 65.2 | — | — | Artificial Analysis |
| LiveBench Data Analysis | Reasoning | 78.2% | 64.6 | — | 25 Jun 2026 | LiveBench |
| LiveBench Reasoning | Reasoning | 77.7% | 63.9 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis IFBench | Instruction | 77.6% | 67.8 | — | — | Artificial Analysis |
| Artificial Analysis MMMU-Pro | Multimodal | 76.3% | 60.3 | — | — | Artificial Analysis |
| LiveBench Language | Knowledge | 73.7% | 59.5 | — | 25 Jun 2026 | LiveBench |
| SWE-bench Verified | Coding | 72.8% | 55.9 | mini-SWE-agent | 1 Sept 2026 | SWE-bench team |
| SWE-bench | Coding | 72.4% | 55.6 | — | 1 Sept 2026 | Vals AI |
| LiveBench Instruction Following | Instruction | 66.4% | 65.0 | — | 25 Jun 2026 | LiveBench |
| SWE-bench Multilingual | Coding | 66.3% | — | mini-SWE-agent | 2 Sept 2026 | SWE-bench team |
| Gert Labs Composite Game Benchmark | Agentic | 51.8% | 58.5 | — | — | Gert Labs |
| LiveBench Agentic Coding | Agentic | 49.4% | 63.2 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis Omniscience Accuracy | Knowledge | 41.1% | 64.7 | — | — | Artificial Analysis |
| Vibe Code Bench v1.1 | Coding | 37.9% | 57.9 | OpenHands | 21 Sept 2026 | Vals AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 35.7% | 63.3 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 28.5% | 58.0 | — | — | Artificial Analysis |
| JobBench | Agentic | 26.0% | 52.0 | — | — | Yuetai Li et al. |
| Critical Physics Tasks | Reasoning | 8.7% | 56.6 | — | — | Artificial Analysis |
20 benchmarks count, from 22 of 23 results. A grey row does not count. Too few models took that benchmark.