InclusionAI
availableShows if the model has enough results for an index.Ling 3.0 Flash
Ling 3.0 Flash is a reasoning model from InclusionAI in the Ling 3.0 family. 28 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.52.4 ±4.6
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.97/s
Input / 1MUS dollars per 1M input tokens.$0.021
Output / 1MUS dollars per 1M output tokens.$0.063
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.N/A
50 is the middle of the board. The range shows the doubt in the index.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
28 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| AIME 2026 | Math | 93.2% | 52.9 | — | — | Qwen |
| Harvard-MIT Mathematics Tournament February 2026 | Math | 87.0% | 55.2 | — | — | Qwen |
| Artificial Analysis GPQA Diamond | Knowledge | 85.5% | 56.2 | — | — | Artificial Analysis |
| Graduate-Level Google-Proof Q&A | Knowledge | 85.0% | 56.7 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 85.0% | 56.7 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 84.8% | 56.6 | — | 1 Sept 2026 | Vals AI |
| GPQA Diamond | Knowledge | 84.8% | 56.6 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 84.0% | 60.7 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 84.0% | 60.7 | — | 1 Sept 2026 | Vals AI |
| IMOAnswerBench | Math | 83.7% | — | — | — | DeepSeek-AI |
| LiveCodeBench v5 | Coding | 82.8% | — | — | — | LiveCodeBench maintainers |
| MMLU Pro | Knowledge | 82.0% | 49.7 | — | 1 Sept 2026 | Vals AI |
| MMLU Pro | Knowledge | 82.0% | 49.7 | — | 1 Sept 2026 | Vals AI |
| Instruction Following Benchmark | Instruction | 74.5% | 52.7 | — | — | Benchmark authors |
| WideResearch | Agentic | 73.6% | 58.9 | — | — | Qwen |
| Berkeley Function Calling Leaderboard v4 | Agentic | 73.0% | 58.0 | — | — | Arcee AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 73.0% | 58.7 | — | — | Artificial Analysis |
| BrowseComp | Agentic | 72.2% | 60.5 | — | — | OpenAI |
| Data Research and Analysis with Complex Operations | Agentic | 70.4% | — | — | — | Anthropic |
| MCP Atlas | Agentic | 65.5% | 57.5 | — | — | OpenAI |
| SWE-bench | Coding | 65.2% | 49.8 | — | 1 Sept 2026 | Vals AI |
| SWE-bench | Coding | 65.2% | 49.8 | — | 1 Sept 2026 | Vals AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 57.0% | 57.7 | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 57.0% | 57.7 | — | — | DeepSeek-AI |
| SWE-bench Pro | Coding | 56.6% | 58.7 | — | — | Xiang Deng et al. |
| Artificial Analysis Coding Index | Coding | 50.6% | 54.7 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 50.2% | 53.7 | — | 21 Sept 2026 | Vals AI |
| Terminal-Bench 2.1 | Agentic | 50.2% | 53.7 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis SciCode | Coding | 42.0% | 51.2 | — | — | Artificial Analysis |
| Scientific Code Benchmark | Coding | 41.2% | 51.2 | — | — | Benchmark authors |
| Artificial Analysis Tau3-Banking | Agentic | 28.0% | 50.8 | — | — | Artificial Analysis |
| SkillsBench | Coding | 27.9% | 46.7 | OpenHands | 11 Sept 2026 | Vals AI |
| SkillsBench | Coding | 27.9% | 46.7 | OpenHands | 11 Sept 2026 | Vals AI |
| Artificial Analysis Intelligence Index | Knowledge | 24.9% | 53.6 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 23.7% | 50.3 | — | — | Artificial Analysis |
| Humanity's Last Exam | Knowledge | 22.7% | 48.0 | — | — | Center for AI Safety et al. |
| Artificial Analysis Agentic Index | Agentic | 21.0% | 54.3 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 18.2% | 36.4 | — | — | Artificial Analysis |
| Vibe Code Bench v1.1 | Coding | 2.9% | 43.4 | OpenHands | 21 Sept 2026 | Vals AI |
| Vibe Code Bench v1.1 | Coding | 2.9% | 43.4 | OpenHands | 21 Sept 2026 | Vals AI |
| Critical Physics Tasks | Reasoning | 1.7% | 41.9 | — | — | Artificial Analysis |
| Code Migration | Coding | 0.0% | 44.9 | — | 21 Sept 2026 | Vals AI |
| Code Migration | Coding | 0.0% | 44.9 | — | 21 Sept 2026 | Vals AI |
28 benchmarks count, from 40 of 43 results. A grey row does not count. Too few models took that benchmark.