Inkling

Inkling is a hybrid model from Thinking Machines Lab. 50 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.56.7 ±2.8
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.119/s
Input / 1MUS dollars per 1M input tokens.$1
Output / 1MUS dollars per 1M output tokens.$4.05
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1440 (#67)

50 is the middle of the board. The range shows the doubt in the index. A free tier is available.

25,922 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
55.0
CodingCode writing and repair.
55.3
ReasoningLogic problems and puzzles.
57.6
MultimodalTasks with images and text.
55.8
KnowledgeFacts and expert knowledge.
59.0
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
64.8
MathMath problems.
55.5

Results

50 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIME 2026Math97.1%55.8Qwen
OTIS Mock AIME 2024-2025Math88.9%61.2xhigh effortEpoch AI
LiveBench MathematicsMath88.4%66.0xhigh effort25 Jun 2026LiveBench
GPQA diamondKnowledge88.3%59.7xhigh effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge87.9%59.4David Rein et al.
GPQA DiamondKnowledge87.9%59.4David Rein et al.
Artificial Analysis GPQA DiamondKnowledge87.2%57.9Artificial Analysis
GPQA DiamondKnowledge87.1%58.71 Sept 2026Vals AI
MMLU ProKnowledge86.3%56.51 Sept 2026Vals AI
LiveCodeBenchCoding85.5%62.11 Sept 2026Vals AI
CharXiv ReasoningMultimodal82.0%59.5CharXiv authors
Instruction Following BenchmarkInstruction79.8%58.9Benchmark authors
ARC-AGI-1 (semi-private)Reasoning79.5%64.8ARC Prize Foundation
MMMU ProMultimodal78.5%55.11 Sept 2026Vals AI
LiveBench ReasoningReasoning78.3%64.8xhigh effort25 Jun 2026LiveBench
CharXiv Reasoning without toolsMultimodal78.1%CharXiv authors
Software Engineering Benchmark VerifiedCoding77.6%59.8Carlos E. Jimenez et al.
SWE-benchCoding77.6%59.81 Sept 2026Vals AI
Artificial Analysis Long Context ReasoningReasoning77.3%61.7Artificial Analysis
BrowseCompAgentic77.1%64.6OpenAI
MCP AtlasAgentic74.1%64.0OpenAI
Massive Multi-discipline Multimodal Understanding ProMultimodal73.5%47.0MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal73.5%56.9Artificial Analysis
LiveBench LanguageKnowledge73.5%59.2xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning72.8%57.0xhigh effort25 Jun 2026LiveBench
LiveBench CodingCoding71.0%55.6xhigh effort25 Jun 2026LiveBench
LiveBench Instruction FollowingInstruction70.1%70.7xhigh effort25 Jun 2026LiveBench
SWE-bench ProCoding54.3%56.5Xiang Deng et al.
Artificial Analysis Coding IndexCoding52.1%55.6Artificial Analysis
LiveBench Agentic CodingAgentic49.4%63.2xhigh effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic47.6%52.121 Sept 2026Vals AI
Artificial Analysis SciCodeCoding47.0%58.1Artificial Analysis
Humanity's Last ExamKnowledge46.0%67.7Center for AI Safety et al.
Artificial Analysis Omniscience AccuracyKnowledge41.6%65.3Artificial Analysis
SimpleQA VerifiedKnowledge40.3%58.9xhigh effortEpoch AI
Artificial Analysis EnterpriseOps-GymAgentic38.0%59.4Artificial Analysis
ARC-AGI-2 (semi-private)Reasoning36.5%58.1ARC Prize Foundation
FrontierMath-Tiers-1-3-v2-PrivateMath33.3%50.2xhigh effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge31.9%59.1Artificial Analysis
Humanity's Last Exam without toolsKnowledge30.0%54.2OpenAI
Artificial Analysis Tau3-BankingAgentic29.1%52.3Artificial Analysis
GDPval-AA normalizedAgentic28.2%57.2Artificial Analysis
SkillsBenchCoding26.1%45.1OpenHands11 Sept 2026Vals AI
τ²-bench BankingAgentic25.0%16.8max effort · Sierra4 Aug 2026Sierra Research
Artificial Analysis Intelligence IndexKnowledge25.0%53.6Artificial Analysis
Artificial Analysis Agentic IndexAgentic24.3%57.0Artificial Analysis
Artificial Analysis AnalystAgentAgentic23.8%58.2Artificial Analysis
Chess PuzzlesReasoning21.0%50.7xhigh effortEpoch AI
Vibe Code Bench v1.1Coding19.2%50.2OpenHands21 Sept 2026Vals AI
IOICoding14.9%52.121 Sept 2026Vals AI
Code MigrationCoding11.8%52.421 Sept 2026Vals AI
Vibe Code Bench 1-100Coding7.3%59.8OpenHands16 Sept 2026Vals AI
Critical Physics TasksReasoning5.4%49.7Artificial Analysis
FrontierMath-Tier-4-v2-PrivateMath4.9%50.3xhigh effortEpoch AI
FrontierSWE v2Coding4.1%56.8Proximal
Agent Arena command recoveryAgentic2.069.615 Sept 2026LMArena
ProgramBenchCoding0.0%21 Sept 2026Vals AI
ProofBench v1.1Math0.0%49.421 Sept 2026Vals AI
Terminal-Bench 4.0Agentic0.0%59.521 Sept 2026Vals AI
Agent Arena steerabilityAgentic-10.555.515 Sept 2026LMArena
Agent Arena task outcomeAgentic-19.345.615 Sept 2026LMArena

50 benchmarks count, from 59 of 61 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useSierra Research, collected directlyMIT — results are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Thinking Machines Lab

Inkling-Small56.9