Mistral Medium 3.5 128B

Mistral Medium 3.5 128B is a reasoning model from Mistral in the Mistral Medium 3.5 family. 19 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.49.8 ±5.2
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.$1.5
Output / 1MUS dollars per 1M output tokens.$7.5
ContextMaximum tokens in one request.256K
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
53.0
CodingCode writing and repair.
53.5
ReasoningLogic problems and puzzles.
47.3
MultimodalTasks with images and text.
46.4
KnowledgeFacts and expert knowledge.
42.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
58.9
MathMath problems.
N/A

Results

19 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic94.2%66.4Victor Barres et al.
τ³-Bench Tool-Agent-User EvaluationAgentic91.4%72.8Sierra Research
Software Engineering Benchmark VerifiedCoding77.6%59.8Carlos E. Jimenez et al.
Artificial Analysis GPQA DiamondKnowledge74.8%45.2Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning69.3%56.2Artificial Analysis
Artificial Analysis Harvey LAB-AAAgentic69.1%43.8Artificial Analysis
Artificial Analysis IFBenchInstruction68.8%58.9Artificial Analysis
Artificial Analysis MMMU-ProMultimodal64.9%46.4Artificial Analysis
Artificial Analysis Coding IndexCoding46.9%52.0Artificial Analysis
Artificial Analysis SciCodeCoding40.2%48.7Artificial Analysis
Gert Labs Composite Game BenchmarkAgentic39.1%47.3Gert Labs
Artificial Analysis EnterpriseOps-GymAgentic33.7%53.6Artificial Analysis
Terminal-Bench HardAgentic33.3%Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge24.7%44.4Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge14.2%40.1Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge13.8%39.5Artificial Analysis
Artificial Analysis AnalystAgentAgentic12.5%50.3Artificial Analysis
GDPval-AA normalizedAgentic12.4%45.1Artificial Analysis
Artificial Analysis Agentic IndexAgentic9.3%44.8Artificial Analysis
Critical Physics TasksReasoning0.0%38.4Artificial Analysis

19 benchmarks count, from 19 of 20 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Mistral

Mistral Medium 3.543.0Mistral Small 438.4Mistral Small 4 (Reasoning)38.5Magistral Medium38.3Magistral Small35.7