Muse Spark 1.2

Muse Spark 1.2 is a reasoning model from Meta in the Muse Spark family. 32 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.69.2 ±3.4
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.126/s
Input / 1MUS dollars per 1M input tokens.$1.25
Output / 1MUS dollars per 1M output tokens.$4.25
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1489 (#11)

50 is the middle of the board. The range shows the doubt in the index.

3,227 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
69.4
CodingCode writing and repair.
67.3
ReasoningLogic problems and puzzles.
70.0
MultimodalTasks with images and text.
67.5
KnowledgeFacts and expert knowledge.
68.5
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
77.3
MathMath problems.
68.3

Results

32 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
LiveBench MathematicsMath91.2%69.8xhigh effort25 Jun 2026LiveBench
Artificial Analysis GPQA DiamondKnowledge90.4%61.2Artificial Analysis
LiveBench ReasoningReasoning90.0%81.0xhigh effort25 Jun 2026LiveBench
MMLU ProKnowledge88.3%59.61 Sept 2026Vals AI
VulcanBench v3Coding87.0%72.7VulcanBench contributors
SWE-benchCoding86.6%67.01 Sept 2026Vals AI
MMMU ProMultimodal86.1%67.51 Sept 2026Vals AI
Terminal-Bench 2.1 (provider run)Agentic82.9%73.0DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic82.9%73.0DeepSeek-AI
Vibe Code Bench v1.1Coding79.1%75.1OpenHands21 Sept 2026Vals AI
Artificial Analysis Long Context ReasoningReasoning79.0%62.9Artificial Analysis
LiveBench LanguageKnowledge78.6%65.4xhigh effort25 Jun 2026LiveBench
LiveBench CodingCoding77.5%66.4xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning76.5%62.2xhigh effort25 Jun 2026LiveBench
LiveBench Instruction FollowingInstruction74.3%77.3xhigh effort25 Jun 2026LiveBench
Artificial Analysis Coding IndexCoding72.2%69.8Artificial Analysis
Terminal-Bench 2.1Agentic69.7%65.221 Sept 2026Vals AI
SimpleQA VerifiedKnowledge60.3%77.4xhigh effortEpoch AI
DeepSWEAgentic59.3%66.9Datacurve AI
LiveBench Agentic CodingAgentic57.6%70.9xhigh effort25 Jun 2026LiveBench
Artificial Analysis SciCodeCoding57.4%72.5Artificial Analysis
SkillsBenchCoding53.0%68.0OpenHands11 Sept 2026Vals AI
IOI v1Coding49.5%68.19 Aug 2026Vals AI
GDPval-AA normalizedAgentic49.1%73.3Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge45.5%73.9Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge45.4%70.0Artificial Analysis
Artificial Analysis Agentic IndexAgentic44.0%73.0Artificial Analysis
ProofBench v1.1Math43.0%66.821 Sept 2026Vals AI
Artificial Analysis Intelligence IndexKnowledge39.6%71.9Artificial Analysis
Code MigrationCoding30.0%64.121 Sept 2026Vals AI
IOICoding21.8%55.121 Sept 2026Vals AI
Critical Physics TasksReasoning17.7%75.5Artificial Analysis
FrontierSWE v2Coding12.0%61.2Proximal
Agent Arena command recoveryAgentic7.375.6xhigh effort15 Sept 2026LMArena
Terminal-Bench 4.0Agentic5.6%63.421 Sept 2026Vals AI
Agent Arena task outcomeAgentic-2.065.2xhigh effort15 Sept 2026LMArena
Agent Arena steerabilityAgentic-4.662.2xhigh effort15 Sept 2026LMArena

32 benchmarks count, from 37 of 37 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · FreeGPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6

More from Meta

Muse Spark 1.375.0Muse Spark 1.167.0Muse Spark61.6Muse Glimmer 30B53.5Llama 4 Maverick33.6