Muse Spark 1.1

Muse Spark 1.1 is a reasoning model from Meta in the Muse Spark family. 40 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.67.0 ±3.1
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.217/s
Input / 1MUS dollars per 1M input tokens.$1.25
Output / 1MUS dollars per 1M output tokens.$4.25
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1480 (#15)

50 is the middle of the board. The range shows the doubt in the index.

27,615 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
64.4
CodingCode writing and repair.
67.6
ReasoningLogic problems and puzzles.
66.4
MultimodalTasks with images and text.
67.5
KnowledgeFacts and expert knowledge.
68.2
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
69.9
MathMath problems.
64.3

Results

40 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
CybenchAgentic92.9%Andy K. Zhang et al.
GPQA DiamondKnowledge91.2%62.41 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge89.8%60.6Artificial Analysis
MMLU ProKnowledge88.7%60.31 Sept 2026Vals AI
CharXiv ReasoningMultimodal88.4%66.8CharXiv authors
MCP AtlasAgentic88.1%74.4OpenAI
LiveBench ReasoningReasoning87.7%77.8xhigh effort25 Jun 2026LiveBench
LiveBench MathematicsMath87.1%64.3xhigh effort25 Jun 2026LiveBench
MMMU ProMultimodal86.6%68.31 Sept 2026Vals AI
LiveCodeBenchCoding85.9%62.41 Sept 2026Vals AI
DeepSearchQAAgentic84.9%65.5Meta AI
SWE-benchCoding82.0%63.31 Sept 2026Vals AI
OSWorld-VerifiedAgentic80.8%69.0Tianbao Xie et al.
Artificial Analysis Long Context ReasoningReasoning77.7%62.0Artificial Analysis
LiveBench CodingCoding77.2%65.8xhigh effort25 Jun 2026LiveBench
BabyVisionMultimodal76.3%Meta AI
ToolathlonAgentic75.6%88.5OpenAI
LiveBench LanguageKnowledge74.3%60.3xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning72.5%56.7xhigh effort25 Jun 2026LiveBench
Vibe Code Bench v1.1Coding72.2%72.2OpenHands21 Sept 2026Vals AI
Artificial Analysis Coding IndexCoding71.3%69.2Artificial Analysis
LiveBench Instruction FollowingInstruction69.6%69.9xhigh effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic69.3%64.921 Sept 2026Vals AI
WebArena-Verified Browser Agent BenchmarkAgentic69.0%Amine El Hattami et al.
Humanity's Last ExamKnowledge62.1%81.4Center for AI Safety et al.
SWE-bench ProCoding61.5%63.5Xiang Deng et al.
HealthBench ProfessionalKnowledge59.3%Rebecca Soskin Hicks et al.
SkillsBenchCoding59.2%73.2OpenHands11 Sept 2026Vals AI
CyberGymAgentic59.0%56.6Zhun Wang et al.
Artificial Analysis SciCodeCoding58.8%74.5Artificial Analysis
LiveBench Agentic CodingAgentic58.5%71.8xhigh effort25 Jun 2026LiveBench
SimpleQA VerifiedKnowledge57.8%75.1Epoch AI
JobBenchAgentic54.7%71.7Yuetai Li et al.
MRCR 1MReasoning54.1%DeepSeek-AI
DeepSWEAgentic53.3%62.5Datacurve AI
Humanity's Last Exam without toolsKnowledge52.2%73.0OpenAI
Artificial Analysis Omniscience AccuracyKnowledge52.1%78.3Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge46.2%74.6Artificial Analysis
τ²-bench BankingAgentic40.5%27.9xhigh effort · Sierra4 Aug 2026Sierra Research
GDPval-AA normalizedAgentic35.4%62.8Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge33.7%64.6Artificial Analysis
Code MigrationCoding31.1%64.821 Sept 2026Vals AI
Artificial Analysis Agentic IndexAgentic27.5%59.6Artificial Analysis
Critical Physics TasksReasoning15.1%70.0Artificial Analysis
OSWorld 2.0Agentic14.2%62.7Mengqi Yuan et al.
ExploitGymAgentic0.8%61.9Zhun Wang et al.
Agent Arena task outcomeAgentic0.067.415 Sept 2026LMArena
ProgramBenchCoding0.0%21 Sept 2026Vals AI
Agent Arena command recoveryAgentic-0.666.715 Sept 2026LMArena
Agent Arena steerabilityAgentic-4.262.615 Sept 2026LMArena

40 benchmarks count, from 44 of 50 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSierra Research, collected directlyMIT — results are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 073164.6 · $0.28

More from Meta

Muse Spark 1.375.0Muse Spark 1.269.2Muse Spark61.6Muse Glimmer 30B53.5Llama 4 Maverick33.6