Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite is a reasoning model from Google. 38 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.54.4 ±2.6
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.73/s
Input / 1MUS dollars per 1M input tokens.$0.3
Output / 1MUS dollars per 1M output tokens.$2.5
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1436 (#79)

50 is the middle of the board. The range shows the doubt in the index.

26,165 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
56.0
CodingCode writing and repair.
55.5
ReasoningLogic problems and puzzles.
48.2
MultimodalTasks with images and text.
63.5
KnowledgeFacts and expert knowledge.
52.6
MultilingualTasks in many languages.
77.8
InstructionTasks with strict rules in the prompt.
66.2
MathMath problems.
47.9

Results

38 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MMLU ProKnowledge85.8%55.81 Sept 2026Vals AI
GPQA DiamondKnowledge83.8%55.61 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge83.8%54.4Artificial Analysis
MMMU ProMultimodal83.6%63.41 Sept 2026Vals AI
GPQA diamondKnowledge83.3%55.2high effortEpoch AI
LiveCodeBenchCoding79.0%56.21 Sept 2026Vals AI
Artificial Analysis MMMU-ProMultimodal79.0%63.6Artificial Analysis
LiveBench CodingCoding76.1%64.0high effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning76.0%60.8Artificial Analysis
SWE-benchCoding75.0%57.71 Sept 2026Vals AI
OSWorld-VerifiedAgentic74.0%62.6Tianbao Xie et al.
LiveBench MathematicsMath73.7%46.3high effort25 Jun 2026LiveBench
MRCRv2Reasoning72.2%OpenAI
LiveBench LanguageKnowledge71.8%57.3high effort25 Jun 2026LiveBench
OTIS Mock AIME 2024-2025Math71.1%51.3high effortEpoch AI
LiveBench Instruction FollowingInstruction67.2%66.2high effort25 Jun 2026LiveBench
EuroEval SwedishMultilingual63.0%81.0EuroEval
EuroEval FrenchMultilingual61.6%79.2EuroEval
EuroEval ItalianMultilingual61.6%79.1EuroEval
EuroEval PortugueseMultilingual61.3%78.8EuroEval
LiveBench ReasoningReasoning60.2%39.5high effort25 Jun 2026LiveBench
EuroEval DutchMultilingual59.6%76.7EuroEval
EuroEval SpanishMultilingual57.1%73.6EuroEval
EuroEval PolishMultilingual54.3%70.1EuroEval
SWE-bench ProCoding54.2%56.4Xiang Deng et al.
ARC-AGI-1 (semi-private)Reasoning53.5%52.5high effortARC Prize Foundation
LiveBench Data AnalysisReasoning53.2%29.9high effort25 Jun 2026LiveBench
EuroEval GermanMultilingual52.4%67.8EuroEval
Terminal-Bench 2.1Agentic50.2%53.721 Sept 2026Vals AI
Artificial Analysis Coding IndexCoding49.3%53.7Artificial Analysis
LiveBench Agentic CodingAgentic45.3%59.3high effort25 Jun 2026LiveBench
Artificial Analysis EnterpriseOps-GymAgentic42.3%65.3Artificial Analysis
Artificial Analysis SciCodeCoding41.3%50.2Artificial Analysis
Vibe Code Bench v1.1Coding37.2%57.6OpenHands21 Sept 2026Vals AI
Artificial Analysis Omniscience AccuracyKnowledge29.5%50.4Artificial Analysis
IOI v1Coding26.2%54.89 Aug 2026Vals AI
FrontierMath-Tiers-1-3-v2-PrivateMath26.0%46.1high effortEpoch AI
GDPval-AA normalizedAgentic23.5%53.6Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge22.2%50.1Artificial Analysis
Chess PuzzlesReasoning22.0%52.0high effortEpoch AI
Mystery Game PuzzlesReasoning19.0%54.1low effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge18.8%44.9Artificial Analysis
Artificial Analysis Agentic IndexAgentic15.9%50.2Artificial Analysis
ARC-AGI-2 (semi-private)Reasoning10.3%44.9high effortARC Prize Foundation
Code MigrationCoding6.1%48.821 Sept 2026Vals AI
Critical Physics TasksReasoning0.0%38.4Artificial Analysis
ProgramBenchCoding0.0%21 Sept 2026Vals AI
FrontierMath-Tier-4-v2-PrivateMath0.0%47.9high effortEpoch AI
Agent Arena steerabilityAgentic-12.853.015 Sept 2026LMArena
Agent Arena task outcomeAgentic-17.647.515 Sept 2026LMArena
Agent Arena command recoveryAgentic-28.735.015 Sept 2026LMArena

38 benchmarks count, from 49 of 51 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Google

Gemini 3.7 Flash72.3Gemini 3.8 Flash72.1Gemini 3.5 Flash67.3Gemini 3.6 Flash66.8Gemini 3.1 Pro64.7