Gemini 3.7 Flash

Gemini 3.7 Flash is a reasoning model from Google. 51 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.72.3 ±2.3
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.80/s
Input / 1MUS dollars per 1M input tokens.$0.75 batch $0.375
Output / 1MUS dollars per 1M output tokens.$3.75 batch $1.88 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1490 (#8)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

5,640 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
70.7
CodingCode writing and repair.
70.1
ReasoningLogic problems and puzzles.
71.0
MultimodalTasks with images and text.
70.3
KnowledgeFacts and expert knowledge.
72.8
MultilingualTasks in many languages.
94.3
InstructionTasks with strict rules in the prompt.
86.0
MathMath problems.
69.8

Results

51 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math97.2%65.9high effortEpoch AI
OpenAI MRCR v2 8-needle 64K-128KReasoning97.0%OpenAI
ARC-AGI-1 (semi-private)Reasoning95.5%72.3high effortARC Prize Foundation
GPQA diamondKnowledge94.8%65.8high effortEpoch AI
Artificial Analysis GPQA DiamondKnowledge94.5%65.4Artificial Analysis
GPQA DiamondKnowledge93.9%65.01 Sept 2026Vals AI
LiveBench MathematicsMath93.5%72.8high effort25 Jun 2026LiveBench
Artificial Analysis Harvey LAB-AAAgentic90.7%73.6Artificial Analysis
MMLU ProKnowledge90.1%62.51 Sept 2026Vals AI
MMMU ProMultimodal89.0%72.11 Sept 2026Vals AI
CharXiv ReasoningMultimodal88.7%67.1CharXiv authors
LiveCodeBenchCoding88.7%64.91 Sept 2026Vals AI
LiveBench ReasoningReasoning87.8%77.9high effort25 Jun 2026LiveBench
BioMysteryBench Human SolvableKnowledge87.1%Anthropic
Terminal-Bench 2.1 (provider run)Agentic85.8%74.7DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic85.8%74.7DeepSeek-AI
Artificial Analysis MMMU-ProMultimodal85.5%71.6Artificial Analysis
LiveBench LanguageKnowledge85.5%73.6high effort25 Jun 2026LiveBench
LVBenchMultimodal85.4%Qwen Team
ARC-AGI-2 (semi-private)Reasoning84.6%82.4high effortARC Prize Foundation
CharXiv Reasoning without toolsMultimodal84.5%CharXiv authors
LABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchKnowledge82.1%Jon M. Laurent et al.
Artificial Analysis Long Context ReasoningReasoning81.7%64.8Artificial Analysis
SWE-benchCoding80.8%62.41 Sept 2026Vals AI
LiveBench Instruction FollowingInstruction79.9%86.0high effort25 Jun 2026LiveBench
LiveBench CodingCoding78.9%68.6high effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic77.5%69.821 Sept 2026Vals AI
EuroEval FrenchMultilingual77.1%95.0EuroEval
EuroEval ItalianMultilingual76.9%95.0EuroEval
EuroEval SwedishMultilingual76.6%95.0EuroEval
Artificial Analysis Coding IndexCoding76.1%72.6Artificial Analysis
EuroEval PortugueseMultilingual74.4%95.0EuroEval
EuroEval DutchMultilingual73.1%93.5EuroEval
FrontierMath-Tiers-1-3-v2-PrivateMath71.6%71.7high effortEpoch AI
Vibe Code Bench v1.1Coding70.4%71.5OpenHands21 Sept 2026Vals AI
SimpleQA VerifiedKnowledge69.2%85.6high effortEpoch AI
EuroEval SpanishMultilingual68.3%87.5EuroEval
LiveBench Data AnalysisReasoning68.0%50.4high effort25 Jun 2026LiveBench
EuroEval PolishMultilingual67.9%87.0EuroEval
IOICoding67.8%75.221 Sept 2026Vals AI
EuroEval GermanMultilingual67.3%86.3EuroEval
SkillsBenchCoding65.9%78.8OpenHands11 Sept 2026Vals AI
DeepSWEAgentic65.3%71.2Datacurve AI
Artificial Analysis AnalystAgentAgentic60.0%83.6Artificial Analysis
LiveBench Agentic CodingAgentic58.3%71.6high effort25 Jun 2026LiveBench
ProofBench v1.1Math58.0%72.921 Sept 2026Vals AI
Artificial Analysis SciCodeCoding57.2%72.2Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge55.3%82.3Artificial Analysis
HLE-VerifiedKnowledge53.6%Weiqi Zhai et al.
OSWorld 2.0Agentic47.9%78.8Mengqi Yuan et al.
Artificial Analysis Humanity's Last ExamKnowledge47.9%76.5Artificial Analysis
Chess PuzzlesReasoning47.0%84.3high effortEpoch AI
GDPval-AA normalizedAgentic43.6%69.1Artificial Analysis
FrontierCode 1.1 MainCoding43.6%72.3Cognition
BioMysteryBench Human DifficultKnowledge43.5%Anthropic
Artificial Analysis Intelligence IndexKnowledge39.1%71.2Artificial Analysis
Mystery Game PuzzlesReasoning37.0%73.3high effortEpoch AI
FrontierMath-Tier-4-v2-PrivateMath36.6%65.5high effortEpoch AI
Artificial Analysis Agentic IndexAgentic36.4%66.9Artificial Analysis
Code MigrationCoding34.8%67.221 Sept 2026Vals AI
AutomationBenchAgentic30.4%66.0Moonshot AI
Furniture AssemblyReasoning26.7%58.5high effortEpoch AI
Agents' Last ExamAgentic26.3%66.0DeepSeek-AI
FrontierSWE v2Coding20.3%65.7Proximal
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic16.0%69.0NeoCognition
Terminal-Bench 3.0Agentic14.9%66.2Ryan Marten et al.
Critical Physics TasksReasoning14.3%68.4Artificial Analysis
Terminal-Bench 4.0.0Agentic11.2%67.4high effort · mini-SWE-agent21 Sept 2026Terminal-Bench
Terminal-Bench 4.0Agentic6.1%63.721 Sept 2026Vals AI
Agent Arena task outcomeAgentic2.970.6high effort15 Sept 2026LMArena
ProgramBenchCoding0.0%21 Sept 2026Vals AI
Agent Arena steerabilityAgentic-0.367.0high effort15 Sept 2026LMArena
Agent Arena command recoveryAgentic-2.165.0high effort15 Sept 2026LMArena

51 benchmarks count, from 65 of 73 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from Google

Gemini 3.8 Flash72.1Gemini 3.5 Flash67.3Gemini 3.6 Flash66.8Gemini 3.1 Pro64.7Gemini 3 Pro60.5