Gemini 3.8 Flash

Gemini 3.8 Flash is a reasoning model from Google. 47 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.72.1 ±2.9
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.86/s
Input / 1MUS dollars per 1M input tokens.$0.75 batch $0.375
Output / 1MUS dollars per 1M output tokens.$3.75 batch $1.88 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1495 (#6)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

5,076 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
73.2
CodingCode writing and repair.
69.2
ReasoningLogic problems and puzzles.
70.6
MultimodalTasks with images and text.
72.0
KnowledgeFacts and expert knowledge.
73.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
88.4
MathMath problems.
66.9

Results

47 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math98.9%66.8high effortEpoch AI
GPQA diamondKnowledge95.4%66.3high effortEpoch AI
Artificial Analysis GPQA DiamondKnowledge95.3%66.2Artificial Analysis
GPQA DiamondKnowledge94.4%65.41 Sept 2026Vals AI
LiveBench MathematicsMath91.6%70.2high effort25 Jun 2026LiveBench
MMLU ProKnowledge90.2%62.71 Sept 2026Vals AI
LiveCodeBenchCoding89.5%65.71 Sept 2026Vals AI
Terminal-Bench 2.1 (provider run)Agentic89.4%76.8DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic89.4%76.8DeepSeek-AI
LiveBench ReasoningReasoning89.3%80.0high effort25 Jun 2026LiveBench
MMMU ProMultimodal89.1%72.31 Sept 2026Vals AI
BioMysteryBench Human SolvableKnowledge88.8%Anthropic
LiveBench LanguageKnowledge87.8%76.4high effort25 Jun 2026LiveBench
LVBenchMultimodal87.1%Qwen Team
CharXiv Reasoning without toolsMultimodal86.2%CharXiv authors
LABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchKnowledge86.2%Jon M. Laurent et al.
Artificial Analysis MMMU-ProMultimodal85.6%71.7Artificial Analysis
LiveBench Instruction FollowingInstruction81.4%88.4high effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning81.3%64.5Artificial Analysis
Terminal-Bench 2.1Agentic81.3%72.021 Sept 2026Vals AI
SWE-benchCoding80.0%61.71 Sept 2026Vals AI
Vibe Code Bench v1.1Coding78.7%74.9OpenHands21 Sept 2026Vals AI
Artificial Analysis Coding IndexCoding76.3%72.7Artificial Analysis
DeepSWEAgentic73.8%77.4Datacurve AI
LiveBench CodingCoding72.5%58.0high effort25 Jun 2026LiveBench
SimpleQA VerifiedKnowledge69.7%86.1high effortEpoch AI
cursorBench32Coding69.2%77.1Benchmark authors
FrontierMath-Tiers-1-3-v2-PrivateMath68.4%70.0high effortEpoch AI
Chess PuzzlesReasoning61.0%95.0high effortEpoch AI
Artificial Analysis AutomationBenchAgentic59.9%71.9Artificial Analysis
OSWorld 2.0Agentic59.0%84.0Mengqi Yuan et al.
SkillsBenchCoding58.0%72.1OpenHands11 Sept 2026Vals AI
IOICoding56.9%70.521 Sept 2026Vals AI
Artificial Analysis SciCodeCoding56.6%71.4Artificial Analysis
BioMysteryBench Human DifficultKnowledge56.5%Anthropic
HLE-VerifiedKnowledge54.9%Weiqi Zhai et al.
Artificial Analysis Omniscience AccuracyKnowledge54.6%81.4Artificial Analysis
LiveBench Agentic CodingAgentic54.2%67.8high effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning54.0%30.9high effort25 Jun 2026LiveBench
ProofBench v1.1Math48.0%68.821 Sept 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge47.8%76.4Artificial Analysis
Mystery Game PuzzlesReasoning47.0%83.9high effortEpoch AI
GDPval-AA normalizedAgentic45.6%70.6Artificial Analysis
Artificial Analysis Tau3-BankingAgentic44.9%73.6Artificial Analysis
Artificial Analysis Agentic IndexAgentic41.1%70.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge40.9%73.6Artificial Analysis
cursorBench40Coding39.6%Benchmark authors
Code MigrationCoding36.5%68.321 Sept 2026Vals AI
Furniture AssemblyReasoning31.7%62.0high effortEpoch AI
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic24.0%73.5NeoCognition
FrontierMath-Tier-4-v2-PrivateMath22.0%58.5high effortEpoch AI
Medical Long Context Reasoning (MLCR-AA)Reasoning21.7%56.6Wisedocs and Artificial Analysis
Artificial Analysis GDP.pdfAgentic21.0%72.2Artificial Analysis
FrontierSWE v2Coding19.6%65.3Proximal
Terminal-Bench 4.0.0Agentic19.1%72.9high effort · mini-SWE-agent21 Sept 2026Terminal-Bench
Vibe Code Bench 1-100Coding18.8%72.4OpenHands16 Sept 2026Vals AI
Critical Physics TasksReasoning18.3%76.7Artificial Analysis
Terminal-Bench 4.0Agentic13.1%68.721 Sept 2026Vals AI
Agent Arena task outcomeAgentic9.377.9high effort15 Sept 2026LMArena
Agent Arena command recoveryAgentic1.969.5high effort15 Sept 2026LMArena
ProgramBenchCoding1.0%21 Sept 2026Vals AI
Agent Arena steerabilityAgentic-1.266.0high effort15 Sept 2026LMArena

47 benchmarks count, from 54 of 62 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from Google

Gemini 3.7 Flash72.3Gemini 3.5 Flash67.3Gemini 3.6 Flash66.8Gemini 3.1 Pro64.7Gemini 3 Pro60.5