DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a reasoning model from DeepSeek in the DeepSeek V4 Flash family. 29 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.64.6 ±5.0
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.221/s
Input / 1MUS dollars per 1M input tokens.$0.14
Output / 1MUS dollars per 1M output tokens.$0.28
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
65.0
CodingCode writing and repair.
64.5
ReasoningLogic problems and puzzles.
68.3
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
61.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
61.1

Results

29 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Harvard-MIT Mathematics Tournament February 2026Math94.8%61.1Qwen
LiveCodeBench Pass@1 with Chain-of-ThoughtCoding91.6%DeepSeek
Artificial Analysis GPQA DiamondKnowledge90.8%61.6Artificial Analysis
VulcanBench v3Coding88.4%75.0VulcanBench contributors
IMOAnswerBenchMath88.4%DeepSeek-AI
Graduate-Level Google-Proof Q&AKnowledge88.1%59.6David Rein et al.
GPQA DiamondKnowledge88.1%59.6David Rein et al.
Massive Multitask Language Understanding ProfessionalKnowledge86.2%56.4Yubo Wang et al.
Apex ShortlistMath85.7%DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic82.7%72.9DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic82.7%72.9DeepSeek-AI
Artificial Analysis Long Context ReasoningReasoning79.7%63.4Artificial Analysis
Software Engineering Benchmark VerifiedCoding79.0%60.9Carlos E. Jimenez et al.
Chinese-SimpleQAKnowledge78.9%DeepSeek-AI
MRCR 1MReasoning78.7%DeepSeek-AI
CyberGymAgentic76.7%68.7Zhun Wang et al.
BrowseCompAgentic73.2%61.3OpenAI
Toolathlon-VerifiedAgentic70.3%69.1Moonshot AI
Artificial Analysis Coding IndexCoding69.1%67.6Artificial Analysis
MCP AtlasAgentic69.0%60.1OpenAI
DeepSeek DSBench FullStackCoding68.7%DeepSeek-AI
CorpusQA 1MReasoning60.5%DeepSeek-AI
DeepSeek DSBench HardCoding59.6%DeepSeek-AI
DeepSWEAgentic54.4%63.3Datacurve AI
NL2RepoCoding54.2%67.6MiniMax
OpenHarmony Bench v1.0Coding53.8%62.7OpenHarmony Bench authors
SWE-bench ProCoding52.6%54.9Xiang Deng et al.
Artificial Analysis SciCodeCoding50.3%62.7Artificial Analysis
ToolathlonAgentic47.8%62.5OpenAI
GDPval-AA normalizedAgentic46.3%71.1Artificial Analysis
Humanity's Last Exam with toolsAgentic45.1%58.3DeepSeek-AI
Artificial Analysis Agentic IndexAgentic41.7%71.2Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge40.4%63.8Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge38.6%66.4Artificial Analysis
Humanity's Last ExamKnowledge34.8%58.3Center for AI Safety et al.
Artificial Analysis Intelligence IndexKnowledge34.3%65.3Artificial Analysis
Measuring Short-Form Factuality in Large Language ModelsKnowledge34.1%Jason Wei et al.
ApexMath33.0%DeepSeek-AI
Agents' Last ExamAgentic25.2%64.9DeepSeek-AI
AutomationBenchAgentic25.1%56.9Moonshot AI
Critical Physics TasksReasoning16.6%73.2Artificial Analysis

29 benchmarks count, from 31 of 41 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from DeepSeek

DeepSeek V4.1 Flash70.6DeepSeek V4 Pro 081368.8DeepSeek V4 Pro 042362.4DeepSeek V4 Flash 042363.0DeepSeek V3.2 (Thinking)49.9