DeepSeek V4 Pro 0423

DeepSeek V4 Pro 0423 is a model from DeepSeek. 27 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.62.4 ±4.6
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.33/s
Input / 1MUS dollars per 1M input tokens.$0.955
Output / 1MUS dollars per 1M output tokens.$1.91
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1451 (#43)

50 is the middle of the board. The range shows the doubt in the index.

54,130 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
61.2
CodingCode writing and repair.
65.6
ReasoningLogic problems and puzzles.
56.4
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
62.9
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
62.7
MathMath problems.
61.3

Results

27 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math96.7%65.6max effortEpoch AI
LiveBench MathematicsMath92.9%72.025 Jun 2026LiveBench
GPQA DiamondKnowledge90.9%62.21 Sept 2026Vals AI
GPQA diamondKnowledge89.6%61.0max effortEpoch AI
LiveCodeBenchCoding87.5%63.91 Sept 2026Vals AI
MMLU ProKnowledge87.1%57.81 Sept 2026Vals AI
SWE-benchCoding86.9%67.21 Sept 2026Vals AI
LiveBench ReasoningReasoning84.3%73.025 Jun 2026LiveBench
LiveBench LanguageKnowledge80.1%67.225 Jun 2026LiveBench
SWE-Bench verifiedCoding77.6%59.8max effortEpoch AI
LiveBench Data AnalysisReasoning76.9%62.825 Jun 2026LiveBench
LiveBench CodingCoding73.6%59.825 Jun 2026LiveBench
Vibe Code Bench v1.1Coding66.1%69.7OpenHands21 Sept 2026Vals AI
LiveBench Instruction FollowingInstruction65.0%62.725 Jun 2026LiveBench
Terminal-Bench 2.0Agentic56.2%62.84 Jun 2026Vals AI
SkillsBenchCoding52.6%67.5OpenHands11 Sept 2026Vals AI
Terminal-Bench 2.1Agentic52.4%55.021 Sept 2026Vals AI
IOICoding51.6%68.121 Sept 2026Vals AI
LiveBench Agentic CodingAgentic48.8%62.625 Jun 2026LiveBench
SimpleQA VerifiedKnowledge47.0%65.1max effortEpoch AI
FrontierMath-Tiers-1-3-v2-PrivateMath45.3%56.9max effortEpoch AI
IOI v1Coding35.8%60.39 Aug 2026Vals AI
Code MigrationCoding33.9%66.621 Sept 2026Vals AI
ProofBench v1.1Math33.0%62.821 Sept 2026Vals AI
Chess PuzzlesReasoning20.0%49.4max effortEpoch AI
Vibe Code Bench 1-100Coding17.5%71.0OpenHands16 Sept 2026Vals AI
Mystery Game PuzzlesReasoning17.0%52.0none effortEpoch AI
Agent Arena command recoveryAgentic3.270.915 Sept 2026LMArena
FrontierMath-Tier-4-v2-PrivateMath2.4%49.1max effortEpoch AI
Terminal-Bench 4.0Agentic1.0%60.221 Sept 2026Vals AI
ProgramBenchCoding0.0%21 Sept 2026Vals AI
Agent Arena task outcomeAgentic-1.765.415 Sept 2026LMArena
Agent Arena steerabilityAgentic-2.165.015 Sept 2026LMArena

27 benchmarks count, from 32 of 33 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from DeepSeek

DeepSeek V4.1 Flash70.6DeepSeek V4 Pro 081368.8DeepSeek V4 Flash 073164.6DeepSeek V4 Flash 042363.0DeepSeek V3.2 (Thinking)49.9