Granite 4.2 8B

Granite 4.2 8B is a reasoning model from IBM in the Granite 4.2 family. 22 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.37.6 ±4.9
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.51/s
Input / 1MUS dollars per 1M input tokens.$0.06
Output / 1MUS dollars per 1M output tokens.$0.25
ContextMaximum tokens in one request.131K
EloLMArena rating and rank.1317 (#225)

50 is the middle of the board. The range shows the doubt in the index.

3,087 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
38.7
CodingCode writing and repair.
36.5
ReasoningLogic problems and puzzles.
39.2
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
34.5
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
58.4
MathMath problems.
41.9

Results

22 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
American Invitational Mathematics Examination 2025Math86.7%45.1Mathematical Association of America
Instruction Following BenchmarkInstruction79.3%58.4Benchmark authors
Harvard-MIT Mathematics Tournament February 2025Math78.3%38.7Qwen
Massive Multitask Language Understanding ProfessionalKnowledge74.0%37.1Yubo Wang et al.
LiveCodeBench v6Coding73.2%43.7LiveCodeBench maintainers
Graduate-Level Google-Proof Q&AKnowledge64.1%37.5David Rein et al.
Artificial Analysis GPQA DiamondKnowledge63.1%33.2Artificial Analysis
τ³-Bench Tool-Agent-User EvaluationAgentic58.1%44.5Sierra Research
Berkeley Function Calling Leaderboard v4Agentic52.4%36.9Arcee AI
Software Engineering Benchmark VerifiedCoding47.7%35.8Carlos E. Jimenez et al.
Artificial Analysis Long Context ReasoningReasoning45.0%39.4Artificial Analysis
Scientific Code BenchmarkCoding36.1%45.7Benchmark authors
Artificial Analysis SciCodeCoding31.5%36.6Artificial Analysis
Artificial Analysis Coding IndexCoding22.4%34.7Artificial Analysis
Terminal-Bench 2.1 (provider run)Agentic20.6%36.2DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic20.6%36.2DeepSeek-AI
SWE-bench ProCoding19.1%22.4Xiang Deng et al.
Artificial Analysis Omniscience AccuracyKnowledge11.2%27.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge11.1%36.2Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge9.7%35.1Artificial Analysis
Artificial Analysis Agentic IndexAgentic3.7%40.2Artificial Analysis
Critical Physics TasksReasoning0.3%39.0Artificial Analysis
GDPval-AA normalizedAgentic0.0%35.6Artificial Analysis

22 benchmarks count, from 23 of 23 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence stated

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from IBM

Granite 4.2 30B41.9Granite 4.2 3B32.8Granite-4.0-H-1B22.7Granite-4.0-H-350M20.6Granite-4.0-350M20.2