Xentir Media Source-backed AI & tech
LIVE INTELLIGENCE · part of the Xentir AI Index

Grok 4 vs o3

Grok 4 · xAI, released 2025-07-09
o3 · OpenAI, released 2025-04-16

“Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks' [online resource].” · epoch.ai/benchmarks · Licensed CC BY 4.0 · data last verified 2026-07-26 · Methodology · Data Sources

Head-to-head

Measured on 6 of 9 tracked benchmarks for both models — Grok 4 leads on 3, o3 leads on 2, tied on 1.

BenchmarkGrok 4 o3Leads
Epoch Capabilities Index (ECI)147.0147.1o3
SWE-bench VerifiedNot measured0.623 (medium)Not measured on both
GPQA Diamond0.8700.818 (high)Grok 4
MATH Level 5Not measured0.978 (high)Not measured on both
OTIS Mock AIME 2024-20250.8400.839 (high)Grok 4
FrontierMath0.1970.187 (high)Grok 4
FrontierMath Tier 40.0210.021 (high)Tied
SimpleQA Verified0.4790.530 (high)o3
Chess Puzzles0.280Not measuredNot measured on both
Benchmark results measure specific tasks and should not be treated as a universal measure of model quality. A model with no row here for a benchmark has not been measured on it — that is left blank, never estimated.

Full profiles

Grok 4 → · o3 → · the full AI Model Rankings table