Anthropic · Status: Active · Released 2025-05-22 · Accessibility: API access
“Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks' [online resource].” · epoch.ai/benchmarks · Licensed CC BY 4.0 · data last verified 2026-07-26 · Methodology · Data Sources
| Benchmark | Score | Source |
|---|---|---|
| Epoch Capabilities Index (ECI) | 142.9 | Epoch AI Benchmarking Hub |
| SWE-bench Verified | 0.707 | Epoch AI Benchmarking Hub |
| GPQA Diamond | 0.763 (16k) | Epoch AI Benchmarking Hub |
| MATH Level 5 | 0.850 | Epoch AI Benchmarking Hub |
| OTIS Mock AIME 2024-2025 | 0.644 (27k) | Epoch AI Benchmarking Hub |
| FrontierMath | 0.045 | Epoch AI Benchmarking Hub |
| FrontierMath Tier 4 | 0.042 (27k) | Epoch AI Benchmarking Hub |
| SimpleQA Verified | Not measured | Epoch AI Benchmarking Hub |
| Chess Puzzles | Not measured | Epoch AI Benchmarking Hub |
No licensed per-token pricing exists for this model yet. Not shown as an estimate — see Data Sources.
History begins 2026-07-26 — one data point so far. A trend needs at least two.
Scores on this page come from Epoch AI Benchmarking Hub, licensed CC BY 4.0. See the full methodology and data sources pages for how rankings, variants and ties are computed, and the full AI Model Rankings table for how Claude Opus 4 compares to every other measured model.