Alibaba · Status: Active · Released 2025-01-25 · Accessibility: API access
“Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks' [online resource].” · epoch.ai/benchmarks · Licensed CC BY 4.0 · data last verified 2026-07-26 · Methodology · Data Sources
| Benchmark | Score | Source |
|---|---|---|
| Epoch Capabilities Index (ECI) | Not measured | Epoch AI Benchmarking Hub |
| SWE-bench Verified | Not measured | Epoch AI Benchmarking Hub |
| GPQA Diamond | 0.481 | Epoch AI Benchmarking Hub |
| MATH Level 5 | 0.653 | Epoch AI Benchmarking Hub |
| OTIS Mock AIME 2024-2025 | 0.178 | Epoch AI Benchmarking Hub |
| FrontierMath | Not measured | Epoch AI Benchmarking Hub |
| FrontierMath Tier 4 | Not measured | Epoch AI Benchmarking Hub |
| SimpleQA Verified | Not measured | Epoch AI Benchmarking Hub |
| Chess Puzzles | Not measured | Epoch AI Benchmarking Hub |
No licensed per-token pricing exists for this model yet. Not shown as an estimate — see Data Sources.
History begins 2026-07-26 — one data point so far. A trend needs at least two.
No head-to-head comparison has been generated for this model yet — see the full AI Model Rankings table for how it stacks up on each benchmark.
No Xentir coverage of this model yet. See the latest AI news.
Scores on this page come from Epoch AI Benchmarking Hub, licensed CC BY 4.0. See the full methodology and data sources pages for how rankings, variants and ties are computed, and the full AI Model Rankings table for how Qwen Plus (01 25) compares to every other measured model.