Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) vs Llama 4 Scout
Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) · Google DeepMind, released 2024-09-24
Llama 4 Scout OPEN WEIGHTS · Meta AI, released 2025-04-05
“Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks' [online resource].” · epoch.ai/benchmarks · Licensed CC BY 4.0 · data last verified 2026-09-09 · Methodology · Data Sources
Llama 4 Scout leads on more of the 5 benchmarks where both are measured — Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) on 1, Llama 4 Scout on 3, tied on 1. The remaining 4 of 9 tracked benchmarks cannot be compared: 0 have a result for only one of the two models, 4 for neither. Nothing on this page is estimated to fill those in.
Head-to-head
Measured on 5 of 9 tracked benchmarks for both models — Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) leads on 1, Llama 4 Scout leads on 3, tied on 1.
| Benchmark | Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) | Llama 4 Scout | Leads |
|---|---|---|---|
| Epoch Capabilities Index (ECI) | 129.4 | 129.6 | Llama 4 Scout |
| SWE-bench Verified | Not measured | Not measured | Not measured on both |
| GPQA Diamond | 0.473 | 0.518 | Llama 4 Scout |
| MATH Level 5 | 0.619 | 0.623 | Llama 4 Scout |
| OTIS Mock AIME 2024-2025 | 0.163 | 0.078 | Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) |
| FrontierMath | 0.000 | 0.000 | Tied |
| FrontierMath Tier 4 | Not measured | Not measured | Not measured on both |
| SimpleQA Verified | Not measured | Not measured | Not measured on both |
| Chess Puzzles | Not measured | Not measured | Not measured on both |
What remains uncertain
- 4 tracked benchmarks have no published result for either model.
A missing result means Xentir has not found one published by the source it cites, not that the model performs badly. See Data Sources.
Full profiles
Gemini 1.5 Flash (Sep 2024, 24 Sep 2024) → · Llama 4 Scout → · the full AI Model Rankings table