Home › AI › Olmo-core 3 Launches With Trillion-Par
AI

Olmo-core 3 Launches With Trillion-Parameter MoE Training

The open-source framework introduces distributed data parallelism and low-precision MXFP8 support to scale mixture-of-experts AI models on GPU clusters.

WHAT YOU NEED TO KNOW
  • Olmo-core 3 replaces fully sharded data parallelism with distributed data parallelism, boosting 47-billion-parameter throughput from 19,400 to 52,000 tokens per second per GPU on eight NVIDIA B300s.
  • The framework supported a 1.2-trillion-parameter model across 512 NVIDIA B300 GPUs with 858 TFLOP/s per GPU throughput.
  • Using the MXFP8 format increased throughput by approximately 21% over BF16 and reduced peak active memory from 103 GiB to 95 GiB.
  • Capacity tests incorporating DeepEP v2 reached 2.38 trillion total parameters.

The Olmo development team released Olmo-core 3 on October 1, 2026, delivering an open training system engineered to scale mixture-of-experts language models into the trillion-parameter range, according to technical documentation published on Hugging Face.

Developers replaced the framework's earlier fully sharded data parallelism design with distributed data parallelism. Rather than gathering and resharding model weights for every batch of data, Olmo-core 3 keeps experts resident on GPUs and routes relevant tokens to those locations. In benchmark testing on eight NVIDIA B300 GPUs, a 47-billion-parameter mixture-of-experts model reached 52,000 tokens per second per GPU, up from 19,400 tokens per second under the prior sharded setup.

Memory and routing optimizations

Engineers organized the framework around three distribution methods to divide training states across hardware:

  • Expert parallelism assigns distinct segments of the expert pool to individual GPUs.
  • Pipeline parallelism divides model layers across separate GPU groups.
  • A distributed optimizer spreads parameter update states across the cluster instead of duplicating them on every GPU.

The system pairs these parallel structures with rowwise expert parallelism to place routed data straight into expert input buffers, GPU-resident routing to keep metadata off the CPU, and grouped GEMM to batch expert computations. In a test scaling an expert pool from 8 to 128 while routing four experts per token, total model capacity expanded from 4.6 billion to 47 billion parameters, active parameters remained steady at roughly 3.2 billion per token, and training throughput declined by less than 5%.

Hardware benchmarks also implemented MXFP8, a reduced-precision numerical format. Evaluated on four NVIDIA B300 GPUs with uniform workload distributions, MXFP8 increased end-to-end training throughput by roughly 21% over a BF16 baseline while lowering peak active memory from 103 GiB to 95 GiB, with most efficiency gains occurring in feed-forward computation and inter-expert data transfers.

Scale and empirical findings

Benchmark tests evaluated a 1.2-trillion-parameter model with 58.36 billion active parameters per token across 512 NVIDIA B300 GPUs using random routing, achieving a top throughput of 858 TFLOP/s per GPU. A separate capacity run using DeepEP v2 communication libraries scaled to 2.38 trillion total parameters.

Testing also exposed several practical training constraints documented in the release. A routing balance metric produced an anomaly the authors called token gerrymandering, where balance scores improved while actual hardware workloads grew more uneven. Reductions to expert learning rates yielded no benefits in the tested model family. Furthermore, overlapping computation and communication across separate GPU streams slowed execution in specific tests, and GPU calculation speeds varied based on numerical input values despite identical matrix shapes.

The training stack will serve as the base for the upcoming generation of Olmo, which will adopt a mixture-of-experts architecture trained across the team's largest dataset and longest context window to date.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →