The Olmo development team released Olmo-core 3 on October 1, 2026, delivering an open training system engineered to scale mixture-of-experts language models into the trillion-parameter range, according to technical documentation published on Hugging Face.
Developers replaced the framework's earlier fully sharded data parallelism design with distributed data parallelism. Rather than gathering and resharding model weights for every batch of data, Olmo-core 3 keeps experts resident on GPUs and routes relevant tokens to those locations. In benchmark testing on eight NVIDIA B300 GPUs, a 47-billion-parameter mixture-of-experts model reached 52,000 tokens per second per GPU, up from 19,400 tokens per second under the prior sharded setup.
Memory and routing optimizations
Engineers organized the framework around three distribution methods to divide training states across hardware:
- Expert parallelism assigns distinct segments of the expert pool to individual GPUs.
- Pipeline parallelism divides model layers across separate GPU groups.
- A distributed optimizer spreads parameter update states across the cluster instead of duplicating them on every GPU.
The system pairs these parallel structures with rowwise expert parallelism to place routed data straight into expert input buffers, GPU-resident routing to keep metadata off the CPU, and grouped GEMM to batch expert computations. In a test scaling an expert pool from 8 to 128 while routing four experts per token, total model capacity expanded from 4.6 billion to 47 billion parameters, active parameters remained steady at roughly 3.2 billion per token, and training throughput declined by less than 5%.
Hardware benchmarks also implemented MXFP8, a reduced-precision numerical format. Evaluated on four NVIDIA B300 GPUs with uniform workload distributions, MXFP8 increased end-to-end training throughput by roughly 21% over a BF16 baseline while lowering peak active memory from 103 GiB to 95 GiB, with most efficiency gains occurring in feed-forward computation and inter-expert data transfers.
Scale and empirical findings
Benchmark tests evaluated a 1.2-trillion-parameter model with 58.36 billion active parameters per token across 512 NVIDIA B300 GPUs using random routing, achieving a top throughput of 858 TFLOP/s per GPU. A separate capacity run using DeepEP v2 communication libraries scaled to 2.38 trillion total parameters.
Testing also exposed several practical training constraints documented in the release. A routing balance metric produced an anomaly the authors called token gerrymandering, where balance scores improved while actual hardware workloads grew more uneven. Reductions to expert learning rates yielded no benefits in the tested model family. Furthermore, overlapping computation and communication across separate GPU streams slowed execution in specific tests, and GPU calculation speeds varied based on numerical input values despite identical matrix shapes.
The training stack will serve as the base for the upcoming generation of Olmo, which will adopt a mixture-of-experts architecture trained across the team's largest dataset and longest context window to date.
