OpenAI deployed GPT-6 Astra Ultrafast on NVIDIA Blackwell graphics processing units, NVIDIA reported on Oct. 1. The model is available immediately through the OpenAI API and to eligible ChatGPT Work and Codex subscribers. The configuration delivers up to 8x faster token generation compared to OpenAI's Astra Standard mode.
NVIDIA reported that the speed increase stems from inference optimizations engineered by OpenAI to tap directly into the capabilities of the Blackwell architecture. These optimizations shorten edit-test-debug cycles for automated coding agents and cut waiting periods between tool calls in agentic workflows, where a system repeatedly writes code, executes a tool, checks the output, and determines its next step.
Philippe Tillet, inference lead at OpenAI, said NVIDIA’s tooling and documentation allowed OpenAI to make its models exceptionally skilled at programming Blackwell and Rubin GPUs. Tillet said Astra translates that capability into high-performance kernels, making NVIDIA chips compelling across latency, throughput, and cost limits as agents tackle multi-step coding tasks.
Uday Ruddarraju, chief technology officer of compute at OpenAI, said the company used internal AI models to optimize the inference software operating on NVIDIA GPUs. Ruddarraju credited the programmability of the platform with enabling the acceleration behind Astra Ultrafast, adding that testing software improvements on deployed infrastructure can yield faster responses over time.
Developers and researchers can repurpose the same programmable NVIDIA platform across model training, inference, and reinforcement learning. This shared setup helps teams adjust compute resources to changing demand without overprovisioning hardware. OpenAI has published implementation details, pricing, and access parameters in its dedicated Ultrafast guide for developers using the API.