Home › AI › OpenAI Launches GPT-6 Astra Ultrafast
AI

OpenAI Launches GPT-6 Astra Ultrafast on NVIDIA Blackwell GPUs

OpenAI released GPT-6 Astra Ultrafast, using NVIDIA Blackwell GPUs to achieve up to eight times faster token generation than standard mode.

WHAT YOU NEED TO KNOW
  • GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available in the OpenAI API, ChatGPT Work, and Codex.
  • The Ultrafast mode delivers up to 8x faster token generation than Astra Standard mode.
  • OpenAI used internal AI models to refine inference kernels and software for NVIDIA Blackwell and Rubin architectures.
EPOCH CAPABILITIES INDEX (ECI)GPT-6 Astra169Claude Opus 5.5167Claude Sonnet 5.5165Claude Fable 5.1165Claude Fable 5164Claude Opus 5163Source: Epoch AI Benchmarking Hub - CC BY 4.0 - as of 2026-10-01

OpenAI deployed GPT-6 Astra Ultrafast on NVIDIA Blackwell graphics processing units, NVIDIA reported on Oct. 1. The model is available immediately through the OpenAI API and to eligible ChatGPT Work and Codex subscribers. The configuration delivers up to 8x faster token generation compared to OpenAI's Astra Standard mode.

NVIDIA reported that the speed increase stems from inference optimizations engineered by OpenAI to tap directly into the capabilities of the Blackwell architecture. These optimizations shorten edit-test-debug cycles for automated coding agents and cut waiting periods between tool calls in agentic workflows, where a system repeatedly writes code, executes a tool, checks the output, and determines its next step.

Philippe Tillet, inference lead at OpenAI, said NVIDIA’s tooling and documentation allowed OpenAI to make its models exceptionally skilled at programming Blackwell and Rubin GPUs. Tillet said Astra translates that capability into high-performance kernels, making NVIDIA chips compelling across latency, throughput, and cost limits as agents tackle multi-step coding tasks.

Uday Ruddarraju, chief technology officer of compute at OpenAI, said the company used internal AI models to optimize the inference software operating on NVIDIA GPUs. Ruddarraju credited the programmability of the platform with enabling the acceleration behind Astra Ultrafast, adding that testing software improvements on deployed infrastructure can yield faster responses over time.

Developers and researchers can repurpose the same programmable NVIDIA platform across model training, inference, and reinforcement learning. This shared setup helps teams adjust compute resources to changing demand without overprovisioning hardware. OpenAI has published implementation details, pricing, and access parameters in its dedicated Ultrafast guide for developers using the API.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →