NVIDIA announced at the Future of Memory and Storage conference in Santa Clara, California, that it is open sourcing its cuFile application programming interfaces and underlying vertical storage software stack. The cuFile software, a component of NVIDIA GPUDirect Storage, allows graphics processing units to read from and write to storage directly without relying solely on central processing units.
In reporting published on its technical site, NVIDIA revealed that Google, Intel, Meta, and NVIDIA are serving as inaugural maintainers for the open API platform. The open availability of cuFile aims to improve interoperability across hardware platforms and support cybersecurity initiatives such as the Open Secure AI Alliance by providing data access within microseconds.
Performance Benchmarks and SCADA
Tests conducted by NVIDIA show that its Vera CPU, integrated into the Vera BlueField-4 STX platform, achieves up to 3.21 times higher throughput than a standard x86 CPU during a two-stage compression and encryption pipeline. The company stated that this performance increase allows storage platforms to process incoming artificial intelligence data with less compute infrastructure.
To further align storage hardware with GPU demands, NVIDIA launched the Storage-Next initiative alongside more than 40 flash and storage vendors, including KIOXIA, Micron, and DDN. As part of this effort, NVIDIA introduced SCADA, a framework that permits GPUs to pull specific data subsets directly from storage into high-speed memory. Storage provider DDN is integrating SCADA into Infinia, its software-defined data intelligence platform.
Modular Rack Architecture
NVIDIA built its STX storage foundation using the Vera Rubin platform, Vera BlueField-4 storage processors, and Spectrum-X Ethernet networking. The architecture uses the unified DOCA security stack to enforce continuous security policies across the data path.
Built on top of the STX platform, NVIDIA also introduced CMX Context Memory Storage. This tool offers a dedicated context tier designed to handle multi-turn, long-context agentic AI inference workloads.
