NVIDIA has opened the source code for the cuFile API and associated GPU-direct storage components. The change gives storage vendors and software developers greater access to the technology used by GPUDirect Storage to move information between NVMe storage and GPU memory without routing every transfer through a conventional CPU-managed system-memory buffer. The objective is to develop lower-latency storage capable of supporting retrieval-augmented generation, agentic AI, large inference datasets and other GPU-heavy workloads. Even a small reduction in those idle periods can improve the economics of large GPU clusters. NVIDIA storage development Details Opened component cuFile API and supporting storage stack Data path NVMe storage to GPU memory through DMA Primary technology GPUDirect Storage Storage-Next participants Approximately 40 companies Initial workloads AI inference, RAG, agents and large datasets Replacement for HBM/GDDR No; operates as a slower capacity tierSource: NVIDIA