GPU workloads no longer have to ask the CPU to handle storage reads and writes.
At the Flash Memory Summit, NVIDIA made the software that lets a GPU initiate storage I/O public: the cuFile APIs and the vertical storage stack behind them. DDN also announced a collaboration to add support for those GPU-initiated workflows into its storage platforms.
Why that matters is simple: data-heavy training and inference often choke not on raw compute but on how fast data can reach the GPU. letting the GPU talk to storage directly removes a common detour through the CPU, which can raise throughput and cut latency for large models.
How it works, in plain terms: picture a kitchen where chefs currently ask a manager to fetch every ingredient. GPU-initiated I/O hands the shopping list to the chef, so they fetch what they need themselves. NVIDIA calls the plumbing that enables this cuFile.
What changes now is practical. Storage vendors can build native support instead of reverse-engineering a workaround, and operators can redesign clusters to reduce CPU load for I/O-heavy jobs. That should translate into faster training runs and snappier inference at scale.
One important reality: this is infrastructure work. Making the software public doesn’t mean you can run it on a laptop. Real benefit requires compatible GPUs, fast storage, and datacenter integration.
What to watch next: will third-party storage systems and independent benchmarks confirm real-world speedups, and how quickly will operators roll this into production?
