CoreWeave announced at its Fully Connected conference in San Francisco that the Nvidia Vera Rubin NVL72 system is now available on its cloud platform, with Cognition as the first customer, having begun using the systems in early September. Cognition reported up to a 4.8x increase in total token throughput for its SWE-2 inference workloads on the new hardware. CoreWeave also announced plans to offer the Vera CPU as a standalone bare-metal product, with some customers expected to begin testing it in coming weeks.
The rapid commercial availability of the Vera Rubin NVL72, which Nvidia claims offers 5x inference performance improvement over Blackwell, signals that next-generation rack-scale GPU infrastructure is moving from demonstration to production faster than prior generations. The addition of a standalone Vera CPU offering marks a strategic expansion for CoreWeave beyond its traditional GPU-focused model, which could reshape how cloud providers position CPU resources for agentic AI workloads.
CoreWeave making Nvidia Vera Rubin NVL72 available with a named early customer (Cognition) is a concrete AI compute infrastructure milestone. Although a related CoreWeave/Forge platform story was previously published, this article appears to cover a distinct availability announcement with a specific customer named, warranting selection.