NVIDIA Announces Advances in AI Factory Efficiency at AI Infra Summit
NVIDIA showcased advancements in AI factory efficiency, including collaborations with Amazon and d-Matrix, and demonstrated performance improvements with NVIDIA DSX MaxLPS and Vera Rubin CPUs.
NVIDIA announced significant advancements in AI factory efficiency at the AI Infra Summit, held at the Santa Clara Convention Center. Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, discussed new collaborations and platform advancements designed to optimize energy usage and increase token generation efficiency.
Amazon’s Annapurna Labs is working with NVIDIA on NVHBM, a custom high-bandwidth memory technology. Additionally, d-Matrix is integrating with NVLink Fusion to combine NVIDIA Vera CPUs with d-Matrix Raptor XPUs, delivering ultra-low-latency inference at scale.
Emerald AI and NVIDIA demonstrated a commercial AI factory flexible-load program in collaboration with Silicon Valley Power. Lambda, an AI cloud provider, improved performance per watt by 23% using NVIDIA DSX MaxLPS on NVIDIA Blackwell servers. DSX MaxLPS continuously monitors power consumption across GPUs and racks, dynamically shifting available power where it’s needed most, thereby increasing cluster-wide token throughput by 24%.
NVIDIA also unveiled the Vera Rubin NVL72, which sets a new benchmark for AI throughput in MLPerf Inference v6.1. The platform is designed to optimize system performance, infrastructure scaling, and software optimization, enabling enterprises to run any model and workload on the same infrastructure.
Startups such as Perplexity and Daytona are also celebrating performance results with the NVIDIA Vera CPU, showing significant gains in agentic execution and sandbox starts.
Source: nvidia-newsroom
