NVIDIA presented new AI factory efficiency results and platform collaborations Tuesday at the AI Infra Summit, emphasizing Vera Rubin systems, NVIDIA DSX power management and NVLink technologies. Lambda reported that DSX MaxLPS raised cluster-wide token throughput by 24% within the same power budget while improving performance per watt by 23%.
Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, framed validated agentic tokens per megawatt as an increasingly important infrastructure metric as agentic AI workloads require greater performance, efficiency and scale. NVIDIA said its full-stack approach spans Vera Rubin systems, Dynamo inference software, NeMo libraries, NVLink scale-up computing, Spectrum-X Ethernet, ConnectX SuperNICs and BlueField products for storage and security.
Lambda’s results were described as the first validation of NVIDIA DSX MaxLPS on NVIDIA Blackwell servers. The AI cloud provider operated 19 nodes within a power budget normally allocated to 16 full-power nodes, increasing throughput from about 4 million to 5 million tokens per second. NVIDIA said DSX MaxLPS, which reallocates available power across GPUs and racks, can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories in the right deployment environments.
NVIDIA also reported that Vera Rubin NVL72 delivered up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on the DeepSeek V4 Pro model in the SemiAnalysis AgentX benchmark, alongside up to 45x lower cost per million tokens. Separately, NVIDIA said a platform combining Vera Rubin NVL72 and Groq 3 LPX can provide up to 35x higher token throughput per megawatt than GB200 NVL72 for models with more than 2 trillion parameters at long context. On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX reached 2,529 output tokens per second per user.
Grid flexibility was another focus. Emerald AI and NVIDIA demonstrated automated load reduction with Silicon Valley Power, responding to hundreds of demand signals while protecting AI workload performance, according to NVIDIA. Emerald AI plans to use NVIDIA DSX Flex in its Conductor software to adjust energy consumption in response to grid signals, temporarily pausing lower-priority work while critical jobs continue.
NVIDIA announced additional ecosystem work around memory and scale-up computing: Amazon’s Annapurna Labs is working with NVIDIA on NVHBM custom high-bandwidth memory technology, while d-Matrix is integrating NVIDIA NVLink Fusion with its Raptor XPUs. NVIDIA also detailed NVLink 6 resiliency features intended to detect, contain and recover from faults as AI factories expand to hundreds of thousands of GPUs, including forward error correction, physical-layer retry, dynamic routing and link rebalancing.

