theCUBE + NYSE Wired

Nvidia's 'scale-in': the network that guards the AI factory

Key points

Key takeaways from theCUBE + NYSE Wired interview with Nvidia senior vice president of networking Gilad Shainer (September 2026):

Scale-in, the new pillar. Shainer describes agentic AI as the most complex workload ever created — an agent is a flow of operations spanning compute, storage, memory and networking, not a single inference call. Nvidia has added scale-in to its AI-factory stack alongside scale-up (NVLink), scale-out (InfiniBand and Spectrum-X Ethernet), scale-across and context-scale infrastructure.

BlueField-4 as the security plane. Every BlueField-4 DPU connects to every ConnectX NIC and manages 7.2 terabits of secure east-west traffic across the factory, versus a 400 gig access pipe in the previous generation. Because it runs out of band, it collects telemetry and enforces policy without degrading GPU or storage performance.

DPU is data centre infrastructure on a chip. Just as CUDA exposes GPU capabilities to applications, the DOCA framework and DPU APIs expose BlueField-4's in-silicon security engines, storage acceleration and telemetry to developers — on a device fully isolated from where agents run.

Security inside the factory, not just at the front door. Early AI factories split into a training-focused back-end network and a disconnected front-end access network; agentic AI demands security across the entire east-west fabric — memory, storage, compute — which is what the scale-in layer provides.

BlueField-4 pairs ConnectX-9 with a Grace CPU. The SuperNIC handles all DMA operations, memory and storage access and telemetry, while the Grace CPU runs the control logic; a GPU may be attached in future so the device can follow exactly what agents are doing. Shainer notes Grace is powerful enough and more cost-effective than a Vera for this role.

OpenShell keeps agents in their sandbox. Nvidia OpenShell is an open-source runtime that sits underneath agents, controlling where they can go and what they may access; BlueField-4 monitors behaviour out of band from a device no agent can influence.

Same arc as browsers and hypervisors. Shainer frames scale-in as the latest step in a familiar pattern: web pages needed browser sandboxes, multi-user servers needed hypervisors and DPUs, and now agents need an isolated infrastructure layer — each time so that the underlying technology can keep accelerating rather than slow down.

AI factories iterate at an annual cadence. Scale-up has grown from 8 GPUs (NVLink) to 72, 576 and now 1,152; scale-out connects hundreds of thousands of GPUs. Nvidia ships a full reference design and digital twin of everything it builds, so partners can run, explore and join the ecosystem.

Read more: theCUBE + NYSE Wired

More top stories