Nvidia ties AI factory economics to inference and efficiency – SiliconANGLE

The economics of AI factories increasingly depend on more than just access to powerful graphics processing units. Because agent systems rely on multiple models, databases and tools, the entire data center must function as a single computer system.

This transition shifts attention from individual chips to the infrastructure that converts computing capacity into useful intelligence. According to Ian Buck (pictured), vice president and general manager of hyperscale and HPC at Nvidia Corp., networks, storage, processors and software must work together at scale while increasing the power produced from each unit of energy.

“Instead of cars or devices or PCs, it’s tokens,” he said. “These assets are not IT; they are not costs. They are actually valuable, revenue-generating, fungible, durable and productive parts of an economy.”

Buck spoke with theCUBE Research’s Dave Vellante and John Furrier at the Fully Connected event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed inference, system-level design, and the changing economics of AI infrastructure. (*Disclosure below.)

Inference is changing the economics of the AI ​​factory

The commercial output of an AI factory is based on inference, in which models used process requests and produce tokens. But inference is not a substitute for training, as companies continually update the models they deploy as data and market conditions change, Buck said.

“All of these services are not just about fire and forget,” he said. “As companies use these models, they refine them, they align them, they add more data to them. Keeping them updated and aware of them – that’s actually a bit of training. We see the work in reinforcement learning and online adaptation.”

Low latency creates another layer of economics for workloads where faster thinking is of greater value. Nvidia’s Groq 3 LPX inference accelerator works with its Vera Rubin platform to increase token rates per user for time-critical workloads.

“If these tokens have value to enable the fastest possible thinking, LPX can be boosted in addition to Vera Rubin to enable this,” Buck said. “We’re seeing a lot of interest in areas like fintech and other areas where things happen in real time.”

Performance makes efficiency a system-level priority

Power capacity ultimately limits how much computing infrastructure a data center can provide. This limitation makes tokens per watt a key measure of AI factory economics and puts pressure on vendors to improve performance with each generation of hardware.

“Data centers have a natural ceiling, and that ceiling is actually their power,” Buck said. “With each GPU generation, we ensure our tokens per watt are more than 10x more efficient. In fact, we saw this with Blackwell – we ended up with a 30x improvement in tokens per watt.”

Change is also changing the scale at which systems must be designed and operated. CoreWeave allows customers to select configurations or leverage higher-level inference services that optimize the balance between throughput and token speed, Buck said.

“CoreWeave can do that for customers,” he said. “You don’t have to feel overwhelmed by all the choices. That’s why our partner ecosystem is so important.”

Here is the full video interview, part of SiliconANGLE and theCUBE’s coverage of the Fully Connected event:

(*Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s coverage, nor other sponsors have editorial control over theCUBE or SiliconANGLE content.)

Photo: SiliconANGLE

Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.

  • Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
  • Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
SiliconANGLE Media is a recognized leader in digital media innovation, combining breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios – with flagship locations in Silicon Valley and the New York Stock Exchange – SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.

Avatar photo
Written by

Mira Edora

Mira Edora is a writer and contributor at CKSOR, creating clear and engaging articles on current topics, technology, science, lifestyle, and stories of interest to readers. She enjoys researching new developments and presenting useful information in a simple, accessible way. Through her writing, Mira aims to keep readers informed with timely, informative, and easy-to-understand content.

Leave a Comment