AI inference will be fully utilized when CoreWeave launches Forge

AI inference is quickly becoming the workload that will determine the economic viability of the AI ​​boom. Training created the first wave of GPU clouds, but deploying models faster and more cost-effectively will define the next wave.

This shift is pushing specialized cloud providers beyond just GPU capacity into storage, networking and software. According to Urvashi Chowdhary (pictured), vice president of product and AI services at CoreWeave Inc., one vendor is layering managed services for training, post-training and inference on top of its infrastructure.

“I think if you look at the AI ​​developer’s journey, they want to solve a problem and do it as quickly as possible with the best performance and scalable cost,” Chowdhary said. “So we’ve really focused on building the layers of our stack and building on top of the reliable infrastructure to develop more managed services, whether it’s for training, retraining or inference.”

Chowdhary spoke with theCUBE Research’s Dave Vellante and John Furrier at the Fully Connected event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed AI inference, managed services across the stack, and CoreWeave RL Rollouts, a new feature to accelerate agent model iteration. (*Disclosure below.)

Optimizing the AI ​​inference stack layer by layer

The demand for inferences is increasing rapidly. A survey of CoreWeave customers and prospects conducted by theCUBE Research found that a healthcare customer’s share of inference workload increased from about 10% in the first year to 40% in the second year, with a share of about 50% expected within 12 months. CoreWeave’s answer is to optimize every layer above the hardware, from the vLLM engine to quantized models and custom speculative decoders, Chowdhary explained.

“One thing that we’ve been very passionate about is using open source tools and technologies to contribute to open systems so customers have flexibility, and then also building our services on top of each other,” she said.

Reinforcement learning increases the pressure. According to Chowdhary, when customers train agent models with rewards and verifiers, inference becomes the bottleneck in adoption. CoreWeave RL Rollouts, a preview feature built on Nvidia Corp.’s Dynamo framework. based, loads new checkpoints into a live deployment. In tests, the feature improved model reload latency by 15x compared to a base configuration.

“When you do RL rollouts, you’re constantly creating new model checkpoints and versions and you want these to be integrated into your inference setup so you can scale it independently,” she said. “And we were able to speed this up by 15x, meaning your training runs fast and your inference scales as you continue to train your model quickly.”

These capabilities are now included in CoreWeave Forge, a platform launched at the event that connects deployment, observability, retraining and assessment. Forge is free to start. Paid tiers provide additional features and expand access to AI development tools and services to individual developers, Chowdhary explained.

“We want to provide access to these leading technologies,” she said. “Even if you’re a single developer logging on alone today, you’ll still get the best performance, the best reliability, and no compromises.”

Here is the full video interview, part of SiliconANGLE and theCUBE’s coverage of the Fully Connected event:

(*Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over theCUBE or SiliconANGLE content.)

Photo: SiliconANGLE

Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.

  • Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
  • Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
SiliconANGLE Media is a recognized leader in digital media innovation, combining breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios – with flagship locations in Silicon Valley and the New York Stock Exchange – SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.

Avatar photo
Written by

Mira Edora

Mira Edora is a writer and contributor at CKSOR, creating clear and engaging articles on current topics, technology, science, lifestyle, and stories of interest to readers. She enjoys researching new developments and presenting useful information in a simple, accessible way. Through her writing, Mira aims to keep readers informed with timely, informative, and easy-to-understand content.

Leave a Comment