The Neocloud market is moving away from its origins as a stopgap solution for scarce graphics processors. AI-native startups now choose their infrastructure based on latency, burst capacity and openness, not just chip availability.
That shift is happening at CoreWeave Inc., which is expanding beyond GPU computing power into networking, storage and software as inference demand grows. At the same time, LlamaIndex Inc. has evolved from an open-source retrieval-enhanced generation framework to a model builder that rents its computing power rather than owning it, according to Jerry Liu (pictured right), co-founder and CEO of LlamaIndex.
“We’re effectively a specialized AI lab right now that focuses solely on building models for parsing and extracting documents. We retrain open-weight models, collect our own datasets, and do really, really well at parsing and reading documents to essentially extract that data,” Liu said. “We are very passionate about ensuring that everything we do can actually align with the Pareto frontier of performance, cost and latency for our customers.”
Liu and Lukas Biewald (left), senior vice president of AI initiatives at CoreWeave, spoke with theCUBE’s John Furrier and Dave Vellante at Fully Connected during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed long-standing agents, governance, and why AI-native startups turn to specialized clouds for inference-intensive workloads. (*Disclosure below.)
Why bursty AI workloads prefer the Neocloud model
The computational effort of LlamaIndex was almost non-existent a year ago. Liu says his workload today is about 75% inference and 25% training, processing millions of document pages daily for financial, legal and insurance clients whose paperwork arrives in bulk. The company does not own a GPU cluster, so guaranteed capacity is more important than owning hardware.
“We serve a wide variety of clients with extremely demanding and heavy workloads,” Liu said. “We really, really need to make sure we have the right capacity to serve our customers without being throttled.”
CoreWeave believes that capacity alone is not the differentiator. Biewald joined the company through the acquisition of Weights & Biases, the AI observability startup he co-founded, and CoreWeave used the event to launch CoreWeave Forge, a development layer that performs training, inference, evaluation and agent development in a connected environment. Coming from the software industry, Biewald noted, he first questioned how important chip configuration could really be.
“I’m telling you, the answer is ‘massive difference,'” Biewald said. “I’m talking about orders of magnitude differences depending on how you network the chips [and] how to do power distribution.”
According to Biewald, openness also distinguishes CoreWeave from hyperscalers. While vendors like Amazon Web Services Inc. rely on proprietary application programming interfaces that make moving workloads difficult, CoreWeave follows Nvidia Corp.’s approach. recommended standard network protocols, providing broader open source support. Analysts have observed that CoreWeave is expanding its portfolio in the same way AWS did in its early days, even as the company resists the “neocloud” label.
“CoreWeave knows that everyone comes from a different cloud,” Biewald said. “Everyone will host their web service on AWS or GCP, not CoreWeave. CoreWeave is fine with that, so CoreWeave works much better with the other clouds.”
This ecosystem suggests a larger shift in who gets to build intelligence, Liu noted. After training, a small open weight model is still today a skill reserved for a small group of specialists. Abundant Neocloud capacity, combined with rapidly improving coding agents, could make this work accessible to far more people.
“Everyone is starting to get really good at defining observability and evaluations and the right metrics to focus on,” Liu said. “I think there will be a world where we basically just automate the entire loop and make it accessible to everyone.”
Here’s the full video interview, part of SiliconANGLE and theCUBE’s coverage of Fully Connected:
(*Disclosure: TheCUBE is a paid media partner for the Fully Connected 2026 event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over theCUBE or SiliconANGLE content.)
Photo: SiliconANGLE
Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.
- Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
- Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.