AI agent memory is part of memory, says Vast Data CTO

AI agent storage places new demands on infrastructure as agents run longer sessions and spread across the enterprise. Maintaining this context and making it available when needed puts pressure on storage capacity and data movement.

These requirements go beyond the context of an individual interaction. According to Alon Horev, (pictured) co-founder and chief technology officer of Vast, corporate agents also need shared knowledge that persists across sessions.

“Storage for agents is a little different,” Horev said. “First, there are multiple types of memory. There is long-term memory, where an agent can see past conversations and interactions and look back and learn from their past experiences.”

Horev spoke with John Furrier and Dave Vellante of theCUBE Research at Fully Connected 2026 during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed AI agent storage, key-value cache offloading, and the shift to data movement as the next bottleneck. (*Disclosure below.)

Why AI agent memory goes beyond the GPU

The pressure first appears in the conclusion. Each long-running session stores its KV cache in the graphics processing unit’s memory, and a session with half a million tokens can take up a tenth to a twentieth of a GPU’s memory, Horev explained.

“It’s also possible that the agent will stop talking to him [large language model] because it’s about compiling code, it’s about testing software, or I as a human want to have a cup of coffee,” he said. “What you’re seeing is that if you could extend the memory wall and basically move those sessions into memory, you can avoid that repeated recomputing.”

Vast’s approach uses storage in layers. First, GPU memory is used, then central processing unit memory on the same machine, then persistent media capable of storing petabytes of KV cache, using Nvidia Corp.’s Dynamo software. orchestrated the process, Horev noted.

“You can move a session from a busy GPU to a less busy GPU and either move the KV cache across the network or read it from Vast,” he said. “If you think about inference as a distributed problem, where you have the ability to use GPU memory, CPU memory, and Vast across a fleet of machines, you have more options and more optimized scheduling.”

The stakes are rising as companies deploy thousands of agents to process sensitive data and act on behalf of customers. These companies need to record everything their agents do and retain it for a specific period of time, which Horev says makes AI agent memory both a governance and performance advantage. Vast has also introduced a confidential computing service for sensitive workloads.

“These conversations that the agent has are also worth their weight in gold,” he said. “It’s the same information that can be used for fine-tuning, training, or building specialized models.”

Here’s the full video interview, part of SiliconANGLE and theCUBE’s coverage of Fully Connected 2026:

(*Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over theCUBE or SiliconANGLE content.)

Photo: SiliconANGLE

Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.

  • Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
  • Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
SiliconANGLE Media is a recognized leader in digital media innovation, combining breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios – with flagship locations in Silicon Valley and the New York Stock Exchange – SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.

Avatar photo
Written by

Mira Edora

Mira Edora is a writer and contributor at CKSOR, creating clear and engaging articles on current topics, technology, science, lifestyle, and stories of interest to readers. She enjoys researching new developments and presenting useful information in a simple, accessible way. Through her writing, Mira aims to keep readers informed with timely, informative, and easy-to-understand content.

Leave a Comment