Ai2 releases Olmo-core 3 to make the development of large, mixed-expert LLMs more efficient

Seattle-based artificial intelligence research firm the Allen Institute for AI announced Thursday a large language model development framework that significantly improves the way large language models are trained with mixed experts.

The new Olmo-core 3 framework enables MoE training to reach the trillion-parameter scale while keeping costs low by maintaining computational efficiency.

Expert mixture models work differently than dense AI models: they split the computation into specific parts of the model each time a token is generated, while dense models activate the entire model. An MoE model can contain many more total parameters, the individual controls that optimize model behavior, but it only counts on a small number of them as it generates each part of a response.

Similarly, training an MoE model to activate only parts of it while simultaneously learning each token allows a small section of text – often a word or part of a word – that an AI reads and generates. This reduces overall computing power; However, the full model still needs to be stored in the memory of the graphics processing unit, and coordinating networking between experts during training still incurs additional costs.

According to Ai2, Olmo-core 3 is designed to bridge the gap between dense models and MoE models, allowing the expert pool to grow from eight to 128 while only selecting four experts per token. With the same infrastructure, LLMs can scale to over a trillion parameters.

Benchmarking the next generation of efficient MoE training

In benchmarks, the company reported that Olmo-core 3 processed 52,000 tokens per second on Nvidia B3000 GPUs for a 47 billion parameter model. Compared to Nvidia Corp.’s Megatron Core training architecture, a well-established option for training large MoEs, which reached around 19,400 tokens per second, this represented a jump of around 2.7x in throughput.

In a white paper published about the project, Ai2 noted that the new architecture uses expert parallelism to distribute experts across multiple GPUs, allowing each card to store only a portion of the entire pool of experts. Additionally, the model’s layers, the successive stages that generate inputs, are split across groups of GPUs to reduce how much of the model each GPU must keep in memory. Finally, a distributed optimizer distributes the optimizer state, additional data used to calculate and apply updates during training, across multiple GPUs rather than storing full copies on each GPU.

Overall, this reduces the amount of memory required when scaling the model, since the entire model and its training state do not have to be saved in memory at once.

In addition, Ai2 supports MXFP8, a number format for LLMs that represents some values ​​with fewer bits. It can reduce the amount of computation and the amount of data moved between GPUs.

The new training architecture and its greater efficiency are part of the company’s vision to provide researchers with the tools to build and train larger models. AI models with trillions of parameters are often beyond the reach of those without access to government and corporate infrastructure. Ai2 added that Olmo-core 3 would also allow researchers to adapt MoE training to different hardware, experiment with routing, concurrency and other parts of the system to build a new ecosystem.

The project and associated systems are currently available to developers and the open source community on GitHub.


Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.

  • Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
  • Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
SiliconANGLE Media is a recognized leader in digital media innovation, combining breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios – with flagship locations in Silicon Valley and the New York Stock Exchange – SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.

Avatar photo
Written by

Mira Edora

Mira Edora is a writer and contributor at CKSOR, creating clear and engaging articles on current topics, technology, science, lifestyle, and stories of interest to readers. She enjoys researching new developments and presenting useful information in a simple, accessible way. Through her writing, Mira aims to keep readers informed with timely, informative, and easy-to-understand content.

Leave a Comment