Microsoft is targeting ultra-realistic voice agents with its first streaming transcription model

Microsoft Corp. today expanded its MAI artificial intelligence model family with its first streaming transcription model, debuting along with two others focused on text-to-speech.

They are designed for developers who want to create voice agents that can listen to people’s voices and respond instantly, similar to how people talk to each other.

The most consequential of the three is MAI Transcribe 2 Streaming, which the company says accepts human speech over a WebSocket and converts it into a transcript that is continually updated as the person continues to speak. Once the person has finished speaking, the transcript is confirmed as final. This allows the model to power applications that display live captions or begin processing a user’s request before they have finished speaking it, Microsoft said.

MAI Transcribe 2 streaming is listed on Microsoft’s Vercel AI Gateway and costs 54 cents per audio hour, the company said. It supports more than 60 languages ​​and can automatically detect which language someone speaks. It delivers its first transcript hypotheses within 320 milliseconds on average, although Microsoft said it cannot guarantee this kind of performance in every scenario because the speed of response is also determined by the network connection and the AI ​​system that generates the response.

Transcription models are not new to Microsoft. Last month, the MAI-Transcribe-2 model was introduced with a price of 10 cents per audio hour, which is more than five times cheaper than the streaming version. This is primarily because the streaming model performs a lot more processing and continually delivers preliminary results while the audio is still arriving. The regular MAI Transcribe 2 model simply waits for the speaker to finish speaking before processing begins.

Aside from transcription, there are two models that focus on generating speech from text and are equally important for conversational AI agents. These include MAI-Voice-2.1, which is said to be the best choice for more expressive and lifelike output, and MAI-Voice-2.1-Flash, which tones down these characteristics for faster response and lower cost.

According to Vercel AI Gateway listings, MAI-Voice-2.1 costs $22 per million characters, while the Flash version costs $15 per million characters. According to Microsoft, both text-to-speech models support 23 languages.

The new disclosures show Microsoft is accelerating its transition away from model providers like OpenAI Group PBC and Anthropic PBC, despite being a major investor in both companies. In July, it was reported that Microsoft AI chief Mustafa Suleyman was becoming increasingly concerned about the costs associated with OpenAI and Anthropic’s powerful frontier models.

In response, he directed Microsoft’s AI researchers to double down on the MAI family of models, with the goal of eventually using these models to support his copilot agents in platforms like Excel and Outlook. “We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate those costs,” Suleyman said in an interview with Bloomberg.

With the introduction of the new models, Microsoft has everything it needs to develop powerful voice AI agents that can communicate with their human users in a natural, human-like manner. To function, voice agents must be able to do three things: They must be able to recognize and understand what someone is saying to them, then decide what to do based on what is said, and finally produce an audible response.

MAI-Transcribe-2 streaming takes care of the first problem and the MAI-Voice models solve the third. What’s in the middle is Microsoft’s powerful reasoning model Mai-Thinking-1, a standard model for large languages ​​that reads through transcripts of what is said to decide what the agent should do. By using three separate models, developers have much more control over the quality, latency and cost of their voice agents.

Image: Microsoft AI

Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.

  • Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
  • Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
SiliconANGLE Media is a recognized leader in digital media innovation, combining breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios – with flagship locations in Silicon Valley and the New York Stock Exchange – SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.

Avatar photo
Written by

Mira Edora

Mira Edora is a writer and contributor at CKSOR, creating clear and engaging articles on current topics, technology, science, lifestyle, and stories of interest to readers. She enjoys researching new developments and presenting useful information in a simple, accessible way. Through her writing, Mira aims to keep readers informed with timely, informative, and easy-to-understand content.

Leave a Comment