Microsoft MAI voice models now span more of a voice-agent loop. On 1 October, Microsoft launched MAI-Transcribe-2-Streaming, its first streaming speech-to-text model, alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash for text-to-speech. The company says the transcription model supports 60 languages, while the voice models cover 23 languages and 26 locales. Microsoft’s announcement frames the three releases as components for building conversational agents.

Microsoft MAI voice models split the work across an agent
MAI-Transcribe-2-Streaming sends incoming audio over a WebSocket and returns provisional text as the speaker talks. It keeps revising those partial transcripts as more audio arrives, then marks the text final when the person stops. Automatic language detection runs continuously, so developers don’t have to select the spoken language before every session, according to Microsoft.
Microsoft says the first transcript hypotheses can appear just over 100 milliseconds after audio begins arriving. The number describes early partial output, not the full time it takes an agent to understand a request and answer it. Network conditions, the reasoning model and the speech-generation step all affect the wait a user actually experiences.
The company also says the model ranks first for accuracy on both final and partial transcripts in Artificial Analysis’ streaming evaluation. That is a vendor-reported benchmark result, not a guarantee for every accent, microphone, or noisy call. The practical use is narrower and clearer: an application can display captions while someone is speaking, or start reasoning and calling a tool before the speaker finishes a request.
The two MAI-Voice models handle the reply. MAI-Voice-2.1 is aimed at more expressive, higher-fidelity speech, and Microsoft says one voice can keep its identity while speaking across 23 languages and 26 locales with a native accent. Flash is the faster, lower-cost variant; it supports the same languages and can generate up to 45 seconds of audio with 150 milliseconds of end-to-end latency, according to Microsoft.
| Model | Main job | Published price |
|---|---|---|
| MAI-Transcribe-2-Streaming | Streaming speech-to-text; 60 languages | $0.54 per audio hour through 2026 |
| MAI-Voice-2.1 | Higher-fidelity text-to-speech; 23 languages, 26 locales | $22 per million characters |
| MAI-Voice-2.1-Flash | Faster text-to-speech; same language coverage | $15 per million characters |
The Microsoft MAI voice models are billed by different units: transcription by audio hour and speech generation by character count. Developers need to compare each rate against the amount of audio and generated speech their service expects to use.
Why streaming transcription costs more than batch
The new streaming model costs $0.54 per audio hour through the end of 2026. Microsoft’s non-streaming MAI-Transcribe-2 launched at $0.10 per hour through the same date, making the streaming rate 5.4 times higher. The earlier model waits for the audio to finish before returning a transcript; the new one keeps sending provisional results while the audio is still arriving. Microsoft’s September launch notes give the batch model’s price, while SiliconANGLE’s report details the new service’s pricing and timing.
That premium is easiest to justify when a transcript can trigger something before a turn ends: live captions, an interruptible assistant or an agent that starts a tool call mid-request. For archived recordings, meeting notes or other jobs that can wait for the speaker to finish, the lower-cost batch model may be sufficient. Developers should compare recognition costs with the cost of separate reasoning and speech-generation steps; the three API rates are billed in different units and aren’t directly interchangeable.
Microsoft is keeping those steps modular. Transcription turns audio into text, an application or language model decides what to do, and the MAI-Voice model produces the spoken answer. That gives developers room to change one component without replacing the others, but it also leaves them responsible for connecting the models and handling the conversation logic. This differs from systems built around live two-way audio, such as OpenAI’s GPT-Live-1 API and Nvidia PersonaPlex.
Where developers can access the MAI models
Microsoft lists the three models through Microsoft Foundry, MAI Playground, Vercel and Azure Voice Live, with LiveKit marked as coming soon. MAI-Voice-2.1 and its Flash variant are also available through OpenRouter. The company built Chatter, a demo in the MAI Playground, to show the transcription and voice models working together in a live agent.
Both speech-generation models also support voice cloning from a few seconds of reference audio. Microsoft says built-in consent guardrails are intended to prevent misuse; developers still need to secure permission to reproduce a person’s voice and check the terms that apply to their deployment. A cloning feature does not settle whether a particular use is authorised.
For UAE developers, the announcement publishes prices in US dollars and does not specify an AED rate or UAE-specific service availability. Teams should confirm regional access and the supported-language list before committing to a production design, particularly if the voice agent must handle a specific Arabic dialect.
Should an agent act on every partial transcript?
Not for high-impact actions. Microsoft says partial text is revised as more audio arrives, so developers can display it provisionally or use it to prefetch low-risk data; purchases, account changes and other consequential tool calls should wait for a stable transcript or user confirmation.
What should a pilot measure beyond transcription accuracy?
Test the whole conversation on the devices and networks customers will use. Include accented and noisy speech, interruptions, partial-versus-final transcript changes, tool-call success and the time from the end of a request to the finished spoken reply; component latency alone will not capture that full path.
How should a UAE team test whether the system handles its target Arabic dialect?
Check the current language and locale lists, then evaluate consented recordings and code-switching patterns customers actually use in those dialects. Compare partial and final transcripts, test the generated voice separately, and measure response time from the team’s intended UAE deployment region; broad language counts don’t establish dialect-level quality.
How should a team estimate the full cost of a voice-agent session?
Model the incoming audio hours, generated speech characters, and any separate reasoning-model or tool calls as different line items. Use representative short and long conversations from the intended workload rather than multiplying one model’s hourly price by the entire session; retries and longer spoken replies can change the total.


















