Microsoft has released MAI-Transcribe-2-Streaming for live speech recognition and MAI-Voice-2.1, including a faster Flash variant, for speech generation. The transcription model covers 60 languages, while the voice models cover 23 languages and 26 locales. Together they offer the input and output ends of a voice agent, but Microsoft’s published timings do not establish the delay of a complete conversational turn.

That distinction is useful for anyone choosing a voice stack. An agent must detect speech, transcribe enough of it to act, run its own reasoning or tools, synthesize a response, deliver audio over a network, and handle an interruption. A fast component can help, but it cannot erase time in the other steps.

A partial transcript can arrive before the user finishes

Microsoft says MAI-Transcribe-2-Streaming emits its first partial transcript just over 100 milliseconds after receiving audio. Partial text can be revised as later speech arrives, and the service eventually commits a final transcript. It continuously detects language and supports 60 languages. Microsoft’s example is an agent starting work before the speaker has finished a sentence.

The timing needs careful reading. An early hypothesis is useful for speculative work such as retrieving likely records. It may be unsafe as the sole basis for an irreversible action: a later word can change “book” into “don’t book,” and partials can be revised. A reliable integration would distinguish provisional text from a final or sufficiently stable command, and measure how often acting early creates wasted work or wrong actions.

Microsoft says the model ranks first for final and partial accuracy on Artificial Analysis’s streaming speech benchmark. The benchmark is independent, but it is a speech-recognition test rather than an end-to-end voice-agent test. Its published methodology combines roughly eight hours of audio from AA-AgentTalk, VoxPopuli and Earnings22, weighted 50%, 25% and 25%. The site’s time-to-final metric starts at detected speech end; that is a different clock from the time between a person’s first word and the agent’s first audible reply.

Flash changes the output budget, with a quality choice

Microsoft’s model page lists roughly 45 milliseconds of model-inference latency for MAI-Voice-2.1-Flash and roughly 550 milliseconds for the standard 2.1 model. The page recommends Flash for latency-sensitive assistants and calls, and standard 2.1 for fidelity-sensitive narration. Both are described as supporting the same 23 languages, cross-language speaker identity and voice matching from a short reference clip.

Those are service-side model figures, not measurements of a user’s full reply. Network delivery, buffering, text generation, speech chunking and playback add time. Microsoft separately says Flash can produce 45 seconds of audio with 150 milliseconds of end-to-end latency, but the announcement does not define that measure well enough to treat it as a 45-second clip fully delivered in 150 milliseconds. For a procurement test, measure time to first audible chunk and time to completed audio separately.

The announced prices are $22 per million input characters for MAI-Voice-2.1 and $15 per million for Flash. Streaming transcription is $0.54 per hour of input audio at an introductory rate through the end of 2026. At those rates, one hour of incoming audio plus 100,000 synthesized characters would be $2.04 with Flash or $2.74 with standard 2.1, before the language model, tools, transport and any other service charges. That example uses the published rates as a simple multiplication; actual billing and future prices should be checked at deployment.

The integration test is a conversation, not a waveform

A practical pilot should replay representative calls with accents, background noise, interruptions and language changes. Log the time of speech onset, first partial, stable transcript, decision, first audio byte, first audible output and complete reply. Record recognition corrections and actions taken from provisional text. This separates latency saved by streaming from latency spent on reasoning and action, while catching failures hidden by a polished demonstration.

The models are available through Microsoft Foundry and MAI Playground, with other listed access routes. Their release gives builders a new option for both sides of a voice interaction. The next evidence that matters is measured task success and turn latency in the actual agent workflow, especially when the first transcript is wrong or the speaker changes course mid-sentence.

Last Update: October 2, 2026