MAI-Transcribe-2-Streaming Hits 2.5% WER in 0.13 Seconds

October 3, 2026 • news
Speech RecognitionVoice AIBenchmarks

Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, according to MarkTechPost. It is Microsoft’s first streaming speech-to-text model and the real-time counterpart to the batch MAI-Transcribe-2 released in September.

The model targets voice agents, live captions and dictation. Artificial Analysis ranks it first for both final and first-partial transcript accuracy among 38 systems on the AA-WER Streaming index, MarkTechPost reports. Its final transcript recorded a 2.5% word error rate (WER) at 0.13 seconds after the detected end of speech. Its first partial also recorded 2.5% WER, at 0.12 seconds.

That equality matters operationally: on this benchmark, the early transcript was as accurate as the committed result. A voice application may therefore be able to start retrieval or tool selection from the partial instead of waiting for a later correction, although teams must test that behavior on their own speakers, microphones and acoustic conditions.

Streaming behavior and integration

MAI-Transcribe-2-Streaming accepts audio continuously and returns text while the speaker is talking. MarkTechPost reports that the model emits its first partial hypotheses just over 100 milliseconds after receiving audio, revises them as context arrives and then commits a stable final transcript.

Microsoft says its internal testing showed words appearing twice as fast as with its closest competitor. That is a vendor-reported result and is separate from the third-party AA-WER latency measurements.

The model supports 60 languages with continuous automatic language detection, according to MarkTechPost. Microsoft documents two integration paths: a Realtime API for applications using an OpenAI Realtime-compatible WebSocket and the Azure Speech SDK, which manages connections, retries and audio streaming. Both return intermediate and final results.

MarkTechPost also lists availability through the MAI Playground, Vercel and Azure Voice Live, with LiveKit support coming later. Microsoft pairs the transcription model with MAI-Voice-2.1-Flash for voice-agent loops. Microsoft describes that text-to-speech model as supporting 23 languages and reports 150-millisecond end-to-end latency for generating 45 seconds of audio.

Benchmark and pricing comparison

The AA-WER Streaming index uses approximately eight hours of audio: 50% AA-AgentTalk, 25% VoxPopuli and 25% Earnings22. Artificial Analysis measures latency from the end of speech as detected by SileroVAD.

ModelFinal WERTime to finalStreaming price/hour
MAI-Transcribe-2-Streaming2.5%0.13s$0.54 introductory
Grok Voice Transcribe 2.02.7%0.49s$0.20
Muse Voice Transcribe3.1%0.16s$0.18
Gemini 3.5 Transcribe Live4.0%Not reportedAbout $0.54 estimated
Cartesia Ink-2, external endpoints4.0%0.07sNot reported

Artificial Analysis data reported by MarkTechPost shows that Cartesia Ink-2 returned a final transcript faster, at 0.07 seconds, but with a higher 4.0% WER. Grok Voice Transcribe 2.0 was cheaper at $0.20 per hour but took 0.49 seconds to return its final transcript. Muse Voice Transcribe cost $0.18 per hour and produced a 3.1% WER at 0.16 seconds.

MarkTechPost reports that MAI-Transcribe-2-Streaming costs an introductory $0.54 per audio hour through the end of 2026, which Artificial Analysis normalizes to $9 per 1,000 minutes. Batch MAI-Transcribe-2 costs $0.10 per hour.

The streaming model is in public preview without a service-level agreement and has no open weights, according to MarkTechPost. Speaker diarization for its streaming output was not stated in the reported release details.

AI Mastery analysis

The standout result is not simply the 2.5% final WER, but that the first partial achieved the same score. Streaming systems often revise early hypotheses as more context arrives; reducing that correction gap can let an agent begin work sooner without building every action around transcript rollbacks.

The latency figures require careful interpretation. Artificial Analysis starts its clock when SileroVAD detects the end of speech, so its 0.12- and 0.13-second results measure turn-closing speed rather than the delay between a speaker’s first phoneme and the first displayed word. Microsoft’s claim of partials appearing just over 100 milliseconds after audio arrives describes that different stage of the pipeline.

Deployment constraints may outweigh a narrow benchmark lead. This is a closed, public-preview service tied to Microsoft-hosted interfaces, not a self-hosted model. Teams requiring cloud portability or on-premises inference face the broader problem that AI portability is often the real bottleneck. Production evaluations should therefore test transcript stability, endpointing, WebSocket recovery and regional behavior—not just leaderboard WER.

Sources

Frequently asked questions

How accurate is MAI-Transcribe-2-Streaming?

Artificial Analysis reports a 2.5% word error rate for both the first partial and final transcript. That ranks MAI-Transcribe-2-Streaming first for accuracy among the 38 systems evaluated on its AA-WER Streaming index.

How fast is MAI-Transcribe-2-Streaming?

Artificial Analysis measured the first partial transcript at 0.12 seconds and the final transcript at 0.13 seconds after SileroVAD detected the end of speech. Microsoft separately says the model begins emitting partial hypotheses just over 100 milliseconds after receiving audio.

How much does MAI-Transcribe-2-Streaming cost?

MarkTechPost reports an introductory price of $0.54 per audio hour through the end of 2026, equivalent to $9 per 1,000 minutes. The batch MAI-Transcribe-2 model costs $0.10 per audio hour.

How many languages does MAI-Transcribe-2-Streaming support?

MarkTechPost reports support for 60 languages with continuous automatic language detection. The model can therefore identify language changes while processing an ongoing audio stream.

How do developers access MAI-Transcribe-2-Streaming?

Microsoft documents an OpenAI Realtime-compatible WebSocket API and the Azure Speech SDK as its two integration paths. MarkTechPost also lists the MAI Playground, Vercel and Azure Voice Live, with LiveKit support coming later.

Free interactive tools for the decisions this piece raises.

Related Reading