MAI-Transcribe-2-Streaming Hits 2.5% WER in 0.13 Seconds
In this article
Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, according to MarkTechPost. It is Microsoft’s first streaming speech-to-text model and the real-time counterpart to the batch MAI-Transcribe-2 released in September.
The model targets voice agents, live captions and dictation. Artificial Analysis ranks it first for both final and first-partial transcript accuracy among 38 systems on the AA-WER Streaming index, MarkTechPost reports. Its final transcript recorded a 2.5% word error rate (WER) at 0.13 seconds after the detected end of speech. Its first partial also recorded 2.5% WER, at 0.12 seconds.
That equality matters operationally: on this benchmark, the early transcript was as accurate as the committed result. A voice application may therefore be able to start retrieval or tool selection from the partial instead of waiting for a later correction, although teams must test that behavior on their own speakers, microphones and acoustic conditions.
Streaming behavior and integration
MAI-Transcribe-2-Streaming accepts audio continuously and returns text while the speaker is talking. MarkTechPost reports that the model emits its first partial hypotheses just over 100 milliseconds after receiving audio, revises them as context arrives and then commits a stable final transcript.
Microsoft says its internal testing showed words appearing twice as fast as with its closest competitor. That is a vendor-reported result and is separate from the third-party AA-WER latency measurements.
The model supports 60 languages with continuous automatic language detection, according to MarkTechPost. Microsoft documents two integration paths: a Realtime API for applications using an OpenAI Realtime-compatible WebSocket and the Azure Speech SDK, which manages connections, retries and audio streaming. Both return intermediate and final results.
MarkTechPost also lists availability through the MAI Playground, Vercel and Azure Voice Live, with LiveKit support coming later. Microsoft pairs the transcription model with MAI-Voice-2.1-Flash for voice-agent loops. Microsoft describes that text-to-speech model as supporting 23 languages and reports 150-millisecond end-to-end latency for generating 45 seconds of audio.
Benchmark and pricing comparison
The AA-WER Streaming index uses approximately eight hours of audio: 50% AA-AgentTalk, 25% VoxPopuli and 25% Earnings22. Artificial Analysis measures latency from the end of speech as detected by SileroVAD.
| Model | Final WER | Time to final | Streaming price/hour |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | 2.5% | 0.13s | $0.54 introductory |
| Grok Voice Transcribe 2.0 | 2.7% | 0.49s | $0.20 |
| Muse Voice Transcribe | 3.1% | 0.16s | $0.18 |
| Gemini 3.5 Transcribe Live | 4.0% | Not reported | About $0.54 estimated |
| Cartesia Ink-2, external endpoints | 4.0% | 0.07s | Not reported |
Artificial Analysis data reported by MarkTechPost shows that Cartesia Ink-2 returned a final transcript faster, at 0.07 seconds, but with a higher 4.0% WER. Grok Voice Transcribe 2.0 was cheaper at $0.20 per hour but took 0.49 seconds to return its final transcript. Muse Voice Transcribe cost $0.18 per hour and produced a 3.1% WER at 0.16 seconds.
MarkTechPost reports that MAI-Transcribe-2-Streaming costs an introductory $0.54 per audio hour through the end of 2026, which Artificial Analysis normalizes to $9 per 1,000 minutes. Batch MAI-Transcribe-2 costs $0.10 per hour.
The streaming model is in public preview without a service-level agreement and has no open weights, according to MarkTechPost. Speaker diarization for its streaming output was not stated in the reported release details.
AI Mastery analysis
The standout result is not simply the 2.5% final WER, but that the first partial achieved the same score. Streaming systems often revise early hypotheses as more context arrives; reducing that correction gap can let an agent begin work sooner without building every action around transcript rollbacks.
The latency figures require careful interpretation. Artificial Analysis starts its clock when SileroVAD detects the end of speech, so its 0.12- and 0.13-second results measure turn-closing speed rather than the delay between a speaker’s first phoneme and the first displayed word. Microsoft’s claim of partials appearing just over 100 milliseconds after audio arrives describes that different stage of the pipeline.
Deployment constraints may outweigh a narrow benchmark lead. This is a closed, public-preview service tied to Microsoft-hosted interfaces, not a self-hosted model. Teams requiring cloud portability or on-premises inference face the broader problem that AI portability is often the real bottleneck. Production evaluations should therefore test transcript stability, endpointing, WebSocket recovery and regional behavior—not just leaderboard WER.
Sources
Frequently asked questions
How accurate is MAI-Transcribe-2-Streaming?
Artificial Analysis reports a 2.5% word error rate for both the first partial and final transcript. That ranks MAI-Transcribe-2-Streaming first for accuracy among the 38 systems evaluated on its AA-WER Streaming index.
How fast is MAI-Transcribe-2-Streaming?
Artificial Analysis measured the first partial transcript at 0.12 seconds and the final transcript at 0.13 seconds after SileroVAD detected the end of speech. Microsoft separately says the model begins emitting partial hypotheses just over 100 milliseconds after receiving audio.
How much does MAI-Transcribe-2-Streaming cost?
MarkTechPost reports an introductory price of $0.54 per audio hour through the end of 2026, equivalent to $9 per 1,000 minutes. The batch MAI-Transcribe-2 model costs $0.10 per audio hour.
How many languages does MAI-Transcribe-2-Streaming support?
MarkTechPost reports support for 60 languages with continuous automatic language detection. The model can therefore identify language changes while processing an ongoing audio stream.
How do developers access MAI-Transcribe-2-Streaming?
Microsoft documents an OpenAI Realtime-compatible WebSocket API and the Azure Speech SDK as its two integration paths. MarkTechPost also lists the MAI Playground, Vercel and Azure Voice Live, with LiveKit support coming later.
Related Reading
Grok Voice Transcribe 2.0 Cuts Short-Phrase WER From 20.6% to 6.8%
SpaceXAI's Grok Voice Transcribe 2.0 claims 2x accuracy over 1.0 at unchanged pricing: $0.10/hr batch, $0.20/hr streaming.
Gradium TTS: 81.0% Hard-Case Accuracy at 216 ms First Audio
Gradium AI's new default TTS model posts 81.0% on a 500-sentence hard-case eval and 216 ms P50 latency with a 30 ms interquartile spread.
ASR Benchmarks Are Gameable: 6 of 11 Top Models Reproduce Audio Errors
Hume AI tested 11 open-source ASR models and found six reproduce VoxPopuli's transcript errors even when audio contradicts them — exposing WER as a gameable metric.