Speech Recognition
5 pieces on Speech Recognition.
News & Analysis
Saaras V4 Covers All 22 Indian Languages With a 3B Hybrid Decoder
Sarvam AI's Saaras V4 handles all 22 scheduled Indian languages plus global English via a 3B hybrid state-space decoder, five output modes, and keyterm prompting at ₹30/hour.
Muse Voice Transcribe: 3.1% WER, One Model for ASR, Diarization, Endpointing
Meta Superintelligence Labs ships a single autoregressive model for streaming ASR, speaker diarization, and endpointing at $3.00 per 1,000 minutes.
Gemini 3.5 Transcribe Removes Filler Words, Covers 85+ Languages
Google's Gemini 3.5 Transcribe strips 'um' and 'uh', supports 85+ languages, and is in public preview via the Gemini API.
Gemini 3.5 Transcribe Hits 2.6% WER, Cuts Latency 70% vs Chirp 3
Google's new speech-to-text model achieves 2.6% WER non-streaming and 4.0% streaming, with 70% faster time-to-final-transcription than Chirp 3.
ASR Benchmarks Are Gameable: 6 of 11 Top Models Reproduce Audio Errors
Hume AI tested 11 open-source ASR models and found six reproduce VoxPopuli's transcript errors even when audio contradicts them — exposing WER as a gameable metric.