Google's New Transcription Model Cuts Word Error Rate to 2.6%
Google's Gemini 3.5 Transcribe cuts word error rate to 2.6%, speeds up transcription 70% over Chirp 3, and supports 85+ languages with real-time and pre-recorded modes.
13 verified stories covering voice AI, product updates and industry developments.
Google's Gemini 3.5 Transcribe cuts word error rate to 2.6%, speeds up transcription 70% over Chirp 3, and supports 85+ languages with real-time and pre-recorded modes.
Qwen-Audio-3.0-Realtime adds tool calling to real-time voice interaction. The piece reviews the verified facts and why the signal matters beyond one announcement.
ChatGPT Voice moves to GPT-Live is reframed for global technology readers, with the key numbers, product claims and open questions kept intact.
Bland says its voice agents handled 175 million calls last year and now processes 3.5 million calls a week.
Vapi raised a $50 million Series B after Amazon Ring chose its voice AI stack over dozens of alternatives for inbound support calls.
Wispr Flow raises new funding for voice AI. The story explains the announcement, the strategic context and the practical risks for the people or companies affected.
ElevenLabs extended its Series D above $550 million at an $11 billion valuation, but the sharper signal is enterprise voice AI revenue growing 43% in one quarter.
OpenAI adds reasoning speech, live translation, and streaming transcription to the Realtime API, with Zillow and Deutsche Telekom among early users.
xAI's grok-voice-think-fast-1.0 scored 67.3% on the τ-voice Bench, leaving Gemini 3.1 Flash Live at 43.8%, and is already handling real customer service calls for Starlink.
xAI launches Grok audio API with low error rates. The story explains the announcement, the strategic context and the practical risks for the people or companies affected.
Alibaba's Qwen team has released Qwen3.5-Omni, its first true all-modal model that processes text, images, audio, and video within a single architecture, featuring real-time voice cloning and audio-video vibe coding.
Mistral released Voxtral TTS, its first voice model, on March 26. The 4B-parameter open-weight model is rated by human evaluators as more natural than ElevenLabs Flash v2.5 and on par with ElevenLabs v3. API pricing is $0.016 per thousand characters, and the model can run on your own server.
Google has released Gemini 3.1 Flash Live, a model that compresses the traditional three-stage voice pipeline into a native audio-to-audio system, reducing latency and enabling features like barge-in and tool calling.