Speech, Voice and Realtime Audio Models — 2026-09-02
Meta Superintelligence Labs has released Muse Voice Transcribe, a unified real-time model for streaming ASR, diarization, and endpointing with a 3.1% WER. Meanwhile, South Korea selected SK Telecom, Kakao, and KT to build free nationwide AI services, marking a major shift in regional voice-agent deployment. Recent benchmarks highlight Cartesia’s 40ms latency lead in TTS and Deepgram’s competitive pricing in STT.
Speech, Voice and Realtime Audio Models — 2026-09-02
Top developments
Meta Releases Muse Voice Transcribe for Real-Time Streaming
Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, 2026, a single real-time model designed to handle streaming Automatic Speech Recognition (ASR), speaker diarization, and endpointing simultaneously. The model achieves a Word Error Rate (WER) of 3.1%, positioning it as a high-accuracy solution for complex multi-speaker environments. This release aims to simplify the architecture for voice agents by consolidating multiple pipeline components into one efficient model, potentially reducing latency and integration complexity for developers using Meta's Model API.

South Korea Selects SKT, Kakao, and KT for National AI Services
On August 28, 2026, the South Korean government selected SK Telecom, Kakao, and KT as the primary partners for the "All People's AI" project. The initiative aims to provide free, unlimited access to AI chatbots and agents for public services by the end of 2026. SK Telecom will utilize its proprietary "A.X K2" foundation model, focusing on an "executing AI" approach that goes beyond simple chat to perform tasks like public service applications. This move is designed to reduce reliance on foreign AI models and improve digital accessibility for elderly and vulnerable populations through phone and text-based interfaces.

Voice Cloning Fraud and Legal Battles Intensify
Recent reports highlight a surge in AI-driven fraud, with voice cloning tools being used to impersonate creators and executives. A TechTimes analysis published on September 1, 2026, noted that voice cloning technology crossed the threshold of being indistinguishable from human speech in 2025, outpacing current deepfake detection methods. Concurrently, Hollywood voice actors are escalating legal battles against AI studios over unauthorized voice replication, citing job displacement concerns. These developments underscore the growing tension between rapid voice synthesis capabilities and regulatory or ethical safeguards.

Benchmarking Latency in Realtime Inference APIs
A new benchmark study published on August 30, 2026, evaluated the lowest-latency inference APIs for voice agents, focusing on Time to First Token (TTFT) and full-pipeline latency budgets. While specific vendor rankings were not detailed in the summary, the study emphasizes that TTFT is now a critical metric for real-time conversational intelligence. This aligns with earlier data showing Cartesia Sonic 4 leading in pure latency at approximately 40ms Time to First Audio (TTFA), setting a high bar for realtime voice agent responsiveness.

Local view
Japan: Japanese media reported on August 27 that NTT Sonority launched "SonoVo AI," a new voice AI solution designed to turn workplace conversations into corporate assets. Additionally, TechnoSpeech announced pre-orders for a new voice library, "Oyasui Motomi," for its VoiSona Talk software, scheduled for release in December 2026.
South Korea: Local outlets like Bloter and ZDNet Korea detailed the "All People's AI" selection, emphasizing that the government is providing NVIDIA GPUs and budget support to ensure domestic models compete effectively against foreign alternatives. The focus is on "action-oriented" AI agents that can handle complex public service procedures rather than just answering questions.
Context & numbers
- Latency Leaders: Cartesia Sonic 4 currently leads in pure latency with ~40ms TTFA, while ElevenLabs v3 leads in expressiveness with coverage of 70+ languages and 5,000+ voices.
- Pricing: Deepgram remains the most affordable cloud STT option at roughly $0.26/hour for batch and $0.46/hour for streaming. ElevenLabs TTS pricing ranges from $0.05 to $0.10 per 1K characters.
- Accuracy Plateau: Top STT providers (Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe) have plateaued within 1-2 percentage points of each other on clean English audio WERs.
On the radar
- SKT Beta Launch: SK Telecom targets an October 2026 beta release for its "All People's AI" service, which will be accessible via phone and text without app installation.
- VoiSona Release: The new "Oyasui Motomi" voice library for VoiSona Talk is scheduled for official release on December 16, 2026, with pre-orders opening August 27, 2026.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.