CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-02

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-02

Speech, Voice and Realtime Audio Models|September 2, 2026(4h ago)3 min read9.1AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

Meta Superintelligence Labs has released Muse Voice Transcribe, a unified real-time model for streaming ASR, diarization, and endpointing with a 3.1% WER. Meanwhile, South Korea selected SK Telecom, Kakao, and KT to build free nationwide AI services, marking a major shift in regional voice-agent deployment. Recent benchmarks highlight Cartesia’s 40ms latency lead in TTS and Deepgram’s competitive pricing in STT.

Speech, Voice and Realtime Audio Models — 2026-09-02


Top developments


Meta Releases Muse Voice Transcribe for Real-Time Streaming

Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, 2026, a single real-time model designed to handle streaming Automatic Speech Recognition (ASR), speaker diarization, and endpointing simultaneously. The model achieves a Word Error Rate (WER) of 3.1%, positioning it as a high-accuracy solution for complex multi-speaker environments. This release aims to simplify the architecture for voice agents by consolidating multiple pipeline components into one efficient model, potentially reducing latency and integration complexity for developers using Meta's Model API.

Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time
Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time

engadget.com

engadget.com


South Korea Selects SKT, Kakao, and KT for National AI Services

On August 28, 2026, the South Korean government selected SK Telecom, Kakao, and KT as the primary partners for the "All People's AI" project. The initiative aims to provide free, unlimited access to AI chatbots and agents for public services by the end of 2026. SK Telecom will utilize its proprietary "A.X K2" foundation model, focusing on an "executing AI" approach that goes beyond simple chat to perform tasks like public service applications. This move is designed to reduce reliance on foreign AI models and improve digital accessibility for elderly and vulnerable populations through phone and text-based interfaces.

SKT, Kakao and KT chosen for South Korea’s public AI service project
SKT, Kakao and KT chosen for South Korea’s public AI service project


Voice Cloning Fraud and Legal Battles Intensify

Recent reports highlight a surge in AI-driven fraud, with voice cloning tools being used to impersonate creators and executives. A TechTimes analysis published on September 1, 2026, noted that voice cloning technology crossed the threshold of being indistinguishable from human speech in 2025, outpacing current deepfake detection methods. Concurrently, Hollywood voice actors are escalating legal battles against AI studios over unauthorized voice replication, citing job displacement concerns. These developments underscore the growing tension between rapid voice synthesis capabilities and regulatory or ethical safeguards.

Morning Show Season 4's Deepfake Science Was Right: Voice Cloning Now Sounds Real
Morning Show Season 4's Deepfake Science Was Right: Voice Cloning Now Sounds Real


Benchmarking Latency in Realtime Inference APIs

A new benchmark study published on August 30, 2026, evaluated the lowest-latency inference APIs for voice agents, focusing on Time to First Token (TTFT) and full-pipeline latency budgets. While specific vendor rankings were not detailed in the summary, the study emphasizes that TTFT is now a critical metric for real-time conversational intelligence. This aligns with earlier data showing Cartesia Sonic 4 leading in pure latency at approximately 40ms Time to First Audio (TTFA), setting a high bar for realtime voice agent responsiveness.

Lowest-Latency Inference APIs for Voice and Realtime Agents
Lowest-Latency Inference APIs for Voice and Realtime Agents

marktechpost.com

marktechpost.com

marktechpost.com

marktechpost.com


Local view

Japan: Japanese media reported on August 27 that NTT Sonority launched "SonoVo AI," a new voice AI solution designed to turn workplace conversations into corporate assets. Additionally, TechnoSpeech announced pre-orders for a new voice library, "Oyasui Motomi," for its VoiSona Talk software, scheduled for release in December 2026.

South Korea: Local outlets like Bloter and ZDNet Korea detailed the "All People's AI" selection, emphasizing that the government is providing NVIDIA GPUs and budget support to ensure domestic models compete effectively against foreign alternatives. The focus is on "action-oriented" AI agents that can handle complex public service procedures rather than just answering questions.


Context & numbers

  • Latency Leaders: Cartesia Sonic 4 currently leads in pure latency with ~40ms TTFA, while ElevenLabs v3 leads in expressiveness with coverage of 70+ languages and 5,000+ voices.
  • Pricing: Deepgram remains the most affordable cloud STT option at roughly $0.26/hour for batch and $0.46/hour for streaming. ElevenLabs TTS pricing ranges from $0.05 to $0.10 per 1K characters.
  • Accuracy Plateau: Top STT providers (Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe) have plateaued within 1-2 percentage points of each other on clean English audio WERs.

On the radar

  • SKT Beta Launch: SK Telecom targets an October 2026 beta release for its "All People's AI" service, which will be accessible via phone and text without app installation.
  • VoiSona Release: The new "Oyasui Motomi" voice library for VoiSona Talk is scheduled for official release on December 16, 2026, with pre-orders opening August 27, 2026.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Meta's Muse Voice Transcribe reduce latency?
  • QWhat specific public services will South Korea's AI offer?
  • QWhat new detection methods combat voice cloning fraud?
  • QWhich API won the recent real-time latency benchmark?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.