Speech, Voice and Realtime Audio Models — 2026-09-08
This week, the voice AI landscape saw significant movement in both open-source benchmarks and national-level strategic deployments. Microsoft’s new MAI-Transcribe-2 model is being touted for high accuracy at low cost, while South Korea officially launched its "Everyone's AI" initiative, selecting SK Telecom, Kakao, and KT to build national-scale voice-enabled agents. Meanwhile, industry experts are increasingly warning that standard speech-to-text benchmarks no longer reflect production realities, urging a shift toward phone-audio testing.
Speech, Voice and Realtime Audio Models — 2026-09-08
South Korea Selects SKT, Kakao, and KT for "Everyone's AI" National Project
On September 4, 2026, the South Korean government confirmed that SK Telecom, Kakao, and KT have been selected as consortium partners for the "Everyone's AI" project. The initiative aims to provide free, nationwide AI services by the end of 2026, with a budget of 250 billion KRW (approx. $185 million) allocated for 2027. The selected companies will leverage their existing infrastructure—SKT’s telephony network, Kakao’s messenger platform, and KT’s content ecosystem—to integrate voice agents into daily citizen services, targeting 15 million users by year-end. This move highlights the growing integration of realtime voice interfaces into public sector digital infrastructure, with a focus on accessibility via familiar channels like phone calls and chat apps.

Microsoft Announces MAI-Transcribe-2 for High-Speed, Low-Cost ASR
Microsoft released details on its new speech recognition model, MAI-Transcribe-2, which it claims is the "fastest, most accurate, and cheapest" in the world. Japanese tech outlet Mado no Mori reported on September 7 that the model aims to lower barriers for real-time transcription applications by significantly reducing inference costs while maintaining high accuracy. While specific latency figures were not detailed in the initial press, the positioning directly challenges existing leaders like Deepgram and AssemblyAI in the enterprise ASR market. This release underscores the intensifying competition in the STT space, where price-performance ratios are becoming key differentiators for large-scale deployments.
Industry Pushback on Standard STT Benchmarks
A report published on September 6 by Famulor argues that traditional word error rate (WER) benchmarks are insufficient for evaluating voice AI performance in production environments. The analysis highlights that while models like Meta Muse and Microsoft MAI-Transcribe-2 achieve low WER on clean audio, they often struggle with the noise and channel variability of real-world phone audio. The article urges enterprise teams to adopt "production-like" testing methodologies rather than relying solely on static datasets. This critique comes as vendors continue to race to claim top spots on leaderboards that may not reflect actual customer experience.
NTT Sonority Launches "SonoVo AI" for Enterprise Asset Management
NTT Sonority announced the launch of "SonoVo AI," a voice AI solution designed to convert field voices into corporate assets. Although the service began in August, coverage continued this week, emphasizing its role in capturing and structuring unstructured voice data from frontline workers. This reflects a broader trend in Asia where voice AI is moving beyond simple transcription to knowledge management and workflow automation within large enterprises.
Local view
In South Korea, ZDNet Korea reported that the three winning consortia—SKT, Kakao, and KT—are prioritizing different entry points: SKT is focusing on telephony integration, Kakao on messaging-based agents, and KT on content-linked services. The competition is framed not just as a technical race but as a battle for user interface dominance in the Korean market. Meanwhile, in Japan, Mado no Mori covered Microsoft's MAI-Transcribe-2 with a focus on its potential to disrupt the local market for cost-sensitive transcription services, noting the model's aggressive pricing strategy.
Context & numbers
The "Everyone's AI" project in South Korea targets 15 million users by the end of 2026, backed by a planned 250 billion KRW budget for next year. In the broader benchmarking context, recent data indicates that hosted APIs like Deepgram Flux and ElevenLabs Scribe v2 maintain competitive word error rates (WER) between 9% and 15% on realistic conversational datasets, though costs remain a barrier for some enterprises compared to self-hosted open-weight models. The FBI previously reported $893 million in losses related to AI voice fraud in 2025, a figure that continues to influence enterprise security protocols for voice authentication.
On the radar
- Regulatory Scrutiny on Voice Cloning: A new report from Recording Law highlights that all 50 US states now have some form of deepfake or voice cloning legislation, complicating compliance for cross-border voice AI startups.
- Hollywood Labor Disputes: The Los Angeles Times reported ongoing clashes between voice actors and studios over AI voice clones, with negotiations likely to impact the licensing terms for commercial TTS models in the entertainment industry.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.
