Speech, Voice and Realtime Audio Models — 2026-09-13
OpenAI officially released the GPT-Live-1 API, enabling full-duplex voice agents that listen while speaking, marking a shift from turn-based to fluid conversational AI. In parallel, Meta leaked details on an 80ms latency custom voice model for Muse, while South Korea’s major telcos (KT, SKT) detailed their "Everyone's AI" strategies involving massive-scale voice interface deployments.
Speech, Voice and Realtime Audio Models — 2026-09-13
Top developments
OpenAI releases GPT-Live-1 API for full-duplex voice agents
On September 10, 2026, OpenAI launched the GPT-Live-1 model as an API, allowing developers to build voice agents that can listen and speak simultaneously. Unlike previous models that required a stop-and-start turn-taking protocol, GPT-Live-1 supports full-duplex communication, enabling more natural interruptions and overlapping speech. This release targets developers building real-time conversational workflows where latency and natural flow are critical, moving beyond the "push-to-talk" paradigm.

Meta leaks 80ms latency custom voice model for Muse
Tech Insider reported on September 13 that Meta is preparing to launch on-demand custom voices for its Muse platform, powered by a new real-time speech model with an estimated 80ms time-to-first-audio (TTFA). The model is reportedly priced at $3 per 1,000 minutes, positioning it as a highly competitive option for low-latency applications compared to current market leaders like Cartesia Sonic 4 (~40ms) and ElevenLabs. This move signals Meta's intent to capture a larger share of the voice-agent infrastructure market by combining ultra-low latency with customizable branding.

South Korean telcos KT and SKT scale "Everyone's AI" voice services
On September 10–12, South Korean telecom giants KT and SK Telecom outlined their participation in the government-backed "Everyone's AI" project, aiming to integrate voice-enabled AI assistants into daily life services for over 60 million touchpoints. KT introduced "Eum Inside," a strategy to embed its AI capabilities into partner services, leveraging NPU efficiency to boost GPU performance by 21%. SK Telecom formed a "Dream Team" with finance and healthcare partners to develop end-to-end executable agents, moving beyond simple chatbots to action-oriented voice interfaces.

NTT TechnoCross launches new AI voice bot features for natural dialogue
On September 10, NTT TechnoCross announced new features for its "CTBASE/SmartCommunicator" AI voice bot, designed to optimize response processes and achieve more natural conversations in Japanese enterprise settings. The update focuses on reducing latency and improving context retention in customer service scenarios, aligning with global trends toward more human-like agent interactions. This release underscores the push by Japanese legacy vendors to modernize their contact center stacks with advanced NLP and speech synthesis integration.

Local view
Japan: Media outlets such as AI Watch and eguweb highlighted the industry shift from sequential processing (listen-transcribe-think-speak) to full-duplex systems, citing OpenAI's GPT-Live-1 as a catalyst. The coverage emphasizes how Japanese enterprises are evaluating these new APIs to replace traditional IVR systems with more fluid, natural language interactions.
South Korea: Financial News and Edaily focused heavily on the national "Everyone's AI" initiative, reporting that KT, SKT, and Kakao are racing to launch free beta versions of voice-enabled super-apps by year-end. The local narrative centers on using these voice agents to solve practical daily problems (finance, shopping) rather than just providing information, with a target of 15 million users initially.
Context & numbers
- Latency Benchmarks: Current market leaders in pure latency include Cartesia Sonic 4 at ~40ms TTFA, while the rumored Meta model aims for 80ms. ElevenLabs v3 leads in expressiveness with 70+ languages and 5,000+ voices.
- Pricing: Standard cloud TTS pricing remains competitive, with Google Cloud and Amazon Polly at $4 per 1M characters. Meta's leaked pricing of $3 per 1,000 minutes for real-time voice is significantly lower than many premium agent platforms.
- Market Growth: The AI voice cloning market is projected to expand to $9.56 billion by 2030, driven by enterprise adoption and fraud prevention needs.
- Word Error Rates (WER): Top STT providers (Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe) have plateaued within 1-2 percentage points of each other on clean English audio, shifting competition to multilingual support and latency.
On the radar
- Voice Agent Disclosure Laws: JustCall published a comprehensive guide on September 8 detailing federal, state, and international rules requiring disclosure when interacting with AI voice agents, which will impact compliance for all major voice platforms.
- Meta Muse Voice Transcribe: Following the September 1 announcement of "Muse Voice Transcribe," developers are awaiting further details on API access and integration with Meta AI for Mac, as reported by AI Watch.
- Korean Beta Launches: SKT, Kakao, and KT are expected to roll out beta versions of their "Everyone's AI" voice services later this year, with specific dates still being finalized by the Ministry of Science and ICT.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.