Speech, Voice and Realtime Audio Models — 2026-09-11
OpenAI launched GPT-Live-1, a full-duplex voice model enabling simultaneous listening and speaking for developers. In Japan, NTT-TX introduced new optimization features for its AI voice bot, while Microsoft’s MAI-Transcribe-2 was highlighted in local tech media for its speed and cost-efficiency. Meanwhile, South Korea’s "AI for Everyone" project progressed with SKT, Kakao, and KT detailing their end-to-end agent strategies.
Speech, Voice and Realtime Audio Models — 2026-09-11
Top developments
OpenAI Releases GPT-Live-1 for Full-Duplex Interaction
On September 11, 2026, OpenAI released GPT-Live-1, a new voice model designed for full-duplex conversations. Unlike previous half-duplex models, GPT-Live-1 allows the AI to listen and speak simultaneously, handling interruptions naturally. The model supports backend delegation and offers 12 distinct voices, aiming to reduce latency in real-time agent applications.

NTT-TX Enhances CTBASE Voice Bot with New Optimization Features
NTT TechnoCross (NTT-TX) announced on September 10, 2026, new features for its CTBASE/SmartCommunicator AI voice bot. The update focuses on optimizing the response process to enable more natural dialogue flows. This move aligns with broader industry trends toward reducing latency and improving contextual awareness in Japanese enterprise voice solutions.

South Korea's "AI for Everyone" Consortia Detail Agent Strategies
On September 10, 2026, SK Telecom announced its strategy for the government-backed "AI for Everyone" project, forming a "dream team" with finance, education, and healthcare partners to build an "end-to-end" executable agent. This follows earlier announcements from the Science Ministry where SKT, Kakao, and KT outlined plans to launch public-facing AI agents by year-end, targeting 15 million users initially.

Microsoft MAI-Transcribe-2 Highlighted in Japanese Tech Media
Japanese outlet Mado no Mori reported on Microsoft's MAI-Transcribe-2 speech recognition model on September 7, 2026. The article highlights Microsoft's claim that the model is the "world's fastest, most accurate, and cheapest" speech recognition solution. While the model itself may have been announced slightly earlier, this coverage underscores its immediate impact on the Asian market's perception of cost-efficient ASR benchmarks.
Local view
In Japan, media attention has shifted toward practical enterprise integrations. AI Watch covered NTT-TX’s specific updates to their voice bot response processes, emphasizing "natural dialogue" as a key competitive differentiator in the domestic market. Additionally, Investing.com Japan reported on OpenAI’s GPT-Live-1 release, framing it within the context of developer tools and potential stock market impacts for tech sectors.
In South Korea, Financial News and ZDNet Korea focused heavily on the "AI for Everyone" consortiums. Stakeholders are scrutinizing how SKT, Kakao, and KT will differentiate their voice and agent capabilities, with SKT emphasizing specialized partnerships in vertical industries like healthcare and finance to create actionable, rather than just conversational, AI agents.
Context & numbers
Recent pricing analyses indicate continued pressure on TTS costs. As of early September 2026, Speechify’s self-serve TTS overages range from $10 down to $6 per 1M characters depending on plan volume. All-in voice-agent minutes are quoted between $0.075 and $0.068. ElevenLabs remains priced at the higher end for quality-focused applications, with rates between $0.05 and $0.10 per 1K characters.
On the radar
- Regulatory Compliance: A new guide from JustCall published on September 8, 2026, details 2026 disclosure laws for AI voice agents across federal, state, and international jurisdictions, noting that non-compliance risks are rising as more states mandate explicit AI caller identification.
- Korean Launch Timeline: The "AI for Everyone" services from SKT, Kakao, and KT are targeted for release by the end of 2026, with a 2.5 trillion KRW budget planned for next year to support nationwide adoption.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.