CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-12

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-12

Speech, Voice and Realtime Audio Models|September 12, 2026(1h ago)3 min read9.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

OpenAI has officially launched the GPT-Live-1 API, enabling developers to build full-duplex voice agents that can listen and speak simultaneously. In parallel, Japanese voice AI provider CoeFont released v4 of its synthesis model, claiming a 61x speed increase, while South Korean tech giants SKT, Kakao, and KT advanced their government-backed "AI for Everyone" voice assistant initiatives.

Speech, Voice and Realtime Audio Models — 2026-09-12


Top developments


OpenAI Releases GPT-Live-1 API for Real-Time Conversations

On September 10, 2026, OpenAI released the API for GPT-Live-1, its latest real-time full-duplex voice model. This update allows software developers to implement applications where AI can listen and speak simultaneously, significantly reducing latency and creating more fluid, natural conversational workflows compared to turn-based systems. The release marks a shift from sequential processing (listen-transcribe-think-speak) to simultaneous interaction, addressing key bottlenecks in realtime voice-agent models.

Screenshot of OpenAI GPT-Live-1 interface
Screenshot of OpenAI GPT-Live-1 interface

theregister.com

theregister.com


CoeFont v4 Boosts TTS Speed by 61x

Japanese voice AI company CoeFont released CoeFont v4 on September 11, 2026. The new version claims a maximum generation speed increase of 61 times compared to previous iterations, targeting high-volume content creation and real-time applications. This development highlights the competitive pressure in the Asian TTS market to reduce time-to-first-audio (TTFA) while maintaining naturalness, directly impacting latency metrics for developers using Japanese-language voice stacks.

CoeFont v4 Release Announcement
CoeFont v4 Release Announcement


South Korea’s "AI for Everyone" Voice Assistants Move to Beta

SK Telecom, Kakao, and KT have finalized their blueprints for the government-backed "AI for Everyone" project, with beta releases targeted for late 2026. SK Telecom announced on September 10 that it is forming an "End-to-End" execution agent team, integrating voice interfaces across finance, education, and healthcare. These initiatives aim to provide free, accessible AI voice assistants to 15 million users by year-end, leveraging existing channels like phone calls and KakaoTalk to lower barriers for elderly and digital-minority populations.

SKT AI Assistant Blueprint
SKT AI Assistant Blueprint


NTT TechnoCross Enhances Voice Bot Response Processes

NTT TechnoCross announced new features for its "CTBASE/SmartCommunicator" AI voice bot on September 10, 2026. The update optimizes response processes to enable more natural dialogue flows, focusing on reducing awkward silences and improving context retention during customer service interactions. This reflects a broader trend among enterprise vendors to refine full-duplex capabilities and interruption handling in production-grade voice agents.


Local view

In Japan, media outlets like AI Watch and Gigazine are highlighting the shift toward "full-duplex" interactions, noting that OpenAI’s GPT-Live-1 allows AI to respond while still listening, mimicking human conversation more closely than previous turn-based models. Eguweb summarized the global AI news on September 12, emphasizing that voice AI is moving away from sequential transcription-inference-synthesis pipelines to simultaneous processing, which drastically reduces perceived latency.

In South Korea, Seoul Finance and ZDNet Korea reported on the competitive dynamics between SKT, Kakao, and KT. The Ministry of Science and ICT is overseeing the "AI for Everyone" consortiums, with a focus on universal accessibility. Reports indicate that SKT will leverage its telecom infrastructure to allow access via voice calls without app installations, aiming to bridge the digital divide.


Context & numbers

Pricing and Latency Benchmarks: Recent comparisons of TTS providers show Cartesia Sonic 4 leading in pure latency with approximately 40ms Time-To-First-Audio (TTFA). ElevenLabs v3 remains a leader in expressiveness, supporting over 70 languages and 5,000+ voices, priced at $0.10 per 1,000 characters. Deepgram Aura-2 offers competitive pricing at $0.030 per 1,000 characters, often integrated directly with their speech-to-text APIs.

Market Growth: The AI voice cloning market is projected to expand to $9.56 billion by 2030, driven by advancements in generative audio and increased adoption in customer service and entertainment.

Regulatory Landscape: New disclosure laws are emerging globally. A recent analysis by JustCall outlines federal, state, and international rules requiring AI voice agents to disclose their non-human status at the start of calls, impacting how developers design opening lines and compliance checks for voice agents.


On the radar

  • SKT Beta Launch: SK Telecom targets an October 2026 beta release for its "AI for Everyone" voice assistant, which will be accessible via phone and SMS without app installation.
  • Full-Duplex Benchmarks: New benchmarks like τ-Voice are emerging to test full-duplex voice agents on real-world domains, moving beyond simple transcription accuracy to evaluate conversational coherence and interruption handling.
  • Deepfake Fraud Trends: Continued rise in voice-clone fraud ("CEO fraud") is prompting stricter enterprise verification protocols, with FBI-reported losses from AI fraud reaching $893 million in 2025.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow much does the GPT-Live-1 API cost?
  • QHow does CoeFont v4 achieve its speed boost?
  • QWhat are South Korea's beta launch details?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.