CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-14

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-14

Speech, Voice and Realtime Audio Models|September 14, 2026(3h ago)3 min read8.8AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

OpenAI’s launch of the GPT-Live-1 API has dominated the week, offering a full-duplex voice model at $0.05 per minute that eliminates the need for cascaded speech pipelines. In parallel, NTT TechnoCross introduced new interruption-handling features for its AI voice bots, while South Korean telecom giants KT and SK Telecom detailed their "Everyone's AI" initiatives to integrate voice assistants into daily apps. The market continues to consolidate around low-latency, natural conversation capabilities, with pricing pressure and regulatory compliance becoming key differentiators.

Speech, Voice and Realtime Audio Models — 2026-09-14


Top developments


OpenAI GPT-Live-1 API Launches at $0.05/Minute

On September 10, 2026, OpenAI opened developer access to GPT-Live-1, a full-duplex voice model capable of listening and speaking simultaneously. Priced at $0.05 per minute for the voice layer, this API powers the 150 million weekly ChatGPT Voice users. Partners report an 80% reduction in false interruptions and a 23,000-line reduction in pipeline code compared to cascaded STT-LLM-TTS architectures. This move forces competitors to match native full-duplex capabilities rather than relying on stitched-together components.

OpenAI GPT-Live-1 interface
OpenAI GPT-Live-1 interface

techtimes.com

techtimes.com


NTT TechnoCross Enhances Voice Bot Interruption Handling

On September 10, 2026, NTT TechnoCross announced new features for its "CTBASE/SmartCommunicator" AI voice bot, focusing on reduced response latency and improved interruption handling. The update allows the bot to process user speech while speaking, enabling more natural, overlapping conversations typical of human interaction. This targets enterprise call centers where natural flow is critical for customer satisfaction, moving beyond turn-taking models.

NTT TechnoCross CTBASE SmartCommunicator
NTT TechnoCross CTBASE SmartCommunicator


KT and SK Telecom Push "Everyone's AI" Voice Integration

KT unveiled its "Inside" strategy on September 11, 2026, aiming to embed its "Everyone's AI" assistant into partner services like Daum, Musinsa, and Zippang, targeting 60 million touchpoints. Meanwhile, SK Telecom is developing end-to-end executable agents with partners in finance and healthcare, aiming for a beta release in October 2026. These initiatives focus on integrating voice and AI into existing high-traffic platforms rather than standalone apps, leveraging local ecosystem dominance.

KT Everyone's AI Strategy
KT Everyone's AI Strategy


Voice Cloning Market Projected to Hit $9.56 Billion by 2030

A new market outlook report released on September 11, 2026, forecasts the AI voice cloning market to expand to $9.56 billion by 2030. This growth is driven by increasing adoption in media, gaming, and customer service, despite rising regulatory scrutiny. The report highlights that while technical barriers drop, trust and verification mechanisms remain critical bottlenecks for enterprise adoption.


Local view

In Japan, eguweb highlighted the shift from cascaded pipelines to full-duplex systems like GPT-Live-1, noting the industry's move toward "listening while responding" architectures. In South Korea, Financial News reported that SK Telecom is forming a "dream team" of specialized partners to build "end-to-end" executable AI agents, emphasizing action-oriented voice assistants over simple chatbots.


Context & numbers

  • Pricing: OpenAI GPT-Live-1 is priced at $0.05 per minute. ElevenLabs remains higher-end at $0.05–$0.10 per 1K characters.
  • Latency: Cartesia Sonic 4 leads in pure Time-to-First-Audio (TTFA) at roughly 40ms, though this data is from May 2026 and serves as a current benchmark baseline.
  • Adoption: GPT-Live-1 supports 150 million weekly ChatGPT Voice users.

On the radar

  • SKT Beta Release: SK Telecom's "Everyone's AI" service is scheduled for a beta launch in October 2026, which will test large-scale voice agent integration in Korea.
  • Regulatory Compliance: JustCall published a comprehensive guide on AI Voice Agent Disclosure Laws on September 8, 2026, highlighting new federal and state requirements for disclosing AI callers, a critical factor for US-based voice agent deployments.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow do competitors plan to match GPT-Live-1 pricing?
  • QWhat security measures target voice cloning risks?
  • QHow do SK Telecom's agents handle complex tasks?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.