CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-11

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-11

Speech, Voice and Realtime Audio Models|September 11, 2026(3h ago)3 min read8.7AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

OpenAI has released GPT-Live-1, a new full-duplex voice model enabling simultaneous listening and speaking for developers. In Japan, NTT TechnoCross launched new natural conversation features for its AI voice bots, while Microsoft's recent MAI-Transcribe-2 model continues to be highlighted in local tech media for its high accuracy and low cost. Meanwhile, South Korea’s "AI for Everyone" consortium, led by SKT, Kakao, and KT, is accelerating development of voice-enabled public service agents targeted for a late 2026 launch.

Speech, Voice and Realtime Audio Models — 2026-09-11


Top developments


OpenAI releases GPT-Live-1 for full-duplex interaction

On September 11, 2026, OpenAI announced the release of GPT-Live-1, a new voice model designed for full-duplex voice agents. The model supports simultaneous listening and speaking, advanced interruption handling, and backend delegation, offering 12 distinct voices for developers. This release aims to significantly reduce latency and improve the naturalness of realtime voice-agent interactions by allowing the AI to process input while generating output, a key requirement for seamless conversational AI.

OpenAI GPT-Live-1 banner
OpenAI GPT-Live-1 banner


NTT TechnoCross enhances CTBASE voice bot capabilities

On September 10, 2026, NTT TechnoCross announced new features for its AI voice bot "CTBASE/SmartCommunicator," aimed at achieving more natural dialogues. The update optimizes response processes to better handle complex user inquiries in Japanese, reflecting the ongoing push by local Asian labs to improve the contextual awareness of speech recognition and synthesis systems for enterprise use. This move highlights the competitive pressure on global providers like Google and OpenAI to offer localized, high-fidelity voice solutions.

NTT CTBASE SmartCommunicator
NTT CTBASE SmartCommunicator


South Korea’s "AI for Everyone" consortium targets voice-enabled agents

In early September 2026, the South Korean government revealed that the consortia led by SK Telecom, Kakao, and KT are developing "end-to-end" executable AI agents for public services. These agents will leverage voice interfaces via phone and messenger apps to handle tasks like reservations and payments, with a nationwide launch targeted for the end of 2026. The project aims to serve 15 million users initially, driving demand for robust Korean-language speech recognition and TTS models from local and international vendors.

SKT AI Dream Team
SKT AI Dream Team


Local view

In Japan, media outlets such as AI Watch and Mado no Mori have been covering the rapid adoption of Microsoft’s MAI-Transcribe-2, which was marketed as the "fastest, most accurate, and cheapest" speech recognition model. Although the initial announcement was slightly earlier, coverage continued through September 7, 2026, focusing on its potential to disrupt the local ASR market dominated by domestic players like NTT and newer entrants. The emphasis is on its cost-efficiency for high-volume transcription tasks.

In South Korea, ZDNet Korea and Financial News are reporting on the strategic alliances formed under the "AI for Everyone" project. They highlight that SKT is focusing on lifestyle services, Kakao on messenger integration, and KT on content delivery, all requiring sophisticated voice agent technologies that can seamlessly switch between text and voice inputs.


Context & numbers

Recent pricing data indicates that ElevenLabs remains positioned in the premium segment for AI voice quality, with costs ranging from $0.05 to $0.10 per 1,000 characters. Competitors like Speechify have adjusted their self-serve TTS overage fees down to $6 per 1 million characters, while all-in voice-agent minutes are dropping to $0.068. These price adjustments reflect intense competition among providers like Cartesia, Deepgram, and OpenAI to secure enterprise contracts for large-scale voice deployments.


On the radar

  • Regulatory Compliance: New state-level disclosure laws for AI voice agents in the US are taking effect, requiring clear identification of AI callers. Developers must integrate compliance checks into their voice-agent architectures immediately to avoid penalties.
  • Hollywood Voice Actor Disputes: Tensions between Hollywood voice actors and AI companies regarding voice cloning rights continue to escalate, potentially influencing future licensing models for high-fidelity TTS voices.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does GPT-Live-1 handle latency?
  • QWhat are the pricing details for GPT-Live-1?
  • QWhen will South Korea's AI agents launch?
  • QHow does MAI-Transcribe-2 perform in Japanese?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.