Speech, Voice and Realtime Audio Models — 2026-09-20
Google has launched Gemini 3.8 Live, a native speech-to-speech model that tops the Artificial Analysis Speech-to-Speech Index with a score of 82.6, marking a significant leap in real-time conversational AI. Meanwhile, voice cloning fraud continues to escalate, with dark web listings for AI voice-cloning tools jumping 7.4 times year-over-year, prompting urgent discussions on enterprise security controls and disclosure laws.
Speech, Voice and Realtime Audio Models — 2026-09-20
Top developments
Google Launches Gemini 3.8 Live with Native Speech-to-Speech Capabilities
On September 15, 2026, Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced native speech-to-speech models to date. The models are available via the Gemini API and Google AI Studio, designed to handle real-time conversations and video inputs with reduced latency and improved reasoning capabilities compared to previous cascaded pipelines. This launch is critical for developers building voice agents, as it moves beyond simple transcription or synthesis to full-duplex conversational understanding, potentially reducing the need for complex orchestration layers between STT and TTS systems.

Dark Web Voice Cloning Listings Surge 7.4x, Fueling Fraud Concerns
A new report highlights a dramatic increase in the availability of AI voice-cloning tools on dark web forums and Telegram channels, with commercial listings jumping 7.4 times between the first seven months of 2024 and the same period in subsequent years. This surge correlates with rising incidents of deepfake CEO fraud and business email compromise (BEC) attacks, where synthetic voices are used to bypass traditional authentication methods. Security experts are urging enterprises to adopt multi-layered verification protocols and identity-disclosure architectures to mitigate these high-velocity AI safety risks.

Voice Agent Disclosure Laws and Compliance Frameworks Gain Momentum
As AI voice agents become ubiquitous in customer service, regulatory scrutiny is intensifying. New analyses of federal, state, and international rules for 2026 emphasize the necessity of clear disclosure when an AI is interacting with users, with specific guidelines emerging for call centers and automated agents. Compliance is no longer optional; businesses must navigate a patchwork of laws that require explicit identification of AI callers to avoid penalties and maintain consumer trust. This regulatory landscape is forcing vendors like ElevenLabs, Deepgram, and OpenAI to embed compliance features directly into their API responses and agent configurations.
NTT Data Refreshes "Voista!" AI Features for Enterprise Voice Applications
In Japan, NTT Data announced a major refresh of its "Voista!" service, updating both the AI functionalities and the underlying service infrastructure. This move signals a push by major Asian telecom and tech players to integrate more sophisticated voice AI into enterprise workflows, competing directly with global platforms. The update focuses on enhancing natural language understanding and synthesis quality, aiming to provide more seamless voice interfaces for business customers in the region.

Local view
In South Korea, the government-led "Everyone's AI" project is seeing significant movement from local tech giants SK Telecom, KT, and Kakao. These companies are collaborating on a national AI assistant service, with SK Telecom highlighting the use of voice and text as primary interaction channels to ensure accessibility for elderly and digitally vulnerable populations. The focus is on practical tasks like hospital reservations and document issuance, leveraging native Korean language models rather than relying solely on global APIs. This initiative aims to launch a beta service in October and a full release by December, targeting 15 million users initially.
Context & numbers
- Benchmark Scores: Google's Gemini 3.8 Live achieved a score of 82.6 on the Artificial Analysis Speech-to-Speech Index, setting a new standard for real-time voice models.
- Fraud Statistics: The FBI logged $893 million in AI fraud losses in 2025, with fewer than 5% of victims reporting the incidents, underscoring the severity of voice cloning risks.
- Latency Leaders: In recent comparisons, Cartesia Sonic 4 leads pure latency at roughly 40ms Time-to-First-Audio (TTFA), while ElevenLabs v3 leads in voice expressiveness with support for 70+ languages.
- WER Plateau: Word error rate (WER) on clean English audio has plateaued among top providers (Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe), with differences now within 1-2 percentage points, shifting competition toward latency, cost, and multilingual support.
On the radar
- Korean "Everyone's AI" Beta: The beta launch of the SKT/KT/Kakao national AI assistant is scheduled for October 2026, which will test large-scale voice interaction capabilities in Korean.
- Deepfake Detection Standards: Industry groups are expected to release updated guidelines for deepfake call detection in Q4 2026, following the recent surge in fraud cases.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.