CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-16

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-16

Speech, Voice and Realtime Audio Models|September 16, 2026(2h ago)4 min read8.5AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

Google launched Gemini 3.8 Live, a native speech-to-speech model that topped the Artificial Analysis Speech-to-Speech Index with a score of 82.6, challenging OpenAI’s recent GPT-Live-1 release. Meanwhile, in South Korea, KT reported achieving a 75% resolution rate for AI voice bots in its contact centers, highlighting the maturation of enterprise voice agents.

Speech, Voice and Realtime Audio Models — 2026-09-16


Top developments


Google Releases Gemini 3.8 Live with "Extended Thinking" for Voice

On September 15, 2026, Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, positioning them as its most advanced native speech-to-speech models to date. The models are available via the Gemini API and Google AI Studio, aiming to eliminate the latency and awkward pauses associated with cascaded STT-LLM-TTS pipelines. Early benchmarks from the Artificial Analysis Speech-to-Speech Index place Gemini 3.8 Live at the top with a score of 82.6, surpassing previous leaders. This release directly competes with OpenAI’s GPT-Live-1, which launched its API earlier this month, signaling a fierce battle for dominance in real-time conversational AI.

Google Gemini 3.8 Live announcement graphic
Google Gemini 3.8 Live announcement graphic

techpluto.com

techpluto.com


KT Achieves 75% Voice Bot Resolution Rate in Contact Centers

KT, one of South Korea’s major telecom operators, reported that its generative AI-enhanced AICC (AI Contact Center) has reached a 75% completion rate for customer inquiries without human intervention. This milestone was highlighted in recent media coverage on September 14–15, 2026, showcasing the effectiveness of their "Eum Inside" strategy to embed AI into existing services. The high resolution rate suggests that current voice-agent models, when properly integrated with enterprise knowledge bases, can handle complex customer service tasks at scale. This development underscores the shift from experimental voice bots to production-grade autonomous agents in regulated industries like telecommunications.

KT AI Contact Center illustration
KT AI Contact Center illustration


Shisa AI Launches High-Precision, Low-Latency Realtime STT API

Japanese startup Shisa AI announced the launch of a new real-time speech recognition (STT) API on September 15, 2026, emphasizing high precision and ultra-low latency. The service targets developers building real-time voice agents where transcription speed is critical for natural conversation flow. While specific word error rate (WER) figures were not detailed in the press release, the focus on "ultra-low latency" positions it against incumbents like Deepgram and AssemblyAI, which currently dominate the sub-300ms latency market segment. This addition to the Japanese market reflects growing demand for localized, high-performance speech infrastructure.


Rising Concerns Over AI Voice Cloning in Financial Fraud

Security firm Neovera published a report on September 15, 2026, detailing how AI voice cloning is increasingly targeting banks and call centers. The report notes that traditional security assumptions—such as the trust inherent in voice conversations—are being eroded by realistic synthetic voices. This trend coincides with broader industry warnings about deepfake CEO fraud, with previous data indicating $893 million in AI fraud losses in 2025. As voice agents become more prevalent, the ability to distinguish between human and AI callers is becoming a critical compliance and security challenge for enterprises.


Local view

Japan: Media outlets such as Impress Watch and AI Watch have focused heavily on Google’s Gemini 3.8 Live launch, analyzing its cost-effectiveness and integration into the Gemini API. A comparative analysis by YAYAFA highlighted the architectural differences between Google’s approach (single model with extended thinking) and OpenAI’s GPT-Live-1, noting how each handles conversational silence and latency differently.

South Korea: The narrative is dominated by the government’s "All People's AI" initiative, where SK Telecom, Kakao, and KT are competing to provide free national AI services by year-end. KT’s achievement of a 75% voice bot resolution rate is being cited as evidence that domestic telcos are successfully leveraging AI to improve customer experience and operational efficiency. Additionally, SKT and KT were recently recognized as top performers in the Call Center Quality Index, largely attributed to their effective integration of AI technologies.


Context & numbers

  • Benchmark Scores: Google’s Gemini 3.8 Live scored 82.6 on the Artificial Analysis Speech-to-Speech Index, setting a new benchmark for native voice models.
  • Latency Standards: Leading TTS providers like Cartesia Sonic 4 continue to push pure latency down to approximately 40ms TTFA (Time to First Audio), while enterprise-grade voice agents aim for sub-300ms end-to-end response times.
  • Fraud Statistics: The FBI previously logged $893 million in AI-related fraud losses in 2025, with voice cloning identified as a primary vector for business email compromise (BEC) and social engineering attacks.
  • Market Landscape: The top STT providers (Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe) now sit within 1-2 percentage points of each other in Word Error Rate (WER) on clean English audio, indicating a plateau in raw accuracy improvements and a shift toward latency and multilingual capabilities.

On the radar

  • Korean "All People's AI" Beta: The beta service for the government-backed "All People's AI" project by Kakao, SKT, and KT is scheduled to launch by the end of September 2026, with full service expected in December. This will be a major test case for large-scale, free-to-use public voice agents.
  • Voice Agent Disclosure Laws: With new federal and state regulations emerging in 2026, enterprises must increasingly comply with strict disclosure requirements for AI-driven phone calls. Legal frameworks are tightening around when and how a bot must identify itself as non-human during voice interactions.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Gemini 3.8 Live's thinking work?
  • QWhat security measures combat voice cloning?
  • QHow does KT's AI handle complex calls?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.