Speech, Voice and Realtime Audio Models — 2026-09-25
Google dominated the week with two new Gemini 3.8 TTS models unveiled September 23, drawing extensive coverage in Japanese-language tech media. OpenAI's ChatGPT mobile app gained voice-based agentic features for paid users, while in South Korea, SK Telecom pushed ahead with voice-call AI agents despite the Chuseok holiday. Voice-cloning fraud and disclosure compliance continued to surface as the sector's main risk themes.
Speech, Voice and Realtime Audio Models — 2026-09-25
Top developments
Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS
On September 23, 2026, Google announced two new text-to-speech models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — available immediately through Gemini API and Google AI Studio, with same-day integration into Gemini Notebook and Google Vids. The models go beyond plain read-aloud: voices can be designed from natural-language prompts specifying character traits and per-line expression, with around 2,000 voice types and multilingual dubbing/translation as headline use cases, positioning them for games, audiobooks and podcasts. This extends Google's September 15 release of Gemini 3.8 Live speech-to-speech models and strengthens its bid against ElevenLabs' expressive-TTS lead.

ChatGPT mobile app adds voice-based agentic features
TechCrunch reported on September 23, 2026 that OpenAI's ChatGPT mobile app now supports voice-based agentic tasks: Pro and Plus users can use the Work tab on their phones to complete tasks initiated by voice. It signals a shift in consumer voice AI from conversation to task execution — a competitive pressure point for real-time voice-agent stacks from Google, ElevenLabs and startups building phone-handling agents.

SK Telecom races voice AI "call on your behalf" through Chuseok holiday
South Korean media reported that SK Telecom's consortium is continuing development of its next-generation A.X K3 model and the "AI for All" service through the Chuseok holiday, with a beta planned for October. In an interview, SKT confirmed that pilot "call on your behalf" ("대신 걸기") and "answer on your behalf" ("대신 받기") voice features launch in October, with full consumer availability targeted for December — a nationwide voice-agent-on-the-phone play aimed at users who struggle with apps, such as the elderly booking train tickets.
AI voice agents enter business phone lines — with a compliance checklist
Telecom Reseller published a practical compliance checklist on September 24, 2026 for AI voice agents replacing IVRs and receptionist workflows within existing PBX call flows. As large-volume calling agents reach production, disclosure requirements (whether callers must be told they're talking to an AI, varying by US state and country) are becoming a procurement gating factor — a major deployment constraint alongside latency and WER metrics.
Local view
Japan: Teach media gave the Gemini 3.8 TTS release prominent play. AI Watch (Impress) reported the September 23 announcement with immediate availability via Google AI Studio and the Gemini API; sbbit.jp highlighted the ability to design 2,000 voice types via natural language; and 일간공업/tech sites noted the voice-clone conditions and pricing details. Separately, Sankei carried a PR Times release on Notta's AI voice recorder "Notta Memo Pro," launching October 13, 2026.
South Korea: Dailyian and Ajunews framed SKT's voice-call AI as a universal-access play — "don't install the app, the AI calls for you" — as part of the government-backed "AI for All" free national service, with beta in October and full launch in December. KT followed with an "AI for All" consortium launch event on September 21 at its Gwanghwamun WEST building, uniting 16 companies ahead of an October beta.
Context & numbers
- ElevenLabs v3 remains the reference point for expressiveness in voice: its own rates are published per-1K characters, and coverage highlights 70+ languages and 5,000+ voices
- Cartesia Sonic 4 leads pure latency at roughly 40ms TTFA (time to first audio), while Hume Octave 2 leads emotional fidelity
- On STT, word error rate on clean English audio has plateaued: Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe, ElevenLabs Scribe v2 and Microsoft MAI-Transcribe-1 sit within 1–2 percentage points of each other
- Voice-cloning fraud remains the sector's dark side: the FBI logged $893 million in AI fraud losses in 2025, with fewer than 5% of victims reporting
On the radar
- Notta Memo Pro (AI voice recorder) launches October 13, 2026 in Japan
- "AI for All" Korean betas: SKT, KT and Kakao consortia all target an October beta and December full launch of nationwide free services, including voice-based call agents
- Deepgram Flux TTS pricing: reported sub-100ms latency with free access through September — pricing after the free window ends is worth watching (earlier report; status may change)
- Compliance tightening: AI voice-agent disclosure laws by US state and country are accumulating, with practitioners converging on a single opening-line disclosure covering multiple regimes
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.