CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-25

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — 2026-09-25

Speech, Voice and Realtime Audio Models|September 25, 2026(1h ago)4 min read8.8AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

Google dominated the week with two new Gemini 3.8 TTS models unveiled September 23, drawing extensive coverage in Japanese-language tech media. OpenAI's ChatGPT mobile app gained voice-based agentic features for paid users, while in South Korea, SK Telecom pushed ahead with voice-call AI agents despite the Chuseok holiday. Voice-cloning fraud and disclosure compliance continued to surface as the sector's main risk themes.

Speech, Voice and Realtime Audio Models — 2026-09-25


Top developments


Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS

On September 23, 2026, Google announced two new text-to-speech models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — available immediately through Gemini API and Google AI Studio, with same-day integration into Gemini Notebook and Google Vids. The models go beyond plain read-aloud: voices can be designed from natural-language prompts specifying character traits and per-line expression, with around 2,000 voice types and multilingual dubbing/translation as headline use cases, positioning them for games, audiobooks and podcasts. This extends Google's September 15 release of Gemini 3.8 Live speech-to-speech models and strengthens its bid against ElevenLabs' expressive-TTS lead.

Announcement coverage of Google's Gemini 3.8 Flash TTS models
Announcement coverage of Google's Gemini 3.8 Flash TTS models


ChatGPT mobile app adds voice-based agentic features

TechCrunch reported on September 23, 2026 that OpenAI's ChatGPT mobile app now supports voice-based agentic tasks: Pro and Plus users can use the Work tab on their phones to complete tasks initiated by voice. It signals a shift in consumer voice AI from conversation to task execution — a competitive pressure point for real-time voice-agent stacks from Google, ElevenLabs and startups building phone-handling agents.

Coverage of ChatGPT's mobile voice agentic features
Coverage of ChatGPT's mobile voice agentic features

techcrunch.com

ChatGPT mobile app gets voice-based agentic features | TechCrunch


SK Telecom races voice AI "call on your behalf" through Chuseok holiday

South Korean media reported that SK Telecom's consortium is continuing development of its next-generation A.X K3 model and the "AI for All" service through the Chuseok holiday, with a beta planned for October. In an interview, SKT confirmed that pilot "call on your behalf" ("대신 걸기") and "answer on your behalf" ("대신 받기") voice features launch in October, with full consumer availability targeted for December — a nationwide voice-agent-on-the-phone play aimed at users who struggle with apps, such as the elderly booking train tickets.


AI voice agents enter business phone lines — with a compliance checklist

Telecom Reseller published a practical compliance checklist on September 24, 2026 for AI voice agents replacing IVRs and receptionist workflows within existing PBX call flows. As large-volume calling agents reach production, disclosure requirements (whether callers must be told they're talking to an AI, varying by US state and country) are becoming a procurement gating factor — a major deployment constraint alongside latency and WER metrics.


Local view

Japan: Teach media gave the Gemini 3.8 TTS release prominent play. AI Watch (Impress) reported the September 23 announcement with immediate availability via Google AI Studio and the Gemini API; sbbit.jp highlighted the ability to design 2,000 voice types via natural language; and 일간공업/tech sites noted the voice-clone conditions and pricing details. Separately, Sankei carried a PR Times release on Notta's AI voice recorder "Notta Memo Pro," launching October 13, 2026.

South Korea: Dailyian and Ajunews framed SKT's voice-call AI as a universal-access play — "don't install the app, the AI calls for you" — as part of the government-backed "AI for All" free national service, with beta in October and full launch in December. KT followed with an "AI for All" consortium launch event on September 21 at its Gwanghwamun WEST building, uniting 16 companies ahead of an October beta.


Context & numbers

  • ElevenLabs v3 remains the reference point for expressiveness in voice: its own rates are published per-1K characters, and coverage highlights 70+ languages and 5,000+ voices
  • Cartesia Sonic 4 leads pure latency at roughly 40ms TTFA (time to first audio), while Hume Octave 2 leads emotional fidelity
  • On STT, word error rate on clean English audio has plateaued: Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe, ElevenLabs Scribe v2 and Microsoft MAI-Transcribe-1 sit within 1–2 percentage points of each other
  • Voice-cloning fraud remains the sector's dark side: the FBI logged $893 million in AI fraud losses in 2025, with fewer than 5% of victims reporting

On the radar

  • Notta Memo Pro (AI voice recorder) launches October 13, 2026 in Japan
  • "AI for All" Korean betas: SKT, KT and Kakao consortia all target an October beta and December full launch of nationwide free services, including voice-based call agents
  • Deepgram Flux TTS pricing: reported sub-100ms latency with free access through September — pricing after the free window ends is worth watching (earlier report; status may change)
  • Compliance tightening: AI voice-agent disclosure laws by US state and country are accumulating, with practitioners converging on a single opening-line disclosure covering multiple regimes

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Gemini 3.8 TTS compare to ElevenLabs?
  • QWhat security risks come with voice-based agents?
  • QHow do SK Telecom's voice features handle privacy?
  • QWhat are the compliance rules for AI phone agents?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.