CrewCrew
FeedSignalsMy Subscriptions
Get Started
Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — October 5, 2026

  1. Signals
  2. /
  3. Speech, Voice and Realtime Audio Models

Speech, Voice and Realtime Audio Models — October 5, 2026

Speech, Voice and Realtime Audio Models|October 5, 2026(2h ago)4 min read8.8AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

Microsoft shipped its first streaming transcription model (MAI-Transcribe-2-Streaming) alongside two multilingual speech synthesis models, closing latency gaps that have plagued real-time voice agents. ElevenLabs doubled its valuation to $22B in a secondary round, while Modulate raised $25M for deepfake detection and compliance tools as voice AI moves from demos into production call centers.

Speech, Voice and Realtime Audio Models — October 5, 2026


Top developments


Microsoft Completes Voice Agent Pipeline With Streaming Transcription

Microsoft AI (MAI) released three speech models on October 1–2: MAI-Transcribe-2-Streaming (its first streaming transcription model), MAI-Voice-2.1, and MAI-Voice-2.1-Flash. The transcription model delivers initial results in an average of 320 milliseconds and costs $0.54 per audio hour on Microsoft Azure AI Speech and Foundry. These models target the silence and latency that still make many voice agents feel unnatural—a critical gap for sub-one-second response times.

Screenshot showing Microsoft's MAI voice model architecture and latency metrics
Screenshot showing Microsoft's MAI voice model architecture and latency metrics

techtimes.com

techtimes.com


ElevenLabs Doubles Valuation to $22B; Secondary Round Signals Investor Confidence

ElevenLabs, the leading AI voice synthesis platform, raised $300 million in a secondary tender offer on September 30, doubling its valuation to $22 billion. The round was co-led by Wellington and T. Rowe Price. This signals sustained investor appetite for production voice AI as call center deployments accelerate globally.


Modulate Raises $25M for Voice Deepfake Detection and Agent Compliance

Boston-based Modulate raised $25 million on September 28 led by Future Ventures, with participation from Hyperplane, to scale its voice analysis suite for deepfake detection and voice agent compliance. Modulate's models flag synthetic voices in real time and audit agent disclosures—critical as call center voice agents proliferate without adequate transparency safeguards. The funding underscores a $5.5 billion annual market forecast for voice detection checks by 2028.

Modulate's voice AI analysis dashboard interface
Modulate's voice AI analysis dashboard interface


Full-Duplex Models Converge on Sub-100ms Latency; Cartesia Sonic-3 Leads Speed

Among full-duplex voice models that listen and talk simultaneously, Cartesia Sonic-3 remains the fastest option at ~85ms latency, suitable for conversational interruptions and WebRTC integrations. OpenAI's GPT-Live-1 and three open models now compete in this space. Meanwhile, Deepgram Aura-2 achieves sub-200ms latency and can be optimized to ~90ms; cloud providers (Google Cloud, Amazon Polly, Microsoft Azure, OpenAI gpt-realtime) typically deliver 150–300ms under optimal conditions.

Latency comparison chart across Cartesia, Deepgram, and OpenAI voice models
Latency comparison chart across Cartesia, Deepgram, and OpenAI voice models

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com


South Korea's "AI for All" Service Launches October Beta; Telcos and Kakao Deploy National AI Stack

SK Telecom, KT, and Kakao, selected as consortium lead for South Korea's government "AI for All" (모두의 AI) initiative, begin beta rollout this October with free national AI service. The deployment integrates voice agents into telecom, messaging, and lifestyle platforms—positioning voice AI as critical national infrastructure.


Local view

Japan (ITmedia, Zaikei Shinbun): ITmedia reported Microsoft's MAI-Transcribe-2-Streaming and MAI-Voice-2.1 lineup as a strategic entry into real-time transcription, filling a gap Japanese telcos had relied on external vendors for. Zaikei Shinbun framed the announcement as part of broader real-time voice AI competition.

South Korea (Yonhap, ZDNet Korea, Daum): Korean media emphasized the October AI for All beta as a nationalistic moment—free, government-backed AI competing against US vendors (OpenAI, Google). Multiple outlets highlighted the SKT–KT–Kakao consortium's deployment across call centers, messaging, and customer service as proof Korea could field production voice agents at scale.


Context & numbers

  • Latency leader: Cartesia Sonic-3 at ~85ms (full-duplex)
  • Transcription cost: Microsoft MAI-Transcribe-2-Streaming at $0.54/audio hour
  • Word error rate parity: Top STT providers (Deepgram Nova-3, AssemblyAI Universal-3 Pro, OpenAI gpt-4o-transcribe, ElevenLabs Scribe v2, Microsoft MAI-Transcribe-1) now sit within 1–2 percentage points on LibriSpeech and FLEURS; clean English WER has plateaued.
  • Open-source representation: Only 16 of 92 TTS models on Artificial Analysis are open-weights (17.4%); NVIDIA's Canary Qwen 2.5B leads Hugging Face Open ASR Leaderboard at 5.63% WER.
  • Voice detection market: Projected to reach $5.5 billion annually by 2028 as deepfake and compliance use cases scale.

On the radar

  • October 5–31, 2026: South Korea's "AI for All" beta window; watch for usage reports and call-center integration milestones
  • Decagon Voice 3 + Chord model: Announced October 1; supports 70+ languages; track for multilingual call-center deployments and latency benchmarks
  • Hugging Face Open TTS Leaderboard: Posted September 30; watch for open-model ranking shifts as NVIDIA, Kokoro, and community models compete
  • Deepfake disclosure regulations: US state and international rules remain fragmented; November policy updates possible as voice agents scale

Sources cited:

  • ElevenLabs $22B valuation
  • Modulate $25M funding
  • South Korea AI for All rollout
  • STT WER parity analysis
  • Open TTS leaderboard
techtimes.com

techtimes.com

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does MAI-Transcribe-2-Streaming compare in cost?
  • QWhat are the main use cases for Modulate's tools?
  • QHow does Cartesia Sonic-3 achieve 85ms latency?
  • QWhat features are included in South Korea's beta?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.