Install
Speech & Audio
Speech recognition, voice analytics and conversation intelligence.
- 79 Tracked terms
- Last 30 days Feed window
What this topic collects on
An article joins this feed when it matches these terms. Each one is also a search of its own.
- +39 more
Related topics
Latest in Speech & Audio
Applied Sciences, Vol. 16, Pages 9314: VQ-CycleDiffusion: Vector Quantized Cycle-Consistent Diffusion Models for Voice Conversion
23+ hour, 23+ min ago (464+ words) This paper proposes a novel hybrid voice conversion framework, termed VQ-CycleDiffusion, which integrates vector quantization (VQ) with CycleDiffusion to address the inherent limitations of continuous latent representations in diffusion-based models. While conventional diffusion-based voice conversion (VC) models provide superior speech quality,…...
I Tried AI Voice Cloning for Content Creation — Here’s What Actually Makes It Useful
1+ day, 1+ hour ago (635+ words) There is a point in content creation when recording your own voice stops being creative and starts feeling …...
How to Choose a British AI Voice for Your Video
1+ day, 7+ hour ago (641+ words) Before publishing client work, check the provider's current usage terms and the permissions for your script and other material. For a performance...
SpaceXAI Releases Grok Voice Transcribe 2.0: A Speech-to-Text API Claiming 2x Accuracy Over 1.0 at $0.10 per Hour
1+ day, 10+ hour ago (491+ words) SpaceXAI has released Grok Voice Transcribe 2.0, its newest speech-to-text (STT) model. The development team claims it to be twice as accurate as Grok Voice Transcribe 1.0 at the same price. The model targets hard audio: noisy phone lines, competing voices, local…...
Google Search Live Now Runs on Gemini 3.8 Live
1+ day, 19+ hour ago (1170+ words) Home » News » Google Search Live Now Runs on Gemini 3.8 Live Google has replaced Search Live’s voice model with a new one. On Sep 15 2026 the company unveiled Gemini 3. 8 Live and Gemini 3. 8 Live Extended Thinking, two live dialogue models that are already…...
xAI Releases Grok Voice Transcribe 2.0 Speech-to-Text Model
1+ day, 12+ hour ago (218+ words) xAI released Grok Voice Transcribe 2.0, its latest speech-to-text model, on September 18, 2026, holding batch pricing at $0.10 per hour of audio while describing the model as twice as accurate as Grok Voice Transcribe 1.0. The official speech-to-text documentation lists 12 supported audio formats, a…...
Introducing Grok Voice Transcribe 2.0 | SpaceXAI
1+ day, 20+ hour ago (345+ words) Today we're releasing Grok Voice Transcribe 2.0, our latest speech-to-text model. Across our real-world evaluations, Grok Voice Transcribe 2.0 is one of the most accurate transcription models available today and twice as accurate as Grok Voice Transcribe 1.0, at the same price. The…...
Salesforce SVP Carlson Details Partner Potential With AIforce, Headless, AgentExchange
1+ day, 17+ hour ago (758+ words) ‘It’s going to be an incredible opportunity for’ Salesforce services partners, CEO and co-founder Marc Benioff said. Salesforce’s new AIforce surface area and Headless Toolkit for building experiences with Salesforce-owned applications and third-party tools are opening up new opportunities for…...
GPT-Realtime-2 and GPT-Live: Choosing a Voice AI Stack
1+ day, 12+ hour ago (19+ words) Compare OpenAI’s GPT-Realtime-2 and GPT-Live voice stack: speech, transcription, translation, reasoning, interruptions, and the costs of a complete session....
Google Gemini 3.8 Live Brings Background Reasoning to Real-Time Voice AI
1+ day, 17+ hour ago (667+ words) Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live-dialogue AI models designed to hold real-time conversations while handling visual context and background tasks. The release matters because it moves voice AI beyond simple question-and-answer exchanges toward systems…...