On July 29, xAI announced Grok Voice Think Fast 2.0, its most advanced speech-to-speech model — and the benchmark numbers stand out: 82.9% on Artificial Analysis's overall voice quality index, ahead of GPT-Realtime-2.1 (79.1%) and Gemini 3.1 Flash (69.5%). The model replaces the previous version as the API default starting August 5.
What Changed
According to xAI's official blog and reporting from BigGo Finance and TestingCatalog, version 2.0 cut time-to-first-audio-response from 1.25 seconds to 0.70 seconds, and reached an agentic score of 56.5% (versus 45.7% for GPT-Realtime-2.1 and 37.7% for Gemini 3.1 Flash). Across 24 tested languages, xAI reports a 1.5-2x improvement over specialized transcription tools like Deepgram Nova 3 and ElevenLabs Scribe v2. Pricing is $0.08 per minute of audio, aimed at developers building voice agents.
Why It Matters
Voice-based service has always been the hardest mode to automate well — noticeable latency, transcription errors, and unnatural conversation flow break the experience quickly. A quality jump on this specific benchmark (response speed nearly halved, higher agentic score) signals voice agents are moving from demo phase into real production viability.
The Impact for Brazil
Phone-based customer service is still a massive sector in Brazil, and most automation to date has stayed limited to chat and traditional IVR. Faster, cheaper voice models open real room for voice agents in call centers, collections, scheduling, and technical support. One caveat: benchmarks published so far are mostly in English — before committing to any voice vendor, it's worth testing directly in Brazilian Portuguese with your sector's specific vocabulary, since benchmark performance doesn't guarantee equivalent performance in another language.
Entercast's Take
The pattern repeats: capability rises, cost falls, and the real decision stops being 'does the technology work?' and becomes 'is my operation ready to use it properly?' Companies that already have process, organized service data, and a defined human-escalation threshold can turn this kind of technical leap into production quickly. Companies without that in place will stay stuck in pilot, no matter how good the voice model gets.