
Introducing Grok Voice Think Fast 2.0 for speech applications
Grok Voice Think Fast 2.0 brings improved reasoning, higher transcription accuracy, and natural conversational pacing to real-time speech applications. The new model processes complex workflows while maintaining low latency and predictable pricing.
Published by Jin · 2 min read · 26 AUG 2026
- $0.08 / min of audio
- August 5, 2026
- 24 languages
A new speech-to-speech voice model has been introduced to improve intelligence, transcription accuracy, and conversational capabilities. The system builds on its predecessor with meaningful gains in speech reasoning, conversational ability, and tool use reliability.
Transcription Accuracy
The model outperforms dedicated transcription tools across multiple languages. In evaluations spanning thousands of short phrases across 24 different languages, it demonstrated a 1.5 to 2.0 times improvement over Deepgram Nova 3 and ElevenLabs Scribe v2. It also achieved a 1.4 times improvement relative to the first version of the model.
In noisy settings, such as environments with background noise and telephone compression, the performance gap widens to roughly 10 times compared to dedicated speech-to-text systems.
Reasoning Efficiency
The system reasons through queries while speaking. This parallel processing makes the model faster and smarter without increasing latency. Version 2.0 uses reasoning tokens more efficiently than its predecessor, allowing tool calls to execute quickly—often before the agent finishes its first sentence.
Conversation Flow
Extensive reinforcement learning—a training method that uses rewards and penalties to guide behavior—encourages the model to mimic natural human conversation. The system speaks in shorter sentences, asks single questions, and avoids unnecessary filler words. This creates a fluid experience while the model guides users through complex workflows.
Migration and Pricing
Source — Original announcement ↗
Worth a read?
Comments · 0