
Grok voice transcribe 2.0 update improves short-phrase speech recognition
The latest speech-to-text update introduces significant accuracy gains for short phrases, alongside new features like word-level timestamps and speaker diarization. Existing API integrations receive these improvements automatically without requiring code changes.
Published by Jin · 2 min read · 3 OCT 2026
- 8 hours
- 50%
- 25%
- 25%
A new speech-to-text release aims to solve common hurdles in audio transcription, particularly when dealing with brief utterances that lack surrounding context.
Accuracy and short phrases
Short phrases, such as in-car commands, often give transcription systems little context to correctly identify the spoken language. On a dedicated short-phrase test set, the word error rate drops from 20.6% to 6.8% in this update.
Features and capabilities
Existing speech-to-text application programming interfaces — the software bridges that let different programs talk to each other — receive the accuracy improvements with no code changes needed for batch or streaming modes. Users can transcribe recorded files and URLs, or process live audio streams in real time.
The system includes several advanced controls:
- Word-level timestamps providing precise start and end times alongside confidence scores.
- Speaker diarization to label individual speakers in a transcript at no additional cost.
- Multichannel transcription supporting up to 8 independent audio channels.
- Key term biasing, which lets developers pass up to 100 domain-specific terms like medical vocabulary per request.
- Automated text formatting that converts numbers, dates, currencies, and emails into written form.
- Filler word removal to omit vocal pauses like "um" and "uh".
- Smart turn detection to identify the end of a speaker's turn for voice-based agents.
Real-world adoption and pricing
Source — Original announcement ↗
Worth a read?

Comments · 0