xAI releases Grok Voice Transcribe 2.0 and makes it default in Speech-to-Text API
Grok Voice Transcribe 2.0 is a production speech-to-text model xAI said improves accuracy over version 1.0 across multiple internal test sets, will be promoted to the default Speech-to-Text API model, and keeps batch/stream pricing identical to v1.0.
In this brief: 3 sections 2 min read
Company claims the model was trained on noisy, multilingual, real-world audio.
xAI says Transcribe 2.0 ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
Internal evaluation sets: telephony support calls, Grok conversations, spoken credentials, short multilingual commands — company reports improvements on all four.
Available via Speech-to-Text API for batch and real-time streaming.
Features: word-level timestamps with confidence, free speaker diarization, up to 8-channel transcription, key-term biasing (100 terms), formatting of numbers/dates/emails, filler-word removal, smart turn detection.
Pricing: batch $0.10 per hour, streaming $0.20 per hour — same as v1.0; v1.0 remains pin-able but will be deprecated in coming weeks.
xAI says Atlassian evaluated Transcribe 2.0 versus its prior solution and now uses it to transcribe every Loom video.
Announcement includes a quoted Atlassian executive endorsing the workflow link between Loom transcripts and downstream tools.
xAI also says its voice models already power broad production usage (tens of thousands of support calls, millions of hours of narration) but those usage figures cover the wider Grok Voice system.