What’s new
More accurate transcription
1.5–2× better word error rate than dedicated STT models across 24 languages. Up to ~10× better in noisy, telephony conditions.
Efficient parallel reasoning
Reasons while speaking with ~60% fewer reasoning tokens than 1.0, so tool calls and responses feel snappier.
More natural conversation
Shorter sentences, one question at a time, less filler—trained to sound like real conversation while still driving complex workflows.
|