Plus Larkspur as an Ensemble LLM provider, Dictate 3 Chinese Mandarin, a healthcare webinar with Cobalt Cloud, and more.
͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ 

View in browser

Sonora Sonora
A dark banner reading Developer Digest beside a white equalizer glyph, with a pill that says Sonora for developers

Sonora Developer Digest

This issue covers telephony voice agents, reusable agent configurations, JavaScript and Python SDK releases, Larkspur LLM support, and Dictate 3 Chinese.

PRODUCT UPDATES

Reusable Agent Configurations for the Ensemble API

Three concentric light arcs beside a card showing a stored agent config id and the words Your agent is live and ready

Store and manage agent configurations through the Ensemble API, then reference them by ID instead of sending the full configuration on every WebSocket session. A/B test voices and prompts, manage per-customer agents, and support multi-agent architectures without a code deploy.

View the docs →

PRODUCT UPDATES

JavaScript SDK v4 and Python SDK v3 are now GA

Two major SDK releases are stable and production-ready. JavaScript SDK v4 and Python SDK v3 include updated APIs, improved type safety, and streamlined authentication. The migration guides list breaking changes and upgrade paths.

Read the blog →

PRODUCT UPDATES

Inbound and outbound telephony voice agents are here

Two production-patterned reference implementations for telephony voice agents on the Ensemble API. Inbound handles calls with function calling and barge-in. Outbound adds answering machine detection and voicemail. Both deploy in one command.

View the telephony docs →

PRODUCT UPDATES

Larkspur now available as an Ensemble LLM provider

Larkspur models can now be used in the Ensemble API. Larkspur Reason 49B handles multi-step agent reasoning, while Larkspur Nano 8B delivers cost-efficient performance for targeted tasks. Both are in the Standard pricing tier. Set the provider type to larkspur and go.

View the provider docs →

PRODUCT UPDATES

Dictate 3 adds Chinese Mandarin: Simplified and Traditional

Dictate 3 now supports Simplified Chinese (zh, zh-CN, zh-Hans) and Traditional Chinese (zh-TW, zh-Hant). Set the model to dictate-3 with the matching language code to transcribe Mandarin audio in both streaming and batch requests.

Try it in the Playground →

EVENTS · WEBINAR

AI-Powered Outbound Dialing in Healthcare

A dark webinar card with the Sonora and Cobalt Cloud names, the title AI-Powered Outbound Dialing in Healthcare, the date Thu Apr 16 at 9 AM PT, and speakers Ines Marlow and Tomas Reyes

Sonora and Cobalt Cloud show how voice agents handle clinical trial screening calls with real conversational flow. See two live reference architectures: Sonora speech models on Cobalt’s managed model platform, and Cobalt’s contact center suite for a fully managed path.

When: Thu, Apr 16 · 9 AM PT / 12 PM ET

Speakers: Ines Marlow, Director of Product at Sonora, and Tomas Reyes, Partner Solutions Architect at Cobalt Cloud

Register now

Quick Hits

Fast links. Big signal.

The Definitive Guide to Voice AI Agents [E-book] – Architecture-level playbook covering the full voice agent stack, four build-approach trade-offs, conversational UX design, performance diagnostics, and compliance architecture.

Low Latency Voice AI: What It Is and How to Achieve It – How to hit sub-300ms voice AI latency using streaming recognition, real-time LLM processing, and enterprise deployment strategies.

Which STT API Handles Production Reality? – Real-world accuracy, latency, pricing, and production scalability compared for enterprise speech-to-text.

WebSocket vs REST for Text-to-Speech: When to Use Which – A decision framework for choosing between WebSocket and REST TTS APIs based on latency requirements, telephony, and voice agent use cases.

Barge-In, Interruptions, and Turn-Taking – When interruption handling holds up in call center deployments, and when noisy audio and concurrency demand custom speech infrastructure.

How AI Contact Centers Detect Caller Intent – How speech recognition, NLU, and classification work together for intent detection, with accuracy benchmarks and latency requirements.

Sonora, 12 Harbor Mill, Brooklyn, NY 11201
Unsubscribe  |  Manage preferences