Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

Cartesia has released Sonic-3.6, the newest version of its real-time text-to-speech model. It arrives roughly three months after Sonic-3.5. The new change is naturalness, and this one is independently checkable. Sonic 3.6 now holds #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on the Provider Voice board and 1,123 on the Controlled Voice board. The second result matters more. That board clones every model onto the same eight reference voices, which isolates the synthesis engine from the voice catalog. Sonic-3.6 leads it, with Sonic-3.5 second and ElevenLabs Eleven v3 third. The model runs on state space models rather than transformers, and Cartesia states sub-90ms time-to-first-audio. It is available in beta.

Is it deployable?

YES, it is available in beta and as a hosted API. Not as self-hosted weights.

Sonic is a closed, commercial model. There are no open weights and no Hugging Face repo. You rent it.

Company level: Solo developers and startups (Free/Pro $5 tiers), scaleups running contact centers (Startup $49 / Scale $299), and regulated enterprises needing DPAs, BAAs, and SSO.

Industries: Financial services, healthcare, retail and e-commerce, logistics, recruiting, SaaS support, consumer companion apps, media localization

Applications: Inbound support agents, outbound qualification calls, IVR replacement, appointment reminders, sales-training simulators, audio localization, in-product voice UI

The Architecture

Sonic runs on state space models rather than transformers. Cartesia’s launch page frames the usual tradeoffs — speed versus naturalness, accuracy versus cost — as architectural, not inevitable.

The practical output is time-to-first-audio. Cartesia states sub-90ms TTS latency, and 100ms transcript latency for its Ink-2 speech-to-text model. Both are vendor-stated model latency, not measured end-to-end round trips.

Interactive explainer