Skip to content

Launching Miso One — The Most Emotive AI Voice Model

YouTube launch video from Miso Labs announcing Miso One, framed as “the most emotive AI voice model.” Posted as a sibling URL to the Miso TTS 8B GitHub release (Miso TTS 8B — Highly Emotive Text-to-Speech Model (Miso Labs)), which is the substantive artifact with architecture details, weights, and inference code. No transcript was available at filing time, so the page is filed as a pointer with detail deferred to the GitHub repo.

  • Marketing title positions the model as “The Most Emotive AI Voice Model” [video title].

Not available — no transcript was extractable at filing time. The companion GitHub repo describes the underlying model as an ~8.2B-parameter RVQ Transformer with a Llama-8B backbone and Llama-300M audio decoder over Mimi codec audio frames, inspired by Sesame CSM (see Miso TTS 8B — Highly Emotive Text-to-Speech Model (Miso Labs)).

Not retrievable from the video alone at filing time. The GitHub repo does not list benchmark numbers either.

The launch frame (“most emotive”) is the same axis being pushed by other recent open TTS drops — Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) (Voxtral TTS — “emotionally expressive natural speech”) and Chatterbox Turbo — 350M Zero-shot TTS (Resemble AI) — suggesting expressivity/affect, not just intelligibility, is becoming the differentiating axis for open TTS in 2026. Worth checking the video for any side-by-side comparisons against Sesame CSM, ElevenLabs, or Voxtral that the GitHub README doesn’t reproduce.