Launching Miso One — The Most Emotive AI Voice Model
YouTube launch video from Miso Labs announcing Miso One, framed as “the most emotive AI voice model.” Posted as a sibling URL to the Miso TTS 8B GitHub release (Miso TTS 8B — Highly Emotive Text-to-Speech Model (Miso Labs)), which is the substantive artifact with architecture details, weights, and inference code. No transcript was available at filing time, so the page is filed as a pointer with detail deferred to the GitHub repo.
Key claims
Section titled “Key claims”- Marketing title positions the model as “The Most Emotive AI Voice Model” [video title].
Method
Section titled “Method”Not available — no transcript was extractable at filing time. The companion GitHub repo describes the underlying model as an ~8.2B-parameter RVQ Transformer with a Llama-8B backbone and Llama-300M audio decoder over Mimi codec audio frames, inspired by Sesame CSM (see Miso TTS 8B — Highly Emotive Text-to-Speech Model (Miso Labs)).
Results
Section titled “Results”Not retrievable from the video alone at filing time. The GitHub repo does not list benchmark numbers either.
Why it’s interesting
Section titled “Why it’s interesting”The launch frame (“most emotive”) is the same axis being pushed by other recent open TTS drops — Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) (Voxtral TTS — “emotionally expressive natural speech”) and Chatterbox Turbo — 350M Zero-shot TTS (Resemble AI) — suggesting expressivity/affect, not just intelligibility, is becoming the differentiating axis for open TTS in 2026. Worth checking the video for any side-by-side comparisons against Sesame CSM, ElevenLabs, or Voxtral that the GitHub README doesn’t reproduce.
See also
Section titled “See also”- Miso TTS 8B — Highly Emotive Text-to-Speech Model (Miso Labs) — sibling GitHub release with the actual architecture, weights, and inference code
- Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) — Voxtral TTS launch on the same “expressive open TTS” axis
- Chatterbox Turbo — 350M Zero-shot TTS (Resemble AI) — Chatterbox Turbo, another expressivity-focused open TTS
- Open foundation-model releases — the cluster this drop joins