Cohere Transcribe — open-source speech-to-text model (announcement)
Cohere’s tweet announces Cohere Transcribe, the company’s first speech-to-text release and an open-source ASR model framed as a new SOTA in open ASR. Headline claim: top spot on Hugging Face’s Open ASR leaderboard for English with a 5.42% WER, validated by human evaluation, and one of the strongest accuracy/speed ratios among comparably sized speech models. The release is pitched as the speech-recognition layer for Cohere’s enterprise agentic platform (North), with the marketing quote emphasizing real-time throughput (“minutes of audio into usable transcripts in seconds”). No technical report, no architecture or parameter count, and no latency hardware are disclosed in the tweet itself.
Key claims
Section titled “Key claims”- Cohere Transcribe is Cohere’s first speech-to-text model release and is positioned as a step toward enterprise speech intelligence inside the North agentic platform [tweet OP].
- The open-source variant tops Hugging Face’s Open ASR leaderboard for English accuracy at 5.42% WER [tweet OP].
- Accuracy is validated by human evaluation in addition to leaderboard WER [tweet OP].
- The model claims one of the strongest accuracy-vs-speed ratios among open speech models in its size class [tweet OP].
- Marketing emphasizes real-time throughput (“minutes of audio into usable transcripts in seconds”), with the implied use case being live products and workflows [tweet OP].
Method
Section titled “Method”The tweet does not disclose architecture, parameter count, training data, tokenizer, decoding strategy, or hardware. It frames Transcribe as the input-side speech layer of Cohere’s broader enterprise AI stack (North, Command R family for LLM, Rerank/Embedding for search), with the only concrete deployment surface mentioned being Hugging Face’s Open ASR leaderboard for the open variant.
Results
Section titled “Results”The only quantitative claim in the tweet is the headline 5.42% WER on the English Open ASR leaderboard, claimed as #1 at the time of the announcement. No multilingual results, no per-domain breakdowns (long-form, telephony, accented English), no latency or throughput numbers with hardware, and no comparison against closed APIs (Whisper-large-v3, Gemini, ElevenLabs Scribe, AssemblyAI). Validation is described as “human evaluation” without a protocol.
Why it’s interesting
Section titled “Why it’s interesting”Direct counterpart to Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) from earlier today: both are open-weight speech announcements from frontier-LLM labs, both ship a single tweet with one headline number and no tech report, and both pitch a speech-to-speech stack (Voxtral pairs TTS with Voxtral Transcribe; Cohere pairs Transcribe with the rest of Command/North). Voxtral’s claim is a head-to-head TTS win over ElevenLabs Flash; Cohere’s claim is the leaderboard-top WER on Open ASR English. Different sides of the speech stack, same announcement playbook.
Fits the Open foundation-model releases pattern of frontier labs adding speech to their open-release surface — Mistral with Voxtral (TTS) on the same day, Qwen with Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation on the TTS side, and now Cohere stepping in on the ASR side. Worth tracking which open ASR model becomes the default baseline for Luma’s multimodal data pipelines (currently Whisper / WhisperX); 5.42% WER would beat large-v3 on the same leaderboard if the claim holds up to independent evaluation. The absence of a tech report or model card details means the leaderboard-top claim should be treated as marketing until reproduced.
See also
Section titled “See also”- Open foundation-model releases — fits the single-model open-release pattern; first Cohere entry on this page (previously all LLM/embedding-side).
- Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) — same-day open speech release from a sibling frontier lab; covers the TTS side of the same stack.
- Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation — adjacent open TTS family with a concrete tech report, the kind of follow-up Cohere Transcribe still owes.
- Tweet: https://twitter.com/cohere/status/2037159129345614174