Mistral: Voxtral TTS, Forge, Leanstral, & what's next for Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Episode
48 min
Read time
2 min
Topics
Design & UX, Software Development, Philosophy & Wisdom
AI-Generated Summary
Key Takeaways
- ✓Autoregressive Flow Matching for TTS: Voxtral TTS uses a flow matching head attached to a 3B Ministral backbone, processing audio at 12.5 Hz latent tokens. Flow matching outperforms depth transformers for audio generation because it models the distribution of possible inflections rather than predicting a blurred mean, and reduces inference to 12–16 steps versus k autoregressive steps per frame.
- ✓Neural Audio Codec Design: The in-house codec converts audio into 12.5 Hz latent tokens, each containing one semantic token plus multiple acoustic tokens. Embeddings are summed at each frame on the input side. This continuous-discrete hybrid design enables streaming-first voice agent deployment, targeting sub-100ms latency for real-time applications rather than batch audio file processing.
- ✓Fine-Tuning on Proprietary Data Yields Outsized Gains: Enterprises using closed-source models via API leave decades of domain-specific data unused. Fine-tuning on proprietary data via Mistral Forge — using the same training infrastructure Mistral's science team uses internally — can produce models 10x cheaper to serve and significantly stronger on domain-specific tasks than any general-purpose model accessed through a shared endpoint.
- ✓LeanStral Uses Lean Formal Proofs as Verifiable RL Reward Signal: Formal proof verification in the Lean language provides a binary, unambiguous reward signal — code either compiles or it does not — solving the reward hacking problem that plagues open-ended mathematical reasoning. This enables reinforcement learning on complex multi-step proofs and transfers reasoning gains to coding and general problem-solving domains.
- ✓Sparse Mixture-of-Experts Architecture for Mistral Small: Mistral Small activates roughly 6B parameters out of a larger sparse MoE architecture, supports a 256k context window, and merges previously separate specialist models — coding, reasoning, vision — into one artifact. The approach keeps per-query compute low while consolidating capabilities, making it practical to deploy on-premise for latency-sensitive or data-privacy-constrained enterprise workloads.
What It Covers
Mistral releases Voxtral TTS, a 3B-parameter text-to-speech model supporting nine languages, built on a novel autoregressive flow matching architecture with an in-house neural audio codec. Guillaume Lample and Pavan Kumar Reddy also cover Mistral Small, the Forge deployment platform, and the LeanStral formal math reasoning project.
Key Questions Answered
- •Autoregressive Flow Matching for TTS: Voxtral TTS uses a flow matching head attached to a 3B Ministral backbone, processing audio at 12.5 Hz latent tokens. Flow matching outperforms depth transformers for audio generation because it models the distribution of possible inflections rather than predicting a blurred mean, and reduces inference to 12–16 steps versus k autoregressive steps per frame.
- •Neural Audio Codec Design: The in-house codec converts audio into 12.5 Hz latent tokens, each containing one semantic token plus multiple acoustic tokens. Embeddings are summed at each frame on the input side. This continuous-discrete hybrid design enables streaming-first voice agent deployment, targeting sub-100ms latency for real-time applications rather than batch audio file processing.
- •Fine-Tuning on Proprietary Data Yields Outsized Gains: Enterprises using closed-source models via API leave decades of domain-specific data unused. Fine-tuning on proprietary data via Mistral Forge — using the same training infrastructure Mistral's science team uses internally — can produce models 10x cheaper to serve and significantly stronger on domain-specific tasks than any general-purpose model accessed through a shared endpoint.
- •LeanStral Uses Lean Formal Proofs as Verifiable RL Reward Signal: Formal proof verification in the Lean language provides a binary, unambiguous reward signal — code either compiles or it does not — solving the reward hacking problem that plagues open-ended mathematical reasoning. This enables reinforcement learning on complex multi-step proofs and transfers reasoning gains to coding and general problem-solving domains.
- •Sparse Mixture-of-Experts Architecture for Mistral Small: Mistral Small activates roughly 6B parameters out of a larger sparse MoE architecture, supports a 256k context window, and merges previously separate specialist models — coding, reasoning, vision — into one artifact. The approach keeps per-query compute low while consolidating capabilities, making it practical to deploy on-premise for latency-sensitive or data-privacy-constrained enterprise workloads.
Notable Moment
Guillaume Lample observed that even native French, Spanish, and German speakers unconsciously slow down and over-articulate when talking to current voice AI, despite those languages having abundant training data — revealing a gap in naturalness that persists well beyond low-resource language limitations and that Mistral treats as a primary unsolved benchmark.
Episode Transcript
Okay. Welcome to Lanespace. We're here in the studio with trustee co host, Bibhu. Welcome. Hey. I added for this one. As well as Guillaume and Pavan from Mistral. Welcome. Excited to be here. Thank you for having us. Pavan, you are leading audio research at Mistral, and, Guillaume, you're a chief scientist. What are we announcing today? We're we're coordinating this release with you guys. Yeah. So we are releasing Vauxhall TTS. So it's our first audio model that generates speech. It's not our first audio model. We had a couple of releases before. We had one in the summer. That was Vauxhall, our first audio model, but it's it was like a transcription model ASR. And, like, a few months later, we released some update on top of this, supporting more languages. Also, a lot of table stack features for our customers, context biasing, position, time stamping, and the transcription. We will start some real time model that can transcribe not just at the end of the it was don't need to fill your entire audio file, but that can also come in real time. On here, this is a natural extension in the audio. So, basically, speech generation. So, yeah, so we support nine languages, and this is a pretty small model, three d model, so very fast and also status. Yeah. Is the same level of the of the best model, but it's much more efficient in terms of cost. And also much in terms of cost, it's also much to go. They're only a fraction of the cost of parking catiters. And we are also releasing the weight of this model. Yeah. Mami linked? Not this time. What yeah. What's the decision Factor him. It's a good question. There'll be more w. Oh. Yeah. Pavan, any other sort of research notes to add on on what it No. We maybe we'll dive into it later in the forecast too, but it's a novel architecture that we developed in house. We iterated on several internal architectures and ended up with a autoregressive flow matching architecture and also have a new in house neural audio codec, which converts this audio into all point by Hertz, latent tokens, semantic and acoustic tokens. And, yeah, that's that's their new part about this model, and we're pretty excited that it it came out with such good quality. And Guillaume was mentioning, yeah, it's a three b model. It's based off of the ministral model that we actually released just a few months back and insert trunk. And it mainly meant for, like, the TTS stuff, but they need text capabilities are also there. Yeah. So there's a lot to cover. I always I love any anything to do with novel encodings and all those things because I think that's obviously increase a lot of efficiency, but also maybe bugs sometimes happen. You were previously at Gemini, and you worked on post training for language models, and maybe a lot of …
Get the full transcript (10,808 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 45-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
Aug 11 · 95 min
The AI Breakdown
AI Optimism Has a Trust Problem
Aug 11
More from Latent Space
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Aug 3 · 101 min
a16z Podcast
ElevenLabs CEO: Why Voice is the Next AI Interface
Nov 5
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
- Mistral ForgeBy guest
by Mistral
“Fine-tuning on proprietary data via Mistral Forge — using the same training infrastructure Mistral's science team uses internally — can produce models 10x cheaper to serve.”
“Formal proof verification in the Lean language provides a binary, unambiguous reward signal — code either compiles or it does not.”
Products
- Voxtral TTSBy guest
by Mistral
“Mistral releases Voxtral TTS, a 3B-parameter text-to-speech model supporting nine languages, built on a novel autoregressive flow matching architecture with an in-house neural audio codec.”
- Mistral SmallBy guest
by Mistral
“Guillaume Lample and Pavan Kumar Reddy also cover Mistral Small, the Forge deployment platform, and the LeanStral formal math reasoning project.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Inside the Model Factory — Eiso Kant, Poolside AI
🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Similar Episodes
Related episodes from other podcasts
The AI Breakdown
Aug 11
AI Optimism Has a Trust Problem
a16z Podcast
Nov 5
ElevenLabs CEO: Why Voice is the Next AI Interface
Software Engineering Daily
Aug 11
SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing
Cognitive Revolution
Aug 10
Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
20VC (20 Minute VC)
Aug 10
20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Software Engineering Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime