Pavan Kumar Reddy leads audio research at Mistral AI. He joins Tim Scarfe for a deep technical tour of Voxtral — and explains why, after everything that has landed in the last two years, the frontier of deployed voice is still a cascade of specialised models rather than one end-to-end system. IN PARTNERSHIP WITH MISTRAL AI: --- This episode was produced in partnership with Mistral AI. Mistral AI: https://mistral.ai/ --- The conversation opens on architecture. Voxtral Chat feeds a 3B Ministral text trunk with continuous embeddings from an audio encoder, fed into the decoder as direct token input rather than through cross-attention as in Whisper, so the model can answer questions about emotion, timing and who spoke when without an intermediate transcript to lose them. The real-time model changes shape into a dual-stream decoder that consumes audio and emits text at the same time, at a configurable target delay down to 160ms, with slower streams running in parallel for anything that can afford to wait for more context. On the generation side, Pavan explains why Voxtral TTS predicts continuous latents rather than discrete codec tokens, traces the lineage from SoundStream through EnCodec to Mimi's split of semantic and acoustic codebooks, and places finite scalar quantisation and flow matching in it. Tim presses on the engineering priors underneath: why a mel spectrogram instead of the raw waveform, what noise augmentation actually buys, and the point at which acoustic overfitting becomes somebody's fine-tuning problem. Then the failure modes. Speaker diarisation is emitted autoregressively as part of the transcript rather than by a separate head, which makes streaming diarisation fragile in a particular way — less context, late speaker changes, invented extra speakers. And because the architecture commits to what it has already predicted, a single out-of-distribution mistake can compound into looping or skipped segments, which is what DPO is there to correct: the negative supervision that pre-training and SFT cannot provide. The last third is the argument Tim keeps returning to. Customers running voice agents over millions of sessions do not describe a solved problem, they describe scaffolding, with a sharp quality drop outside the top few languages. Cascades survive because each component stays separately adaptable, observable and constrainable — fine-tune the ASR for your acoustics, log what compliance demands, keep it on your own hardware. The interface has its own limit: voice alone is cognitive debt, because absorbing information and deciding in one serial audio stream is much harder than glancing at a menu. Voice becomes ubiquitous beside a screen, not instead of one. --- --- REFERENCES: paper: [00:01:42] Mistral 7B https://arxiv.org/abs/2310.06825 [00:09:38] Voxtral https://arxiv.org/abs/2507.13264 [00:14:41] Robust Speech Recognition via Large-Scale Weak Supervision (Whisper) https://arxiv.org/abs/2212.04356 [00:19:11] Voxtral Realtime https://arxiv.org/abs/2602.11298 [00:21:52] Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling https://arxiv.org/abs/2509.08753 [00:30:52] Voxtral TTS https://arxiv.org/abs/2603.25551 [00:32:38] SoundStream: An End-to-End Neural Audio Codec https://arxiv.org/abs/2107.03312 [00:34:59] Flow Matching for Generative Modeling https://arxiv.org/abs/2210.02747 [00:37:03] High Fidelity Neural Audio Compression (EnCodec) https://arxiv.org/abs/2210.13438 [00:37:42] Moshi: a speech-text foundation model for real-time dialogue (Mimi) https://arxiv.org/abs/2410.00037 [00:39:05] Finite Scalar Quantization: VQ-VAE Made Simple https://arxiv.org/abs/2309.15505 [01:03:33] Direct Preference Optimization: Your Language Model is Secretly a Reward Model https://arxiv.org/abs/2305.18290 benchmark: [01:15:40] ElevenLabs v3 and Flash comparison in Voxtral TTS https://arxiv.org/abs/2603.25551 dataset: [00:46:14] Mozilla Common Voice datasets https://commonvoice.mozilla.org/en/datasets organization: [00:01:24] Mistral AI https://mistral.ai/ [00:50:47] Hugging Face https://huggingface.co/