Key ideas
Scroll to read ↓Reliable speech recognition depends on matching encoders, decoders, speaker models, and customization tools to the audio and deployment constraints of each application.
- What must a speech system handle?0:32 ↗
- Choose the decoder for the workload2:42 ↗
- Shorten the sequence before decoding4:26 ↗
- Parakeet and Canary divide the work5:27 ↗
- Connect speaker timestamps to transcript tokens6:57 ↗
- Make recognized speech usable8:26 ↗
- Recognize lyrics over music9:50 ↗
- Build coverage into the training data10:53 ↗
- Keep training fed and validation broad12:01 ↗
- Serve streaming and offline workloads13:22 ↗
- Adapt the layer responsible for the error14:27 ↗
- Find a model, then deploy or fine-tune it15:32 ↗
