← All speakers

Bio, Work & Ideas

Jaehan

Conference affiliation: NVIDIA · 2025

Jaehan was historically affiliated with NVIDIA at a recorded 2025 event. The session examined streaming speech recognition, speaker diarization techniques, and GPU-efficient audio processing.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Reliable speech recognition depends on matching encoders, decoders, speaker models, and customization tools to the audio and deployment constraints of each application.

  • What must a speech system handle?
    0:32 ↗
  • Choose the decoder for the workload
    2:42 ↗
  • Shorten the sequence before decoding
    4:26 ↗
  • Parakeet and Canary divide the work
    5:27 ↗
  • Connect speaker timestamps to transcript tokens
    6:57 ↗
  • Make recognized speech usable
    8:26 ↗
  • Recognize lyrics over music
    9:50 ↗
  • Build coverage into the training data
    10:53 ↗
  • Keep training fed and validation broad
    12:01 ↗
  • Serve streaming and offline workloads
    13:22 ↗
  • Adapt the layer responsible for the error
    14:27 ↗
  • Find a model, then deploy or fine-tune it
    15:32 ↗

References