Context
Child of #2118. #1863 added a substantial audio-tagging pipeline, but its only Fern change was to the speaker-separation page. The main guidance remains in tutorials/audio/tagging/README.md.
APIs to cover
PNCwithvLLMInferenceStage / CleanLLMOutputStage
- Reusable
VLLMInference lifecycle
- WER, CER, SQUIM, and bandwidth stages
PrepareModuleSegmentsStage
- ITN, Arabic diacritic removal, and Chinese conversion
- ASR and TTS pipeline configurations
Requirements
- Explain stage ordering, expected manifest schema, columns added/removed, and model/dependency setup.
- Document first-pass/second-pass PnC, CER validation, fallback behavior, and operational limits.
- Include an end-to-end runnable example and a configuration reference for both ASR and TTS paths.
- Cross-link speaker diarization/separation and audio quality filtering.
Acceptance criteria
Related PR: #1863
Supplemental code-freeze attribution
#1679 introduced the foundational audio tagging pipeline, including resampling, PyAnnote diarization, WhisperX VAD, NeMo ASR alignment, segment merging, YAML configuration, and the initial tutorial. Treat #1679 and #1863 as one implementation history and document the final API on main.
Related PRs: #1679, #1863
Context
Child of #2118. #1863 added a substantial audio-tagging pipeline, but its only Fern change was to the speaker-separation page. The main guidance remains in
tutorials/audio/tagging/README.md.APIs to cover
PNCwithvLLMInferenceStage/CleanLLMOutputStageVLLMInferencelifecyclePrepareModuleSegmentsStageRequirements
Acceptance criteria
Related PR: #1863
Supplemental code-freeze attribution
#1679 introduced the foundational audio tagging pipeline, including resampling, PyAnnote diarization, WhisperX VAD, NeMo ASR alignment, segment merging, YAML configuration, and the initial tutorial. Treat #1679 and #1863 as one implementation history and document the final API on
main.Related PRs: #1679, #1863