Third ASR backend option alongside faster-whisper and qwen3. Builds
microsoft/VibeVoice from source (pinned to a specific commit, since it's
custom modeling code not in transformers' Auto* registry) into its own venv,
inheriting the base image's torch/CUDA. Comes with built-in speaker
diarization (VibeVoiceASRProcessor.post_process_transcription returns
per-segment speaker ids directly, no separate pyannote pass needed).
Verified end-to-end against a real recording: 200 OK, correct Korean
transcription, speaker labels populated.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>