Add VibeVoice-ASR backend

Third ASR backend option alongside faster-whisper and qwen3. Builds
microsoft/VibeVoice from source (pinned to a specific commit, since it's
custom modeling code not in transformers' Auto* registry) into its own venv,
inheriting the base image's torch/CUDA. Comes with built-in speaker
diarization (VibeVoiceASRProcessor.post_process_transcription returns
per-segment speaker ids directly, no separate pyannote pass needed).

Verified end-to-end against a real recording: 200 OK, correct Korean
transcription, speaker labels populated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
du5t
2026-07-23 16:54:34 +09:00
parent a32b0acb09
commit ce064d6894
7 changed files with 202 additions and 2 deletions

View File

@@ -19,6 +19,7 @@ DEFAULT_LANGUAGE = env_str("DEFAULT_LANGUAGE", "ko")
FASTER_WHISPER_URL = env_str("FASTER_WHISPER_URL", "http://127.0.0.1:8001")
QWEN3_URL = env_str("QWEN3_URL", "http://127.0.0.1:8004")
VIBEVOICE_URL = env_str("VIBEVOICE_URL", "http://127.0.0.1:8006")
PYANNOTE_HF_TOKEN = env_str("PYANNOTE_HF_TOKEN", "")