du5t ce064d6894 Add VibeVoice-ASR backend
Third ASR backend option alongside faster-whisper and qwen3. Builds
microsoft/VibeVoice from source (pinned to a specific commit, since it's
custom modeling code not in transformers' Auto* registry) into its own venv,
inheriting the base image's torch/CUDA. Comes with built-in speaker
diarization (VibeVoiceASRProcessor.post_process_transcription returns
per-segment speaker ids directly, no separate pyannote pass needed).

Verified end-to-end against a real recording: 200 OK, correct Korean
transcription, speaker labels populated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 16:54:34 +09:00
2026-07-23 16:54:34 +09:00
2026-06-18 15:25:56 +09:00
2026-07-23 16:54:34 +09:00
Description
ASR v2 — faster-whisper + Qwen3-ASR, multi-venv, GPU
217 KiB
Languages
Python 49.1%
JavaScript 26.2%
HTML 14.7%
CSS 5.9%
Shell 2.2%
Other 1.9%