Self-restart worker process when model loading fails
Discovered vibevoice sitting at 3.2GB resident GPU memory with loaded_models: [] and no idle-unload ever firing for it again. Root cause: a load attempt had OOM'd partway through (competing with an unrelated ollama process on the same GPU), so _MODEL_CACHE was never populated — the idle-unload loop only clears that cache, so it had nothing to act on, even though the partially-constructed model had already left memory allocated. gc.collect()+empty_cache() don't reliably reclaim memory from an interrupted from_pretrained() call. Confirmed a plain process restart does fully reclaim it, so each worker now treats any load failure as fatal: log it and os._exit(1), letting supervisord's autorestart=true respawn a clean process immediately. Verified with a bogus model id — worker exits, respawns, and passes the smoke test right after. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
@@ -109,15 +109,25 @@ def _load_model(model_id: str) -> Tuple[Any, Any]:
|
||||
_log(f"loading model {model_id}")
|
||||
dtype = torch.float16 if DEVICE == "cuda" else torch.float32
|
||||
|
||||
processor = AutoProcessor.from_pretrained(model_id, cache_dir=str(MODEL_CACHE))
|
||||
model = AutoModelForSpeechSeq2Seq.from_pretrained(
|
||||
model_id,
|
||||
dtype=dtype,
|
||||
low_cpu_mem_usage=True,
|
||||
cache_dir=str(MODEL_CACHE),
|
||||
)
|
||||
model.to(DEVICE)
|
||||
model.eval()
|
||||
try:
|
||||
processor = AutoProcessor.from_pretrained(model_id, cache_dir=str(MODEL_CACHE))
|
||||
model = AutoModelForSpeechSeq2Seq.from_pretrained(
|
||||
model_id,
|
||||
dtype=dtype,
|
||||
low_cpu_mem_usage=True,
|
||||
cache_dir=str(MODEL_CACHE),
|
||||
)
|
||||
model.to(DEVICE)
|
||||
model.eval()
|
||||
except Exception as e:
|
||||
# A load interrupted partway (e.g. OOM) can leave CUDA memory
|
||||
# fragmented/leaked in ways gc.collect()+empty_cache() don't reliably
|
||||
# reclaim, and since _MODEL_CACHE never got populated the idle-unload
|
||||
# loop has nothing to clean up either. Restarting the whole process
|
||||
# is the only guaranteed way to get that memory back — supervisord's
|
||||
# autorestart=true respawns it immediately.
|
||||
_log(f"model load failed, restarting process to reclaim GPU memory: {type(e).__name__}: {e}")
|
||||
os._exit(1)
|
||||
|
||||
_MODEL_CACHE[model_id] = (model, processor)
|
||||
_log(f"model {model_id} loaded")
|
||||
|
||||
Reference in New Issue
Block a user