What happens
If a model download is cut short, the partial file stays on disk and is treated as valid on every subsequent run. The feature that depends on it is silently degraded from then on, permanently, with no error the user will ever see.
Confirmed in the wild: a reporter's install has been running this way since setup. Their log shows both models failing 25 seconds into warmup, far too fast to be a download attempt.
WARMUP_FAILED beat_this ('Could not load the checkpoint given the provided name', 'final0')
WARMUP_FAILED vocal_split Unsupported Model File: parameters for MD5 hash 844a5bf2... could not be found
PytorchStreamReader failed reading zip archive: failed finding central directory
Why it never heals
Existence is the entire validity check. In torch.hub, which is what beat_this loads through:
if not os.path.exists(cached_file):
download_url_to_file(url, cached_file, hash_prefix, progress=progress)
and download_url_to_file reads until the body ends, then moves the result into place with no comparison against Content-Length. check_hash defaults to False and beat_this does not pass it. audio-separator has the same shape, which is why a bad file there surfaces as an MD5 matching nothing.
So a connection that drops mid-body and closes cleanly is promoted to a complete checkpoint.
Why nobody notices
app/pipeline/beat_detect.py caches the failure for the process and falls back to librosa. That fallback is indistinguishable, from the outside, from a machine that never had the model. The user gets a permanently worse beat grid and nothing anywhere connects the two.
Worse, both app/pipeline/warmup.py and desktop/ui/setup.js state that a failed model "falls back to lazy-download-on-first-use". That is true for an empty cache and false for a wrong one. The retry never fires, because the bad file is still there.
Constraints
check_hash=True is not ours to pass. beat_this calls into the hub itself, and the URL carries no hash.
- Pinning our own checksums means tracking upstream reissues forever.
- Vendoring a patched
beat_this is worse than the problem.
- Any retry has to be bounded, and must not fire when the cache was simply empty. A machine that is offline should not pay the same network timeout twice to reach the same answer.
Surfaced by the logs on #502.
What happens
If a model download is cut short, the partial file stays on disk and is treated as valid on every subsequent run. The feature that depends on it is silently degraded from then on, permanently, with no error the user will ever see.
Confirmed in the wild: a reporter's install has been running this way since setup. Their log shows both models failing 25 seconds into warmup, far too fast to be a download attempt.
Why it never heals
Existence is the entire validity check. In
torch.hub, which is whatbeat_thisloads through:and
download_url_to_filereads until the body ends, then moves the result into place with no comparison againstContent-Length.check_hashdefaults toFalseandbeat_thisdoes not pass it. audio-separator has the same shape, which is why a bad file there surfaces as an MD5 matching nothing.So a connection that drops mid-body and closes cleanly is promoted to a complete checkpoint.
Why nobody notices
app/pipeline/beat_detect.pycaches the failure for the process and falls back to librosa. That fallback is indistinguishable, from the outside, from a machine that never had the model. The user gets a permanently worse beat grid and nothing anywhere connects the two.Worse, both
app/pipeline/warmup.pyanddesktop/ui/setup.jsstate that a failed model "falls back to lazy-download-on-first-use". That is true for an empty cache and false for a wrong one. The retry never fires, because the bad file is still there.Constraints
check_hash=Trueis not ours to pass.beat_thiscalls into the hub itself, and the URL carries no hash.beat_thisis worse than the problem.Surfaced by the logs on #502.