Install a torchvision that matches the torch it ships beside - #615
Merged
Merged
Conversation
The karaoke split failed on every attempt after a cu128 install: vocal split failed: operator torchvision::nms does not exist torchvision ships compiled ops registered against one torch ABI. Put it next to a different torch and the registration silently does not happen, so the failure arrives far away from its cause, from whatever first calls one of those ops. The lockfile resolves torch 2.6.0 with torchvision 0.21.0, a matched pair. The CUDA passes install torch and torchaudio with --no-deps, so pip touches only what is named, and torchvision was not named. On the cu128 line, and only there, torch moves to 2.8.0 and leaves torchvision behind at 0.21.0. audio-separator pulls onnx2torch, which needs torchvision, so the karaoke split is where it shows up. Demucs never touches torchvision, which is why separation kept working and this looked like a karaoke bug rather than a broken install. Both install paths now name a torchvision pinned to their torch line, and the pairing sits beside torch_version_for_tag because the two have to agree. --no-deps stays: it is what makes the swap fast, and naming the package is the fix rather than resolving dependencies. restore_cpu_torch gets the same treatment. It happened to heal the mismatch already, since putting torch back to 2.6.0 re-pairs it with 0.21.0, but that was luck and not intent. Versions checked against the index rather than remembered: torchvision 0.23.0+cu128 and 0.21.0+cpu both exist for cp312. Worth recording plainly: this only ever bit installs that land on cu128, which used to mean Blackwell alone. #602 widened it by selecting cu128 for any CUDA 13 driver, which is right for the driver and is what put the reporter on that line. The change that fixed their install is what exposed this. Two tests. One pins the pairing and that every other tag stays on the locked pair; one asserts every CUDA pass names torchvision at all, which is the actual bug, since --no-deps means anything unnamed is simply not touched. Verified by dropping torchvision from the passes: the second fails. Checked on both targets, since cuda_wheel_needs_runtime_deps is Linux-gated: 91 tests on Windows, 100 under WSL. Closes #614
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #614
Reported in #502.
The failure
Every attempt, permanently, after a cu128 install. Separation kept working, which is what made it look like a karaoke bug rather than a broken install.
Why
torchvision ships compiled ops registered against one torch ABI. Next to a different torch the registration silently does not happen, and the failure surfaces far from its cause.
--no-deps, so pip touches only what is named.onnx2torch-py313, which needs torchvision. The karaoke split calls into it and finds nonmsregistered.Demucs never touches torchvision, hence separation being fine.
Worth saying plainly
This only bit installs landing on cu128, which used to mean Blackwell alone. #602 widened it by selecting cu128 for any CUDA 13 driver, which is correct for the driver and is exactly what put the reporter on that line. The change that fixed their install is what exposed this one.
The fix
Both install paths now name a torchvision pinned to their torch line, and the mapping sits directly beside
torch_version_for_tagbecause the two have to agree.--no-depsstays. It is what makes the swap fast and stops pip rebuilding the world; naming the package is the fix, not resolving dependencies.restore_cpu_torchgets the same treatment. It happened to heal the mismatch already, since putting torch back to 2.6.0 re-pairs it with 0.21.0, but that was luck rather than intent.Versions checked against the index, not remembered:
torchvision-0.23.0+cu128-cp312andtorchvision-0.21.0+cpu-cp312both exist.Tests
One pins the pairing and that every non-cu128 tag stays on the locked pair. One asserts every CUDA pass names torchvision at all, which is the actual bug:
--no-depsmeans anything unnamed is simply not touched. Verified by dropping torchvision from the passes, at which point the second fails.Run on both targets, since
cuda_wheel_needs_runtime_depsis Linux-gated: 91 on Windows, 100 under WSL.What this does not cover
I cannot reproduce the original failure locally: it needs a real cu128 install on a CUDA 13 machine. The reasoning is from the lockfile, the install args and the reporter's traceback, and the pinned versions are verified against the index. Worth a confirmation from @raniaamina once it ships.