Skip to content

Install a torchvision that matches the torch it ships beside - #615

Merged
thcp merged 1 commit into
mainfrom
fix/torchvision-pair
Sep 9, 2026
Merged

thcp merged 1 commit into
mainfrom
fix/torchvision-pair

Conversation

@thcp

@thcp thcp commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator

Closes #614
Reported in #502.

The failure

vocal split failed: operator torchvision::nms does not exist

Every attempt, permanently, after a cu128 install. Separation kept working, which is what made it look like a karaoke bug rather than a broken install.

Why

torchvision ships compiled ops registered against one torch ABI. Next to a different torch the registration silently does not happen, and the failure surfaces far from its cause.

  1. The lockfile resolves torch 2.6.0 with torchvision 0.21.0, a matched pair.
  2. The CUDA passes install torch and torchaudio with --no-deps, so pip touches only what is named.
  3. torchvision was not named, so it stayed at 0.21.0.
  4. On the cu128 line, and only there, torch moves to 2.8.0.
  5. audio-separator pulls onnx2torch-py313, which needs torchvision. The karaoke split calls into it and finds no nms registered.

Demucs never touches torchvision, hence separation being fine.

Worth saying plainly

This only bit installs landing on cu128, which used to mean Blackwell alone. #602 widened it by selecting cu128 for any CUDA 13 driver, which is correct for the driver and is exactly what put the reporter on that line. The change that fixed their install is what exposed this one.

The fix

Both install paths now name a torchvision pinned to their torch line, and the mapping sits directly beside torch_version_for_tag because the two have to agree.

--no-deps stays. It is what makes the swap fast and stops pip rebuilding the world; naming the package is the fix, not resolving dependencies.

restore_cpu_torch gets the same treatment. It happened to heal the mismatch already, since putting torch back to 2.6.0 re-pairs it with 0.21.0, but that was luck rather than intent.

Versions checked against the index, not remembered: torchvision-0.23.0+cu128-cp312 and torchvision-0.21.0+cpu-cp312 both exist.

Tests

One pins the pairing and that every non-cu128 tag stays on the locked pair. One asserts every CUDA pass names torchvision at all, which is the actual bug: --no-deps means anything unnamed is simply not touched. Verified by dropping torchvision from the passes, at which point the second fails.

Run on both targets, since cuda_wheel_needs_runtime_deps is Linux-gated: 91 on Windows, 100 under WSL.

What this does not cover

I cannot reproduce the original failure locally: it needs a real cu128 install on a CUDA 13 machine. The reasoning is from the lockfile, the install args and the reporter's traceback, and the pinned versions are verified against the index. Worth a confirmation from @raniaamina once it ships.

The karaoke split failed on every attempt after a cu128 install:

  vocal split failed: operator torchvision::nms does not exist

torchvision ships compiled ops registered against one torch ABI. Put it next
to a different torch and the registration silently does not happen, so the
failure arrives far away from its cause, from whatever first calls one of
those ops.

The lockfile resolves torch 2.6.0 with torchvision 0.21.0, a matched pair. The
CUDA passes install torch and torchaudio with --no-deps, so pip touches only
what is named, and torchvision was not named. On the cu128 line, and only
there, torch moves to 2.8.0 and leaves torchvision behind at 0.21.0.
audio-separator pulls onnx2torch, which needs torchvision, so the karaoke
split is where it shows up. Demucs never touches torchvision, which is why
separation kept working and this looked like a karaoke bug rather than a
broken install.

Both install paths now name a torchvision pinned to their torch line, and the
pairing sits beside torch_version_for_tag because the two have to agree.
--no-deps stays: it is what makes the swap fast, and naming the package is the
fix rather than resolving dependencies.

restore_cpu_torch gets the same treatment. It happened to heal the mismatch
already, since putting torch back to 2.6.0 re-pairs it with 0.21.0, but that
was luck and not intent.

Versions checked against the index rather than remembered:
torchvision 0.23.0+cu128 and 0.21.0+cpu both exist for cp312.

Worth recording plainly: this only ever bit installs that land on cu128, which
used to mean Blackwell alone. #602 widened it by selecting cu128 for any CUDA
13 driver, which is right for the driver and is what put the reporter on that
line. The change that fixed their install is what exposed this.

Two tests. One pins the pairing and that every other tag stays on the locked
pair; one asserts every CUDA pass names torchvision at all, which is the
actual bug, since --no-deps means anything unnamed is simply not touched.
Verified by dropping torchvision from the passes: the second fails.

Checked on both targets, since cuda_wheel_needs_runtime_deps is Linux-gated:
91 tests on Windows, 100 under WSL.

Closes #614
@thcp thcp mentioned this pull request Sep 9, 2026
2 tasks done
@thcp
thcp merged commit 7d60953 into main Sep 9, 2026
12 of 13 checks passed
@thcp
thcp deleted the fix/torchvision-pair branch September 9, 2026 21:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Karaoke split dies with "operator torchvision::nms does not exist" after a cu128 install

1 participant