Skip to content

Support empty JaggedTensor structures and validate raw constructors - #812

Open
swahtz wants to merge 15 commits into
openvdb:mainfrom
swahtz:fix/jagged-empty-structures
Open

swahtz wants to merge 15 commits into
openvdb:mainfrom
swahtz:fix/jagged-empty-structures

Conversation

@swahtz

@swahtz swahtz commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #89, fixes #722, fixes #802, fixes #803, fixes #805, fixes #811. Also covers the JaggedTensor launch sites of #627 (JaggedTensorIndex.cu, JCat0.cu). The rest of #627 stays open.

Summary

A JaggedTensor with zero tensors, zero elements, or an empty outer list now constructs correctly and works with every JaggedTensor operation. Before this PR, such tensors either could not be built or crashed their readers. The PR defines the structural invariants once, enforces them in the public raw constructors, and fixes the readers that assumed every outer list had a tensor.

Root causes

Issue Symptom Cause
#89 JaggedTensor([]) raises The list constructors rejected empty input or called torch::cat on an empty list.
#802 Empty outer lists dropped by lshape, unbind, from_data_offsets_and_list_ids and CPU indexing Readers derived outer lists from list ids (max + 1, or "the outer id changed") instead of num_outer_lists.
#805 lshape segfaults list_ids.numel() == 0 bypassed the row-count check, and the reader then read past the list ids.
#803 Gsplat contributor ids segfault when no pixel has a contributor The op's empty path built the #805 structure.
#722 Malformed offsets cause out-of-bounds reads in later ops The raw constructors never checked offset values.
#811 CUDA jcat(dim=0) writes a wrong jidx The kernel wrote at the in-tensor position and overran the offsets by one.

Behavior changes

  • New num_outer_lists keyword. from_data_offsets_and_list_ids and from_data_indices_and_list_ids take an optional num_outer_lists, because list ids cannot express trailing empty outer lists. Without it, the count is the largest outer id + 1.
  • The public raw constructors now raise ValueError on malformed structure. This applies to from_data_offsets_and_list_ids, from_data_indices_and_list_ids, from_data_and_offsets and from_data_and_indices. Malformed structure includes:
    • offsets that don't start at 0, decrease, or don't end at data.shape[0];
    • indices that are unsorted or out of range;
    • list ids whose row count doesn't match the tensor count, or whose ids are unsorted or not numbered 0, 1, …
  • Integer structure tensors are cast and moved instead of rejected. For example, int64 list ids are cast to the canonical dtype, and host-built offsets move to the data's device.
  • The nested lsizes constructor no longer takes total_tensors. This changes the C++ signature JaggedTensor(lsizes, total_tensors, data) to JaggedTensor(lsizes, data). The Python API is unaffected.
  • An empty list holds an empty CPU float32 tensor. JaggedTensor([]) and JaggedTensor([[], []]) have no tensor to take a device, dtype or element shape from, so they use these defaults. To get an empty structure on a specific device or dtype, use from_data_offsets_and_list_ids(..., num_outer_lists=N) or the from_* lsizes builders.

Also fixed

Bugs found along the way that no open issue covered:

  • set_jdata rejected zero-element nested JaggedTensors, such as ray-op outputs with zero hits.
  • jsum, jmin and jmax at dim 0 treated "no elements" as "one tensor": the sum had the wrong row count, and min/max raised.
  • A pickle round-trip dropped trailing empty outer lists.
  • CPU slices with a step dropped empty outer lists and renumbered the rest; CUDA kept them.
  • fvdb/attention.py dropped the query's list ids from its output.

Performance

Per-call time on an RTX PRO 6000 with the GPU idle. "main" is a 2026-09-29 build; the code timed here has not changed on main since 2026-08-06. Lower is better.

Call main This PR Change
from_data_and_offsets, ldim 1, offsets on GPU, B=1 3.9 µs 15.7 µs ~12 µs slower: one validation launch and sync
from_data_and_offsets, ldim 1, offsets on GPU, B=64 8.0 µs 19.3 µs ~11 µs slower: same
from_data_and_offsets, ldim 1, offsets on host, B=1 3.3 µs 18.7 µs ~15 µs slower: the offsets are now copied to the data device (main left CPU offsets on CUDA data); validation itself runs on the host with no sync
from_data_and_offsets, ldim 1, offsets on host, B=64 88 µs 25 µs ~3.5× faster
from_data_and_indices, ldim 1, B=1 / B=64 139 / 185 µs 29 / 33 µs ~5× faster: the fused check replaces unique_dim
from_data_and_indices, 200k elements, 4 tensors 193 µs 36 µs ~5× faster
from_data_offsets_and_list_ids, ldim 2, 64 × 64 35 µs 27 µs slightly faster, while also validating every list id
jflatten(dim=1), ldim 2, 64 × 64 205 µs 40 µs ~5× faster: no unique_dim
jt[i] integer index, ldim 2, 64 × 64 50 µs 42 µs slightly faster: one binary-search kernel
jagged_like (no-sync reference) 2.1 µs 1.9 µs unchanged
  • What gets slower. Only from_data_and_offsets and from_data_offsets_and_list_ids with ldim 1 and small offsets. Main did no sync there; validation needs one launch and one sync on the GPU, or a host-to-device copy when the offsets start on the host.
  • What is unchanged. Op outputs (grid, mesh, gsplat) and jagged_like skip validation through the internal unchecked constructors. Their code path is the same as main.
  • Cost inside a busy training step. I also timed each call with ~2 ms of GPU work queued ahead of it. The run-to-run noise there was about ±50 µs, larger than these differences, so those numbers are not shown. Expect a GPU-side validation to add one sync's pipeline bubble when the host would otherwise run ahead.
  • Reality-capture. The sparse rasterizer's backward rebuilds its pixel JaggedTensor with from_data_offsets_and_list_ids every step. Keep the pixel JaggedTensor on ctx in the sparse rasterizer backward fvdb-reality-capture#344 keeps the original on ctx instead. It doesn't depend on this PR, so it can merge first.

CI note

The unit-test jobs fail only on test_viz_shutdown.py::test_viewer_shutdown_with_gpu_buffers, with Failed to create Vulkan instance. VkResult: -9 (VK_ERROR_INCOMPATIBLE_DRIVER). That test came in with #810, whose unit-test jobs were skipped. This is the first PR to run it on the GPU runners, which have no usable Vulkan driver. It is unrelated to this change, and main will hit it too. It needs a skip when Vulkan device creation fails, or a Vulkan driver on the runners.

Test plan

./build.sh install
cd tests && pytest unit -v
cd tests && pytest unit/test_jagged_tensor.py unit/test_gaussian_splat_functional.py -v
./build.sh install gtests   # then run JaggedTensorTest via ctest
  • pytest unit: 3119 passed, 10 skipped.
  • TestEmptyStructures and the updated tests in test_jagged_tensor.py, run on CPU and CUDA.
  • ContributingGaussianIdsEmptyTests in test_gaussian_splat_functional.py covers the rasterize_contributing_gaussian_ids_sparse segfaults when no requested pixel has a contributor #803 repro cases plus multi-camera, non-uniform and dense all-background cases.
  • C++ gtests that exercise JaggedTensor: JaggedTensorTest (including new cases), PackedJaggedAccessorTest, GaussianProjectionJaggedTest, GaussianComputeSparseInfoTest, GaussianComputeNanInfMaskTest, GaussianRasterizeForwardTest and GaussianRasterizeBackwardTest. All pass.
  • Each issue's reproducer, run in its own process, now behaves as the issue expected.
  • black and clang-format 18 are clean.
  • The full C++ gtest suite has not been run. Only the JaggedTensor-related targets above were built and run.
  • The intermediate commits have not been built one by one. Only the final tree has been built and tested.

🤖 Generated with Claude Code

swahtz and others added 7 commits October 9, 2026 10:52
The raw constructors from_data_offsets_and_list_ids and
from_data_indices_and_list_ids accepted structures that later crashed
their readers: an empty list_ids with tensors present (openvdb#805), and
decreasing, negative, or out-of-range offsets (openvdb#722). The list and
lsizes constructors rejected empty input outright (openvdb#89).

- Document the structural invariants on JaggedTensor.
- Validate offsets, indices, and list ids in the public raw
  constructors. All value checks share one device-to-host readback.
- Add an optional num_outer_lists argument (keyword-only in Python),
  since list ids cannot express trailing empty outer lists (openvdb#802).
- Add from_data_offsets_and_list_ids_unsafe and
  from_data_indices_and_list_ids_unsafe for internal callers whose
  structure is correct by construction, and move the grid, mesh, and
  gsplat op outputs to them so op outputs pay no validation sync.
- Accept JaggedTensor([]), nested lists whose outer lists are empty,
  and lsizes with empty lists.
- jreshape passes the new structure's tensor count to the lsizes
  constructor instead of the old one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Readers derived the outer list structure from list ids, which have no
row for an empty outer list. lshape and unbind dropped empty outer
lists, and lshape read past the list ids when they had fewer rows than
tensors (openvdb#802, openvdb#805).

- recompute_lsizes_if_dirty sizes the ldim-2 result to num_outer_lists
  and buckets each tensor by its outer id, with a range check.
- ldim() and check_valid() require exactly num_tensors list id rows for
  ldim 2.
- set_jdata compares the new data with the current element count. Its
  offsets check assumed ldim 1 and rejected zero-element ldim-2 tensors.
- jflatten(dim=1) builds the outer offsets from the outer ids, so it
  keeps empty outer lists and no longer needs jidx or a unique_dim sync.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Integer indexing of an empty outer list raised on CPU, and on CUDA it
read an uninitialized buffer that its mask kernel never wrote (openvdb#802).
CPU slicing with a step dropped empty outer lists and renumbered the
rest, while CUDA kept them.

- Replace the CPU mask path and the CUDA getJOffsetsIndexMask kernel
  with one searchsorted over the sorted outer ids. An empty outer list
  returns a zero-tensor JaggedTensor on the input's device. Dispatch on
  ldim() instead of on an empty jlidx.
- Integer indexing of a list of tensors returns an empty jlidx and
  creates jidx on the input device.
- CPU nested slices with a step number outer lists as CUDA does and
  keep empty ones.
- ldim-1 slices use the unchecked constructor, so they add no sync.
- Skip zero-size kernel launches in the CUDA slice and JaggedTensor
  index paths (part of openvdb#627).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
- jsum, jmin, and jmax at dim 0 took their single-tensor shortcut
  whenever jidx was empty, which is also the case for any JaggedTensor
  with no elements. Three empty tensors summed to one row, and jmin and
  jmax raised. Empty data now takes the scatter path, which already
  handles empty groups, and the index broadcast no longer reshapes an
  empty tensor with -1.
- jcat at dim 0 skips the CUDA index kernel for inputs with no elements
  and initializes the CPU output offsets when there are no tensors.
- jcat over lists of ldim-1 JaggedTensors returns empty list ids, since
  inputs may mix empty and explicit list ids.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Unpickling rebuilt the JaggedTensor from its offsets and list ids and
ignored the stored outer list count, so trailing empty outer lists were
lost. Its numOuterLists check also compared against the offsets length
instead of the tensor count. Pass the stored count to the validating
constructor, which checks it.

Add TestEmptyStructures, which covers construction, validation,
indexing, slicing, reductions, jcat, jflatten, set_jdata, and pickling
of empty structures on CPU and CUDA (openvdb#89, openvdb#722, openvdb#802, openvdb#805).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
When no requested pixel had a contributing Gaussian, the contributor
id and weight queries built their result with empty list ids and one
tensor per pixel. lshape on that result read past the list ids and
segfaulted, and the result had no camera nesting (openvdb#803).

Remove the special empty path. The general path already builds offsets,
list ids, and the camera count for every pixel, and now skips the copy
when there is nothing to copy. Guard the zero-camera case against a
division by zero and an index into empty camera offsets.

Fixes openvdb#803.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
computeIndexPutArg wrote each element's tensor index at its position
inside its input tensor instead of its position in the output, so the
output jidx was wrong and partly uninitialized for ordinary inputs.
computeTensorSizes ran one thread per offset instead of one per tensor,
so the last thread read and wrote one element past the end of the
offsets. Write jidx at the output position, launch one thread per
tensor into zeroed offsets, and skip the launch when there are no
tensors.

Fixes openvdb#811.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
@swahtz
swahtz requested a review from a team as a code owner October 8, 2026 21:54
@swahtz swahtz added bug Something isn't working core library Core fVDB library. i.e. anything in the _Cpp module (C++) or fvdb python module Python API labels Oct 8, 2026
@swahtz swahtz self-assigned this Oct 8, 2026
swahtz and others added 7 commits October 9, 2026 13:05
Follow-up from review of this PR.

- Store an empty jidx whenever there is at most one tensor. The nested
  list constructor built a full-length jidx for a single tensor, which
  it now allows among empty outer lists, and from_data_indices kept a
  caller's all-zero indices. Element-wise ops compare jidx shapes, so
  JaggedTensor([[], [t]]) + an identical tensor built from offsets or a
  pickle raised. The nested constructor now derives jidx from offsets,
  which also replaces a torch::full per tensor and a cat.
- Move offsets, indices, and list ids to the data device instead of
  rejecting them. Host-built offsets with CUDA data worked before
  validation was added.
- Stop checking ldim-1 list id values. Nothing reads them, and pickles
  of older slices hold un-rebased ids.
- from_data_indices_and_list_ids computes offsets from the validated,
  sorted indices with searchsorted, so it syncs once instead of twice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
The nested lsizes constructor took the total tensor count as a
parameter even though it is the sum of the inner list lengths. A wrong
count overflowed its offset and list id buffers, which is how jreshape
wrote out of bounds before this PR. Drop the parameter and count
inside the constructor. jreshape and the jrand/jzeros/jones/jempty
builders no longer compute it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
The attention output rebuilt its JaggedTensor with from_data_and_offsets,
which now validates and syncs, and dropped the query's list ids and
outer list count. The output has the query's structure by construction,
so use query.jagged_like, which does neither.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
- jcat over lists builds shifted list ids only for ldim 2. For ldim 1
  it discarded them.
- Integer indexing checks bounds once before dispatching to the
  PrivateUse1 path, and an orphaned comment from the removed mask
  kernel is gone.
- Rewrap the num_outer_lists parameter docs that clang-format split.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Validation in the checked raw constructors ran about two dozen small
torch ops (diff, compare, searchsorted, stack, ...) before its single
readback. That cost 95 us for ldim-1 offsets and 230 us for ldim-2 list
ids on an RTX PRO 6000, against 4-8 us and 37 us on main.

Add checkJaggedStructure, which checks offsets, indices, and list ids in
one pass and returns a bitmask of violated invariants plus the last
outer id. The CUDA path is one kernel and one device-to-host copy; the
CPU path is a loop over accessors. Both share one per-position check.

Per-call cost is now 29-32 us for ldim-1 offsets, 41-44 us for indices
(139-184 us on main), and 32 us for ldim-2 list ids (37 us on main).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Integer indexing of an ldim-2 JaggedTensor ran a column select, a
contiguous copy, a host-built search tensor, searchsorted, a gather,
and a cat before its readback: 87 us against 50 us on main. Binary
search the sorted outer ids directly, in a single-thread kernel on CUDA
and a host loop on CPU. It now costs 38 us.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
The fused structure check still cleared its result buffer with a
separate fill kernel and read it back with a pageable copy, and it
validated host-built structure tensors on the GPU after copying them
there.

- Validate structure tensors on the host when they all start there,
  before moving them to the data device. That needs no kernel and no
  sync.
- Check inputs of up to 65536 entries with one block, which combines
  its result in shared memory and writes it once, so no buffer needs
  clearing.
- Write that result straight into pinned host memory the device can
  address, so the readback is just the stream sync. Larger inputs, or
  memory the device cannot address, use a cleared device buffer and an
  async copy instead.

On an RTX PRO 6000 the ldim-1 offsets constructor drops from 28-32 us
to 16-19 us (4-8 us on main). Host-built offsets drop from about 45 us
to 19-25 us. from_data_and_indices drops to 29-33 us (139-185 us on
main).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
swahtz added a commit that referenced this pull request Oct 9, 2026
## Summary


`tests/unit/test_viz_shutdown.py::test_viewer_shutdown_with_gpu_buffers`,
added in #810, needs a Vulkan device. On the GPU runners it fails with
`Failed to create Vulkan instance. VkResult: -9`
(`VK_ERROR_INCOMPATIBLE_DRIVER`). #810's unit-test jobs were skipped, so
#812 is the first PR to hit it. Every PR based on current main will fail
the same way.

This PR gives the three GPU unit-test jobs a working NVIDIA Vulkan
driver.

## Cause

The runner AMI is fine. Its 580.65.06 driver includes the Vulkan driver,
and `vulkaninfo` on the host lists the L4. The unit-test containers
could not load that driver, for two reasons:

1. **The CUDA images set `NVIDIA_DRIVER_CAPABILITIES=compute,utility`.**
The container toolkit therefore does not mount the Vulkan driver
(`libGLX_nvidia` and `/etc/vulkan/icd.d/nvidia_icd.json`).
2. **The driver depends on libraries the slim images lack.**
   - `libGLX_nvidia` links against `libXext.so.6`.
   - It also opens `libEGL.so.1` at run time.
- Without either one, the driver loads but returns no
`vkCreateInstance`, and instance creation fails with -9.

## Change

The fvdb Unit Tests jobs in `tests.yml` (conda), `cu130.yml` and
`cu132.yml` now set up Vulkan in their containers:
- **All three jobs:** set `NVIDIA_DRIVER_CAPABILITIES: all`.
- **pip jobs:** install `libvulkan1 libxext6 libegl1`.
- **conda job:**
  - install `libXext libglvnd-egl`;
- set `VK_DRIVER_FILES=/etc/vulkan/icd.d/nvidia_icd.json`, in case the
conda-provided Vulkan loader doesn't search `/etc/vulkan`.

The gtest and docs jobs are unchanged.

## Test plan

- [x] Diagnosed by hand on a runner instance from the CI AMI (Ubuntu
24.04, driver 580.65.06, L4, container toolkit 1.17.8):
- The stock `nvidia/cuda:13.0.2-cudnn-runtime-ubuntu22.04` container has
no NVIDIA Vulkan driver files.
- With `NVIDIA_DRIVER_CAPABILITIES=all` but without `libxext6`, the
loader skips the driver (`libXext.so.6: cannot open`).
- With `libxext6` but without `libegl1`, the driver loads but returns no
`vkCreateInstance`.
- With `libvulkan1 libxext6 libegl1`, `vulkaninfo --summary` lists
`NVIDIA L4`.
- [ ] CI: these workflows run on `pull_request_target`, which uses
`main`'s workflow definitions. This PR's own unit-test jobs will
therefore still fail on the viewer test. The fix takes effect once it
merges; the next PR, or a rerun of #812, should then pass
`test_viz_shutdown`.
- [ ] The conda job's `VK_DRIVER_FILES` setting is untested. If that job
still fails after merge, its loader output will show whether the
variable is needed or not enough.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working core library Core fVDB library. i.e. anything in the _Cpp module (C++) or fvdb python module Python API

Projects

None yet

1 participant