Repository navigation
Conversation
The raw constructors from_data_offsets_and_list_ids and from_data_indices_and_list_ids accepted structures that later crashed their readers: an empty list_ids with tensors present (openvdb#805), and decreasing, negative, or out-of-range offsets (openvdb#722). The list and lsizes constructors rejected empty input outright (openvdb#89). - Document the structural invariants on JaggedTensor. - Validate offsets, indices, and list ids in the public raw constructors. All value checks share one device-to-host readback. - Add an optional num_outer_lists argument (keyword-only in Python), since list ids cannot express trailing empty outer lists (openvdb#802). - Add from_data_offsets_and_list_ids_unsafe and from_data_indices_and_list_ids_unsafe for internal callers whose structure is correct by construction, and move the grid, mesh, and gsplat op outputs to them so op outputs pay no validation sync. - Accept JaggedTensor([]), nested lists whose outer lists are empty, and lsizes with empty lists. - jreshape passes the new structure's tensor count to the lsizes constructor instead of the old one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Readers derived the outer list structure from list ids, which have no row for an empty outer list. lshape and unbind dropped empty outer lists, and lshape read past the list ids when they had fewer rows than tensors (openvdb#802, openvdb#805). - recompute_lsizes_if_dirty sizes the ldim-2 result to num_outer_lists and buckets each tensor by its outer id, with a range check. - ldim() and check_valid() require exactly num_tensors list id rows for ldim 2. - set_jdata compares the new data with the current element count. Its offsets check assumed ldim 1 and rejected zero-element ldim-2 tensors. - jflatten(dim=1) builds the outer offsets from the outer ids, so it keeps empty outer lists and no longer needs jidx or a unique_dim sync. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Integer indexing of an empty outer list raised on CPU, and on CUDA it read an uninitialized buffer that its mask kernel never wrote (openvdb#802). CPU slicing with a step dropped empty outer lists and renumbered the rest, while CUDA kept them. - Replace the CPU mask path and the CUDA getJOffsetsIndexMask kernel with one searchsorted over the sorted outer ids. An empty outer list returns a zero-tensor JaggedTensor on the input's device. Dispatch on ldim() instead of on an empty jlidx. - Integer indexing of a list of tensors returns an empty jlidx and creates jidx on the input device. - CPU nested slices with a step number outer lists as CUDA does and keep empty ones. - ldim-1 slices use the unchecked constructor, so they add no sync. - Skip zero-size kernel launches in the CUDA slice and JaggedTensor index paths (part of openvdb#627). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
- jsum, jmin, and jmax at dim 0 took their single-tensor shortcut whenever jidx was empty, which is also the case for any JaggedTensor with no elements. Three empty tensors summed to one row, and jmin and jmax raised. Empty data now takes the scatter path, which already handles empty groups, and the index broadcast no longer reshapes an empty tensor with -1. - jcat at dim 0 skips the CUDA index kernel for inputs with no elements and initializes the CPU output offsets when there are no tensors. - jcat over lists of ldim-1 JaggedTensors returns empty list ids, since inputs may mix empty and explicit list ids. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Unpickling rebuilt the JaggedTensor from its offsets and list ids and ignored the stored outer list count, so trailing empty outer lists were lost. Its numOuterLists check also compared against the offsets length instead of the tensor count. Pass the stored count to the validating constructor, which checks it. Add TestEmptyStructures, which covers construction, validation, indexing, slicing, reductions, jcat, jflatten, set_jdata, and pickling of empty structures on CPU and CUDA (openvdb#89, openvdb#722, openvdb#802, openvdb#805). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
When no requested pixel had a contributing Gaussian, the contributor id and weight queries built their result with empty list ids and one tensor per pixel. lshape on that result read past the list ids and segfaulted, and the result had no camera nesting (openvdb#803). Remove the special empty path. The general path already builds offsets, list ids, and the camera count for every pixel, and now skips the copy when there is nothing to copy. Guard the zero-camera case against a division by zero and an index into empty camera offsets. Fixes openvdb#803. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
computeIndexPutArg wrote each element's tensor index at its position inside its input tensor instead of its position in the output, so the output jidx was wrong and partly uninitialized for ordinary inputs. computeTensorSizes ran one thread per offset instead of one per tensor, so the last thread read and wrote one element past the end of the offsets. Write jidx at the output position, launch one thread per tensor into zeroed offsets, and skip the launch when there are no tensors. Fixes openvdb#811. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
3 tasks done
Follow-up from review of this PR. - Store an empty jidx whenever there is at most one tensor. The nested list constructor built a full-length jidx for a single tensor, which it now allows among empty outer lists, and from_data_indices kept a caller's all-zero indices. Element-wise ops compare jidx shapes, so JaggedTensor([[], [t]]) + an identical tensor built from offsets or a pickle raised. The nested constructor now derives jidx from offsets, which also replaces a torch::full per tensor and a cat. - Move offsets, indices, and list ids to the data device instead of rejecting them. Host-built offsets with CUDA data worked before validation was added. - Stop checking ldim-1 list id values. Nothing reads them, and pickles of older slices hold un-rebased ids. - from_data_indices_and_list_ids computes offsets from the validated, sorted indices with searchsorted, so it syncs once instead of twice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
The nested lsizes constructor took the total tensor count as a parameter even though it is the sum of the inner list lengths. A wrong count overflowed its offset and list id buffers, which is how jreshape wrote out of bounds before this PR. Drop the parameter and count inside the constructor. jreshape and the jrand/jzeros/jones/jempty builders no longer compute it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
The attention output rebuilt its JaggedTensor with from_data_and_offsets, which now validates and syncs, and dropped the query's list ids and outer list count. The output has the query's structure by construction, so use query.jagged_like, which does neither. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
- jcat over lists builds shifted list ids only for ldim 2. For ldim 1 it discarded them. - Integer indexing checks bounds once before dispatching to the PrivateUse1 path, and an orphaned comment from the removed mask kernel is gone. - Rewrap the num_outer_lists parameter docs that clang-format split. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Validation in the checked raw constructors ran about two dozen small torch ops (diff, compare, searchsorted, stack, ...) before its single readback. That cost 95 us for ldim-1 offsets and 230 us for ldim-2 list ids on an RTX PRO 6000, against 4-8 us and 37 us on main. Add checkJaggedStructure, which checks offsets, indices, and list ids in one pass and returns a bitmask of violated invariants plus the last outer id. The CUDA path is one kernel and one device-to-host copy; the CPU path is a loop over accessors. Both share one per-position check. Per-call cost is now 29-32 us for ldim-1 offsets, 41-44 us for indices (139-184 us on main), and 32 us for ldim-2 list ids (37 us on main). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Integer indexing of an ldim-2 JaggedTensor ran a column select, a contiguous copy, a host-built search tensor, searchsorted, a gather, and a cat before its readback: 87 us against 50 us on main. Binary search the sorted outer ids directly, in a single-thread kernel on CUDA and a host loop on CPU. It now costs 38 us. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
The fused structure check still cleared its result buffer with a separate fill kernel and read it back with a pageable copy, and it validated host-built structure tensors on the GPU after copying them there. - Validate structure tensors on the host when they all start there, before moving them to the data device. That needs no kernel and no sync. - Check inputs of up to 65536 entries with one block, which combines its result in shared memory and writes it once, so no buffer needs clearing. - Write that result straight into pinned host memory the device can address, so the readback is just the stream sync. Larger inputs, or memory the device cannot address, use a cleared device buffer and an async copy instead. On an RTX PRO 6000 the ldim-1 offsets constructor drops from 28-32 us to 16-19 us (4-8 us on main). Host-built offsets drop from about 45 us to 19-25 us. from_data_and_indices drops to 29-33 us (139-185 us on main). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
1 of 3 tasks
swahtz
added a commit
that referenced
this pull request
Oct 9, 2026
## Summary `tests/unit/test_viz_shutdown.py::test_viewer_shutdown_with_gpu_buffers`, added in #810, needs a Vulkan device. On the GPU runners it fails with `Failed to create Vulkan instance. VkResult: -9` (`VK_ERROR_INCOMPATIBLE_DRIVER`). #810's unit-test jobs were skipped, so #812 is the first PR to hit it. Every PR based on current main will fail the same way. This PR gives the three GPU unit-test jobs a working NVIDIA Vulkan driver. ## Cause The runner AMI is fine. Its 580.65.06 driver includes the Vulkan driver, and `vulkaninfo` on the host lists the L4. The unit-test containers could not load that driver, for two reasons: 1. **The CUDA images set `NVIDIA_DRIVER_CAPABILITIES=compute,utility`.** The container toolkit therefore does not mount the Vulkan driver (`libGLX_nvidia` and `/etc/vulkan/icd.d/nvidia_icd.json`). 2. **The driver depends on libraries the slim images lack.** - `libGLX_nvidia` links against `libXext.so.6`. - It also opens `libEGL.so.1` at run time. - Without either one, the driver loads but returns no `vkCreateInstance`, and instance creation fails with -9. ## Change The fvdb Unit Tests jobs in `tests.yml` (conda), `cu130.yml` and `cu132.yml` now set up Vulkan in their containers: - **All three jobs:** set `NVIDIA_DRIVER_CAPABILITIES: all`. - **pip jobs:** install `libvulkan1 libxext6 libegl1`. - **conda job:** - install `libXext libglvnd-egl`; - set `VK_DRIVER_FILES=/etc/vulkan/icd.d/nvidia_icd.json`, in case the conda-provided Vulkan loader doesn't search `/etc/vulkan`. The gtest and docs jobs are unchanged. ## Test plan - [x] Diagnosed by hand on a runner instance from the CI AMI (Ubuntu 24.04, driver 580.65.06, L4, container toolkit 1.17.8): - The stock `nvidia/cuda:13.0.2-cudnn-runtime-ubuntu22.04` container has no NVIDIA Vulkan driver files. - With `NVIDIA_DRIVER_CAPABILITIES=all` but without `libxext6`, the loader skips the driver (`libXext.so.6: cannot open`). - With `libxext6` but without `libegl1`, the driver loads but returns no `vkCreateInstance`. - With `libvulkan1 libxext6 libegl1`, `vulkaninfo --summary` lists `NVIDIA L4`. - [ ] CI: these workflows run on `pull_request_target`, which uses `main`'s workflow definitions. This PR's own unit-test jobs will therefore still fail on the viewer test. The fix takes effect once it merges; the next PR, or a rerun of #812, should then pass `test_viz_shutdown`. - [ ] The conda job's `VK_DRIVER_FILES` setting is untested. If that job still fails after merge, its loader output will show whether the variable is needed or not enough. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Jonathan Swartz <jonathan@jswartz.info> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #89, fixes #722, fixes #802, fixes #803, fixes #805, fixes #811. Also covers the JaggedTensor launch sites of #627 (
JaggedTensorIndex.cu,JCat0.cu). The rest of #627 stays open.Summary
A JaggedTensor with zero tensors, zero elements, or an empty outer list now constructs correctly and works with every JaggedTensor operation. Before this PR, such tensors either could not be built or crashed their readers. The PR defines the structural invariants once, enforces them in the public raw constructors, and fixes the readers that assumed every outer list had a tensor.
Root causes
JaggedTensor([])raisestorch::caton an empty list.lshape,unbind,from_data_offsets_and_list_idsand CPU indexingmax + 1, or "the outer id changed") instead ofnum_outer_lists.lshapesegfaultslist_ids.numel() == 0bypassed the row-count check, and the reader then read past the list ids.jcat(dim=0)writes a wrongjidxBehavior changes
num_outer_listskeyword.from_data_offsets_and_list_idsandfrom_data_indices_and_list_idstake an optionalnum_outer_lists, because list ids cannot express trailing empty outer lists. Without it, the count is the largest outer id + 1.ValueErroron malformed structure. This applies tofrom_data_offsets_and_list_ids,from_data_indices_and_list_ids,from_data_and_offsetsandfrom_data_and_indices. Malformed structure includes:data.shape[0];int64list ids are cast to the canonical dtype, and host-built offsets move to the data's device.total_tensors. This changes the C++ signatureJaggedTensor(lsizes, total_tensors, data)toJaggedTensor(lsizes, data). The Python API is unaffected.JaggedTensor([])andJaggedTensor([[], []])have no tensor to take a device, dtype or element shape from, so they use these defaults. To get an empty structure on a specific device or dtype, usefrom_data_offsets_and_list_ids(..., num_outer_lists=N)or thefrom_*lsizes builders.Also fixed
Bugs found along the way that no open issue covered:
set_jdatarejected zero-element nested JaggedTensors, such as ray-op outputs with zero hits.jsum,jminandjmaxat dim 0 treated "no elements" as "one tensor": the sum had the wrong row count, and min/max raised.fvdb/attention.pydropped the query's list ids from its output.Performance
Per-call time on an RTX PRO 6000 with the GPU idle. "main" is a 2026-09-29 build; the code timed here has not changed on main since 2026-08-06. Lower is better.
from_data_and_offsets, ldim 1, offsets on GPU, B=1from_data_and_offsets, ldim 1, offsets on GPU, B=64from_data_and_offsets, ldim 1, offsets on host, B=1from_data_and_offsets, ldim 1, offsets on host, B=64from_data_and_indices, ldim 1, B=1 / B=64unique_dimfrom_data_and_indices, 200k elements, 4 tensorsfrom_data_offsets_and_list_ids, ldim 2, 64 × 64jflatten(dim=1), ldim 2, 64 × 64unique_dimjt[i]integer index, ldim 2, 64 × 64jagged_like(no-sync reference)from_data_and_offsetsandfrom_data_offsets_and_list_idswith ldim 1 and small offsets. Main did no sync there; validation needs one launch and one sync on the GPU, or a host-to-device copy when the offsets start on the host.jagged_likeskip validation through the internal unchecked constructors. Their code path is the same as main.from_data_offsets_and_list_idsevery step. Keep the pixel JaggedTensor on ctx in the sparse rasterizer backward fvdb-reality-capture#344 keeps the original onctxinstead. It doesn't depend on this PR, so it can merge first.CI note
The unit-test jobs fail only on
test_viz_shutdown.py::test_viewer_shutdown_with_gpu_buffers, withFailed to create Vulkan instance. VkResult: -9(VK_ERROR_INCOMPATIBLE_DRIVER). That test came in with #810, whose unit-test jobs were skipped. This is the first PR to run it on the GPU runners, which have no usable Vulkan driver. It is unrelated to this change, and main will hit it too. It needs a skip when Vulkan device creation fails, or a Vulkan driver on the runners.Test plan
pytest unit: 3119 passed, 10 skipped.TestEmptyStructuresand the updated tests intest_jagged_tensor.py, run on CPU and CUDA.ContributingGaussianIdsEmptyTestsintest_gaussian_splat_functional.pycovers the rasterize_contributing_gaussian_ids_sparse segfaults when no requested pixel has a contributor #803 repro cases plus multi-camera, non-uniform and dense all-background cases.JaggedTensorTest(including new cases),PackedJaggedAccessorTest,GaussianProjectionJaggedTest,GaussianComputeSparseInfoTest,GaussianComputeNanInfMaskTest,GaussianRasterizeForwardTestandGaussianRasterizeBackwardTest. All pass.🤖 Generated with Claude Code