Skip to content

Keep the pixel JaggedTensor on ctx in the sparse rasterizer backward - #344

Open
swahtz wants to merge 1 commit into
openvdb:mainfrom
swahtz:perf/sparse-raster-reuse-pixels
Open

swahtz wants to merge 1 commit into
openvdb:mainfrom
swahtz:perf/sparse-raster-reuse-pixels

Conversation

@swahtz

@swahtz swahtz commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

_RasterizeScreenSpaceGaussiansSparseFn saved the pixel selection's jdata, joffsets and jlidx, then rebuilt the JaggedTensor in backward with JaggedTensor.from_data_offsets_and_list_ids. This PR keeps the original JaggedTensor on ctx and reuses it in backward.

Why

openvdb/fvdb-core#812 (fixing openvdb/fvdb-core#722) makes fvdb-core's checked JaggedTensor constructors validate their structure. Validation reads the structure tensors on the host, so after that PR the rebuild would add one device-to-host sync to every sparse backward step.

Measured on an RTX PRO 6000, the checked ldim-1 constructor costs about 16–19 µs per call with an idle GPU, against 4–8 µs on current main. Its sync also stops the host from running ahead of queued GPU work.

The pixel JaggedTensor from forward is already valid, so backward doesn't need to rebuild it.

This change is preventive. With current fvdb-core the ldim-1 rebuild does not sync, so today it saves only a small jidx kernel and binding overhead. It only affects sparse rendering; dense training never goes through this path.

Notes

  • No new fvdb-core API. The change uses nothing new from core, so it works with current and future fvdb-core versions.
  • Holding it on ctx is safe. A tensor held as a ctx attribute skips saved-tensor hooks and the in-place modification check. Pixel coordinates are integer, need no gradient, and are not modified in place, so neither check matters here.

Test plan

  • Unit tests pass:
    pytest tests/unit/test_functional_gaussian_splatting.py
    pytest tests/unit/test_gaussian_splat_3d.py -k "sparse or pixels"
    
    Results: 23 passed, and 43 passed with 9 subtests.
  • Gradient parity against upstream/main with the same seed, on the sparse scene from test_sparse_matches_oo_and_backward_runs:
    • The max absolute difference was at most 3e-8 for means, quats and log_scales, which is atomicAdd ordering noise.
    • The difference was exactly 0 for logit_opacities, sh0 and shN.
  • black --check --target-version=py311 --line-length=120 passes.

🤖 Generated with Claude Code

The sparse screen-space rasterizer saved the pixel selection's jdata,
joffsets and jlidx and rebuilt the JaggedTensor in backward with
JaggedTensor.from_data_offsets_and_list_ids. fvdb-core's checked
constructors validate their structure (openvdb/fvdb-core#722), which
costs a device-to-host sync, so the rebuild would add one sync to every
sparse backward step. The JaggedTensor is already valid, so hold it on
ctx and reuse it. Integer pixel coordinates need no gradient and are not
modified in place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
@swahtz
swahtz requested a review from a team as a code owner October 8, 2026 21:40
@swahtz
swahtz requested review from blackencino and matthewdcong and removed request for a team October 8, 2026 21:40
@swahtz swahtz added the enhancement New feature or request label Oct 8, 2026
@swahtz swahtz self-assigned this Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant