Skip to content

Composable differentiable Gaussian splatting API over fvdb.functional - #338

Merged
swahtz merged 33 commits into
openvdb:mainfrom
swahtz:feature/functional-gaussian-splatting
Oct 8, 2026
Merged

swahtz merged 33 commits into
openvdb:mainfrom
swahtz:feature/functional-gaussian-splatting

Conversation

@swahtz

@swahtz swahtz commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Builds the composable, differentiable Gaussian splatting pipeline on top of the flat kernel surface fvdb-core publishes in fvdb.functional (openvdb/fvdb-core#797, openvdb/fvdb-core#799), and makes GaussianSplat3d a wrapper that composes it. This repo no longer imports fvdb._fvdb_cpp or unwraps JaggedTensor._impl.

Fixes #337. Completes the fvdb-reality-capture half of openvdb/fvdb-core#590.

Requires a fvdb-core build containing openvdb/fvdb-core#799. The pin is raised to fvdb-core>=0.7.0dev0; CI fails at collection (cannot import name 'CameraModel' from 'fvdb') until that PR is merged into fvdb-core main, which the test workflow clones.

Scope. The port, the stage contracts, and the bugs the port exposed at those boundaries.

Changes

Enums. CameraModel and RollingShutterType are re-exported from fvdb as the same objects; the local copies are gone. ProjectionMethod stays owned here, with the same members and values, because no fvdb kernel takes it: resolve_projection_method picks between fvdb's analytic and unscented functions. GaussianRenderMode (FEATURES, DEPTH, FEATURES_AND_DEPTH) is new and owned here.

fvdb_reality_capture.functional. Four stages passing frozen dataclasses, plus analysis:

Stage Functions Types
1. Projection project_gaussians, resolve_projection_method, requires_distortion_coeffs, check_distortion_coeffs ProjectedGaussians
2. Features evaluate_gaussian_sh, sh_degree_from_coefficients, compute_gaussian_opacities
3. Tile intersection intersect_gaussian_tiles, intersect_gaussian_tiles_sparse, deduplicate_pixels, as_pixel_jagged (re-exported from fvdb), check_tiles_match GaussianTileIntersection, SparseGaussianTileIntersection
4. Rasterization rasterize_screen_space_gaussians, rasterize_world_space_gaussians, rasterize_screen_space_gaussians_sparse, pixel_mask_to_tile_mask, validate_crop, apply_crop, apply_pixel_mask, pad_crop Crop
Analysis rasterize_num_contributing_gaussians, rasterize_contributing_gaussian_ids and their _sparse variants

The six torch.autograd.Function classes moved from radiance_fields/_gaussian_autograd.py to functional/_autograd.py and call fvdb.functional. Stages after projection take per-camera opacities ([C, N], from compute_gaussian_opacities) rather than logits, which is the contract the kernel boundary had before. Analysis functions keep the gsplat-aligned names decided in #337.

GaussianSplat3d. Every render, projection, analysis, MCMC and PLY method is a few lines over the stage functions; the file goes from 4928 to 4062 lines with all docstrings kept. Existing signatures are unchanged; the additions are keyword arguments: render_from_projected_gaussians(tiles=..., tile_size=None) for repeated crops from one projection, and crop/crop_masks on the three *_from_world methods. ProjectedGaussianSplats exposes its ProjectedGaussians as projected_gaussians, so a projection made through the class can continue in the functional pipeline; it takes its opacities at construction, alongside the features, so the projection is a consistent snapshot of the model, and tile_intersection(tile_size) computes on demand and caches nothing.

Docs and tests. New docs/api/functional.rst; the enums page points at fvdb through intersphinx for the shared enums and documents ProjectionMethod and GaussianRenderMode locally. tests/unit/test_functional_gaussian_splatting.py checks each stage's contract and the pipeline against GaussianSplat3d (forward, backward, depth modes, world space through the unscented projection, sparse with duplicates, crops, analysis, empty selections, deduplication order, pixel validation, opacity snapshot). The API ownership test asserts CameraModel and RollingShutterType are shared with fvdb and that ProjectionMethod and GaussianRenderMode are not in fvdb.

Behaviour changes

Each of these differs from main; everything else renders the same pixels.

  • Crops render correctly. main sized the tile grid to the crop but indexed tile offsets by absolute position, so any crop not at the origin read out of bounds (Dense rasterizers cannot render a crop into crop-sized buffers: window origin and tile-grid contract are inconsistent fvdb-core#800). Crops now mask tiles outside the window and slice the full-size render, and the result equals that region of the uncropped image. The edge contract is chosen rather than preserved, since main had none that worked: render_from_projected_gaussians returns the requested size with the outside as background at zero alpha (a crop entirely outside is all background), the stage functions return the clipped part (empty for a crop entirely outside), a zero crop size raises, and the outputs stay differentiable. Raster buffers are still allocated at full image size (fvdb-core#800), so a crop does not save that memory.
  • Masks have one coordinate frame per argument. Stage functions take masks in image coordinates and, with a crop, crop_masks in crop coordinates (requested or clipped size), never both; render_from_projected_gaussians takes a crop-space mask. Any other shape, or a mask on another device, raises instead of being read from the top-left corner. The sparse rasterizer takes per-tile tile_masks, since its pixels are explicit, and checks their shape and device; GaussianSplat3d.sparse_render_* keep their per-tile masks argument.
  • Pixel deduplication. deduplicate_pixels returns unique pixels in first-requested order, as its docstring said, where main returned sorted row-major order, and it never merges an out-of-image pixel with a valid one: on main, (0, W) shared a key with (1, 0). Sparse stage functions return results in requested-pixel order, duplicates included, as the class methods already did.
  • Plain-tensor sparse renders. A [C, P, 2] tensor of pixels returns stacked [C, P, D] tensors, as the docstrings said; main returned the flat [C * P, D] data. A [P, 2] tensor is rejected (as_pixel_jagged, shared with fvdb's own sparse wrappers, wants [C, P, 2]), where main accepted it for one camera. In-repo callers pass jagged input and are unaffected.
  • Validation that used to be a kernel fault or silent garbage. Tile intersections are checked against the projection they are rendered with (image size, camera count, and an identity token the projection carries, so tiles from before an optimizer step or refinement are rejected; dataclasses.replace keeps the token). Opacities must be on the projection's device and are made contiguous. Distortion coefficients must be a contiguous [C, 12] on the right device for distortion models and are ignored for pinhole and orthographic cameras. Pixel selections reject non-integer coordinates instead of truncating them. Precomputed tiles set the tile size; an explicit tile_size that disagrees raises.
  • Depth channel. evaluate_gaussian_sh computes the depth channel from the Gaussian centers rather than reading the projection's, so it is the same function of its inputs under either projection method and is differentiable through the unscented projection; it equals the projection's depth, zero for culled Gaussians included.
  • Sparse contributor lists keep a trailing camera that requests no pixels when duplicates trigger the expansion (JaggedTensor.from_data_offsets_and_list_ids would drop it, JaggedTensor: empty outer lists are dropped by from_data_offsets_and_list_ids, hidden from lshape/unbind, and fail CPU indexing fvdb-core#802). A request where no pixel anywhere has a contributor segfaults inside fvdb (rasterize_contributing_gaussian_ids_sparse segfaults when no requested pixel has a contributor fvdb-core#803); pre-existing, not worked around.

Test plan

cd tests && pytest -v unit/test_functional_gaussian_splatting.py unit/test_gaussian_splat_3d.py unit/test_gaussian_splat_api.py unit/test_rasterize_from_world.py
cd tests && pytest -v unit
black --check --target-version=py311 --line-length=120 .
sphinx-build -W --keep-going -a ./docs build/sphinx

Run locally against fvdb-core feature/gsplat-functional-surface at 06047779:

Suite Result
test_functional_gaussian_splatting.py 23 passed
test_gaussian_splat_3d.py, test_gaussian_splat_api.py, test_rasterize_from_world.py 173 passed, 1 skipped
Remaining unit tests 461 passed, 1 skipped; the 56 failures are all the tests that need the gettysburg or mipnerf360 datasets (53) or S3 credentials (3), none available locally. Those run in CI
One-epoch training on safety_park, image_space and world_space, crops_per_image 1 and 2 completes with main's training loop over the composed model
tests/benchmarks/test_3dgs.py intersects tiles every iteration, so its numbers stay comparable with earlier runs
Black clean

🤖 Generated with Claude Code

swahtz and others added 3 commits September 23, 2026 00:57
…rom fvdb

Add the composable, differentiable Gaussian splatting pipeline: four stages
(project_gaussians, evaluate_gaussian_sh, intersect_gaussian_tiles and the
sparse variant, and the screen-space, world-space and sparse rasterizers)
passing the frozen dataclasses ProjectedGaussians, GaussianTileIntersection
and SparseGaussianTileIntersection, plus the contributing-Gaussian analysis
functions. The autograd Functions live in functional/_autograd.py and call the
flat kernel wrappers in fvdb.functional; nothing here imports fvdb._fvdb_cpp.

CameraModel, ProjectionMethod and RollingShutterType are now re-exported from
fvdb as the same objects. GaussianRenderMode is new and owned here. The
fvdb-core pin rises to the release that publishes fvdb.functional's Gaussian
surface (openvdb/fvdb-core#797).

Part of openvdb#337.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Every render, projection, analysis, MCMC and PLY method is now a few lines
over the stage functions; the public API and docstrings are unchanged.
ProjectedGaussianSplats wraps a ProjectedGaussians plus features and exposes
it as projected_gaussians. The old radiance_fields/_gaussian_autograd.py is
removed in favour of functional/_autograd.py.

Crops now render the full image with tiles outside the crop masked off and
slice the result. The previous crop-sized tile grid combined with a rasterizer
origin offset indexed tile offsets out of bounds for nonzero origins.

Part of openvdb#337.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Tests check each stage's contract and the pipeline against GaussianSplat3d
for forward, backward, depth modes, world-space rendering through the
unscented projection, sparse rendering with duplicate pixels, crops, analysis
and empty selections, plus a short training loop. The API ownership test now
asserts the camera enums are shared with fvdb.

Docs gain an API page for fvdb_reality_capture.functional; the enums page
points at fvdb's enums through intersphinx, since fvdb is mocked in this
repo's docs build.

Fixes openvdb#337.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
@swahtz
swahtz requested a review from a team as a code owner September 22, 2026 12:59
@swahtz swahtz added enhancement New feature or request Tech Debt The future cost of today's velocity. labels Sep 22, 2026
@swahtz
swahtz requested review from harrism and matthewdcong and removed request for a team September 22, 2026 12:59
@swahtz swahtz added enhancement New feature or request Tech Debt The future cost of today's velocity. labels Sep 22, 2026
…eline

- Crops derive the per-tile render mask from the crop rectangle directly and
  skip the per-pixel post-pass when no pixel mask is given, so a crop no
  longer allocates and multiplies a full-image pixel mask. Output buffers are
  still full size; the docstring says so.
- render_from_projected_gaussians clips the crop to the image before embedding
  a crop-space mask, so edge crops no longer raise or misalign the mask.
- rasterize_world_space_gaussians raises for OpenCV camera models without
  distortion coefficients instead of substituting zeros, matching projection.
- The sparse rasterizer's mask parameter is now tile_masks, since it is per
  tile; the dense ones remain per pixel and pool internally.
- Opacities are computed once per pipeline call: the tile-culling copy runs
  under no_grad, the analysis functions share one computation, and
  ProjectedGaussianSplats caches its opacities property.
- The pipeline dataclasses use identity equality (eq=False); the generated
  element-wise tensor comparison raised on == and hashing.
- as_pixel_jagged is exported from functional and GaussianSplat3d uses it;
  the class's duplicate helper and its unused _resolve_projection_method are
  removed.
- Masks are coerced to bool before being combined with the crop window.
- The gradient-parity test tolerance matches the forward tolerances, since
  both paths run the same nondeterministic atomicAdd kernels.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
swahtz and others added 9 commits September 23, 2026 01:30
- Sparse contributor expansion derives each pixel's camera from the jagged
  offsets; jidx is empty for a single camera, which flattened the nesting to
  one outer list per pixel.
- render_from_projected_gaussians validates the crop before clipping it, so an
  origin at or past the image edge raises the out-of-bounds error instead of a
  negative size or a mask shape mismatch. validate_crop is public.
- ProjectedGaussianSplats caches tile intersections per tile size, so several
  crops from one projection intersect once. Full-size raster buffers remain
  (openvdb/fvdb-core#800).
- Opacities are computed once per pipeline call: every stage that takes
  logit_opacities accepts precomputed opacities, resolve_opacities validates
  them, and GaussianSplat3d passes one tensor through intersection, rasterization
  and analysis. The projected-splats opacities cache feeds its render path.
- The class's unused _deduplicate_pixels wrapper is removed; the dedup tests
  call fvdb_reality_capture.functional.deduplicate_pixels directly.
- Tests for single-camera sparse analysis nesting, precomputed-opacity parity
  and validation, tile caching, and out-of-image crop origins.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…e's edges

The training loop called the render backend once per crop, so with
crops_per_image > 1 the image-space backend re-projected the Gaussians,
re-evaluated spherical harmonics and re-intersected tiles for every crop,
and the world-space backend re-rendered the whole image. forward_train now
does the crop-independent work once and returns a TrainingView whose
render_crop reuses it, so the tile and opacity caches on
ProjectedGaussianSplats finally apply during training. Raster buffers are
still full size (openvdb/fvdb-core#800); the crops_per_image docstring says
so.

Smaller findings from the same review, each verified against the tree:

- ProjectedGaussianSplats.opacities recomputes when the cached copy was made
  under no_grad but a graph is wanted, instead of feeding a graph-less tensor
  to a training render.
- deduplicate_pixels gives out-of-image pixels their own keys so they never
  alias a valid pixel, and returns unique pixels in first-request order as
  its docstring always claimed (stable sort plus a rank remap).
- as_pixel_jagged rejects empty camera batches, malformed (row, col) data and
  non-integer coordinates, matching fvdb's own normalization.
- resolve_opacities checks the device and materializes non-contiguous input
  rather than letting the kernel fail on it.
- The dense rasterizers slice to the crop before the pixel-mask pass and take
  the tile grid from the intersection instead of recomputing it.
- rasterize_screen_space_gaussians documents that only analytic projections
  carry a gradient, and ImageSpaceRenderBackend.validate_scene_cameras
  refuses cameras that resolve to the forward-only unscented projection,
  pointing at render_backend="world_space".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Tile intersection, the three rasterizers and the four analysis functions
took logit_opacities and derived the [C, N] kernel opacities internally,
then also accepted a precomputed opacities override. Two parameters for
one quantity, with nothing checking that they agreed. They now take only
opacities, which restores the contract the kernel boundary had before the
functional refactor: compute_gaussian_opacities builds them once per render
from the logits and the same tensor is passed through every stage.

intersect_gaussian_tiles and intersect_gaussian_tiles_sparse keep the
argument optional, since None selects bounding-box culling. resolve_opacities
is gone; the shape, device and contiguity checks it carried live in a
private check_opacities that every stage runs. GaussianSplat3d and its
ProjectedGaussianSplats already computed once and passed through, so the
class changes only in what it forwards. The docs example and the tests call
compute_gaussian_opacities explicitly.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…t opacities at projection

Refusing image-space training on cameras that need the unscented projection
was the right diagnosis but the wrong remedy: it made a COLMAP scene with
distortion fail under the default config and blocked resuming checkpoints
trained before this branch. GaussianSplatReconstruction now resolves its
backend against the scene's cameras in the constructor and, when image
space cannot train them, logs a warning and renders through the world-space
backend instead. The warning also says that densification statistics are
not accumulated on that path, which has always been the case for the
unscented projection; the optimizer already skips insertion and reports it.
The image-space backend keeps its own check for direct callers.

ProjectedGaussianSplats computes its opacities when it is constructed,
together with the features, instead of from a live reference to the model's
logits at render time. An optimizer step between projection and render no
longer mixes new opacities with old geometry, and a projection made without
a graph renders without one for both quantities.

Also from the same review: masks are checked against the crop (class) or
the full image (stage) instead of being read from their top-left corner;
sparse contributor expansion builds one gather plan for ids and weights
with a single sync; Crop and apply_crop are exported from the functional
package and shared by the training backend; the backend imports the
projection resolver from the public package.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…g it

fvdb.functional now exports the pixel-selection normalizer its own sparse
wrappers use (openvdb/fvdb-core#799), so this package re-exports it rather
than keeping a line-for-line copy that could drift from the kernels' checks.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…depth

Scenes with distortion cameras now train through the world-space backend
by default, so its gaps matter. The depth feature came from the unscented
projection, which has no backward pass, so depth supervision moved the
geometry only through alpha. evaluate_gaussian_sh now recomputes the depth
as the view-space z of each center when a gradient is wanted and the
projection cannot supply one; depth is linear in the center, so the value is
unchanged and the gradient is exact. The world-space training view also
marks that every crop is a slice of one render, and the training loop sums
the crop losses and runs that render's backward once instead of a
full-image backward per crop.

The refinement guard checked the 2D-gradient accumulators for None, but the
model allocates them as zeros whenever accumulation is enabled and the
unscented projection never writes to them, so the check never fired and
refinement was silently inert. It now checks that anything was accumulated,
warns once, and names the screen-space size thresholds it also disables.

Sparse depth points are in full-image pixels; inside the crop loop they
indexed a crop-sized depth map and re-normalized the target every crop.
Points are now filtered to the crop and indexed in crop coordinates, and the
target is normalized into a per-crop view.

Also: requires_distortion_coeffs is the single test for distortion camera
models, used by projection, rasterization and the backends, so a new model
in fvdb is handled consistently; the unscented branch clears its own
compensations; the two forward-only messages share their explanation; and
the render benchmark intersects tiles on every iteration again, since the
per-projection cache would otherwise drop that work from the timing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…views die per step

Training views now render crops from detached copies of their shared work
(projection and features for image space, the full render for world space).
The loop runs loss.backward() per crop, which frees that crop's raster and
loss graphs and accumulates into the copies, then finish_backward() runs the
shared backward once. The projection and feature backwards no longer repeat
per crop, the densification statistics see the whole image's gradient once
exactly as with a single crop (tested against the whole-image path), and no
crop's loss graph outlives its own backward. The view is deleted at the end
of each step so two projections never coexist across steps.

The image-space backend decides per training batch whether the camera needs
the unscented projection and renders only those batches through the
world-space path, instead of switching the whole reconstruction. Pinhole
views in a mixed scene keep image space and keep feeding densification
statistics, as on main. Validation logs the affected camera models once and
probes them through the path they will take; the shared probe loop replaces
two copies of it. resolve_training_backend is gone.

The refinement guard also requires a non-zero accumulated 2D-mean gradient:
world-space rendering with an analytic projection reaches the projection
backward through depth or antialiasing compensation with a zero mean
gradient, which bumped the step counts and passed the previous check.

render_from_projected_gaussians keeps the requested crop size when the crop
runs past the image edge, filling the outside with the background at zero
alpha, as main's callers expect; the stage function still clips. The
sh_degree property uses sh_degree_from_coefficients. The docs describe the
re-exported as_pixel_jagged in prose rather than autodoc'ing the mocked fvdb
function.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…ntract, drop the tile cache

The end-of-step cleanup deleted the training view but not the rendered
crop, which is a view into the full-size raster buffer, so the buffer lived
on until the next step rebound the name. It is released with the view now.

render_from_projected_gaussians has one edge contract, stated once: the
output always has the requested crop size. Where the crop runs past the
image the inside is rendered and the outside is background at zero alpha,
and a crop entirely outside the image is all background rather than an
error. main's docstring promised clipping while its code handed a crop-sized
tile grid and an origin to the kernel, which fvdb-core#800 shows produced
wrong pixels, so there was no correct edge behavior to preserve. The
padding lives in a functional pad_crop, which shares its background fill
with the pixel-mask pass and its window slicing with apply_crop and the
mask crop, so each exists once.

ProjectedGaussianSplats no longer caches tile intersections; the training
views keep their own, and a long-lived projection holds no tile buffers.
deduplicate_pixels returns an empty inverse index when nothing was
deduplicated, since expand_to_requested never reads it then. A TODO about
spherical-harmonics work the contributor-ids path stopped doing is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…p radius-based splitting

Each crop's loss is a mean over its own pixels or depth points, so with the
shared backward four crops summed to four times the full-image gradient and
the densification statistics saw a 4x sample. Every per-crop loss is now
weighted by the crop's share of the image (or of the depth points) before
its backward, so the crops sum to the full-image loss and training matches
crops_per_image=1 exactly; the equivalence test uses mean losses weighted
the same way. Regularization and the pose penalty do not depend on the crop
and are applied once per view. crop_loss_weight lives next to
crop_image_batch.

The routing that sends forward-only camera batches to world space moves out
of ImageSpaceRenderBackend, which is pure again and rejects such cameras,
into RoutedRenderBackend, which make_render_backend returns for
"image_space". It routes training and evaluation alike, so validation
renders through the renderer being optimized, and its validation probes
each camera model through the path it will take.

The refinement guard no longer returns early when no 2D-gradient statistics
were accumulated: gradient-driven duplication and splitting are skipped,
and splitting on screen-space radius and deletion proceed as on main. The
warning says so.

Also: check_tiles_match verifies that a tile intersection belongs to the
projection it is rasterized with, in every stage and analysis function, so
stale tiles raise instead of indexing the kernel out of range;
render_from_projected_gaussians takes precomputed tiles for repeated crops;
the four analysis methods share _project_and_opacities; the docs list Crop
and as_pixel_jagged; the benchmark comment no longer describes a cache.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
swahtz and others added 6 commits September 24, 2026 09:11
…butor expansion

JaggedTensor.from_data_offsets_and_list_ids takes the camera count from the
largest list id, so when duplicates trigger the expansion a camera after the
last one with a requested pixel was dropped from the result (openvdb/fvdb-core#802).
_expand_contributions now falls back to the nested constructor in that case,
which keeps one outer list per camera. A test covers a duplicate pixel in one
camera and no pixels in the next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…crops differentiable

The per-crop loss and backward now live in _train_crop, which returns
detached terms, so the tensors holding a crop's graph die when it returns.
Before, crop_loss and friends survived into the next forward pass and, through
their AccumulateGrad nodes, kept the view's detached copies of the shared
work and their gradients alive; in world space that is a full image and its
gradient. _SharedWork.backward drops the copies' gradients once it has
propagated them. Depth supervision for a view is bundled in _DepthTargets.

A crop that lies entirely outside the image is now sliced from the projected
features instead of allocated, so its background output stays connected to
the projection with zero gradient and backward through it no longer raises.

Tests cover the lifetime of the copies across a two-view sequence and the
gradient of an out-of-image crop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…ks, resolve cameras once

World-space rendering does not differentiate through the projected means, so
_render_dense now projects without the densification accumulators and its
views leave the step counts untouched; the optimizer guard reduces to "no view
fed the statistics". Both render paths use _project_and_opacities.

render_from_projected_gaussians rejects tiles binned at another tile_size and
validates the mask shape before the out-of-image early return. The routed
backend resolves a batch's camera model once and calls the chosen backend's
body directly. Validation probes build each camera batch from the scene
metadata instead of decoding an image.

The crop-equivalence claim is narrowed: per-pixel terms sum to the full-image
loss, SSIM differs along crop seams, and leftover rows or columns from a
non-divisible image size go unsupervised. Documented at crops_per_image,
crop_loss_weight, TrainingView and the equivalence test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
… background per view

Projection and world-space rasterization share check_distortion_coeffs
(exported), so a tensor that is not a contiguous [C, 12] on the right device
is rejected before a kernel reads twelve coefficients per camera. Masks are
checked for device where their shape is checked, in the stage functions and in
render_from_projected_gaussians.

The training loop draws a random background once per view so every crop is
supervised against the same target. The sparse-depth term is a masked sum over
all of the view's points, which removes the per-crop host syncs.

render_from_projected_gaussians takes its tile size from precomputed tiles
(tile_size defaults to None; an explicit size that disagrees raises). The
dense rasterizers share their crop-and-mask epilogue, and the three backends
share the camera-batching guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
… sparse tile masks

Validation warns when pose optimization is on for cameras that render in world
space, since that rasterizer returns no gradient for the camera matrices. The
crops_per_image docstring adds the per-crop scale-and-shift fit of the relative
depth loss to the list of differences from whole-image training, and a crop
count that would make any crop smaller than the SSIM window is rejected up
front. The sparse rasterizer checks tile_masks against the tile grid's shape
and device.

The depth channel is recomputed from the centers whenever a gradient is
wanted, on either projection, as an elementwise product and sum, so pinhole
world-space training no longer runs the analytic projection's full backward
for one row of the rotation. A crop-space mask is applied to the sliced
render through the exported apply_pixel_mask instead of being embedded in a
full-image mask; the stage and class docstrings cross-reference their
different treatment of crops outside the image. The docs give
as_pixel_jagged's real parameter name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…jection identity for tiles

The three render backends share _ModelRenderBackend, which owns validation
and the forward passes; a backend provides _probe, _train_view and
_eval_render for a resolved camera model and one dict of render arguments.
render_backend="auto" (the new default) is the routed backend, and
"image_space" is the pure one again. World-space views render each crop
directly through a new crop argument on the *_from_world methods instead of
slicing a full render held across crops; the camera matrices are the shared
input, so they go through detached copies and finish_backward.

Every ProjectedGaussians carries an identity token that its tile
intersections record, so check_tiles_match rejects tiles from before an
optimizer step or a refinement, not only from another image size or camera
count. replace() keeps the token.

The accumulators stay allocated on every path so refinement can read the
radius accumulator after world-space training; only the gradient ones are
withheld from world-space projections. The depth channel is always computed
from the centers, so it does not depend on the projection method or autograd
state. A zero crop size raises. Validation probes and the crop-size check use
the size the dataset delivers, honouring patch_size.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
swahtz and others added 9 commits September 24, 2026 22:01
…contract, crop masks skip tiles

The world-space training view projects, evaluates features and intersects
tiles once and rasterizes each crop from that through
rasterize_world_space_gaussians(crop=...); features, opacities and the camera
matrices are the shared inputs and go through the detached-copy mechanism.

validate_crop clips a crop entirely outside the image to nothing instead of
raising, so the stage returns an empty render and the class method pads it;
the class no longer fabricates an empty render, and a zero crop size raises
for any origin. With a crop, the dense stages accept a mask of the clipped
crop's size, pool the tile mask from it so masked-out tiles are skipped, and
apply it to the sliced render.

check_distortion_coeffs returns None for camera models that ignore the
coefficients and validates only the distortion models. The depth channel is
zero for culled Gaussians, as the projection and the feature kernel give
them. Camera probes read the dataset's per-image arrays.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
A run renders every batch with the same backend; only the projection method
may differ between batches when a scene holds several camera models, as on
main. render_backend="auto" now resolves once in the constructor through
resolve_render_backend: image space unless some camera needs the unscented
projection, in which case the whole run renders in world space and a log
names the camera, the missing densification statistics and, when poses are
optimized, the missing camera-matrix gradient. RoutedRenderBackend and the
per-batch dispatch are gone; make_render_backend builds the pure backends
only. The render_backend docstring states the rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…k contract, empty crops skip the kernel

The world-space training view projects through _project_for with
accumulate_statistics off, so with antialiasing on the opacity gradient no
longer bumps the step counts and the "no statistics" warning fires as it
should. The camera matrices are passed to the rasterizer detached: it returns
no gradient for them, and passing them with their graph made each crop's
backward free the pose-adjustment graph the shared features still needed.

The dense stages accept a crop mask of the requested or the clipped size, so
render_from_projected_gaussians passes its mask straight through and the
class-side slicing is gone. A crop clipped to nothing returns a connected
empty render without launching the rasterizer. Three docstrings that still
promised a raise for all-outside crops now describe the empty render, and a
spliced sentence in the screen-space mask docstring is repaired.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…istics, crop masks win on shape collision

The analytic kernel records the radius accumulator only inside its gradient
pass, so a projection made without the gradient accumulators recorded no
radii and radius-based splitting and deletion were dead in world space.
_project now records max(rx, ry) per Gaussian from the forward in that case,
and the resolution log says radius-based refinement proceeds.

accumulate_statistics is a keyword on the three public project_gaussians_for_*
methods, so the world-space backend and the tests use the public surface. The
private projection chain is called with keyword arguments throughout, so a
transposed pair of floats can no longer type-check. With a crop, a mask whose
shape matches both the crop and the image is read as a crop mask. The
crops_per_image docstring and the PR body state the regularization weighting
change and the tensor-input sparse render shape change relative to main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…asks by coordinate frame, derive opacities once

Radii are recorded from the forward only when the kernel's gradient pass will
not see the projection (gradient accumulators withheld or disabled, or the
unscented projection) and grad is enabled, so evaluation renders and probes
no longer write the accumulator and accumulate_max_2d_radii works without
accumulate_mean_2d_gradients. The dense stages take masks in image coordinates
and a separate crop_masks in crop coordinates, never both, so no shape is
guessed; render_from_projected_gaussians passes crop_masks. ProjectedGaussianSplats
receives its opacities from the projection helper. The crop docstring no
longer promises a raise for mixed negative values.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…d-space renders, constructor docs

record_radii is a keyword on the three public projection methods and the
private chain; the training backends pass it, so radii are recorded in the
forward for every training projection (world space, frozen geometry and the
unscented projection included) and never from autograd state, so evaluation,
probes and viewers leave the accumulator alone. accumulate_statistics keeps
its meaning of wiring the gradient accumulators into the backward. The three
*_from_world methods take crop_masks. ProjectedGaussianSplats requires its
opacities and documents them in the constructor; the flag docstrings state
the real rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…ning-backends

This PR is the functional API port: fvdb_reality_capture.functional, the
GaussianSplat3d composed from it, and the stage contracts and boundary bugs
the port exposed. The training-side work that grew out of the review rounds
(backend resolution, the crop-based TrainingView, per-view regularization,
sparse depth under crops, densification statistics on the world-space path,
view lifetime) is independent of the port and moves to its own branch on top
of this one. _gaussian_rendering.py, the training loop, the optimizer guard,
_private/utils.py and test_training.py are back to their main versions, the
accumulate_statistics and record_radii keywords are gone, and the backend
tests go with the code they test. The index-set test expectation returns to
main's.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
fvdb's build_sparse_gaussian_tile_layout now requires image_width and
image_height so it can reject pixels that fall inside the tile padding
but outside the image. Passing them also gives deduplicated selections
the exact bounds check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
No fvdb kernel takes a ProjectionMethod; the choice between the analytic
and unscented projections is made here by resolve_projection_method,
which calls a different fvdb function for each. fvdb no longer exports
the enum, so it is defined here with the same members and values.
CameraModel and RollingShutterType stay shared with fvdb because its
kernels accept them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>

@harrism harrism left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @swahtz, this is a big cleanup, and the functional stages look clean! The crop fix and the validation listed under "Behaviour changes" look right to me, including the full-size raster buffers until fvdb-core#800 is fixed.

Requesting changes for a few docs and comments that don't match the code (inline).

Comment thread fvdb_reality_capture/radiance_fields/gaussian_splatting.py Outdated
Comment thread fvdb_reality_capture/radiance_fields/gaussian_splatting.py Outdated
Comment thread fvdb_reality_capture/radiance_fields/gaussian_splatting.py Outdated
Comment thread fvdb_reality_capture/radiance_fields/gaussian_splatting.py Outdated
swahtz and others added 5 commits October 8, 2026 19:16
…mulator wiring as it is

The crop and crop_masks entries sat in the docstrings of sparse_render_images
and render_num_contributing_gaussians, which take neither; they now live with
render_images_from_world, and the two other *_from_world methods point at it.
The _project comment claimed world-space projections skip the gradient
accumulators, which this branch does not do (it wires them as main did); the
comment now says so and names the follow-up. tile_intersection's docstring no
longer refers to training views.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Resolves gaussian_splatting.py and test_gaussian_splat_api.py in favour of
this branch: main adopted the fvdb-core kernel renames at call sites that no
longer exist here (every kernel call goes through fvdb_reality_capture.functional),
and main's enum test tolerated fvdb exposing the camera enums where this
branch asserts they are the same objects. Takes main's exponential momentum
scaling (openvdb#341) and the Python 3.10 typing fix (openvdb#343) as is.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Found by an audit of every docstring this branch touched. The depth channel
from evaluate_gaussian_sh equals the projection's depth for kept Gaussians and
is zero for culled ones, whose projected depth is not defined. The three
project_gaussians_for_* docstrings referred to a render_projected_gaussians
method that does not exist, and the depth example used an undefined variable.
render_num_contributing_gaussians returns (C, H, W) counts and alphas, not
(C, H, W, 1). inv_covar_2d is (C, N, 3), not (C, N, D). Sparse pixel
selections are (row, col), not (x, y). accumulated_max_2d_radii says when the
kernel updates it. requires_distortion_coeffs no longer mentions training
backends.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
… that broke the Sphinx build

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…nt tense

Drops the two remaining references to the training follow-up (the _project
comment, the check_tiles_match note about views), fixes a test comment that
misdescribed culled depths, says the step-count accumulator is updated only
alongside the norm accumulator rather than required with it, notes that
pixels_to_render is normalized by as_pixel_jagged, qualifies two cross-
references so they resolve from the functional docs page, documents the
ValueError raises of pad_crop and evaluate_gaussian_sh, moves the
deduplicate_pixels detail into its body so napoleon renders the Returns block,
and drops a historicizing clause in the benchmark.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>

@harrism harrism left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the doc fixes!

@swahtz
swahtz merged commit a123a88 into openvdb:main Oct 8, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request Tech Debt The future cost of today's velocity.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Composable differentiable Gaussian splatting API over fvdb.functional

2 participants