Repository navigation
Composable differentiable Gaussian splatting API over fvdb.functional - #338
Merged
swahtz merged 33 commits intoOct 8, 2026
Merged
Conversation
…rom fvdb Add the composable, differentiable Gaussian splatting pipeline: four stages (project_gaussians, evaluate_gaussian_sh, intersect_gaussian_tiles and the sparse variant, and the screen-space, world-space and sparse rasterizers) passing the frozen dataclasses ProjectedGaussians, GaussianTileIntersection and SparseGaussianTileIntersection, plus the contributing-Gaussian analysis functions. The autograd Functions live in functional/_autograd.py and call the flat kernel wrappers in fvdb.functional; nothing here imports fvdb._fvdb_cpp. CameraModel, ProjectionMethod and RollingShutterType are now re-exported from fvdb as the same objects. GaussianRenderMode is new and owned here. The fvdb-core pin rises to the release that publishes fvdb.functional's Gaussian surface (openvdb/fvdb-core#797). Part of openvdb#337. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Every render, projection, analysis, MCMC and PLY method is now a few lines over the stage functions; the public API and docstrings are unchanged. ProjectedGaussianSplats wraps a ProjectedGaussians plus features and exposes it as projected_gaussians. The old radiance_fields/_gaussian_autograd.py is removed in favour of functional/_autograd.py. Crops now render the full image with tiles outside the crop masked off and slice the result. The previous crop-sized tile grid combined with a rasterizer origin offset indexed tile offsets out of bounds for nonzero origins. Part of openvdb#337. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Tests check each stage's contract and the pipeline against GaussianSplat3d for forward, backward, depth modes, world-space rendering through the unscented projection, sparse rendering with duplicate pixels, crops, analysis and empty selections, plus a short training loop. The API ownership test now asserts the camera enums are shared with fvdb. Docs gain an API page for fvdb_reality_capture.functional; the enums page points at fvdb's enums through intersphinx, since fvdb is mocked in this repo's docs build. Fixes openvdb#337. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
swahtz
requested review from
harrism and
matthewdcong
and removed request for
a team
September 22, 2026 12:59
…eline - Crops derive the per-tile render mask from the crop rectangle directly and skip the per-pixel post-pass when no pixel mask is given, so a crop no longer allocates and multiplies a full-image pixel mask. Output buffers are still full size; the docstring says so. - render_from_projected_gaussians clips the crop to the image before embedding a crop-space mask, so edge crops no longer raise or misalign the mask. - rasterize_world_space_gaussians raises for OpenCV camera models without distortion coefficients instead of substituting zeros, matching projection. - The sparse rasterizer's mask parameter is now tile_masks, since it is per tile; the dense ones remain per pixel and pool internally. - Opacities are computed once per pipeline call: the tile-culling copy runs under no_grad, the analysis functions share one computation, and ProjectedGaussianSplats caches its opacities property. - The pipeline dataclasses use identity equality (eq=False); the generated element-wise tensor comparison raised on == and hashing. - as_pixel_jagged is exported from functional and GaussianSplat3d uses it; the class's duplicate helper and its unused _resolve_projection_method are removed. - Masks are coerced to bool before being combined with the crop window. - The gradient-parity test tolerance matches the forward tolerances, since both paths run the same nondeterministic atomicAdd kernels. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
- Sparse contributor expansion derives each pixel's camera from the jagged offsets; jidx is empty for a single camera, which flattened the nesting to one outer list per pixel. - render_from_projected_gaussians validates the crop before clipping it, so an origin at or past the image edge raises the out-of-bounds error instead of a negative size or a mask shape mismatch. validate_crop is public. - ProjectedGaussianSplats caches tile intersections per tile size, so several crops from one projection intersect once. Full-size raster buffers remain (openvdb/fvdb-core#800). - Opacities are computed once per pipeline call: every stage that takes logit_opacities accepts precomputed opacities, resolve_opacities validates them, and GaussianSplat3d passes one tensor through intersection, rasterization and analysis. The projected-splats opacities cache feeds its render path. - The class's unused _deduplicate_pixels wrapper is removed; the dedup tests call fvdb_reality_capture.functional.deduplicate_pixels directly. - Tests for single-camera sparse analysis nesting, precomputed-opacity parity and validation, tile caching, and out-of-image crop origins. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…e's edges The training loop called the render backend once per crop, so with crops_per_image > 1 the image-space backend re-projected the Gaussians, re-evaluated spherical harmonics and re-intersected tiles for every crop, and the world-space backend re-rendered the whole image. forward_train now does the crop-independent work once and returns a TrainingView whose render_crop reuses it, so the tile and opacity caches on ProjectedGaussianSplats finally apply during training. Raster buffers are still full size (openvdb/fvdb-core#800); the crops_per_image docstring says so. Smaller findings from the same review, each verified against the tree: - ProjectedGaussianSplats.opacities recomputes when the cached copy was made under no_grad but a graph is wanted, instead of feeding a graph-less tensor to a training render. - deduplicate_pixels gives out-of-image pixels their own keys so they never alias a valid pixel, and returns unique pixels in first-request order as its docstring always claimed (stable sort plus a rank remap). - as_pixel_jagged rejects empty camera batches, malformed (row, col) data and non-integer coordinates, matching fvdb's own normalization. - resolve_opacities checks the device and materializes non-contiguous input rather than letting the kernel fail on it. - The dense rasterizers slice to the crop before the pixel-mask pass and take the tile grid from the intersection instead of recomputing it. - rasterize_screen_space_gaussians documents that only analytic projections carry a gradient, and ImageSpaceRenderBackend.validate_scene_cameras refuses cameras that resolve to the forward-only unscented projection, pointing at render_backend="world_space". Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Tile intersection, the three rasterizers and the four analysis functions took logit_opacities and derived the [C, N] kernel opacities internally, then also accepted a precomputed opacities override. Two parameters for one quantity, with nothing checking that they agreed. They now take only opacities, which restores the contract the kernel boundary had before the functional refactor: compute_gaussian_opacities builds them once per render from the logits and the same tensor is passed through every stage. intersect_gaussian_tiles and intersect_gaussian_tiles_sparse keep the argument optional, since None selects bounding-box culling. resolve_opacities is gone; the shape, device and contiguity checks it carried live in a private check_opacities that every stage runs. GaussianSplat3d and its ProjectedGaussianSplats already computed once and passed through, so the class changes only in what it forwards. The docs example and the tests call compute_gaussian_opacities explicitly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…t opacities at projection Refusing image-space training on cameras that need the unscented projection was the right diagnosis but the wrong remedy: it made a COLMAP scene with distortion fail under the default config and blocked resuming checkpoints trained before this branch. GaussianSplatReconstruction now resolves its backend against the scene's cameras in the constructor and, when image space cannot train them, logs a warning and renders through the world-space backend instead. The warning also says that densification statistics are not accumulated on that path, which has always been the case for the unscented projection; the optimizer already skips insertion and reports it. The image-space backend keeps its own check for direct callers. ProjectedGaussianSplats computes its opacities when it is constructed, together with the features, instead of from a live reference to the model's logits at render time. An optimizer step between projection and render no longer mixes new opacities with old geometry, and a projection made without a graph renders without one for both quantities. Also from the same review: masks are checked against the crop (class) or the full image (stage) instead of being read from their top-left corner; sparse contributor expansion builds one gather plan for ids and weights with a single sync; Crop and apply_crop are exported from the functional package and shared by the training backend; the backend imports the projection resolver from the public package. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…g it fvdb.functional now exports the pixel-selection normalizer its own sparse wrappers use (openvdb/fvdb-core#799), so this package re-exports it rather than keeping a line-for-line copy that could drift from the kernels' checks. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…depth Scenes with distortion cameras now train through the world-space backend by default, so its gaps matter. The depth feature came from the unscented projection, which has no backward pass, so depth supervision moved the geometry only through alpha. evaluate_gaussian_sh now recomputes the depth as the view-space z of each center when a gradient is wanted and the projection cannot supply one; depth is linear in the center, so the value is unchanged and the gradient is exact. The world-space training view also marks that every crop is a slice of one render, and the training loop sums the crop losses and runs that render's backward once instead of a full-image backward per crop. The refinement guard checked the 2D-gradient accumulators for None, but the model allocates them as zeros whenever accumulation is enabled and the unscented projection never writes to them, so the check never fired and refinement was silently inert. It now checks that anything was accumulated, warns once, and names the screen-space size thresholds it also disables. Sparse depth points are in full-image pixels; inside the crop loop they indexed a crop-sized depth map and re-normalized the target every crop. Points are now filtered to the crop and indexed in crop coordinates, and the target is normalized into a per-crop view. Also: requires_distortion_coeffs is the single test for distortion camera models, used by projection, rasterization and the backends, so a new model in fvdb is handled consistently; the unscented branch clears its own compensations; the two forward-only messages share their explanation; and the render benchmark intersects tiles on every iteration again, since the per-projection cache would otherwise drop that work from the timing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…views die per step Training views now render crops from detached copies of their shared work (projection and features for image space, the full render for world space). The loop runs loss.backward() per crop, which frees that crop's raster and loss graphs and accumulates into the copies, then finish_backward() runs the shared backward once. The projection and feature backwards no longer repeat per crop, the densification statistics see the whole image's gradient once exactly as with a single crop (tested against the whole-image path), and no crop's loss graph outlives its own backward. The view is deleted at the end of each step so two projections never coexist across steps. The image-space backend decides per training batch whether the camera needs the unscented projection and renders only those batches through the world-space path, instead of switching the whole reconstruction. Pinhole views in a mixed scene keep image space and keep feeding densification statistics, as on main. Validation logs the affected camera models once and probes them through the path they will take; the shared probe loop replaces two copies of it. resolve_training_backend is gone. The refinement guard also requires a non-zero accumulated 2D-mean gradient: world-space rendering with an analytic projection reaches the projection backward through depth or antialiasing compensation with a zero mean gradient, which bumped the step counts and passed the previous check. render_from_projected_gaussians keeps the requested crop size when the crop runs past the image edge, filling the outside with the background at zero alpha, as main's callers expect; the stage function still clips. The sh_degree property uses sh_degree_from_coefficients. The docs describe the re-exported as_pixel_jagged in prose rather than autodoc'ing the mocked fvdb function. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…ntract, drop the tile cache The end-of-step cleanup deleted the training view but not the rendered crop, which is a view into the full-size raster buffer, so the buffer lived on until the next step rebound the name. It is released with the view now. render_from_projected_gaussians has one edge contract, stated once: the output always has the requested crop size. Where the crop runs past the image the inside is rendered and the outside is background at zero alpha, and a crop entirely outside the image is all background rather than an error. main's docstring promised clipping while its code handed a crop-sized tile grid and an origin to the kernel, which fvdb-core#800 shows produced wrong pixels, so there was no correct edge behavior to preserve. The padding lives in a functional pad_crop, which shares its background fill with the pixel-mask pass and its window slicing with apply_crop and the mask crop, so each exists once. ProjectedGaussianSplats no longer caches tile intersections; the training views keep their own, and a long-lived projection holds no tile buffers. deduplicate_pixels returns an empty inverse index when nothing was deduplicated, since expand_to_requested never reads it then. A TODO about spherical-harmonics work the contributor-ids path stopped doing is gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…p radius-based splitting Each crop's loss is a mean over its own pixels or depth points, so with the shared backward four crops summed to four times the full-image gradient and the densification statistics saw a 4x sample. Every per-crop loss is now weighted by the crop's share of the image (or of the depth points) before its backward, so the crops sum to the full-image loss and training matches crops_per_image=1 exactly; the equivalence test uses mean losses weighted the same way. Regularization and the pose penalty do not depend on the crop and are applied once per view. crop_loss_weight lives next to crop_image_batch. The routing that sends forward-only camera batches to world space moves out of ImageSpaceRenderBackend, which is pure again and rejects such cameras, into RoutedRenderBackend, which make_render_backend returns for "image_space". It routes training and evaluation alike, so validation renders through the renderer being optimized, and its validation probes each camera model through the path it will take. The refinement guard no longer returns early when no 2D-gradient statistics were accumulated: gradient-driven duplication and splitting are skipped, and splitting on screen-space radius and deletion proceed as on main. The warning says so. Also: check_tiles_match verifies that a tile intersection belongs to the projection it is rasterized with, in every stage and analysis function, so stale tiles raise instead of indexing the kernel out of range; render_from_projected_gaussians takes precomputed tiles for repeated crops; the four analysis methods share _project_and_opacities; the docs list Crop and as_pixel_jagged; the benchmark comment no longer describes a cache. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
This was referenced Sep 23, 2026
…butor expansion JaggedTensor.from_data_offsets_and_list_ids takes the camera count from the largest list id, so when duplicates trigger the expansion a camera after the last one with a requested pixel was dropped from the result (openvdb/fvdb-core#802). _expand_contributions now falls back to the nested constructor in that case, which keeps one outer list per camera. A test covers a duplicate pixel in one camera and no pixels in the next. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…crops differentiable The per-crop loss and backward now live in _train_crop, which returns detached terms, so the tensors holding a crop's graph die when it returns. Before, crop_loss and friends survived into the next forward pass and, through their AccumulateGrad nodes, kept the view's detached copies of the shared work and their gradients alive; in world space that is a full image and its gradient. _SharedWork.backward drops the copies' gradients once it has propagated them. Depth supervision for a view is bundled in _DepthTargets. A crop that lies entirely outside the image is now sliced from the projected features instead of allocated, so its background output stays connected to the projection with zero gradient and backward through it no longer raises. Tests cover the lifetime of the copies across a two-view sequence and the gradient of an out-of-image crop. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…ks, resolve cameras once World-space rendering does not differentiate through the projected means, so _render_dense now projects without the densification accumulators and its views leave the step counts untouched; the optimizer guard reduces to "no view fed the statistics". Both render paths use _project_and_opacities. render_from_projected_gaussians rejects tiles binned at another tile_size and validates the mask shape before the out-of-image early return. The routed backend resolves a batch's camera model once and calls the chosen backend's body directly. Validation probes build each camera batch from the scene metadata instead of decoding an image. The crop-equivalence claim is narrowed: per-pixel terms sum to the full-image loss, SSIM differs along crop seams, and leftover rows or columns from a non-divisible image size go unsupervised. Documented at crops_per_image, crop_loss_weight, TrainingView and the equivalence test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
… background per view Projection and world-space rasterization share check_distortion_coeffs (exported), so a tensor that is not a contiguous [C, 12] on the right device is rejected before a kernel reads twelve coefficients per camera. Masks are checked for device where their shape is checked, in the stage functions and in render_from_projected_gaussians. The training loop draws a random background once per view so every crop is supervised against the same target. The sparse-depth term is a masked sum over all of the view's points, which removes the per-crop host syncs. render_from_projected_gaussians takes its tile size from precomputed tiles (tile_size defaults to None; an explicit size that disagrees raises). The dense rasterizers share their crop-and-mask epilogue, and the three backends share the camera-batching guard. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
… sparse tile masks Validation warns when pose optimization is on for cameras that render in world space, since that rasterizer returns no gradient for the camera matrices. The crops_per_image docstring adds the per-crop scale-and-shift fit of the relative depth loss to the list of differences from whole-image training, and a crop count that would make any crop smaller than the SSIM window is rejected up front. The sparse rasterizer checks tile_masks against the tile grid's shape and device. The depth channel is recomputed from the centers whenever a gradient is wanted, on either projection, as an elementwise product and sum, so pinhole world-space training no longer runs the analytic projection's full backward for one row of the rotation. A crop-space mask is applied to the sliced render through the exported apply_pixel_mask instead of being embedded in a full-image mask; the stage and class docstrings cross-reference their different treatment of crops outside the image. The docs give as_pixel_jagged's real parameter name. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…jection identity for tiles The three render backends share _ModelRenderBackend, which owns validation and the forward passes; a backend provides _probe, _train_view and _eval_render for a resolved camera model and one dict of render arguments. render_backend="auto" (the new default) is the routed backend, and "image_space" is the pure one again. World-space views render each crop directly through a new crop argument on the *_from_world methods instead of slicing a full render held across crops; the camera matrices are the shared input, so they go through detached copies and finish_backward. Every ProjectedGaussians carries an identity token that its tile intersections record, so check_tiles_match rejects tiles from before an optimizer step or a refinement, not only from another image size or camera count. replace() keeps the token. The accumulators stay allocated on every path so refinement can read the radius accumulator after world-space training; only the gradient ones are withheld from world-space projections. The depth channel is always computed from the centers, so it does not depend on the projection method or autograd state. A zero crop size raises. Validation probes and the crop-size check use the size the dataset delivers, honouring patch_size. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…contract, crop masks skip tiles The world-space training view projects, evaluates features and intersects tiles once and rasterizes each crop from that through rasterize_world_space_gaussians(crop=...); features, opacities and the camera matrices are the shared inputs and go through the detached-copy mechanism. validate_crop clips a crop entirely outside the image to nothing instead of raising, so the stage returns an empty render and the class method pads it; the class no longer fabricates an empty render, and a zero crop size raises for any origin. With a crop, the dense stages accept a mask of the clipped crop's size, pool the tile mask from it so masked-out tiles are skipped, and apply it to the sliced render. check_distortion_coeffs returns None for camera models that ignore the coefficients and validates only the distortion models. The depth channel is zero for culled Gaussians, as the projection and the feature kernel give them. Camera probes read the dataset's per-image arrays. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
A run renders every batch with the same backend; only the projection method may differ between batches when a scene holds several camera models, as on main. render_backend="auto" now resolves once in the constructor through resolve_render_backend: image space unless some camera needs the unscented projection, in which case the whole run renders in world space and a log names the camera, the missing densification statistics and, when poses are optimized, the missing camera-matrix gradient. RoutedRenderBackend and the per-batch dispatch are gone; make_render_backend builds the pure backends only. The render_backend docstring states the rule. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…k contract, empty crops skip the kernel The world-space training view projects through _project_for with accumulate_statistics off, so with antialiasing on the opacity gradient no longer bumps the step counts and the "no statistics" warning fires as it should. The camera matrices are passed to the rasterizer detached: it returns no gradient for them, and passing them with their graph made each crop's backward free the pose-adjustment graph the shared features still needed. The dense stages accept a crop mask of the requested or the clipped size, so render_from_projected_gaussians passes its mask straight through and the class-side slicing is gone. A crop clipped to nothing returns a connected empty render without launching the rasterizer. Three docstrings that still promised a raise for all-outside crops now describe the empty render, and a spliced sentence in the screen-space mask docstring is repaired. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…istics, crop masks win on shape collision The analytic kernel records the radius accumulator only inside its gradient pass, so a projection made without the gradient accumulators recorded no radii and radius-based splitting and deletion were dead in world space. _project now records max(rx, ry) per Gaussian from the forward in that case, and the resolution log says radius-based refinement proceeds. accumulate_statistics is a keyword on the three public project_gaussians_for_* methods, so the world-space backend and the tests use the public surface. The private projection chain is called with keyword arguments throughout, so a transposed pair of floats can no longer type-check. With a crop, a mask whose shape matches both the crop and the image is read as a crop mask. The crops_per_image docstring and the PR body state the regularization weighting change and the tensor-input sparse render shape change relative to main. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…asks by coordinate frame, derive opacities once Radii are recorded from the forward only when the kernel's gradient pass will not see the projection (gradient accumulators withheld or disabled, or the unscented projection) and grad is enabled, so evaluation renders and probes no longer write the accumulator and accumulate_max_2d_radii works without accumulate_mean_2d_gradients. The dense stages take masks in image coordinates and a separate crop_masks in crop coordinates, never both, so no shape is guessed; render_from_projected_gaussians passes crop_masks. ProjectedGaussianSplats receives its opacities from the projection helper. The crop docstring no longer promises a raise for mixed negative values. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…d-space renders, constructor docs record_radii is a keyword on the three public projection methods and the private chain; the training backends pass it, so radii are recorded in the forward for every training projection (world space, frozen geometry and the unscented projection included) and never from autograd state, so evaluation, probes and viewers leave the accumulator alone. accumulate_statistics keeps its meaning of wiring the gradient accumulators into the backward. The three *_from_world methods take crop_masks. ProjectedGaussianSplats requires its opacities and documents them in the constructor; the flag docstrings state the real rule. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…ning-backends This PR is the functional API port: fvdb_reality_capture.functional, the GaussianSplat3d composed from it, and the stage contracts and boundary bugs the port exposed. The training-side work that grew out of the review rounds (backend resolution, the crop-based TrainingView, per-view regularization, sparse depth under crops, densification statistics on the world-space path, view lifetime) is independent of the port and moves to its own branch on top of this one. _gaussian_rendering.py, the training loop, the optimizer guard, _private/utils.py and test_training.py are back to their main versions, the accumulate_statistics and record_radii keywords are gone, and the backend tests go with the code they test. The index-set test expectation returns to main's. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
fvdb's build_sparse_gaussian_tile_layout now requires image_width and image_height so it can reject pixels that fall inside the tile padding but outside the image. Passing them also gives deduplicated selections the exact bounds check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
No fvdb kernel takes a ProjectionMethod; the choice between the analytic and unscented projections is made here by resolve_projection_method, which calls a different fvdb function for each. fvdb no longer exports the enum, so it is defined here with the same members and values. CameraModel and RollingShutterType stay shared with fvdb because its kernels accept them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
harrism
requested changes
Oct 8, 2026
harrism
left a comment
Contributor
There was a problem hiding this comment.
Thanks @swahtz, this is a big cleanup, and the functional stages look clean! The crop fix and the validation listed under "Behaviour changes" look right to me, including the full-size raster buffers until fvdb-core#800 is fixed.
Requesting changes for a few docs and comments that don't match the code (inline).
…mulator wiring as it is The crop and crop_masks entries sat in the docstrings of sparse_render_images and render_num_contributing_gaussians, which take neither; they now live with render_images_from_world, and the two other *_from_world methods point at it. The _project comment claimed world-space projections skip the gradient accumulators, which this branch does not do (it wires them as main did); the comment now says so and names the follow-up. tile_intersection's docstring no longer refers to training views. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Resolves gaussian_splatting.py and test_gaussian_splat_api.py in favour of this branch: main adopted the fvdb-core kernel renames at call sites that no longer exist here (every kernel call goes through fvdb_reality_capture.functional), and main's enum test tolerated fvdb exposing the camera enums where this branch asserts they are the same objects. Takes main's exponential momentum scaling (openvdb#341) and the Python 3.10 typing fix (openvdb#343) as is. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
Found by an audit of every docstring this branch touched. The depth channel from evaluate_gaussian_sh equals the projection's depth for kept Gaussians and is zero for culled ones, whose projected depth is not defined. The three project_gaussians_for_* docstrings referred to a render_projected_gaussians method that does not exist, and the depth example used an undefined variable. render_num_contributing_gaussians returns (C, H, W) counts and alphas, not (C, H, W, 1). inv_covar_2d is (C, N, 3), not (C, N, D). Sparse pixel selections are (row, col), not (x, y). accumulated_max_2d_radii says when the kernel updates it. requires_distortion_coeffs no longer mentions training backends. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
… that broke the Sphinx build Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
…nt tense Drops the two remaining references to the training follow-up (the _project comment, the check_tiles_match note about views), fixes a test comment that misdescribed culled depths, says the step-count accumulator is updated only alongside the norm accumulator rather than required with it, notes that pixels_to_render is normalized by as_pixel_jagged, qualifies two cross- references so they resolve from the functional docs page, documents the ValueError raises of pad_crop and evaluate_gaussian_sh, moves the deduplicate_pixels detail into its body so napoleon renders the Returns block, and drops a historicizing clause in the benchmark. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Jonathan Swartz <jonathan@jswartz.info>
harrism
approved these changes
Oct 8, 2026
harrism
left a comment
Contributor
There was a problem hiding this comment.
Thanks for the doc fixes!
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Builds the composable, differentiable Gaussian splatting pipeline on top of the flat kernel surface fvdb-core publishes in
fvdb.functional(openvdb/fvdb-core#797, openvdb/fvdb-core#799), and makesGaussianSplat3da wrapper that composes it. This repo no longer importsfvdb._fvdb_cppor unwrapsJaggedTensor._impl.Fixes #337. Completes the fvdb-reality-capture half of openvdb/fvdb-core#590.
Requires a fvdb-core build containing openvdb/fvdb-core#799. The pin is raised to
fvdb-core>=0.7.0dev0; CI fails at collection (cannot import name 'CameraModel' from 'fvdb') until that PR is merged into fvdb-core main, which the test workflow clones.Scope. The port, the stage contracts, and the bugs the port exposed at those boundaries.
Changes
Enums.
CameraModelandRollingShutterTypeare re-exported fromfvdbas the same objects; the local copies are gone.ProjectionMethodstays owned here, with the same members and values, because no fvdb kernel takes it:resolve_projection_methodpicks between fvdb's analytic and unscented functions.GaussianRenderMode(FEATURES,DEPTH,FEATURES_AND_DEPTH) is new and owned here.fvdb_reality_capture.functional. Four stages passing frozen dataclasses, plus analysis:project_gaussians,resolve_projection_method,requires_distortion_coeffs,check_distortion_coeffsProjectedGaussiansevaluate_gaussian_sh,sh_degree_from_coefficients,compute_gaussian_opacitiesintersect_gaussian_tiles,intersect_gaussian_tiles_sparse,deduplicate_pixels,as_pixel_jagged(re-exported from fvdb),check_tiles_matchGaussianTileIntersection,SparseGaussianTileIntersectionrasterize_screen_space_gaussians,rasterize_world_space_gaussians,rasterize_screen_space_gaussians_sparse,pixel_mask_to_tile_mask,validate_crop,apply_crop,apply_pixel_mask,pad_cropCroprasterize_num_contributing_gaussians,rasterize_contributing_gaussian_idsand their_sparsevariantsThe six
torch.autograd.Functionclasses moved fromradiance_fields/_gaussian_autograd.pytofunctional/_autograd.pyand callfvdb.functional. Stages after projection take per-cameraopacities([C, N], fromcompute_gaussian_opacities) rather than logits, which is the contract the kernel boundary had before. Analysis functions keep the gsplat-aligned names decided in #337.GaussianSplat3d. Every render, projection, analysis, MCMC and PLY method is a few lines over the stage functions; the file goes from 4928 to 4062 lines with all docstrings kept. Existing signatures are unchanged; the additions are keyword arguments:render_from_projected_gaussians(tiles=..., tile_size=None)for repeated crops from one projection, andcrop/crop_maskson the three*_from_worldmethods.ProjectedGaussianSplatsexposes itsProjectedGaussiansasprojected_gaussians, so a projection made through the class can continue in the functional pipeline; it takes its opacities at construction, alongside the features, so the projection is a consistent snapshot of the model, andtile_intersection(tile_size)computes on demand and caches nothing.Docs and tests. New
docs/api/functional.rst; the enums page points at fvdb through intersphinx for the shared enums and documentsProjectionMethodandGaussianRenderModelocally.tests/unit/test_functional_gaussian_splatting.pychecks each stage's contract and the pipeline againstGaussianSplat3d(forward, backward, depth modes, world space through the unscented projection, sparse with duplicates, crops, analysis, empty selections, deduplication order, pixel validation, opacity snapshot). The API ownership test assertsCameraModelandRollingShutterTypeare shared with fvdb and thatProjectionMethodandGaussianRenderModeare not in fvdb.Behaviour changes
Each of these differs from
main; everything else renders the same pixels.mainsized the tile grid to the crop but indexed tile offsets by absolute position, so any crop not at the origin read out of bounds (Dense rasterizers cannot render a crop into crop-sized buffers: window origin and tile-grid contract are inconsistent fvdb-core#800). Crops now mask tiles outside the window and slice the full-size render, and the result equals that region of the uncropped image. The edge contract is chosen rather than preserved, sincemainhad none that worked:render_from_projected_gaussiansreturns the requested size with the outside as background at zero alpha (a crop entirely outside is all background), the stage functions return the clipped part (empty for a crop entirely outside), a zero crop size raises, and the outputs stay differentiable. Raster buffers are still allocated at full image size (fvdb-core#800), so a crop does not save that memory.masksin image coordinates and, with a crop,crop_masksin crop coordinates (requested or clipped size), never both;render_from_projected_gaussianstakes a crop-space mask. Any other shape, or a mask on another device, raises instead of being read from the top-left corner. The sparse rasterizer takes per-tiletile_masks, since its pixels are explicit, and checks their shape and device;GaussianSplat3d.sparse_render_*keep their per-tilemasksargument.deduplicate_pixelsreturns unique pixels in first-requested order, as its docstring said, wheremainreturned sorted row-major order, and it never merges an out-of-image pixel with a valid one: onmain,(0, W)shared a key with(1, 0). Sparse stage functions return results in requested-pixel order, duplicates included, as the class methods already did.[C, P, 2]tensor of pixels returns stacked[C, P, D]tensors, as the docstrings said;mainreturned the flat[C * P, D]data. A[P, 2]tensor is rejected (as_pixel_jagged, shared with fvdb's own sparse wrappers, wants[C, P, 2]), wheremainaccepted it for one camera. In-repo callers pass jagged input and are unaffected.dataclasses.replacekeeps the token). Opacities must be on the projection's device and are made contiguous. Distortion coefficients must be a contiguous[C, 12]on the right device for distortion models and are ignored for pinhole and orthographic cameras. Pixel selections reject non-integer coordinates instead of truncating them. Precomputedtilesset the tile size; an explicittile_sizethat disagrees raises.evaluate_gaussian_shcomputes the depth channel from the Gaussian centers rather than reading the projection's, so it is the same function of its inputs under either projection method and is differentiable through the unscented projection; it equals the projection's depth, zero for culled Gaussians included.JaggedTensor.from_data_offsets_and_list_idswould drop it, JaggedTensor: empty outer lists are dropped by from_data_offsets_and_list_ids, hidden from lshape/unbind, and fail CPU indexing fvdb-core#802). A request where no pixel anywhere has a contributor segfaults inside fvdb (rasterize_contributing_gaussian_ids_sparse segfaults when no requested pixel has a contributor fvdb-core#803); pre-existing, not worked around.Test plan
Run locally against fvdb-core
feature/gsplat-functional-surfaceat06047779:test_functional_gaussian_splatting.pytest_gaussian_splat_3d.py,test_gaussian_splat_api.py,test_rasterize_from_world.pygettysburgormipnerf360datasets (53) or S3 credentials (3), none available locally. Those run in CIsafety_park,image_spaceandworld_space,crops_per_image1 and 2main's training loop over the composed modeltests/benchmarks/test_3dgs.py🤖 Generated with Claude Code