Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
0df544d
feat: add fvdb_reality_capture.functional and take the camera enums f…
swahtz Sep 22, 2026
76ceaef
refactor: compose GaussianSplat3d from fvdb_reality_capture.functional
swahtz Sep 22, 2026
953ba0b
test,docs: cover the functional Gaussian splatting pipeline
swahtz Sep 22, 2026
7cb10c3
fix: address review findings on the functional Gaussian splatting pip…
swahtz Sep 22, 2026
dab7896
fix: second review round on the functional Gaussian splatting pipeline
swahtz Sep 22, 2026
0bc986c
fix: project once per training view and harden the functional pipelin…
swahtz Sep 22, 2026
04a72c8
refactor: stage functions take per-camera opacities instead of logits
swahtz Sep 23, 2026
f36862c
fix: fall back to world-space training for unscented cameras, snapsho…
swahtz Sep 23, 2026
574617d
refactor: take as_pixel_jagged from fvdb.functional instead of copyin…
swahtz Sep 23, 2026
420d29f
fix: make the world-space training path whole, and crop-aware sparse …
swahtz Sep 23, 2026
22f2708
fix: crops share one backward, forward-only cameras route per batch, …
swahtz Sep 23, 2026
b0effdf
fix: release the rendered crop with its view, settle the edge-crop co…
swahtz Sep 23, 2026
3a8e072
fix: weight crop losses by pixel share, route eval like training, kee…
swahtz Sep 23, 2026
5d686fb
fix: keep a trailing camera with no requested pixels in sparse contri…
swahtz Sep 23, 2026
06b49c8
fix: release each crop's loss graph before the next step, keep empty …
swahtz Sep 24, 2026
7e1cc7c
fix: project without accumulators in world space, check tiles and mas…
swahtz Sep 24, 2026
7202946
fix: check distortion coefficients and mask device in every path, one…
swahtz Sep 24, 2026
560dbb5
fix: warn on pose optimization in world space, bound crop size, check…
swahtz Sep 24, 2026
36fdf4b
refactor: one backend base class, per-crop world-space rendering, pro…
swahtz Sep 24, 2026
7d28ca7
fix: share the projection across world-space crops, one outside-crop …
swahtz Sep 24, 2026
7e3f5da
refactor: one render backend per run, resolved from the scene
swahtz Sep 24, 2026
006c191
fix: world-space training projects without accumulators, one crop-mas…
swahtz Sep 24, 2026
5ba86d2
fix: record radii for world-space projections, expose accumulate_stat…
swahtz Sep 24, 2026
4319046
fix: record radii only in training forwards the kernel skips, split m…
swahtz Sep 24, 2026
c9f2b77
fix: record radii on explicit training intent, crop masks on the worl…
swahtz Sep 24, 2026
59143bc
refactor: move the training-loop and backend work out to feature/trai…
swahtz Sep 24, 2026
769bcba
fix: pass the exact image size to the sparse tile layout
swahtz Oct 6, 2026
956c565
refactor: own ProjectionMethod here instead of re-exporting it from fvdb
swahtz Oct 6, 2026
1b35310
docs: put the crop docs on the methods that take them, state the accu…
swahtz Oct 8, 2026
1543323
Merge upstream main into feature/functional-gaussian-splatting
swahtz Oct 8, 2026
d3e9ff8
docs: correct docstrings that disagreed with the code they describe
swahtz Oct 8, 2026
969228c
docs: fix an over-indented line in the evaluate_gaussian_sh docstring…
swahtz Oct 8, 2026
e5128fb
docs: second pass over docstrings and comments for accuracy and prese…
swahtz Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 10 additions & 7 deletions docs/api/enums.rst
Original file line number Diff line number Diff line change
@@ -1,15 +1,18 @@
Enums
=====

The camera and Gaussian-splatting enums are provided by
``fvdb_reality_capture``. Their member values remain compatible with the
underlying compiled fVDB kernels.
The camera enums that fvdb kernels accept are owned by ``fvdb`` and re-exported by
``fvdb_reality_capture`` as the same objects, so values pass between the two packages without
conversion:

.. autoclass:: fvdb_reality_capture.RollingShutterType
:members:
- :class:`fvdb.CameraModel` (also available as ``fvdb_reality_capture.CameraModel``)
- :class:`fvdb.RollingShutterType` (also available as ``fvdb_reality_capture.RollingShutterType``)

.. autoclass:: fvdb_reality_capture.CameraModel
:members:
:class:`ProjectionMethod` and :class:`GaussianRenderMode` select stages of the composable rendering
pipeline in :mod:`fvdb_reality_capture.functional` and are defined here.

.. autoclass:: fvdb_reality_capture.ProjectionMethod
:members:

.. autoclass:: fvdb_reality_capture.GaussianRenderMode
:members:
142 changes: 142 additions & 0 deletions docs/api/functional.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,142 @@
Functional Gaussian Splatting
=============================

.. module:: fvdb_reality_capture.functional

:mod:`fvdb_reality_capture.functional` exposes Gaussian splat rendering as four composable stages
that pass small frozen dataclasses between them. :class:`~fvdb_reality_capture.GaussianSplat3d`
is a thin wrapper that composes these stages; use them directly when you need to insert your own
logic between projection and rasterization, reuse a projection for several renders, or build a
training loop over plain tensors.

The kernels themselves live in :mod:`fvdb.functional` as flat, non-differentiable forward and
backward functions. The stages here attach autograd to them, so gradients flow through every stage
except tile intersection.

.. code-block:: python

import torch
import fvdb_reality_capture.functional as F
from fvdb_reality_capture import CameraModel, GaussianRenderMode

# Plain tensors: means [N, 3], quats [N, 4], log_scales [N, 3], logit_opacities [N],
# sh0 [N, 1, 3], shN [N, K - 1, 3], world_to_cam [C, 4, 4], K [C, 3, 3]

# Stage 1: project the 3D Gaussians into every camera
projected = F.project_gaussians(
means, quats, log_scales, world_to_cam, K, image_width=640, image_height=480
)

# Stage 2: view-dependent features from spherical harmonics
features = F.evaluate_gaussian_sh(
means, sh0, shN, world_to_cam, projected, render_mode=GaussianRenderMode.FEATURES
)

# Per-camera opacities, computed once and shared by stages 3 and 4
opacities = F.compute_gaussian_opacities(logit_opacities, projected)

# Stage 3: bin Gaussians into image tiles (opacities enable tighter culling)
tiles = F.intersect_gaussian_tiles(projected, opacities)

# Stage 4: alpha-blend into images
images, alphas = F.rasterize_screen_space_gaussians(projected, features, opacities, tiles)

loss = torch.nn.functional.l1_loss(images, target_images)
loss.backward() # gradients reach means, quats, log_scales, logit_opacities, sh0 and shN

The sparse path renders an arbitrary set of pixels per camera. Swap stages 3 and 4 for
:func:`intersect_gaussian_tiles_sparse` and :func:`rasterize_screen_space_gaussians_sparse`; results
come back as :class:`~fvdb.JaggedTensor` in the order of the requested pixels, duplicates included.
The world-space path, :func:`rasterize_world_space_gaussians`, evaluates the 3D Gaussians along
per-pixel rays and is the training path for the unscented projection, whose kernel has no backward
pass.


Types
-----

.. autoclass:: ProjectedGaussians
:members:

.. autoclass:: GaussianTileIntersection
:members:

.. autoclass:: SparseGaussianTileIntersection
:members:


Stage 1: Projection
-------------------

.. autofunction:: project_gaussians

.. autofunction:: resolve_projection_method

.. autofunction:: requires_distortion_coeffs
.. autofunction:: check_distortion_coeffs


Stage 2: Features
-----------------

.. autofunction:: evaluate_gaussian_sh

.. autofunction:: sh_degree_from_coefficients


Stage 3: Tile Intersection
--------------------------

.. autofunction:: intersect_gaussian_tiles

.. autofunction:: intersect_gaussian_tiles_sparse

.. autofunction:: deduplicate_pixels

.. autofunction:: check_tiles_match

.. py:function:: as_pixel_jagged(value)

``fvdb.functional.as_pixel_jagged``, re-exported so the sparse pipeline here applies the same
pixel-selection checks as fvdb's own sparse kernels. Normalizes a ``[C, P, 2]`` tensor or a
:class:`~fvdb.JaggedTensor` of ``(row, col)`` integer pixels to a JaggedTensor with one list per
camera, raising ``TypeError`` for non-integer coordinates and ``ValueError`` for malformed shapes.


Stage 4: Rasterization
----------------------

.. autofunction:: rasterize_screen_space_gaussians

.. autofunction:: rasterize_world_space_gaussians

.. autofunction:: rasterize_screen_space_gaussians_sparse

.. autofunction:: compute_gaussian_opacities

.. py:data:: Crop

``tuple[int, int, int, int]``: a crop window as ``(origin_w, origin_h, width, height)`` in pixels.

.. autofunction:: validate_crop

.. autofunction:: apply_crop
.. autofunction:: apply_pixel_mask

.. autofunction:: pad_crop

.. autofunction:: pixel_mask_to_tile_mask


Analysis
--------

These do not build an autograd graph.

.. autofunction:: rasterize_num_contributing_gaussians

.. autofunction:: rasterize_contributing_gaussian_ids

.. autofunction:: rasterize_num_contributing_gaussians_sparse

.. autofunction:: rasterize_contributing_gaussian_ids_sparse
5 changes: 3 additions & 2 deletions docs/api/gaussian_splatting.rst
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,9 @@ Gaussian Splatting
==================

The high-level Gaussian splatting API is provided by
``fvdb_reality_capture``. The underlying rendering kernels and supporting
tensor types remain in ``fvdb``.
``fvdb_reality_capture``. :class:`GaussianSplat3d` composes the stages of
:mod:`fvdb_reality_capture.functional`; the underlying rendering kernels and
supporting tensor types remain in ``fvdb``.

.. autoclass:: fvdb_reality_capture.ProjectedGaussianSplats
:members:
Expand Down
11 changes: 10 additions & 1 deletion docs/conf.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,16 @@
# Add any Sphinx extension module names here, as strings. They can be
# extensions coming with Sphinx (named 'sphinx.ext.*') or your custom
# ones.
extensions = ["sphinx.ext.autodoc", "sphinx.ext.viewcode", "sphinx.ext.napoleon", "myst_parser"]
extensions = [
"sphinx.ext.autodoc",
"sphinx.ext.intersphinx",
"sphinx.ext.viewcode",
"sphinx.ext.napoleon",
"myst_parser",
]

# fvdb is mocked during the docs build; resolve references to it against its published docs.
intersphinx_mapping = {"fvdb": ("https://fvdb-core.readthedocs.io/latest/", None)}

myst_enable_extensions = [
"amsmath",
Expand Down
1 change: 1 addition & 0 deletions docs/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,7 @@ A common reality capture pipeline typically resembles the figure below:

api/enums
api/gaussian_splatting
api/functional
api/radiance_fields
api/sfm_scene
api/tools
Expand Down
5 changes: 4 additions & 1 deletion fvdb_reality_capture/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,8 @@
tools,
transforms,
)
from .enums import CameraModel, ProjectionMethod, RollingShutterType
from . import functional
from .enums import CameraModel, GaussianRenderMode, ProjectionMethod, RollingShutterType
from .radiance_fields import (
GaussianSplat3d,
ProjectedGaussianSplats,
Expand All @@ -39,6 +40,8 @@
"RollingShutterType",
"CameraModel",
"ProjectionMethod",
"GaussianRenderMode",
"functional",
"checkpoints",
"dev",
"foundation_models",
Expand Down
104 changes: 26 additions & 78 deletions fvdb_reality_capture/enums.py
Original file line number Diff line number Diff line change
@@ -1,99 +1,47 @@
# Copyright Contributors to the OpenVDB Project
# SPDX-License-Identifier: Apache-2.0
#
"""Enums used by the Gaussian splatting API.

from enum import IntEnum

__all__ = ["RollingShutterType", "CameraModel", "ProjectionMethod"]


class RollingShutterType(IntEnum):
"""
Rolling shutter policy for camera projection / ray generation.

Rolling shutter models treat different image rows/columns as having different exposure times.
FVDB uses this to interpolate between per-camera start/end poses when generating rays.
"""

NONE = 0
"""
No rolling shutter: the start pose is used for all pixels.
"""

VERTICAL = 1
"""
Vertical rolling shutter: exposure time varies with image row (y).
"""

HORIZONTAL = 2
"""
Horizontal rolling shutter: exposure time varies with image column (x).
"""


class CameraModel(IntEnum):
"""
Camera model for projection / ray generation.

Notes:
The camera enums that fvdb kernels accept are owned by :mod:`fvdb` and re-exported here unchanged,
so :class:`fvdb_reality_capture.CameraModel` is the same object as :class:`fvdb.CameraModel`.
:class:`ProjectionMethod` and :class:`GaussianRenderMode` select stages of the composable rendering
pipeline in :mod:`fvdb_reality_capture.functional` and are defined here.
"""

- ``PINHOLE`` and ``ORTHOGRAPHIC`` ignore distortion coefficients.
- ``OPENCV_*`` variants use pinhole intrinsics plus OpenCV-style distortion. When distortion
coefficients are provided, FVDB expects a packed layout:
from enum import IntEnum

``[k1,k2,k3,k4,k5,k6,p1,p2,s1,s2,s3,s4]``
from fvdb import CameraModel, RollingShutterType

Unused coefficients for a given model should be set to 0.
"""
__all__ = ["RollingShutterType", "CameraModel", "ProjectionMethod", "GaussianRenderMode"]

PINHOLE = 0
"""
Ideal pinhole camera model (no distortion).
"""

OPENCV_RADTAN_5 = 1
"""
OpenCV radial-tangential distortion with 5 parameters (k1,k2,p1,p2,k3).
"""

OPENCV_RATIONAL_8 = 2
class ProjectionMethod(IntEnum):
"""
OpenCV rational radial-tangential distortion with 8 parameters (k1..k6,p1,p2).
Which fvdb projection kernel :func:`fvdb_reality_capture.functional.project_gaussians` calls.
"""

OPENCV_RADTAN_THIN_PRISM_9 = 3
"""
OpenCV radial-tangential + thin-prism distortion with 9 parameters (k1,k2,p1,p2,k3,s1..s4).
"""
AUTO = 0
"""Choose the default implementation for the selected camera model."""

OPENCV_THIN_PRISM_12 = 4
"""
OpenCV rational radial-tangential + thin-prism distortion with 12 parameters
(k1..k6,p1,p2,s1..s4).
"""
ANALYTIC = 1
"""Use the analytic (EWA) projection path."""

ORTHOGRAPHIC = 5
"""
Orthographic camera model (no distortion).
"""
UNSCENTED = 2
"""Use the unscented-transform projection path."""


class ProjectionMethod(IntEnum):
class GaussianRenderMode(IntEnum):
"""
Projection implementation selector for Gaussian splatting camera models.
Which per-Gaussian features :func:`fvdb_reality_capture.functional.evaluate_gaussian_sh` produces
for rasterization.
"""

AUTO = 0
"""
Choose the default implementation for the selected camera model.
"""
FEATURES = 0
"""Spherical-harmonics evaluated features only, ``[C, N, D]``."""

ANALYTIC = 1
"""
Use the analytic projection path.
"""
DEPTH = 1
"""View-space depth only, ``[C, N, 1]``. No spherical harmonics are evaluated."""

UNSCENTED = 2
"""
Use the unscented projection path.
"""
FEATURES_AND_DEPTH = 2
"""Spherical-harmonics features with depth appended as the last channel, ``[C, N, D + 1]``."""
Loading
Loading