Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions cuvslam-skills/cuvslam-ci/reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,11 +68,13 @@ Repository secrets, split read from write so fork-reachable jobs never hold a ke
## Dataset registry and layout

- `DATASETS` in `tools/python_tools/cuvslam_tools/dataset_registry.py` is the single source of truth. A `DatasetSpec` holds the ID, the preparation module, and its `EvalSpec` records; an `EvalSpec` holds the reporter config filename, the `cuvslam_app` flags, suite membership, and gating. Provisionable means the dataset is present; eval-enabled means it has at least one `EvalSpec`. KITTI, EuRoC, TUM, ICL-NUIM and M3ED-SPOT are eval-enabled; `tartan` and `coda` are provisionable only. Smoke runs KITTI, EuRoC and ICL-NUIM, which covers stereo, stereo-inertial and RGB-D; TUM is full-only because it is the larger RGB-D corpus and ICL-NUIM already covers the modality pre-merge, and M3ED-SPOT is full-only because at 56 GiB and 57k frames in two modes it is the most expensive record in the suite.
- One dataset can carry several records. KITTI, TUM and ICL-NUIM each have a second, full-only `informational` record that replays the same reporter config in `multisensor` mode, so the `MSF` KPI rows compare the cuNLS solver against the default one on identical frames. EuRoC and M3ED-SPOT have none: the cuNLS solver projects through a pinhole camera and only warns on the fisheye and polynomial models those two use, so a record there would report numbers computed from the wrong geometry. `multisensor` needs `USE_CUNLS=ON`, which every CI configuration has.
- The module is standard library only and imports converters lazily, so shell wrappers call it with `PYTHONPATH=tools/python_tools` inside `cuvslam-ci:local` before anything is installed. `datasets_config.sh` wraps it as `dataset_registry`; `run_eval.sh` defines its own shim because the S3 variables `datasets_config.sh` requires are absent in the eval container.
- Subcommands: `validate [--dataset] [--suite]`, `list [--eval] [--suite]`, `eval-records [--suite]` (tab-separated `id`, KPI prefix, config path, flags), `kpi-keys [--suite]`, `prepare-module`, `prepare --root-file`, `verify-staged --root`.
- Suites: `EVAL_SUITE` selects records in `stage_eval_datasets.sh`, `check_eval_prerequisites.sh`, and `run_eval.sh`, and `eval_cuvslam_in_docker.sh` forwards it into the container. Unset means every record; validation requires every `EvalSpec` to belong to `full`, so unset and `full` agree. `validate --suite` additionally rejects a suite that would select nothing.
- `kpi-keys --suite` prints the `<PREFIX>_<METRIC>_<TYPE>_<MODE>` keys a suite can produce, so a consumer can tell a key that is legitimately absent from one that went missing. It lists both `ODOM` and `SLAM` for every record because which of the two a run emits depends on `sequence_title` inside the reporter config, which ships in the tarball. `KPI_METRICS` and `ODOMETRY_MODE_TYPES` in the registry mirror `REQUIRED_METRICS` and `odometry_mode_to_type` in `scripts/cuvslam_kpi_report.py` and have to be edited together; compare them with `python3 -m cuvslam_tools.dataset_registry kpi-keys --suite full` against the keys in `scripts/kpi_baseline_ranges.json` after changing either side.
- Derived, never declared: `<id>.tar`, the staged directory, the `/sequences` mount, and the KPI prefix (first hyphen-delimited token of the config filename, upper-cased). Validation rejects two records that derive the same prefix and `--odometry_mode`.
- The reporter writes to `$CUVSLAM_OUTPUT/<config stem>-<odometry mode>/<timestamp>/`. The mode is part of the directory because the KPI collector reads only the newest run under each directory, so two records sharing one config would otherwise overwrite each other; the prefix is unaffected because it is the first hyphen-delimited token. KPI types are `MCAM`, `MONO`, `VIO`, `RGBD` and `MSF`, one per odometry mode, and an unrecognized mode is rejected rather than filed under another mode's keys.
- Tarball: uncompressed `<id>.tar` at `<S3_DATASETS_BUCKET>/<id>.tar`, whose root is the directory `prepare()` returned. Staged to `<RUNNER_LOCAL_DATASETS_ROOT>/datasets/vslam/<id>/` and mounted read-only into the eval container at `/datasets`. An ETag file skips re-download when the cache is current. After extraction, `verify-staged` checks the shipped config's `dataset_folder` equals `<id>/`.

## KPI outputs
Expand Down
72 changes: 43 additions & 29 deletions scripts/cuvslam_kpi_report.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@


def display_dataset_key(key):
"""TARTAN_FLAKY-STEREO_ODOM -> TARTAN_F-STEREO_ODOM (display only)."""
"""TARTAN_FLAKY-MCAM_ODOM -> TARTAN_F-MCAM_ODOM (display only)."""
for full, short in DATASET_DISPLAY_ALIASES.items():
if key.startswith(full + "-"):
return short + key[len(full):]
Expand Down Expand Up @@ -78,28 +78,43 @@ def parse_all_stats_json(json_path):
return None


# One KPI type per odometry mode. Keep in step with ODOMETRY_MODE_TYPES in
# tools/python_tools/cuvslam_tools/dataset_registry.py, which names the keys a
# suite is expected to produce: a mismatch reads as a missing KPI. A type must
# not contain an underscore, because parse_kpi_key splits keys on it.
ODOMETRY_MODE_TYPES = {
'multicamera': 'MCAM',
'mono': 'MONO',
'inertial': 'VIO',
'rgbd': 'RGBD',
'multisensor': 'MSF',
}


def odometry_mode_to_type(odometry_mode):
"""Convert odometry_mode string to dataset type.

Args:
odometry_mode: String like "OdometryMode.Multicamera". Matching is
case-insensitive, so command-line values like "multicamera" map the
same way. None or non-string values fall back to STEREO.
odometry_mode: String like "OdometryMode.Multicamera". The enum qualifier
is optional and matching is case-insensitive, so command-line values
like "multicamera" map the same way.

Returns:
str: Dataset type (MONO, STEREO, VIO, RGBD)
str: Dataset type (MCAM, MONO, VIO, RGBD, MSF)

Raises:
ValueError: If the mode is unrecognized. Defaulting would file the run
under another mode's KPI keys and overwrite them.
"""
normalized = str(odometry_mode).lower()
if 'multicamera' in normalized:
return 'STEREO'
elif 'mono' in normalized:
return 'MONO'
elif 'inertial' in normalized:
return 'VIO'
elif 'rgbd' in normalized:
return 'RGBD'
else:
return 'STEREO'
# Exact lookup on the unqualified name, not substring containment: a value
# such as "notmultisensor" has to stay unmapped so the caller skips the run
# instead of recording it as MSF.
normalized = str(odometry_mode).rsplit('.', 1)[-1].lower()
if normalized in ODOMETRY_MODE_TYPES:
return ODOMETRY_MODE_TYPES[normalized]
raise ValueError(
f"unknown odometry_mode {odometry_mode!r}; expected one of {', '.join(ODOMETRY_MODE_TYPES)}"
)


def load_baseline_ranges(path):
Expand Down Expand Up @@ -210,20 +225,19 @@ def process_dataset_folder(dataset_folder_path):
print(f'Warning: failed to parse all_stats.json in {stats_folder}')
return None

if all_stats and 'odometry_mode' in all_stats[0]:
# The mode is the only trustworthy source of the KPI type. Guessing it from
# the folder name mislabels every config whose name mentions another mode,
# such as kitti-vio_slam_gt run as multicamera.
if 'odometry_mode' not in all_stats[0]:
print(f'Warning: odometry_mode not found in {all_stats_json}; cannot name this run\'s KPI keys')
return None

try:
dataset_type = odometry_mode_to_type(all_stats[0]['odometry_mode'])
print(f' Detected dataset type: {dataset_type} (from odometry_mode: {all_stats[0]["odometry_mode"]})')
else:
print(f' Warning: odometry_mode not found in JSON, falling back to folder name parsing')
folder_name = os.path.basename(dataset_folder_path).lower()
if 'mono' in folder_name:
dataset_type = 'MONO'
elif 'vio' in folder_name or 'imu' in folder_name:
dataset_type = 'VIO'
elif 'rgbd' in folder_name or 'depth' in folder_name:
dataset_type = 'RGBD'
else:
dataset_type = 'STEREO'
except ValueError as exc:
print(f'Warning: {exc} in {all_stats_json}')
return None
print(f' Detected dataset type: {dataset_type} (from odometry_mode: {all_stats[0]["odometry_mode"]})')

odom_stats = [s for s in all_stats if 'ODOM' in s.get('sequence_title', '').upper()]
slam_stats = [s for s in all_stats if 'SLAM' in s.get('sequence_title', '').upper()]
Expand Down
70 changes: 50 additions & 20 deletions scripts/kpi_baseline_ranges.json
Original file line number Diff line number Diff line change
Expand Up @@ -12,16 +12,16 @@
"trusted nightly kpi_<date>.json into 'expected' below and pick a tolerance."
],
"kpis": {
"KITTI_ATE_STEREO_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_ARE_STEREO_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_Kabsch_STEREO_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_TrackingLosts_STEREO_ODOM": {"expected": null, "tol_abs": 1},
"KITTI_FPS_STEREO_ODOM": {"expected": null, "tol_pct": 15},
"KITTI_ATE_STEREO_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_ARE_STEREO_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_Kabsch_STEREO_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_TrackingLosts_STEREO_SLAM": {"expected": null, "tol_abs": 1},
"KITTI_FPS_STEREO_SLAM": {"expected": null, "tol_pct": 15},
"KITTI_ATE_MCAM_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_ARE_MCAM_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_Kabsch_MCAM_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_TrackingLosts_MCAM_ODOM": {"expected": null, "tol_abs": 1},
"KITTI_FPS_MCAM_ODOM": {"expected": null, "tol_pct": 15},
"KITTI_ATE_MCAM_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_ARE_MCAM_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_Kabsch_MCAM_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_TrackingLosts_MCAM_SLAM": {"expected": null, "tol_abs": 1},
"KITTI_FPS_MCAM_SLAM": {"expected": null, "tol_pct": 15},
"EUROC_ATE_VIO_ODOM": {"expected": null, "tol_pct": 10},
"EUROC_ARE_VIO_ODOM": {"expected": null, "tol_pct": 10},
"EUROC_Kabsch_VIO_ODOM": {"expected": null, "tol_pct": 10},
Expand Down Expand Up @@ -52,15 +52,45 @@
"ICL_NUIM_Kabsch_RGBD_SLAM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_TrackingLosts_RGBD_SLAM": {"expected": null, "tol_abs": 1},
"ICL_NUIM_FPS_RGBD_SLAM": {"expected": null, "tol_pct": 15},
"M3ED_SPOT_ATE_STEREO_ODOM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_ARE_STEREO_ODOM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_Kabsch_STEREO_ODOM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_TrackingLosts_STEREO_ODOM": {"expected": null, "tol_abs": 1},
"M3ED_SPOT_FPS_STEREO_ODOM": {"expected": null, "tol_pct": 15},
"M3ED_SPOT_ATE_STEREO_SLAM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_ARE_STEREO_SLAM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_Kabsch_STEREO_SLAM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_TrackingLosts_STEREO_SLAM": {"expected": null, "tol_abs": 1},
"M3ED_SPOT_FPS_STEREO_SLAM": {"expected": null, "tol_pct": 15}
"M3ED_SPOT_ATE_MCAM_ODOM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_ARE_MCAM_ODOM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_Kabsch_MCAM_ODOM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_TrackingLosts_MCAM_ODOM": {"expected": null, "tol_abs": 1},
"M3ED_SPOT_FPS_MCAM_ODOM": {"expected": null, "tol_pct": 15},
"M3ED_SPOT_ATE_MCAM_SLAM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_ARE_MCAM_SLAM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_Kabsch_MCAM_SLAM": {"expected": null, "tol_pct": 10},
"M3ED_SPOT_TrackingLosts_MCAM_SLAM": {"expected": null, "tol_abs": 1},
"M3ED_SPOT_FPS_MCAM_SLAM": {"expected": null, "tol_pct": 15},
"KITTI_ATE_MSF_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_ARE_MSF_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_Kabsch_MSF_ODOM": {"expected": null, "tol_pct": 10},
"KITTI_TrackingLosts_MSF_ODOM": {"expected": null, "tol_abs": 1},
"KITTI_FPS_MSF_ODOM": {"expected": null, "tol_pct": 15},
"KITTI_ATE_MSF_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_ARE_MSF_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_Kabsch_MSF_SLAM": {"expected": null, "tol_pct": 10},
"KITTI_TrackingLosts_MSF_SLAM": {"expected": null, "tol_abs": 1},
"KITTI_FPS_MSF_SLAM": {"expected": null, "tol_pct": 15},
"TUM_ATE_MSF_ODOM": {"expected": null, "tol_pct": 10},
"TUM_ARE_MSF_ODOM": {"expected": null, "tol_pct": 10},
"TUM_Kabsch_MSF_ODOM": {"expected": null, "tol_pct": 10},
"TUM_TrackingLosts_MSF_ODOM": {"expected": null, "tol_abs": 1},
"TUM_FPS_MSF_ODOM": {"expected": null, "tol_pct": 15},
"TUM_ATE_MSF_SLAM": {"expected": null, "tol_pct": 10},
"TUM_ARE_MSF_SLAM": {"expected": null, "tol_pct": 10},
"TUM_Kabsch_MSF_SLAM": {"expected": null, "tol_pct": 10},
"TUM_TrackingLosts_MSF_SLAM": {"expected": null, "tol_abs": 1},
"TUM_FPS_MSF_SLAM": {"expected": null, "tol_pct": 15},
"ICL_NUIM_ATE_MSF_ODOM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_ARE_MSF_ODOM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_Kabsch_MSF_ODOM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_TrackingLosts_MSF_ODOM": {"expected": null, "tol_abs": 1},
"ICL_NUIM_FPS_MSF_ODOM": {"expected": null, "tol_pct": 15},
"ICL_NUIM_ATE_MSF_SLAM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_ARE_MSF_SLAM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_Kabsch_MSF_SLAM": {"expected": null, "tol_pct": 10},
"ICL_NUIM_TrackingLosts_MSF_SLAM": {"expected": null, "tol_abs": 1},
"ICL_NUIM_FPS_MSF_SLAM": {"expected": null, "tol_pct": 15}
}
}
61 changes: 53 additions & 8 deletions tools/python_tools/cuvslam_tools/dataset_registry.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,16 +61,17 @@
KPI_METRICS = ("ATE", "ARE", "Kabsch", "TrackingLosts", "FPS")
KPI_MODES = ("ODOM", "SLAM")
ODOMETRY_MODE_TYPES = {
"multicamera": "STEREO",
"multicamera": "MCAM",
"mono": "MONO",
"inertial": "VIO",
"rgbd": "RGBD",
"multisensor": "MSF",
}

# Values cuvslam_app accepts for --odometry_mode. cuvslam_kpi_report.py maps each
# to a distinct KPI type, and anything unrecognized there falls back to STEREO,
# so an unknown mode here would silently mislabel a KPI key.
ODOMETRY_MODES = ("multicamera", "mono", "inertial", "rgbd")
# Values cuvslam_app accepts for --odometry_mode, which are exactly the modes
# cuvslam_kpi_report.py can name a KPI type for. It rejects anything else rather
# than filing the run under another mode's keys.
ODOMETRY_MODES = tuple(ODOMETRY_MODE_TYPES)

_DATASET_ID = re.compile(r"^[a-z0-9_]+$")
_KPI_PREFIX = re.compile(r"^[A-Z0-9_]+$")
Expand Down Expand Up @@ -111,8 +112,18 @@ def odometry_mode(self) -> str:

@property
def kpi_type(self) -> str:
"""Dataset type the KPI collector will derive from the odometry mode."""
return ODOMETRY_MODE_TYPES.get(self.odometry_mode, "STEREO")
"""Dataset type the KPI collector will derive from the odometry mode.

Raises:
RegistryError: when the mode has no type, which ``validate`` also
rejects. Falling back to one would name another mode's keys.
"""
try:
return ODOMETRY_MODE_TYPES[self.odometry_mode]
except KeyError:
raise RegistryError(
f"eval '{self.config}': no KPI type for odometry mode {self.odometry_mode!r}"
) from None

def kpi_keys(self) -> tuple[str, ...]:
"""KPI keys this record can produce, as `<PREFIX>_<METRIC>_<TYPE>_<MODE>`.
Expand Down Expand Up @@ -176,6 +187,16 @@ def _rgbd_args() -> tuple[str, ...]:
return ("--odometry_mode=rgbd", "--async_sba=false", "--use_segments")


def _multisensor_args(*extra: str) -> tuple[str, ...]:
"""Flags for the unified multi-sensor mode.

Only registered for rigs the cuNLS solver can model. It projects through a
pinhole camera and merely warns on any other model, so EuRoC (fisheye) and
M3ED-SPOT (polynomial) would report numbers computed from the wrong geometry.
"""
return ("--odometry_mode=multisensor", *extra, "--async_sba=false", "--use_segments")


DATASETS: dict[str, DatasetSpec] = {
"kitti": DatasetSpec(
dataset_id="kitti",
Expand All @@ -186,6 +207,16 @@ def _rgbd_args() -> tuple[str, ...]:
args=_stereo_args("--rectified_stereo_camera=true"),
suites=frozenset(SUITES),
),
# Same config, same frames, same ground truth as the record above, so
# the two KPI rows compare the cuNLS solver against the default one
# directly. The rig carries no depth; multi-sensor qualifies here on
# the overlapping stereo pair alone.
EvalSpec(
config="kitti-vio_slam_gt.cfg",
args=_multisensor_args("--rectified_stereo_camera=true", "--multicam_mode=moderate"),
suites=frozenset({FULL_SUITE}),
gating="informational",
),
),
),
"euroc": DatasetSpec(
Expand Down Expand Up @@ -216,6 +247,12 @@ def _rgbd_args() -> tuple[str, ...]:
args=_rgbd_args(),
suites=frozenset({FULL_SUITE}),
),
EvalSpec(
config="tum-rgbd_slam.cfg",
args=_multisensor_args(),
suites=frozenset({FULL_SUITE}),
gating="informational",
),
),
),
# One config in both suites, not one per suite: a second would derive its own
Expand All @@ -229,6 +266,14 @@ def _rgbd_args() -> tuple[str, ...]:
args=_rgbd_args(),
suites=frozenset(SUITES),
),
# Full only while the mode is experimental, even though the dataset
# is cheap enough for smoke. Promote once the MSF keys are calibrated.
EvalSpec(
config="icl_nuim-rgbd_slam.cfg",
args=_multisensor_args(),
suites=frozenset({FULL_SUITE}),
gating="informational",
),
),
),
# Full only, and by a wide margin the most expensive record: 56 GiB to stage
Expand Down Expand Up @@ -312,7 +357,7 @@ def _validate_eval(spec: DatasetSpec, record: EvalSpec) -> None:
if modes[0] not in ODOMETRY_MODES:
raise RegistryError(
f"{where}: --odometry_mode must be one of {', '.join(ODOMETRY_MODES)}; "
"an unrecognized mode silently becomes a STEREO KPI type"
"the KPI collector has no type for anything else"
)
if not record.suites:
raise RegistryError(f"{where}: must belong to at least one suite")
Expand Down
Loading
Loading