NVIDIA-accelerated, deep learned model support for object detection including DetectNet.
Isaac ROS Object Detection contains ROS 2 packages to perform object
detection.
isaac_ros_rtdetr, isaac_ros_detectnet, isaac_ros_yolov8, and isaac_ros_grounding_dino each provide a method for spatial
classification using bounding boxes with an input image. Classification
is performed by a GPU-accelerated model of the appropriate architecture:
isaac_ros_rtdetr: RT-DETR modelsisaac_ros_detectnet: DetectNet modelsisaac_ros_yolov8: YOLOv8 modelsisaac_ros_grounding_dino: Grounding DINO models
The output prediction can be used by perception functions to understand the presence and spatial location of an object in an image.
Each Isaac ROS Object Detection package is used in a graph of nodes to provide a bounding box detection array with object classes from an input image. A trained model of the appropriate architecture is required to produce the detection array.
Input images may need to be cropped and resized to maintain the aspect ratio and match the
input resolution of the specific object detection model; image resolution may be reduced to
improve DNN inference performance, which typically scales directly with
the number of pixels in the image. isaac_ros_dnn_image_encoder
provides DNN encoder utilities to process the input image into Tensors for the
object detection models.
Prediction results are decoded in model-specific ways,
often involving clustering and thresholding to group multiple detections
on the same object and reduce spurious detections.
Output is provided as a detection array with object classes.
DNNs have a minimum number of pixels that need to be visible on the object to provide a classification prediction. If a person cannot see the object in the image, it’s unlikely the DNN will. Reducing input resolution to reduce compute may reduce what is detected in the image. For example, a 1920x1080 image containing a distant person occupying 1k pixels (64x16) would have 0.25K pixels (32x8) when downscaled by 1/2 in both X and Y. The DNN may detect the person with the original input image, which provides 1K pixels for the person, and fail to detect the same person in the downscaled resolution, which only provides 0.25K pixels for the person.
Object detection classifies a rectangle of pixels as containing an object, whereas image segmentation provides more information and uses more compute to produce a classification per pixel. Object detection is used to know if, and where in a 2D image, the object exists. If a 3D spacial understanding or size of an object in pixels is required, use image segmentation.
This package uses rosidl::Buffer, a feature built into ROS 2 Lyrical, to
avoid unnecessary copies of large payloads between CPU and accelerator
memory. The CUDA buffer backend builds on this native ROS 2 feature to provide
CUDA memory storage and transport. Most applications can use standard ROS
messages and conversion packages without depending directly on a buffer
backend. See rosidl::Buffer and Buffer Backends for details.
| Sample Graph |
Input Size |
AGX Thor T5000 |
AGX Thor T4000 |
AGX Orin |
Orin Nano Super 8GB |
DGX Spark |
x86_64 w/ RTX 5090 |
x86_64 w/ RTX 5070 |
|---|---|---|---|---|---|---|---|---|
| DetectNet Object Detection Graph |
544p |
216 fps 7.7 ms @ 30Hz |
157 fps 8.4 ms @ 30Hz |
73.5 fps 15 ms @ 30Hz |
30.4 fps 37 ms @ 30Hz |
120 fps 8.9 ms @ 30Hz |
291 fps 4.7 ms @ 30Hz |
166 fps 7.1 ms @ 30Hz |
| Grounding DINO Object Detection Graph |
544p |
25.3 fps |
17.5 fps |
13.6 fps |
– |
17.3 fps |
156 fps 8.0 ms @ 30Hz |
65.6 fps 17 ms @ 30Hz |
| RT-DETR Object Detection Graph SyntheticaDETR |
720p |
195 fps 6.7 ms @ 30Hz |
177 fps 8.3 ms @ 30Hz |
87.3 fps 13 ms @ 30Hz |
40.0 fps 27 ms @ 30Hz |
181 fps 6.8 ms @ 30Hz |
718 fps 2.3 ms @ 30Hz |
374 fps 3.6 ms @ 30Hz |
Please visit the Isaac ROS Documentation to learn how to use this repository.
Update 2026-09-21: Migrated the object detection nodes from NITROS to rosidl::Buffer with the CUDA buffer backend



