We are seeing DCGM_FR_IMEX_UNHEALTHY being set by NVSentinel (v1.18) via the DCGM monitor on DRA enabled (MNNVL) clusters.
For context, nvbandwidth workloads are successfully deployed on this cluster, which validates that compute domain pods (and IMEX channels) are forming correctly. I wonder if the DCGM monitor is polling nvidia-imex service running on the hosts, and whether these are false positives as IMEX management is done by ephemeral compute domain (DRA) pods?
@lalith, the NVSentinel maintainer confirmed this to be a known issue.
We are seeing
DCGM_FR_IMEX_UNHEALTHYbeing set by NVSentinel (v1.18) via the DCGM monitor on DRA enabled (MNNVL) clusters.For context, nvbandwidth workloads are successfully deployed on this cluster, which validates that compute domain pods (and IMEX channels) are forming correctly. I wonder if the DCGM monitor is polling nvidia-imex service running on the hosts, and whether these are false positives as IMEX management is done by ephemeral compute domain (DRA) pods?
@lalith, the NVSentinel maintainer confirmed this to be a known issue.