Health Checks

Proactive health monitoring identifies issues before they impact workloads through both active diagnostic testing and passive continuous monitoring that automatically remediates common problems.

Passive: GPU Health Monitoring

  • ✓
    DCGM background health checks enabledDCGM diagnostics
  • ✓
    GPUs falling off the bus monitoring (XID 79, NVML_ERROR_GPU_IS_LOST)NVIDIA XID errors
  • ✓
    GPU and CPU memory ECC errors (SBE/DBE volatile and aggregate)

Passive: Network and Hardware Monitoring

  • ✓
    PCIe errors via NVML and DCGM counters (replay/error thresholds)DCGM error injection
  • ✓
    Ethernet and InfiniBand link flaps (ethtool, ibdiagnet, ibportstate)
  • ✓
    GPU temperature monitoring (DCGM_FI_DEV_GPU_TEMP)DCGM field IDs

Passive: Error Detection and Performance

  • ✓
    Uncorrectable NVIDIA XID and SXID error code monitoringNVIDIA XID errors
  • ✓
    Stalled NCCL/RCCL job detection (GPU_UTIL vs power consumption)
  • ✓
    InfiniBand health (PKey consistency, link specs, error counters)

Active: GPU Diagnostics and Performance

  • ✓
    NVIDIA DCGM diag level 3 with Extensive Testing (EUD)DCGM diagnostics
  • ✓
    DtoH and HtoD bandwidth testing for PCIe performance validationnvbandwidth
  • ✓
    gpu-burn/gpu-fryer for validating GPU under loadgpu-burn

Active: Communication and Network Testing

  • ✓
    Local NCCL all reduce tests for NVLink/NVSwitch/NVLS performancenccl-tests
  • ✓
    Local InfiniBand all reduce test (with NCCL_P2P_DISABLE=1)nccl-tests
  • ✓
    Pairwise GPU ib_write_bw and ib_write_latency bidirectional testsRDMA perftest

Active: Hardware Validation and AI Workload Testing

  • ✓
    NVIDIA TinyMeg2 for hardware correctness and SDC-free validation
  • ✓
    Megatron or TorchTitan tests for TFLOP/s/GPU performance and loss convergenceTorchTitan

Automation

  • ✓
    Weekly scheduled active health checks on idle nodes
  • ✓
    NCCL and scheduler topology health validation (Slurm topology.yaml/topology.conf; K8s topology-aware scheduling, gang scheduling, bin packing)
  • ✓
    GPU/node health detection, e.g. Node Problem Detector, GPUd, or equivalent, with automated cordon/drain and repair/replace, e.g. Draino, NVSentinel, or equivalentNode Problem DetectorGPUdDrainoNVSentinel
  • ✓
    NVLink connectivity and error tracking (critical for NVL72)
  • ✓
    AI/ML-based prediction of failures

General Expectations

  • ✓
    Console, dashboard, CLI and/or API available to manage resources
  • ✓
    24x7 support availability
  • ✓
    Process for security fixes and upgrades exists, proactive notifications are clear
  • ✓
    Integration with comprehensive monitoring and alerting systems

All expectations