GPU overview
Each node lists its 8 GPUs with utilization, VRAM, temperature, power, and a per-GPU health badge. See Health checks for what each badge state means.
Time range
Pick a time range from 1h, 6h, 12h, 1d, 7d, or 30d. For deeper analysis, open the linked Grafana dashboard.Metric charts
Metrics are grouped into 6 sections with 25 charts total.Utilization (2 charts)
Track whether GPUs are actively executing tasks and how much memory each one holds.
System (7 charts)
Spot bottlenecks outside the GPU — vCPU saturation, memory pressure, disk, network, or the VM’s overall health.

Temperature & power (3 charts)
Watch for thermal throttling and power-related instability.
Memory & clock detail (3 charts)
Identify throttling and HBM activity issues.
ECC & errors (4 charts)
Catch GPU hardware faults early — before they take down a training run.
InfiniBand detail (6 charts)
Monitor the node-to-node fabric used by multi-node distributed training — throughput, errors, and link stability. Each chart is plotted per HCA (Host Channel Adapter).

