Skip to main content
Each node exposes a per-GPU overview plus time-series charts for GPU, system, network, and InfiniBand metrics. Open a node’s detail page from the Active cluster → Node management tab → click a node.

GPU overview

Each node lists its 8 GPUs with utilization, VRAM, temperature, power, and a per-GPU health badge. See Health checks for what each badge state means.
Node detail page header and the GPU overview, showing all 8 GPUs with utilization, VRAM, temperature, and power

Time range

Pick a time range from 1h, 6h, 12h, 1d, 7d, or 30d. For deeper analysis, open the linked Grafana dashboard.

Metric charts

Metrics are grouped into 6 sections with 25 charts total.

Utilization (2 charts)

Track whether GPUs are actively executing tasks and how much memory each one holds.
Utilization section: GPU Utilization and GPU Memory Used time-series charts

System (7 charts)

Spot bottlenecks outside the GPU — vCPU saturation, memory pressure, disk, network, or the VM’s overall health.
System section (top): CPU Usage, Load Average, System Memory Usage, and Root Disk Usage charts
System section (continued): Network RX, Network TX, and Node Uptime charts

Temperature & power (3 charts)

Watch for thermal throttling and power-related instability.
Temperature & power section: GPU temperature, memory temperature, and power usage charts

Memory & clock detail (3 charts)

Identify throttling and HBM activity issues.
Memory & clock detail section: memory utilization, memory clock, and SM clock charts

ECC & errors (4 charts)

Catch GPU hardware faults early — before they take down a training run.
ECC & errors section: ECC SBE/DBE and remapped rows charts

InfiniBand detail (6 charts)

Monitor the node-to-node fabric used by multi-node distributed training — throughput, errors, and link stability. Each chart is plotted per HCA (Host Channel Adapter).
InfiniBand detail section: per-HCA throughput, receive, and symbol error charts
InfiniBand detail section: IB Link Error Recovery and IB Link Downed charts