Unified Inventory
Every node in one list: state, inline CPU/RAM summary, and a tabbed detail view.
Live Metrics
CPU, RAM, disk, network and InfiniBand time series — per node and cluster-wide.
GPU Monitoring
NVIDIA GPU inventory, utilization and temperature metrics via DCGM.
Health Checks
Service and node health probes, recorded automatically on the unified event timeline.
Alert Rules
Threshold-based rules; SMTP, webhook and in-panel notifications; silences and maintenance windows.
Central Logging
Node logs stream to the head over RELP+mTLS; search from the panel with LogsQL — 90-day retention.

