Unified Cluster Dashboard

The Whole Cluster on One Screen

Node health, workload, resource consumption and open alerts converge on a single board — it refreshes itself over WebSocket, so you never have to go looking for a problem.

A board that refreshes itself

Metrics and states stream in live; no page reloads, no stale numbers.

Problems surface first

Nodes with failing health checks and firing alerts are collected at the top of the board.

What's Inside

Status at a Glance, Depth on a Click

The dashboard doesn't just summarize — every card is a door into the management screen behind it.

Cluster Status

Node status distribution — ready, installing, down and needs-attention — summarized in one chart.

Live Time Series

CPU, memory and load curves; pick a node to drill into a single machine's resource profile.

Workload Summary

Running, pending and completed job counts read live from Slurm.

Open Alerts

Firing alerts and failing health checks are listed with a link straight to the node concerned.

Quick Access

One-click cards into inventory, provisioning, monitoring, Slurm and alerting.

Role-Aware View

RBAC scope applies to the dashboard too — everyone sees only the data they are entitled to.

The Data Path Behind the Board

Every number on the board has a known origin; this is a mirror of the live system, not a showcase.

Dashboard data streams over WebSocket rather than periodic polling; if the connection is blocked it falls back to HTTP polling automatically.

Slurm (slurmrestd), Prometheus metrics, inventory and the alerting engine converge on the same screen — no tab-switching between systems.

Health checks are broken down per node; which machine failed, when and why is on the record.

Cluster events collect into a central timeline; findings from the nightly audits are written here too.

%