Node Monitoring

Your Cluster's Pulse on One Screen

Instead of hopping between nodes over SSH, see the real state of your cluster instantly with Prometheus-based metrics, health checks and alerts.

SSH-free metric collection

Node agents stream metrics into Prometheus — scales naturally to 100+ nodes.

Observability built into the panel

Charts, alerts and logs inside the panel — without installing and maintaining Grafana.

What's Inside

End-to-End Monitoring from Metric to Alert

From inventory to GPUs, from health checks to central logging — the whole monitoring stack ships with the installation.

Unified Inventory

Every node in one list: state, inline CPU/RAM summary, and a tabbed detail view.

Live Metrics

CPU, RAM, disk, network and InfiniBand time series — per node and cluster-wide.

GPU Monitoring

NVIDIA GPU inventory, utilization and temperature metrics via DCGM.

Health Checks

Service and node health probes, recorded automatically on the unified event timeline.

Alert Rules

Threshold-based rules; SMTP, webhook and in-panel notifications; silences and maintenance windows.

Central Logging

Node logs stream to the head over RELP+mTLS; search from the panel with LogsQL — 90-day retention.

From Inventory to Action

Watching isn't enough — TULPAR lets you act on what you see, from the same screen.

Metrics flow from node_exporter and tulpar-agent into Prometheus; the panel reads via PromQL — none of the scaling limits of SSH-based collection.

Node, job and service events merge into one chronological stream, filterable by type and severity.

Rule → firing → notification → ack; time-boxed silences and maintenance windows keep alert fatigue away.

Drain, resume and reboot straight from the node detail view; the automation engine can bind the same actions to rules.

%