Grafana Dashboards
Klio ships a Grafana dashboard template that visualizes the metrics it
exposes. The dashboard is generated from code with the
Grafana Foundation SDK
and committed to the repository at
observability/grafana/klio-dashboard.json.
The dashboard queries the Prometheus export of Klio's OpenTelemetry metrics, so it complements the OpenTelemetry setup rather than replacing it.
What the dashboard shows
The dashboard is a single dashboard split into row sections:
- Client / Plugin — the backup lifecycle as seen by the plugin sidecar running in each PostgreSQL pod: backups in progress, time since the last backup started, succeeded and failed, the latest backup duration, the p50/p90/p99 backup duration distribution, backup run and failure rates, and the backup success ratio. Also the WAL streaming client the sidecar supervises as a child process: the PostgreSQL timeline it is currently streaming and the p50/p90/p99 latency of sending a WAL block to the server.

- Server — the state of the Klio server StatefulSet: uptime, WAL ingest throughput and freshness per tier, the latest written LSN, backup verification totals per cluster, the base snapshot inventory (including file and directory counts), p50/p90/p99 WAL block/get/upload duration by path, stage and tier (each as both a rolling-window and a since-restart total variant), tier-2 relay and maintenance run totals, the retention window of the physical PostgreSQL backups (counts by tier, latest/oldest backup age, start/end LSN and PostgreSQL timeline per cluster), and the embedded NATS JetStream queue.

- WAL replication lag — how far behind Klio is in copying WAL:
- Tier-1 replication lag (bytes) and (seconds): how much WAL, and how much time, Klio is behind the PostgreSQL primary. The write line is WAL Klio has received; the flush line is WAL Klio has safely saved to disk. These two panels need CloudNativePG monitoring (see the prerequisites below).
- Tier-2 archival lag (bytes): how much WAL is on Klio's local disk (tier 1) but not yet copied to remote storage (tier 2).

Some panels need extra context to interpret correctly. Two are derived from the alerting guidance in OpenTelemetry:
- Time since last WAL written surfaces the staleness signal described under Alerting on stalled WAL processing: a stale tier-1 value means PostgreSQL is no longer shipping WALs, while a stale tier-2 value means the remote backend is no longer receiving them.
- Tier-2 archival lag (bytes) plots the LSN difference between tier 1 (local disk) and tier 2 (remote storage). Read together with the staleness panel, it tells a slow pipeline (timestamps advancing, gap growing) apart from a stalled one (timestamps and LSN both frozen).
Prerequisites
The dashboard reads the Prometheus names of Klio's metrics (for example
klio_plugin_backup_runs_total and klio_server_wal_written_total). You
therefore need Prometheus scraping those metrics. Any of the export paths
described in OpenTelemetry works:
- An OpenTelemetry Collector with a Prometheus exporter that Prometheus scrapes.
- The Klio Prometheus exporter (
OTEL_METRICS_EXPORTER=prometheus), scraped directly.
The WAL replication lag row additionally reads CloudNativePG's
cnpg_pg_stat_replication_* metrics. To populate it, scrape the
CloudNativePG cluster monitoring (its PodMonitor) into the same
Prometheus. The rest of the dashboard works without it.
The LSN table panels (Latest backup LSN, Oldest backup LSN, Latest written LSN) need Grafana 13 or later: they use SQL Expressions, available starting in that version. Without it, these three panels show an error instead of a table.
When you route metrics through an OpenTelemetry Collector, enable
resource_to_telemetry_conversion on the Prometheus exporter so that
resource attributes such as the pod and namespace become Prometheus labels.
The sample collector under
operator/config/samples/opentelemetry/base/otel_collector.yaml already does
this.
Importing the dashboard
- In Grafana, go to Dashboards → New → Import.
- Upload
observability/grafana/klio-dashboard.jsonor paste its contents. - When prompted, select your Prometheus data source for the
datasourcevariable.
The dashboard declares a datasource template variable, so it is portable
across Grafana installations and is not tied to a specific data source UID.
Three more template variables at the top filter the panels:
namespace: the Kubernetes namespace, matched againstk8s.namespace.name. It scopes the Client / Plugin and WAL Replication Lag panels, whose metrics are emitted from the PostgreSQL pods and therefore carry the cluster's namespace.server: the Klio server, matched against the OpenTelemetryservice.name(not the pod host name, which two servers of the same name in different namespaces would share). It scopes the Server panels.cluster: the PostgreSQL cluster, matched againstcluster_name. It scopes every per-cluster panel across all sections. Because the server-side metrics carry the server's own namespace, per-cluster Server panels are filtered byserverandclusterrather thannamespace, so a cluster backed up by a server in another namespace is still attributed correctly.
Every aggregation groups by the identifying label (cluster, server, tier), so multiple clusters or servers are never folded into a single misleading value.
Give each Klio server a distinct OTEL_SERVICE_NAME (as the sample server
manifests do). The server variable identifies servers by their OpenTelemetry
service.name; if it is left unset, every server reports the SDK default
(unknown_service:klio) and they collapse into a single, indistinguishable
entry.
Example: kube-prometheus-stack
This example follows the same flow as the CloudNativePG quickstart, using the kube-prometheus-stack chart. Adapt it to your own Prometheus/Grafana if you run a different setup — only the metric prerequisites described above are required.
Install Prometheus and Grafana:
helm repo add prometheus-community \
https://prometheus-community.github.io/helm-charts
helm upgrade --install \
-f https://raw.githubusercontent.com/cloudnative-pg/cloudnative-pg/main/docs/src/samples/monitoring/kube-stack-config.yaml \
prometheus-community prometheus-community/kube-prometheus-stack
Ensure Prometheus scrapes Klio's metrics by deploying a ServiceMonitor (or a
PodMonitor, if the collector's Service has no labels) for the
OpenTelemetry collector's Prometheus exporter — see
operator/config/samples/opentelemetry/base/otel_collector_svc_monitor.yaml.
Port-forward Grafana and log in with admin / prom-operator:
kubectl port-forward svc/prometheus-community-grafana 3000:80
Open http://localhost:3000/ and import
observability/grafana/klio-dashboard.json via Dashboards → New → Import,
selecting your Prometheus data source.
Alternatively, load it automatically through the Grafana dashboard sidecar
with a labeled ConfigMap:
kubectl create configmap klio-grafana-dashboard \
--from-file=klio-dashboard.json=observability/grafana/klio-dashboard.json
kubectl label configmap klio-grafana-dashboard grafana_dashboard=1
Regenerating the dashboard
The committed JSON is generated from the Go program under
observability/grafana/. To regenerate it after changing the generator (or
after a grafana-foundation-sdk bump), run from the repository root:
task grafana:gen
CI runs task grafana:uncommitted, which regenerates the dashboard and
fails if the committed JSON has drifted from the generator output. Commit
the regenerated file whenever it changes.