Skip to main content

Monitoring & Observability

The deployment ships with a lightweight, high-density observability stack so your operations team has full visibility into both the platform and the application.

The stack​

Observability stack
OpenShift namespaces
Application pods
metrics · stdout/stderr
vmagentscrape metrics
VictoriaMetricsvmstorage · insert · select
Vectorharvest logs
VictoriaLogs
Grafana
Operations team

vmagent scrapes metrics into VictoriaMetrics and Vector harvests logs into VictoriaLogs; Grafana reads both and serves the operations team.

ConcernToolRole
Metricsvmagent → VictoriaMetricsScrapes performance endpoints across namespaces into a distributed store (vmstorage, vminsert, vmselect).
LogsVector → VictoriaLogsHarvests stdout/stderr and forwards them for structured search.
DashboardsGrafanaConnects to VictoriaMetrics and VictoriaLogs as data sources.

What to monitor​

  • Application health — request rate, error rate, and latency of the Cockpit frontend and API backend.
  • OpenBao — seal status, Raft leadership, and request throughput. An unsealed, healthy quorum is critical.
  • PostgreSQL — replication lag, connections, and disk usage on primary and replicas.
  • Redis — cluster health and master/replica failover events.
  • Platform — node CPU/memory, pod restarts, and PersistentVolume capacity.
AlertCondition
OpenBao sealedAny OpenBao node reports sealed = true
PostgreSQL replica lagLag > 30s for more than 5 min
Pod crash loopAny DuoKey pod restarts > 3 times in 10 min
Ingress 5xx5xx rate > 1% over 5 min
PV nearly fullAny PersistentVolume > 85% used
Backup failedVelero scheduled backup did not complete successfully

Accessing dashboards​

Grafana is exposed via an OpenShift route in the duokey-observability namespace:

oc get route grafana -n duokey-observability

Import the DuoKey dashboards included in your delivery package, or build your own on top of the VictoriaMetrics/VictoriaLogs data sources.