Monitoring & Observability
The deployment ships with a lightweight, high-density observability stack so your operations team has full visibility into both the platform and the application.
The stack
OpenShift namespaces
Application pods
metrics · stdout/stderr
vmagentscrape metrics
VictoriaMetricsvmstorage · insert · select
Vectorharvest logs
VictoriaLogs
Grafana
Operations team
vmagent scrapes metrics into VictoriaMetrics and Vector harvests logs into VictoriaLogs; Grafana reads both and serves the operations team.
| Concern | Tool | Role |
|---|---|---|
| Metrics | vmagent → VictoriaMetrics | Scrapes performance endpoints across namespaces into a distributed store (vmstorage, vminsert, vmselect). |
| Logs | Vector → VictoriaLogs | Harvests stdout/stderr and forwards them for structured search. |
| Dashboards | Grafana | Connects to VictoriaMetrics and VictoriaLogs as data sources. |
What to monitor
- Application health — request rate, error rate, and latency of the Cockpit frontend and API backend.
- OpenBao — seal status, Raft leadership, and request throughput. An unsealed, healthy quorum is critical.
- PostgreSQL — replication lag, connections, and disk usage on primary and replicas.
- Redis — cluster health and master/replica failover events.
- Platform — node CPU/memory, pod restarts, and PersistentVolume capacity.
Recommended alerts
| Alert | Condition |
|---|---|
| OpenBao sealed | Any OpenBao node reports sealed = true |
| PostgreSQL replica lag | Lag > 30s for more than 5 min |
| Pod crash loop | Any DuoKey pod restarts > 3 times in 10 min |
| Ingress 5xx | 5xx rate > 1% over 5 min |
| PV nearly full | Any PersistentVolume > 85% used |
| Backup failed | Velero scheduled backup did not complete successfully |
Accessing dashboards
Grafana is exposed via an OpenShift route in the duokey-observability namespace:
oc get route grafana -n duokey-observability
Import the DuoKey dashboards included in your delivery package, or build your own on top of the VictoriaMetrics/VictoriaLogs data sources.