Health Checks & Metrics
This page gives your operations team concrete checks and the key metrics to watch. It complements the Monitoring & Observability stack overview.
Daily health check (5 minutes)
Run these commands to confirm the platform is healthy:
# 1. All DuoKey pods Running/Ready and spread across nodes
oc get pods -n duokey -o wide
# 2. Cluster operators healthy
oc get clusteroperators | grep -vi "True.*False.*False"
# 3. Ingress route serving TLS
oc get route -n duokey
curl -I https://cockpit.duokey.example.local
# 4. Secrets being injected (no plaintext on disk)
oc get externalsecret -n duokey
# 5. GitOps in sync
argocd app get duokey | grep -E "Sync Status|Health Status"
# 6. PersistentVolume usage
oc get pvc -A
External / VM tier:
# OpenBao: must be unsealed with an active Raft leader
bao status
# PostgreSQL: primary up, replicas streaming, low lag
oc exec -n duokey <pg-primary-pod> -- psql -c "SELECT client_addr, state, replay_lag FROM pg_stat_replication;"
Key metrics (SLIs)
These are the signals that matter most. Build Grafana panels and alerts on them.
Application
| Metric | Why it matters | Healthy range |
|---|---|---|
| Request rate (req/s) | Load / capacity | Baseline-dependent |
| Error rate (5xx %) | Reliability | < 1% |
| P95 / P99 latency | User experience | Within your SLO |
| Active sessions | Usage | Baseline-dependent |
Secret engine (OpenBao / Vault)
| Metric | Healthy |
|---|---|
| Seal status | Unsealed on quorum |
| Raft leader present | Exactly one leader |
| Request latency | Stable, low |
PostgreSQL
| Metric | Healthy |
|---|---|
| Replication lag | < 30s |
| Active connections | Below max_connections |
| Disk usage | < 85% |
| Primary availability | Always one writable primary |
Platform
| Metric | Healthy |
|---|---|
| Node CPU / memory | Headroom maintained |
| Pod restarts | No crash loops |
| PV capacity | < 85% used |
| etcd health | All members healthy |
Suggested SLOs
| Service-level objective | Target |
|---|---|
| API availability | ≥ 99.9% monthly |
| API P95 latency | ≤ your agreed threshold |
| Successful backups | 100% of scheduled runs |
| RTO (disaster recovery) | See Backup & DR |
| RPO (data loss window) | See Backup & DR |
Synthetic / liveness probes
- Kubernetes liveness and readiness probes are configured on every DuoKey pod so OpenShift restarts or removes unhealthy instances automatically.
- Add an external synthetic check (from your monitoring system or LB health monitor) against the Cockpit URL to detect end-to-end failures.
Alerts
The recommended alert set is documented in Monitoring & Observability → Recommended alerts. Route alerts to your on-call channel (Alertmanager → email / Slack / PagerDuty / OpsGenie) and to your SIEM for correlation.