Skip to main content

Health Checks & Metrics

This page gives your operations team concrete checks and the key metrics to watch. It complements the Monitoring & Observability stack overview.

Daily health check (5 minutes)​

Run these commands to confirm the platform is healthy:

# 1. All DuoKey pods Running/Ready and spread across nodes
oc get pods -n duokey -o wide

# 2. Cluster operators healthy
oc get clusteroperators | grep -vi "True.*False.*False"

# 3. Ingress route serving TLS
oc get route -n duokey
curl -I https://cockpit.duokey.example.local

# 4. Secrets being injected (no plaintext on disk)
oc get externalsecret -n duokey

# 5. GitOps in sync
argocd app get duokey | grep -E "Sync Status|Health Status"

# 6. PersistentVolume usage
oc get pvc -A

External / VM tier:

# OpenBao: must be unsealed with an active Raft leader
bao status

# PostgreSQL: primary up, replicas streaming, low lag
oc exec -n duokey <pg-primary-pod> -- psql -c "SELECT client_addr, state, replay_lag FROM pg_stat_replication;"

Key metrics (SLIs)​

These are the signals that matter most. Build Grafana panels and alerts on them.

Application​

MetricWhy it mattersHealthy range
Request rate (req/s)Load / capacityBaseline-dependent
Error rate (5xx %)Reliability< 1%
P95 / P99 latencyUser experienceWithin your SLO
Active sessionsUsageBaseline-dependent

Secret engine (OpenBao / Vault)​

MetricHealthy
Seal statusUnsealed on quorum
Raft leader presentExactly one leader
Request latencyStable, low

PostgreSQL​

MetricHealthy
Replication lag< 30s
Active connectionsBelow max_connections
Disk usage< 85%
Primary availabilityAlways one writable primary

Platform​

MetricHealthy
Node CPU / memoryHeadroom maintained
Pod restartsNo crash loops
PV capacity< 85% used
etcd healthAll members healthy

Suggested SLOs​

Service-level objectiveTarget
API availability≥ 99.9% monthly
API P95 latency≤ your agreed threshold
Successful backups100% of scheduled runs
RTO (disaster recovery)See Backup & DR
RPO (data loss window)See Backup & DR

Synthetic / liveness probes​

  • Kubernetes liveness and readiness probes are configured on every DuoKey pod so OpenShift restarts or removes unhealthy instances automatically.
  • Add an external synthetic check (from your monitoring system or LB health monitor) against the Cockpit URL to detect end-to-end failures.

Alerts​

The recommended alert set is documented in Monitoring & Observability → Recommended alerts. Route alerts to your on-call channel (Alertmanager → email / Slack / PagerDuty / OpsGenie) and to your SIEM for correlation.