Troubleshooting
A symptom-driven guide for the most common issues. Start with the daily health check to localize the problem, then jump to the relevant section.
First triage
oc get pods -n duokey -o wide # what is not Running/Ready?
oc describe pod <pod> -n duokey # events at the bottom
oc logs <pod> -n duokey --tail=200 # application errors
Pods not starting
| Symptom | Likely cause | Resolution |
|---|---|---|
Pending | Insufficient resources / no schedulable node | Check oc describe pod; add capacity or relax anti-affinity |
CrashLoopBackOff | App error or missing secret | Check oc logs; verify the ExternalSecret synced |
CreateContainerConfigError | Secret/ConfigMap missing | Confirm ESO produced the Secret (below) |
ImagePullBackOff | Registry unreachable / unsigned image | Check mirror registry and image signature policy |
Secrets not injected
oc get externalsecret -n duokey
oc describe externalsecret <name> -n duokey # see SecretSyncedError reasons
oc get secretstore -n duokey
| Cause | Resolution |
|---|---|
| Auth to secret manager failing | Verify the Kubernetes auth role and ServiceAccount mapping in Vault/OpenBao/CyberArk |
| OpenBao sealed | Unseal the quorum: bao operator unseal |
| Path / key not found | Confirm the secret path in the SecretStore and the source value exists |
See Secret Manager Integration.
Database issues
# Replication health
oc exec -n duokey <pg-primary> -- psql -c "SELECT * FROM pg_stat_replication;"
| Symptom | Cause | Resolution |
|---|---|---|
| Replica lag growing | Network / IO pressure | Check bandwidth and disk IO on replicas |
| No primary / read-only | Failover in progress | Let the operator promote a replica; verify quorum |
| Connection refused | Firewall / max_connections | Check FW rule (5432) and connection pool |
GitOps / ArgoCD
argocd app get duokey
argocd app sync duokey # force reconcile
argocd app diff duokey # see drift
| Symptom | Cause | Resolution |
|---|---|---|
OutOfSync | Manual change / drift | Re-sync; investigate who changed cluster state |
Degraded | Bad manifest / failing pod | Inspect the resource ArgoCD flags as unhealthy |
| Cannot reach repo | GitLab unreachable | Check FW (443/22) and credentials |
Ingress / TLS
| Symptom | Cause | Resolution |
|---|---|---|
| 503 from router | No healthy backend pods | Fix the backend pods first |
| Certificate error | Expired / wrong cert | Renew/replace the route certificate |
| Connection reset | TLS profile / cipher mismatch | Align tlsSecurityProfile and LB/WAF settings (Network Security) |
| Blocked requests | WAF false positive | Tune the WAF rule; check WAF logs |
Performance
- Check key metrics for saturation (CPU/memory, DB connections, latency).
- Confirm pods are spread across nodes/zones (anti-affinity).
- Review HPA / replica counts if request rate exceeds capacity.
Disaster recovery
For full or partial restore procedures, see Backup & Disaster Recovery.
Collecting a support bundle
When escalating to DuoKey support, attach:
oc adm must-gather # cluster-wide diagnostics
oc get events -n duokey --sort-by=.lastTimestamp
oc logs <failing-pod> -n duokey --previous
argocd app get duokey -o yaml
Send to [email protected] with a description of the symptom, timeline, and any recent changes.