Skip to main content

Troubleshooting

A symptom-driven guide for the most common issues. Start with the daily health check to localize the problem, then jump to the relevant section.

First triage
oc get pods -n duokey -o wide # what is not Running/Ready?
oc describe pod <pod> -n duokey # events at the bottom
oc logs <pod> -n duokey --tail=200 # application errors

Pods not starting​

SymptomLikely causeResolution
PendingInsufficient resources / no schedulable nodeCheck oc describe pod; add capacity or relax anti-affinity
CrashLoopBackOffApp error or missing secretCheck oc logs; verify the ExternalSecret synced
CreateContainerConfigErrorSecret/ConfigMap missingConfirm ESO produced the Secret (below)
ImagePullBackOffRegistry unreachable / unsigned imageCheck mirror registry and image signature policy

Secrets not injected​

oc get externalsecret -n duokey
oc describe externalsecret <name> -n duokey # see SecretSyncedError reasons
oc get secretstore -n duokey
CauseResolution
Auth to secret manager failingVerify the Kubernetes auth role and ServiceAccount mapping in Vault/OpenBao/CyberArk
OpenBao sealedUnseal the quorum: bao operator unseal
Path / key not foundConfirm the secret path in the SecretStore and the source value exists

See Secret Manager Integration.

Database issues​

# Replication health
oc exec -n duokey <pg-primary> -- psql -c "SELECT * FROM pg_stat_replication;"
SymptomCauseResolution
Replica lag growingNetwork / IO pressureCheck bandwidth and disk IO on replicas
No primary / read-onlyFailover in progressLet the operator promote a replica; verify quorum
Connection refusedFirewall / max_connectionsCheck FW rule (5432) and connection pool

GitOps / ArgoCD​

argocd app get duokey
argocd app sync duokey # force reconcile
argocd app diff duokey # see drift
SymptomCauseResolution
OutOfSyncManual change / driftRe-sync; investigate who changed cluster state
DegradedBad manifest / failing podInspect the resource ArgoCD flags as unhealthy
Cannot reach repoGitLab unreachableCheck FW (443/22) and credentials

Ingress / TLS​

SymptomCauseResolution
503 from routerNo healthy backend podsFix the backend pods first
Certificate errorExpired / wrong certRenew/replace the route certificate
Connection resetTLS profile / cipher mismatchAlign tlsSecurityProfile and LB/WAF settings (Network Security)
Blocked requestsWAF false positiveTune the WAF rule; check WAF logs

Performance​

  • Check key metrics for saturation (CPU/memory, DB connections, latency).
  • Confirm pods are spread across nodes/zones (anti-affinity).
  • Review HPA / replica counts if request rate exceeds capacity.

Disaster recovery​

For full or partial restore procedures, see Backup & Disaster Recovery.

Collecting a support bundle​

When escalating to DuoKey support, attach:

oc adm must-gather # cluster-wide diagnostics
oc get events -n duokey --sort-by=.lastTimestamp
oc logs <failing-pod> -n duokey --previous
argocd app get duokey -o yaml

Send to [email protected] with a description of the symptom, timeline, and any recent changes.