Skip to main content

Backup & Disaster Recovery

To deliver a reliable Recovery Time Objective (RTO) and Recovery Point Objective (RPO), the architecture uses a dual-layer disaster-recovery strategy: one layer protects the active OpenShift cluster, the other protects the external VM tier.

Objectives & coverage matrix​

Targets are tuned per customer; the table below is the reference baseline.

ComponentMethodFrequencyRPO (baseline)
PostgreSQLFull + incremental + VM snapshotContinuous WAL + daily full≤ 5 min
OpenBao / secret engineRaft snapshotHourly≤ 1 h
Application state (PVs)Velero CSI snapshotDaily≤ 24 h
K8s manifests / configVelero metadata backupDaily≤ 24 h
GitLab (source)Built-in backupDaily≤ 24 h
HSM materialStays in HSM (replicated per vendor)n/an/a
ScenarioTarget RTO (baseline)
Single pod / node failureAutomatic, seconds (self-healing)
Stateful component restore< 1 hour
Full site failover (secondary cluster)< 4 hours

Disaster-recovery lifecycle​

Disaster-recovery lifecycle
1. ScheduleVelero cron policies
2. Snapshotconsistent volumes + manifests
3. Off-site pushencrypted object storage
4. Replicationto secondary site

Scheduled Velero snapshots are pushed off-site to encrypted object storage and replicated to the secondary site.


Layer 1 — Backup operations (active cluster)​

  • Orchestration — the Velero operator runs scheduled, automated snapshots inside the OpenShift cluster.
  • Data state — Velero coordinates with ODF or the CSI plugin to take crash-consistent block snapshots of the volumes backing stateful workloads.
  • Manifests & metadata — Velero simultaneously captures Kubernetes API objects: namespaces, ServiceAccounts, and ArgoCD applications.
  • Target — all backup metadata and volume payloads are encrypted and written out of the cluster into S3-compatible object storage.
# Example: a daily scheduled backup of the duokey namespace
velero schedule create duokey-daily \
--schedule="0 2 * * *" \
--include-namespaces duokey \
--snapshot-volumes

Layer 2 — Infrastructure backup (external VMs)​

  • OpenBao — backed up independently by triggering automated Raft storage snapshots. The encrypted snapshot files are stored off-node in object storage.
  • GitLab — uses its built-in backup tasks to archive repositories, databases, and configuration to object storage.
  • PostgreSQL — as the most critical data layer, it uses multiple complementary methods: VM snapshots plus scheduled full and incremental database backups.
# OpenBao Raft snapshot (illustrative)
bao operator raft snapshot save openbao-$(date +%F).snap

Recovery procedures (secondary / passive site)​

In the event of a catastrophic failure affecting the active cluster:

  1. Provision the target cluster — activate or initialize a secondary, minimal OpenShift cluster in the recovery site / availability zone.
  2. Restore the secrets engine — the secondary OpenBao instance unseals and loads the replicated Raft snapshot, restoring key capabilities.
  3. Run Velero restore — Velero connects to the replicated object storage in the recovery site and executes a global restore.
  4. Re-inflate volumes — the CSI driver restores block snapshots and binds them to freshly spawned PostgreSQL and Redis StatefulSet pods.
  5. GitOps reconciliation — ArgoCD reconnects to GitLab over HTTPS/SSH, scans the manifests, and reconciles any drift back to compliance.
  6. Restore PostgreSQL — restore from the latest database backup in object storage, or fall back to a full VM snapshot if backups cannot be replayed.
# Restore from the most recent Velero backup
velero restore create --from-backup duokey-daily-<timestamp>

Testing your DR plan​

Don't wait for a real disaster

Rehearse the recovery procedure on a regular cadence (at least quarterly). Verify RTO/RPO targets are actually met and that runbooks are current.

ItemFrequency
Verify backups completedDaily (automated alert)
Restore test (single component)Monthly
Full DR failover rehearsalQuarterly
Review RTO/RPO targetsAnnually