Zum Hauptinhalt springen

VM-Only Deployment (without OpenShift)

Not every site runs Kubernetes. When you only have a hypervisor (VMware, Proxmox, Nutanix, KVM/OpenStack) and prefer plain virtual machines, DuoKey can be deployed on a VM-only topology — the same components as the OpenShift reference architecture, packaged as containers on VMs (Podman/Docker) instead of an OpenShift cluster.

The stack is small: a React frontend (UI), a single backend — the Cockpit API (serving the DKE / KMS / KMIP / ACME-EST / XKS endpoints together) — PostgreSQL, Redis, OpenBao for secrets, plus GitOps (ArgoCD / GitLab), an optional Harbor registry, VictoriaMetrics / VictoriaLogs for observability, and — where a deployment needs it — the optional DuoKey MPC KMS cluster (3+ nodes) as a dedicated key-custody backend.

Key custody options

Key custody defaults to the built-in Software Vault — no extra cluster to deploy. Add the optional DuoKey MPC KMS (DuoKey's own multi-party-computation cluster; 3+ nodes, key split into shares, no single node ever holds a complete key) or a Securosys HSM when your deployment requires a dedicated or hardware-backed key-custody tier. See Prerequisites → Key custody.

What the optional DuoKey MPC stack does not include

Because the DuoKey MPC KMS is DuoKey's own cluster (not a third-party Sepior TSM), adding it brings no MongoDB, no Audit Node and no KMaaS portal (those are Sepior-only components). The frontend is React, not Angular.

Which model should I choose?
  • OpenShift — recommended for large / enterprise scale, self-healing, autoscaling and GitOps-driven operations. See the Reference Architecture.
  • VM-only — a lighter footprint for pilot, standard, and mid-size deployments, or sites with no Kubernetes platform. You own HA at the VM/LB layer instead of relying on the cluster. This page covers that model.

Both models share the same network requirements, DNS records and security.

Single-site architecture​

The stateless tiers (UI and API) sit behind a load-balancer VIP and run N+1 replicas across separate hypervisor hosts; the stateful tiers run their own HA (PostgreSQL primary+replicas, OpenBao Raft, Redis Sentinel).

Single-site architecture
Users / Office / M365
Load balancer / reverse proxy VIPkeepalived + HAProxy / Caddy / F5
Cockpit UI (React)2+ VMs
Cockpit API3+ VMsDKE · KMS · KMIP
PostgreSQL cluster1 primary + 2 replicas
RedisSentinel
OpenBao3 VMs · Raft
DuoKey MPC KMSoptional · 3+ nodes
Securosys HSMoptional
ArgoCD (GitOps)deploys API
VictoriaMetrics · VictoriaLogsscrape / logs
Edge / proxyUIAPIDataSecrets / key custodyOpsOptional

Stateless UI and API run N+1 behind a load-balancer VIP; the API talks to key custody, PostgreSQL, Redis, and OpenBao, with ArgoCD and observability alongside.

VM roles & sizing​

Each row is the size per VM and the minimum count for a highly available deployment. The figures mirror the per-component sizing in Prerequisites.

RoleCount (HA)vCPU / VMRAM / VMDiskNotes
Load balancer / reverse proxy224 GB20 GBActive/standby VIP (keepalived); HAProxy / Nginx / F5
Cockpit UI (React frontend)224 GB20 GBStateless SPA
Cockpit API backend348 GB20 GBCockpit management + DKE / KMS / KMIP key endpoints — the hot path
DuoKey MPC KMS node (optional)348 GB20 GBOnly if added as your key-custody backend; CPU-bound joint computation — keep same-site
PostgreSQL (Cockpit + optional MPC)3616 GB100 GB SSD1 primary + 2 replicas; one process, separate DBs/users
Redis348 GB50 GBSentinel (or cluster); pass-through cache
OpenBao324 GB20 GB SSDRaft integrated storage
VictoriaMetrics (monitoring)1412 GB200 GBMetrics + Grafana; primary DC only
VictoriaLogs (logging)1412 GB600 GBCentral logs / audit; primary DC only
Harbor (registry, optional)118 GB200 GBAir-gapped image mirror; primary DC only
GitOps — ArgoCD / GitLab (optional)114 GB100 GBReconciles manifests from your Git source
Keep roles off the same physical host

Spread each HA set (PostgreSQL, OpenBao, Redis, the MPC nodes and the API replicas) across separate hypervisor hosts / anti-affinity groups so a single host failure never takes down a quorum or all replicas of a tier.

LIGHT profile assumptions

This is the LIGHT reference sizing, shown with the optional DuoKey MPC KMS included: one PostgreSQL process serves both the Cockpit database and the MPC database when present (separate DBs/users); PostgreSQL is CPU-light (a simple data store — no stored procedures); the DuoKey MPC nodes need less CPU than a Sepior TSM; Redis is a pass-through cache; and logging retention is trimmed. If you stay on the default Software Vault, drop the MPC row entirely. Confirm the final numbers with DuoKey.

Sizing tiers by user count​

Indicative VM counts to make planning easier — not a contractual mapping. Final sizing (document activity, peak concurrency, availability target) is validated with DuoKey.

TierUsersUI VMsAPI VMsMPC nodes (optional)Data tierNotes
Pilot / PoC≤ 1,000220 or 3PG×3, Bao×3, Redis×3Minimum HA, reduced replicas
Standard≤ 10,000230 or 3PG×3, Bao×3, Redis×3Reference VM HA
Large≤ 50,000360 or 3–5PG×3 (scaled), Redis clusterScale the API (+ MPC, if added) first
Enterprise100,000+CustomCustomCustomMulti-siteTypically active/passive across DCs
Indicative figures — confirm with DuoKey

As with the OpenShift sizing tiers, these are starting points. The Cockpit API is the hot path for DKE key operations — scale it first under heavy load. If you've added the optional DuoKey MPC KMS as your key-custody backend, its nodes are CPU-bound the same way and scale alongside the API.

High availability within a site​

TierHA mechanism
Load balancer / reverse proxyTwo VMs with a floating VIP (keepalived/VRRP); active/standby
Cockpit UI / APIStateless, N+1 replicas behind the VIP; LB health-checks remove bad backends
DuoKey MPC KMS (optional)3+ nodes on separate hosts; the cluster tolerates node loss without ever exposing a complete key
PostgreSQL1 primary + 2 replicas with automatic failover (Patroni/repmgr); VIP follows the primary
OpenBao3-node Raft quorum; auto-unseal recommended
RedisSentinel (or cluster) for automatic primary election

Multi-site HA with Geo load balancing​

For datacenter-level resilience, run an independent stack per datacenter across three sites and put a Geo load balancer (GSLB) in front. The GSLB health-checks each site's local LB VIP and steers users to a healthy site; inside each site the local load balancer distributes across the UI and API VMs. PostgreSQL replicates across the three sites, and — where the optional DuoKey MPC KMS is your key-custody backend — it runs as a per-site cluster with one active datacenter at a time (active/passive).

Multi-site HA across three datacenters
Users / Office / M365
Geo load balancer (GSLB)active / passive · health-checked
DC 1 · active
Local LB VIP
UI · Cockpit API
PostgreSQL · Redis · OpenBao
DuoKey MPC cluster (3)optional
Monitoring · Logging · Harborshared ops
DC 2 · passive
Local LB VIP
UI · Cockpit API
PostgreSQL · Redis · OpenBao
DuoKey MPC cluster (3)optional
DC 3 · passive
Local LB VIP
UI · Cockpit API
PostgreSQL · Redis · OpenBao
DuoKey MPC cluster (3)optional
PostgreSQL replication across all sites · sync ≤5 ms RTT, else async

A health-checked GSLB steers users to the active datacenter; each of the three sites runs an independent stack, PostgreSQL replicates across all sites, and the optional DuoKey MPC KMS (if added) runs as a per-site cluster (one active DC at a time).

Topology choices

ModeHow it worksWhen to use
Active / passive (3-DC)One datacenter serves; the other two are warm standbys with replicated PostgreSQL/OpenBao. The GSLB fails traffic over to the next healthy DC.Default for three DCs, especially when inter-site links are slow or long-distance — see the latency guidance
Active / active (metro)Multiple sites serve; PostgreSQL uses synchronous replication over a ≤ 5 ms metro linkOnly on fast, reliable metro links

Design rules

  • MPC placement (if added): where the optional DuoKey MPC KMS is your key-custody backend, run a per-site cluster (one active DC at a time). Only stretch one cluster across sites on a verified low-latency link (≤ 2 ms) — every key op is computed jointly across the nodes, so cross-site placement is slow otherwise. See Network Requirements.
  • Shared observability lives in DC 1: the Monitoring, Logging, Harbor and GitOps VMs run in the primary datacenter only; DC 2 and DC 3 run the core stack. Replicate/back them up off-site.
  • Database replication: synchronous only if primary↔replica RTT is ≤ 5 ms (zero RPO); otherwise asynchronous and accept a small, documented RPO.
  • DKE endpoint: front it with the GSLB and keep the same hostname + TLS certificate across sites; combine with split-horizon DNS so on-net and off-net users reach the nearest healthy site.
  • Securosys HSM (optional): if you add a hardware HSM, give each site low-latency reachability to it — see Key custody.
  • Failover targets: agree an RTO/RPO with DuoKey and rehearse site failover; document which tier fails over automatically vs manually.

Per-datacenter sizing (3-site LIGHT profile)​

Indicative rollup following DuoKey's LIGHT reference profile, shown with the optional DuoKey MPC KMS included. DC 1 additionally carries the shared observability (Monitoring, Logging, Harbor, GitOps); DC 2 and DC 3 run the core stack only. If you stay on the default Software Vault, subtract the MPC nodes from these figures. Confirm final sizing with DuoKey.

DatacenterRAMvCPUStorageContents
DC 1 (active + shared ops)~160 GB~60~1.7 TBCore stack + Monitoring, Logging, Harbor, GitOps
DC 2 (passive)~120 GB~48~580 GBCore stack
DC 3 (passive)~120 GB~48~580 GBCore stack
Overall~400 GB~156~2.9 TBThree-site total
Cockpit HA data flow (per datacenter)
User
Load balancer
Cockpit UI (React)
Cockpit API
Redis cluster
PostgreSQL clusterreplicated across sites

Users reach the load balancer, which fans out to the stateless UI + API VMs in each datacenter; the APIs share the site's Redis and PostgreSQL clusters.

Load balancing details​

ServiceLB typeHealth checkTLSAffinity
DKE key endpoint (Cockpit API)L7 (or L4 pass-through)GET on the key endpoint over TLSPass-through preferred (keep client→API TLS intact)None
Cockpit UI / APIL7HTTP(S) health pathTerminate or pass-throughNone (stateless)
KMIP endpoint (optional)L4TCP 5696Pass-throughNone
PostgreSQL VIPL4Primary-only probe (follows failover)n/aPrimary pinning
Never SSL-inspect the DKE endpoint

A TLS-terminating or SSL-inspecting middlebox on the DKE endpoint breaks the client↔API trust the DKE flow relies on. Use pass-through (L4) there, and set LB idle timeouts longer than the longest key operation.

See also​