Aller au contenu principal

DuoKey MPC Cluster — 3 sites, 3 nodes per site

The DuoKey MPC KMS is DuoKey's own key-custody cluster: every key is split into shares across the nodes and each cryptographic operation is computed jointly across them, so no single node ever holds a complete key. For datacenter-level resilience, deploy 3 nodes in each of 3 datacenters — 9 nodes total.

Optional key-custody backend

The DuoKey MPC KMS is an optional vault backend — key custody defaults to the built-in Software Vault. Add the MPC cluster when your deployment needs it. See Prerequisites → Key custody. This is DuoKey's own MPC — no third-party Sepior TSM, and therefore no MongoDB, no Audit Node, no KMaaS portal and no message queue. The nodes are reached by the Cockpit API over OAuth2 and keep their metadata in PostgreSQL.

Topology at a glance​

DuoKey MPC — 9 nodes across 3 datacenters
Cockpit APIOAuth2 · resolve / key op
DC 1
MPC node 1-1
MPC node 1-2
MPC node 1-3
PostgreSQLkey metadata
DC 2
MPC node 2-1
MPC node 2-2
MPC node 2-3
PostgreSQLkey metadata
DC 3
MPC node 3-1
MPC node 3-2
MPC node 3-3
PostgreSQLkey metadata
distributed: joint computation across sites · active/passive: PostgreSQL replication

The Cockpit API drives the MPC cluster; three nodes run in each of the three datacenters, each site backed by a PostgreSQL cluster for key metadata.

Two topologies​

How you connect the 9 nodes depends on how far apart the datacenters are.

1. Distributed cluster — one 9-node cluster​

  • A single logical MPC cluster, with 3 nodes in each datacenter.
  • Every key operation is computed jointly across the participating nodes, so the inter-site link is on the hot path. Only viable on low-latency metro/campus links (≤ 2 ms RTT).
  • Survives a full-site loss (6 of 9 nodes remain across the other two sites), subject to the cluster threshold.
  • Use it when the three datacenters are close together (same metro).

2. Per-site cluster — active/passive (3 × 3 nodes)​

  • Each datacenter runs an independent 3-node cluster; one DC is active at a time.
  • Key metadata is replicated across sites (PostgreSQL); a Geo load balancer / LB fails traffic over to the next healthy datacenter.
  • Key operations stay same-site, so this tolerates higher inter-site latency (regional WAN links).
  • Use it when the datacenters are regionally separated.
Distributed (9-node)Per-site active/passive (3×3)
ClustersOne cluster of 9 nodesThree clusters of 3 nodes
Active sitesAll three participateOne at a time
Inter-site latency≤ 2 ms (on the hot path)Higher OK (replication only)
Survives a site lossYes (threshold permitting)Yes (failover)
Best forSame metro / campusRegionally separated DCs
Which one?

If your three datacenters are on a fast metro ring, the distributed cluster gives the strongest single-cluster resilience. If they are far apart, use per-site active/passive so key operations never cross the WAN. Confirm the choice with DuoKey against your latency measurements — see Network Requirements.

Sizing​

Per the LIGHT profile. Nine MPC nodes total (three per site), plus a PostgreSQL cluster per site for key metadata.

RoleCountvCPU / nodeRAM / nodeDisk
DuoKey MPC node9 (3 × 3)48 GB20 GB
PostgreSQL (key metadata, per site)3 / site616 GB100 GB SSD
MPC nodes are CPU-bound

Each key operation is computed jointly across the nodes, so the MPC tier is the CPU-bound hot path. Scale node vCPU (or add nodes) before other tiers under heavy DKE load.

Network​

FlowSourceDestinationPort
Cockpit API → MPC clusterCockpit APIMPC nodes (per site)443/TCP (OAuth2)
MPC node ↔ nodeMPC nodesMPC nodesper DuoKey release (mTLS/TCP)
MPC node → databaseMPC nodesPostgreSQL5432/TCP
PostgreSQL replicationPrimary ↔ replicasPostgreSQL nodes (cross-site)5432/TCP
  • Latency between MPC nodes is critical — see Network Requirements → Inter-datacenter latency.
  • Keep MPC traffic on a dedicated, hardened VLAN (the secure zone).
  • Synchronize clocks (NTP) across all nodes — a drifting node can be rejected from the joint computation.

High availability​

  • Threshold custody — no single node, and no single administrator, ever holds a complete key. Losing a node does not expose key material.
  • Distributed — losing a whole datacenter leaves the cluster running on the remaining two sites (subject to the threshold).
  • Active/passive — losing the active datacenter fails over to a passive one; key metadata is already replicated.
  • Spread the three nodes in a site across separate hypervisor hosts / anti-affinity groups.

See also​