Cassandra TDE
Cluster-wide SSTable and commitlog encryption keyed by DuoKey over KMIP — with a rotation design built for a replicated, peer-to-peer cluster.
Overview
Apache Cassandra can encrypt its SSTables and commitlog segments through a pluggable key-provider extension point (the same one DataStax Enterprise's Transparent Data Encryption and third-party KMIP key-provider jars use). This app configures that key provider to fetch its key from the DuoKey Cockpit over KMIP, and generates the cassandra.yaml and CQL needed to bring a keyspace under encryption.
Every node loads the same key_alias from the same KMIP server, so streamed data and commitlog replay stay decryptable across the cluster.
| Property | Value |
|---|---|
| Encryption model | Engine-native TDE (SSTable + commitlog, per-node) |
| Key transport | KMIP-backed key provider (fixed provider — always DuoKey KMIP) |
| Onboarding | App detail page — no wizard entry |
| Proof status | Generated artifacts (config, CQL, rotation runbook) are real and directly usable; no live cluster driver session yet — see below |
Why Cassandra is different from the other TDE engines
Every other TDE engine in this catalog talks to what is effectively a single logical instance — even Percona PostgreSQL with read replicas has all replicas pulling from the same KMIP principal key at once. Rotating the key there means: create a new key, point the engine at it, done.
Apache Cassandra is a peer-to-peer, replicated cluster. With replication factor N, every row lives on N nodes, and nodes constantly exchange raw data with each other — streaming during bootstrap/decommission/repair, hinted handoff, and each node replaying its own commitlog on restart. SSTable and commitlog encryption happens locally, per node, using whatever key that node's own cassandra.yaml currently points at. Because it is the same symmetric key, not a per-node key wrapped by a KEK, the key configured across the cluster must be byte-identical for streamed data and commitlog replay to work.
Creating a new key and deactivating the old one immediately, then rolling nodes one at a time, strands every node that hasn't yet picked up the new config: it can no longer decrypt its own commitlog or data streamed from nodes still on the old key. There is always a window where node A has the new key and node B does not — and repair, streaming, and hinted handoff must keep working throughout that window.
Configuration
| Field | Purpose | |
|---|---|---|
contact_points | native_port | Seed nodes and native-protocol port (default 9042) used to reach the cluster. |
cluster_name | Must match `cluster_name` in `cassandra.yaml`. | |
keyspace | Keyspace this endpoint manages encryption for. | |
replication_strategy | `SimpleStrategy` (single datacenter) or `NetworkTopologyStrategy` (multi-datacenter). | |
replication_factor | Replication factor for `SimpleStrategy`. | |
datacenter_replication | Per-datacenter replication factors (dc_name:rf) for `NetworkTopologyStrategy` — every node in every listed datacenter must apply the same principal key. | |
sstable_encryption | commitlog_encryption | Toggle SSTable-level and commitlog segment encryption independently. |
linked_key_id | The DuoKey vault key backing the cluster-wide principal key. |
Deployment artifacts
# every node in the cluster MUST use the SAME key_alias so
# streamed/replicated data stays decryptable cluster-wide.
transparent_data_encryption_options:
enabled: true
cipher: AES/CBC/PKCS5Padding
key_alias: <app-name>
key_provider:
- class_name: org.apache.cassandra.security.KmipKeyProviderFactory
parameters:
- kmip_host: <dke-cockpit-kmip-host>
kmip_port: 5696
key_label: <app-name>CREATE KEYSPACE IF NOT EXISTS "<keyspace>"
WITH replication = {'class': 'NetworkTopologyStrategy', 'dc1': 3, 'dc2': 2};
-- or, for SimpleStrategy:
-- WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 3};ALTER TABLE <keyspace>.<table> WITH compression = {
'sstable_compression': 'EncryptingLZ4Compressor',
'cipher_algorithm': 'AES/CBC/PKCS5Padding',
'secret_key_strength': 128,
'key_provider': 'KmipKeyProviderFactory'
};commitlog_compression:
- class_name: EncryptingLZ4Compressor
parameters:
cipher_algorithm: AES/CBC/PKCS5Padding
secret_key_strength: 128Replication-aware key rotation
Because the cluster-wide key must stay byte-identical during a rolling rollout, rotation is a five-phase runbook, not an instant cutover. The invariant that holds it together: the old key is never destroyed as a side effect of creating the new one — it stays active, then moves to deactivated (decrypt-only), and is only ever destroyed by a separate, explicitly authorized operator action after a retention window.
The old key is only ever destroyed in phase 5 — an explicit, manual, audited action taken after the retention window.
Create the new key — dual-active window opens
A new principal key is created and activated on the KMIP server. The old key stays active. Both keys are simultaneously servable, so no in-flight decrypt request can fail regardless of which key a given node still has configured.
Rolling per-node rollout
Update the key_alias in cassandra.yaml on one node, restart it, wait for it to rejoin (UN in nodetool status), confirm via the health-check endpoint, then move to the next node — never more than one node per rack/datacenter at a time, the same quorum-preserving rule as any other Cassandra rolling restart.
Cluster-wide confirmation
Poll the health check / self-test until every contact point reports the new key_alias. No node is assumed migrated without a positive health signal.
Deactivate the old key
Once every node confirms, the old key moves to deactivated (decrypt-only, never destroyed). New writes only ever use the new key from this point; existing SSTables and commitlog segments still encrypted under the old key remain decryptable until compaction rewrites them under the new key.
Destroy after a retention window
Only after min_old_key_retention_days (default 90 — chosen to comfortably exceed a typical gc_grace_seconds / major-compaction cycle) and an operator has manually verified via nodetool/sstablemetadata that no SSTable still references the old key_alias, the old key is destroyed. This step is deliberately manual and audited — it is never automated.
Destroying a key that any on-disk SSTable still references causes permanent, unrecoverable data loss. Phase 5 is intentionally never automated, regardless of future automation of phases 2–3.
What is real vs. what is a runbook today
| Capability | Status |
|---|---|
| cassandra.yaml key-provider snippet, CQL keyspace/table statements, rotation plan | Generated for real — an operator can act on these directly |
| Per-node rollout (phases 2–3): config push, restart, nodetool status polling | Not yet automated — today this is an operator runbook, not a driven state machine; the rotation plan reports no per-node confirmation status rather than fabricating it |
| Key destruction (phase 5) | Never automated, by design — a safety property, not a maturity gap |
Cassandra TDE shares its current fidelity level with MySQL TDE, MongoDB CSFLE and SQL Always Encrypted: generated setup instructions and a simulated health/self-test rather than a live database session. Automating the rolling rollout end-to-end is a natural next step, requiring either an on-host agent or a native-protocol driver with cluster admin credentials.