Databricks
Guided onboarding of format-preserving tokenization into a Databricks Lakehouse workspace.
Overview
The Databricks integration is a guided onboarding path into DuoKey's format-preserving tokenization service (see Tokenization for the service itself). Rather than protecting a whole disk or VM, it protects individual columns in a Lakehouse table: deploying it generates ready-to-run Spark SQL that registers tokenize / detokenize Unity Catalog user-defined functions (UDFs). Those UDFs call DuoKey's tokenization service over HTTPS at query time — the key-encryption-key behind tokenization never leaves DuoKey and is never handled by the Databricks cluster.
The UDF is a thin HTTPS client; DuoKey resolves the tenant's tokenization key and performs the format-preserving encryption itself.
Databricks appears both as its own entry in the Apps catalog (this page) and as one of the three variants under the Tokenization integration. Both onboard the same underlying tokenize / detokenize service and produce equivalent Spark UDFs — pick whichever path you land on; you do not need to configure both for the same workspace.
Configuration
| Field | Purpose |
|---|---|
| Workspace URL | The Databricks workspace to onboard (for example `https://<workspace-id>.cloud.databricks.com`). Validated at save time to ensure it does not resolve to a private/internal address. |
| Unity Catalog catalog | The catalog the tokenize/detokenize UDFs are registered into. |
| Unity Catalog schema | The schema (database) the UDFs are registered into. |
| Warehouse / cluster id | Optional, informational — the SQL warehouse or cluster expected to run the UDF registration DDL. |
| Data-type hint | The default data type (for example `creditcard`, `ssn`, `email`) used to pick the tokenization alphabet and domain-separation tweak for this deployment. |
| Vault | Optional. Recorded against the integration for reference — the tokenize/detokenize key itself is derived deterministically per tenant (see "What happens on a tokenize call" below), not stored in or served from this vault. |
DuoKey checks that the workspace URL you provide does not resolve to a private or internal network address before storing it, as a guard against misconfiguration pointing the integration at an internal target instead of your actual Databricks workspace.
Deploying the integration
Provide workspace details
Enter the workspace URL, target catalog and schema, and the data-type hint for the columns you plan to tokenize.
Review the generated UDF SQL
DuoKey returns ready-to-run SQL that creates the catalog/schema if needed and registers tokenize and detokenize UDFs pointed at DuoKey's tokenization endpoint.
Store a DuoKey API token as a Databricks secret
Create a Databricks secret holding a DuoKey API token for a user with tokenization access, and expose it to Spark via cluster configuration.
Run the UDF registration SQL
Execute the generated DDL in the target catalog/schema.
Verify end-to-end
Run a query that tokenizes then detokenizes a sample value and confirm you get the original value back, in the same format.
databricks secrets create-scope dke
databricks secrets put-secret dke api_tokenspark.conf.set("dke.api_token", dbutils.secrets.get("dke", "api_token"))CREATE CATALOG IF NOT EXISTS <catalog>;
USE CATALOG <catalog>;
CREATE SCHEMA IF NOT EXISTS <schema>;
CREATE OR REPLACE FUNCTION <catalog>.<schema>.tokenize(input STRING)
RETURNS STRING LANGUAGE PYTHON AS $$
-- calls DuoKey's tokenization endpoint over HTTPS,
-- authenticated with the token stored above
$$;
CREATE OR REPLACE FUNCTION <catalog>.<schema>.detokenize(input STRING)
RETURNS STRING LANGUAGE PYTHON AS $$
-- calls DuoKey's detokenization endpoint over HTTPS
$$;SELECT <catalog>.<schema>.detokenize(<catalog>.<schema>.tokenize('4111111111111111'));What happens on a tokenize call
The four steps below are what runs end to end for a single tokenize (or detokenize) call from a Spark query, once the UDFs are registered:
The same four steps run in reverse for detokenize. See Tokenization for the full FF1 mechanics inside step 3.
Enable, disable, health and self-test
Once deployed, the integration can be enabled or disabled, and exposes a health check and a self-test. Both report the integration's configuration and readiness — including that the UDF template is in place and that the key-encryption-key stays in DuoKey, never in Databricks — rather than probing the Databricks workspace live. For a live connectivity round-trip against a linked vault, plus a live workspace-reachability probe, deploy through the Tokenization page's Databricks variant instead — its health check and self-test exercise both live.
Tokenized values keep the same length and character shape as the original — a 16-digit card number tokenizes to another 16-digit sequence, an email keeps its @ and domain separators — so downstream Databricks queries, joins and format validation keep working against tokenized columns.
Removing the Databricks integration disables the UDFs but does not retroactively detokenize any data already written with tokenized values. Detokenize any data you need in the clear before removing the integration, or keep the underlying tokenization service available.
Prerequisites
Prérequis
- A Databricks workspace with Unity Catalog enabled
- Permission to create catalogs/schemas and register UDFs in the target workspace
- A DuoKey API token for a user with tokenization access, stored as a Databricks secret
- A DuoKey vault available for the integration's connectivity checks