Tokenization
Format-preserving tokenization for sensitive columns and fields, exposed as a single service with three onboarding paths.
Overview
Tokenization replaces a sensitive value — a card number, a national ID, an email — with a token that has the same length and character shape, so downstream systems keep working (queries, joins, format validation) without ever seeing the original value. DuoKey's tokenization service performs format-preserving encryption (FF1, per NIST SP 800-38G): tokenizing and detokenizing are true inverse operations under a tokenization key, not a lookup table, so there is no separate token vault to manage or synchronize.
Three onboarding paths front the same service:
Databricks
Spark / Unity-Catalog UDFs that tokenize and detokenize columns at query time.
Snowflake
External functions that tokenize and detokenize columns from Snowflake SQL.
Tokenization Secret
A standalone tokenization key/secret in a linked vault, for applications that call the service directly.
The Databricks variant is documented in full, including generated setup SQL, on the Databricks page. This page covers the tokenization service itself, plus the Snowflake and generic-secret variants.
Every variant calls the same tokenize / detokenize endpoints and the same FF1 engine — only the setup artifacts (UDFs, external functions, or a direct secret) differ.
How tokenization works
Format-preserving
A token has the same length and character set as the original value: a 16-digit card number tokenizes to another 16-digit sequence, an email keeps its @ and dot separators.
Data-type aware
A data-type hint (creditcard, ssn, phone, email, …) selects the right character alphabet automatically, or you can pin an explicit format.
Domain-separated
The data type doubles as part of the cryptographic tweak, so the same plaintext value tokenizes differently in two different columns/domains — a card number and a lookalike account number never collide.
No stored token vault
Tokenization keys are derived deterministically per tenant and per named key, so detokenizing a value never depends on retrieving previously stored key material. Rotating protection for a field is a matter of choosing a new key name going forward.
| Data-type hint | Alphabet used |
|---|---|
| creditcard, ssn, phone, account, numeric | Digits only (0–9) — preserves numeric formats like card numbers and phone numbers. |
| email, iban, name, custom / anything else | Mixed-case alphanumeric — preserves separators such as @ and . while tokenizing the rest. |
If the default alphabet inferred from the data-type hint is not what you need, you can pin an explicit format (digits, alphanumeric, upper-alpha, lower-alpha, and variants) instead of relying on the data-type inference.
Named tokenization keys
Every tokenization call names a key (defaults to "default") that is scoped to your tenant. Different named keys are cryptographically independent, which gives you two practical controls:
- Segmentation — use different named keys for different applications or environments so their tokenized values are never comparable to each other.
- Rotation — start tokenizing new values under a new key name to rotate protection for a field going forward, without needing to re-tokenize existing data on a deadline.
A value tokenized under one key name can only be detokenized by requesting the same key name. Keep a clear mapping of which key name protects which column/field, especially if you introduce a new key name for rotation.
How an FPE tokenize call works
This is the real sequence behind every tokenize / detokenize call, whichever variant it arrives from — the same steps the Databricks page's sequence diagram summarizes as "DuoKey resolves the key and performs FPE":
Detokenize runs the identical sequence with the decrypt operation in step 4 — FF1 is a true inverse, not a lookup.
The three variants
| Variant | Consumed from | Distinguishing config |
|---|---|---|
| Databricks | Spark / Unity-Catalog UDFs | Workspace URL, catalog, schema — see the dedicated Databricks page. |
| Snowflake | External functions | Snowflake account, region, warehouse, database and schema, plus an API integration binding Snowflake to DuoKey. |
| Tokenization Secret | Direct API calls from your own application | A named secret created in a linked vault, with its own key type, size and exportability setting. |
Deploying any variant links a DuoKey vault (and optionally a key inside it) used for connectivity health checks, self-tests and audit; the vault and key you choose are recorded against the integration, but do not need to be re-selected per tokenize/detokenize call — those are authorized purely by API permission and the named key you pass.
Snowflake variant
Deploying the Snowflake variant generates the SQL to create a Snowflake API integration pointed at DuoKey's tokenization endpoints, and TOKENIZE / DETOKENIZE external functions in your chosen database and schema that call it.
CREATE OR REPLACE API INTEGRATION <integration-name>
API_PROVIDER = aws_api_gateway
API_ALLOWED_PREFIXES = ('<dke-cockpit-origin>/api/apps/tokenization/')
ENABLED = TRUE;
CREATE OR REPLACE DATABASE <database>;
USE DATABASE <database>;
CREATE OR REPLACE SCHEMA <schema>;
CREATE OR REPLACE EXTERNAL FUNCTION <database>.<schema>.TOKENIZE(input VARCHAR) RETURNS VARCHAR
API_INTEGRATION = <integration-name> AS '<dke-cockpit-origin>/api/apps/tokenization/tokenize';
CREATE OR REPLACE EXTERNAL FUNCTION <database>.<schema>.DETOKENIZE(input VARCHAR) RETURNS VARCHAR
API_INTEGRATION = <integration-name> AS '<dke-cockpit-origin>/api/apps/tokenization/detokenize';This Snowflake tokenization variant is not the same feature as DuoKey for Snowflake (Tri-Secret Secure). Tri-Secret Secure protects an entire Snowflake account's data at rest via a customer-managed composite master key in AWS KMS. This tokenization variant protects individual columns by replacing their values with format-preserving tokens, called from Snowflake SQL as external functions. The two are complementary — you can run Tri-Secret Secure account-wide encryption and column-level tokenization for specific sensitive fields at the same time — but they solve different problems and are configured independently.
Tokenization Secret variant
The generic Tokenization Secret variant creates a named tokenization key in a vault you choose, without binding it to Databricks or Snowflake. Use it when your own application calls the tokenization service directly rather than through a data-platform UDF or external function.
| Field | Purpose |
|---|---|
| Secret name | The name of the tokenization key/secret created in the linked vault. |
| Key type / size | The underlying key material type and size (defaults to AES-256). |
| Exportable | Whether the key may be exported from the vault. Defaults to non-exportable. |
| Data-type hint | The default data type used to pick the tokenization alphabet. |
| Custom metadata | Optional free-form key/value tags recorded against the secret. |
Health and self-test
Every variant exposes the same health check and self-test surface, resolved against the linked vault: the health check reports vault reachability (and authentication, when determinable); the self-test additionally performs a real round-trip against the vault — and, for Databricks, a live workspace-reachability probe — rather than a simulated result.