إنتقل إلى المحتوى الرئيسي
ينطبق على:
DuoKey Cockpit v2Format-preserving tokenization (NIST SP 800-38G, FF1)Databricks · Snowflake · generic application data

Overview​

Tokenization replaces a sensitive value — a card number, a national ID, an email — with a token that has the same length and character shape, so downstream systems keep working (queries, joins, format validation) without ever seeing the original value. DuoKey's tokenization service performs format-preserving encryption (FF1, per NIST SP 800-38G): tokenizing and detokenizing are true inverse operations under a tokenization key, not a lookup table, so there is no separate token vault to manage or synchronize.

Three onboarding paths front the same service:

Databricks

Spark / Unity-Catalog UDFs that tokenize and detokenize columns at query time.

Snowflake

External functions that tokenize and detokenize columns from Snowflake SQL.

Tokenization Secret

A standalone tokenization key/secret in a linked vault, for applications that call the service directly.

Databricks has its own page

The Databricks variant is documented in full, including generated setup SQL, on the Databricks page. This page covers the tokenization service itself, plus the Snowflake and generic-secret variants.

Three onboarding paths, one tokenization service
DatabricksSpark / Unity Catalog UDFs
SnowflakeExternal functions
Tokenization SecretDirect API calls
POST tokenize / detokenize — same service
DuoKey tokenization serviceFormat-preserving encryption (FF1, NIST SP 800-38G)Key derived per tenant and named key from the platform master key

Every variant calls the same tokenize / detokenize endpoints and the same FF1 engine — only the setup artifacts (UDFs, external functions, or a direct secret) differ.

How tokenization works​

Format-preserving

A token has the same length and character set as the original value: a 16-digit card number tokenizes to another 16-digit sequence, an email keeps its @ and dot separators.

Data-type aware

A data-type hint (creditcard, ssn, phone, email, …) selects the right character alphabet automatically, or you can pin an explicit format.

Domain-separated

The data type doubles as part of the cryptographic tweak, so the same plaintext value tokenizes differently in two different columns/domains — a card number and a lookalike account number never collide.

No stored token vault

Tokenization keys are derived deterministically per tenant and per named key, so detokenizing a value never depends on retrieving previously stored key material. Rotating protection for a field is a matter of choosing a new key name going forward.

Data-type hintAlphabet used
creditcard, ssn, phone, account, numericDigits only (0–9) — preserves numeric formats like card numbers and phone numbers.
email, iban, name, custom / anything elseMixed-case alphanumeric — preserves separators such as @ and . while tokenizing the rest.
Explicit format override

If the default alphabet inferred from the data-type hint is not what you need, you can pin an explicit format (digits, alphanumeric, upper-alpha, lower-alpha, and variants) instead of relying on the data-type inference.

Named tokenization keys​

Every tokenization call names a key (defaults to "default") that is scoped to your tenant. Different named keys are cryptographically independent, which gives you two practical controls:

  • Segmentation — use different named keys for different applications or environments so their tokenized values are never comparable to each other.
  • Rotation — start tokenizing new values under a new key name to rotate protection for a field going forward, without needing to re-tokenize existing data on a deadline.
Detokenization needs the same key name

A value tokenized under one key name can only be detokenized by requesting the same key name. Keep a clear mapping of which key name protects which column/field, especially if you introduce a new key name for rotation.

How an FPE tokenize call works​

This is the real sequence behind every tokenize / detokenize call, whichever variant it arrives from — the same steps the Databricks page's sequence diagram summarizes as "DuoKey resolves the key and performs FPE":

Inside a tokenize / detokenize call
1. Resolve the alphabetAn explicit format override wins; otherwise it is inferred from the data-type hint
2. Derive the FF1 keyHMAC-SHA256 of the platform master key, scoped to the tenant and the named key — nothing stored
3. Build the tweakThe data type, plus an optional caller tweak, binds the token to its column or domain
4. Encrypt or decrypt with FF1Tokenize and detokenize are the same operation run in opposite directions under the same key and tweak
5. Return the token or original valueSame length and character shape as the input

Detokenize runs the identical sequence with the decrypt operation in step 4 — FF1 is a true inverse, not a lookup.

The three variants​

VariantConsumed fromDistinguishing config
DatabricksSpark / Unity-Catalog UDFsWorkspace URL, catalog, schema — see the dedicated Databricks page.
SnowflakeExternal functionsSnowflake account, region, warehouse, database and schema, plus an API integration binding Snowflake to DuoKey.
Tokenization SecretDirect API calls from your own applicationA named secret created in a linked vault, with its own key type, size and exportability setting.

Deploying any variant links a DuoKey vault (and optionally a key inside it) used for connectivity health checks, self-tests and audit; the vault and key you choose are recorded against the integration, but do not need to be re-selected per tokenize/detokenize call — those are authorized purely by API permission and the named key you pass.

Snowflake variant​

Deploying the Snowflake variant generates the SQL to create a Snowflake API integration pointed at DuoKey's tokenization endpoints, and TOKENIZE / DETOKENIZE external functions in your chosen database and schema that call it.

Generated Snowflake setup (shape)SQL
CREATE OR REPLACE API INTEGRATION <integration-name>
API_PROVIDER = aws_api_gateway
API_ALLOWED_PREFIXES = ('<dke-cockpit-origin>/api/apps/tokenization/')
ENABLED = TRUE;

CREATE OR REPLACE DATABASE <database>;
USE DATABASE <database>;
CREATE OR REPLACE SCHEMA <schema>;

CREATE OR REPLACE EXTERNAL FUNCTION <database>.<schema>.TOKENIZE(input VARCHAR) RETURNS VARCHAR
API_INTEGRATION = <integration-name> AS '<dke-cockpit-origin>/api/apps/tokenization/tokenize';
CREATE OR REPLACE EXTERNAL FUNCTION <database>.<schema>.DETOKENIZE(input VARCHAR) RETURNS VARCHAR
API_INTEGRATION = <integration-name> AS '<dke-cockpit-origin>/api/apps/tokenization/detokenize';
Different from Snowflake TSS

This Snowflake tokenization variant is not the same feature as DuoKey for Snowflake (Tri-Secret Secure). Tri-Secret Secure protects an entire Snowflake account's data at rest via a customer-managed composite master key in AWS KMS. This tokenization variant protects individual columns by replacing their values with format-preserving tokens, called from Snowflake SQL as external functions. The two are complementary — you can run Tri-Secret Secure account-wide encryption and column-level tokenization for specific sensitive fields at the same time — but they solve different problems and are configured independently.

Tokenization Secret variant​

The generic Tokenization Secret variant creates a named tokenization key in a vault you choose, without binding it to Databricks or Snowflake. Use it when your own application calls the tokenization service directly rather than through a data-platform UDF or external function.

FieldPurpose
Secret nameThe name of the tokenization key/secret created in the linked vault.
Key type / sizeThe underlying key material type and size (defaults to AES-256).
ExportableWhether the key may be exported from the vault. Defaults to non-exportable.
Data-type hintThe default data type used to pick the tokenization alphabet.
Custom metadataOptional free-form key/value tags recorded against the secret.

Health and self-test​

Every variant exposes the same health check and self-test surface, resolved against the linked vault: the health check reports vault reachability (and authentication, when determinable); the self-test additionally performs a real round-trip against the vault — and, for Databricks, a live workspace-reachability probe — rather than a simulated result.