Contents

Data Engineering › Data Governance & Privacy

Data Masking

Hiding sensitive values while keeping data usable.

Also known as: masking, PII masking, dynamic data masking, static data masking, obfuscation

Data masking means hiding or altering sensitive values while keeping the data usable, so people and systems that don’t need the real values never see them. It’s one of the main controls for protecting personal and confidential data in analytics, test environments and support tools.

Original:   Ana Putri | ana.putri@example.com | 4111 1111 1111 1111 | 1990-03-14
Masked:     A*** P*** | a***@example.com      | **** **** **** 1111 | 1990-**-**

Common techniques

TechniqueWhat it doesExample
Redaction / nullingRemoves the valueNULL
Partial maskingShows only partLast 4 digits of a card
SubstitutionReplaces with realistic fake valuesA random but plausible name
HashingReplaces with a one-way hash, keeping equal values equal (so joins and counts still work)SHA-256(salt + email)
TokenizationReplaces with a token, with the real value stored in a secure vault (tokenization)tok_8f3a2c
GeneralizationReduces precisionExact age → age band, birthday → year
ShufflingMoves real values among rowsBreak the link to the person

Static vs dynamic

  • Static masking creates a masked copy of the data, such as a test or development database, or a dataset for analysts. The sensitive values never leave production.
  • Dynamic masking applies masks at query time based on who is asking: an analyst sees a***@example.com, while an authorized support agent sees the full email (column-level security).

Where it’s used

  • Development and test data: use masked or synthetic data, never raw production copies (dev and prod environments).
  • Analytics where individuals’ identities aren’t needed (data minimization).
  • Sharing data with partners or vendors.
  • Support and operations tools that need partial visibility.

Things to get right

  • Mask consistently, in every system where the data lands: warehouse, lake, backups, logs, exports. One forgotten copy defeats the control.
  • Preserve what the use case needs: formats (so validation still works), referential integrity (the same person masks to the same token everywhere), and distributions.
  • Hashing isn’t anonymization. Low-entropy values (emails, phone numbers, national IDs) can often be reversed by guessing and hashing candidates. Use salted or keyed hashing, tokenization, or true anonymization where the risk requires it (anonymization).
  • Re-identification risk: combinations of “harmless” fields (postcode, birth date, gender) can identify people.
  • Classify first: you can’t mask what you haven’t found (data classification, PII).
  • Control access to unmasked data, and log it.

Masking reduces exposure. It doesn’t replace access control, minimization and retention limits.