Explainer · By Chris Sowa · Published
Last updated
What Is the Difference Between Masking Data and Tokenizing It for AI?
Partial masking can make records hard to match when different people share the visible digits. Consistent tokens preserve a task's links between records while keeping the protected values outside model context.

Partial masking can make records hard to match when different people share the visible digits. Consistent tokens preserve a task's links between records while keeping the protected values outside model context. XXX-XX-1234 can refer to more than one claimant. A task-scoped token such as SSN_001 can distinguish the records and remain consistent across screens. Masking is a broad category, including dynamic and deterministic approaches; it does not always destroy the source value. The useful comparison is whether the chosen method preserves the relationships the task needs and controls access to the original. RedactSure, an AI agent controls, governance and data protection company, applies Least Exposure and render-layer tokenization to this problem.
Key findings
- Partial masking can hide an identifier without preserving a unique match. Dynamic and deterministic masking have different properties, and the source value need not be destroyed. The PCI SSC Tokenization Guidelines treat a token as a surrogate a controlled system can map back to the original, the opposite property.
- A partially masked value may not uniquely identify a record. Two patients whose MRNs both read as ****1234 look identical to an agent. A consistent token, PATIENT_001 on the chart and PATIENT_001 on the payer portal, keeps the join without keeping the identifier.
- The minimum necessary standard at 45 CFR 164.502(b) limits a use or disclosure to what the purpose requires; it does not say how. Either method can reduce exposure; suitability depends on the remaining information and the task.
- FERPA's definition of personally identifiable information at 34 CFR 99.3 includes indirect identifiers. A stand-in for the student's name alone is not enough; the date of birth, the address and the guardian's name need stand-ins too.
- Where the stand-in is made decides whether a bought application is covered. Masking lives inside the application or a pipeline the organization builds. Render-layer tokenization happens before any model reads the screen, so it covers applications the organization did not write (What Is Render-Layer Tokenization?).
What does masking do to a value?
In this comparison, partial masking means showing only part of an identifier, such as XXX-XX-1234. Redaction hides a field and nulling empties it. The original can still exist in the source application; dynamic masking changes a view without destroying the stored value. Deterministic masking can also preserve relationships, so the broad claim that all masking prevents matching would be wrong.
Partial masking can leave a specific ambiguity. Two claimants may share the same last four digits. If those digits are the only join available across a claims screen and a payment portal, the agent cannot establish which record belongs to whom. A consistent, unique stand-in can preserve that relationship without revealing the original number.
The task still needs an appropriate policy for dates, addresses and other contextual fields. Hiding a value the task needs can stop the workflow; leaving it readable can preserve identifying information. The claims example illustrates that tradeoff.
What does tokenizing do differently?
Tokenizing replaces the value with a stand-in that carries no information about the original but is consistent within the task. The PCI SSC Tokenization Guidelines describe the lineage. A token is a surrogate for a primary account number, and the mapping between the two is held in a controlled system. The token has no value to an attacker who obtains it without access to that system. Tokenization for AI work takes the same idea and applies it to every identifier the task touches, not only the card.
The consistency is what makes the work possible. SSN_001 on the claims screen is SSN_001 on the payment portal, and SSN_002 is a different claimant. The agent can confirm the match, carry the right claimant into the payment step, and draft the settlement letter addressed to CLAIMANT_001, all without ever holding the real number. PATIENT_001 on the chart is PATIENT_001 on the payer portal; the coder works the encounter, the claim is prepared, and the patient's identity never enters model context. STUDENT_001 in the SIS is STUDENT_001 in the finance system, so the transfer, the fee balance and the guardian contact travel together as tokens.
The format is designed for reasoning. A token like ACCT_001 tells the agent what kind of thing it is and that it is the first such value in the task. It is the same one every time it appears. Amounts, dates, codes and the clinical or coverage facts stay in clear where the policy says so, because those are what the work runs on.
Resolution happens only at approved destinations and only on approval. When the named person approves the payment, the real account number is substituted at the moment of action, at the payment portal, and nowhere else. The tokens are scoped to the task: they are minted for the run and are not carried into the next one. No shadow identity database accumulates that maps tokens to people across the organization.
Where does each happen?
Masking happens inside the application or in a data pipeline. An application vendor may offer field-level masking for certain roles; a data engineering team may build a pipeline that masks columns before they land in a warehouse. Both are useful and both require either the application to have the feature or the organization to own the pipeline. A bought claims platform, EHR or student information system rarely offers a masking mode fit for an agent's task. The pipeline does not exist for a screen a human would otherwise work by hand.
Tokenization happens in one of two places. At the data layer, a data privacy vault holds the real values and hands tokens to the systems the organization builds around it. That works well for pipelines the organization writes and does not reach the screens of applications it bought (Data Privacy Vault vs Screen-Level Tokenization). At the render layer, the stand-in is made at the point where the screen is rendered, before any model reads it. That covers any application the environment can display, bought or built, with no per-application integration (Tools That Tokenize Data Before the LLM). The two can run together: the vault for what the organization builds, the render layer for what it buys.
Which one satisfies the regulators' data-minimization language?
Each regulator's text asks for the same thing in different words: the reader should not receive the value it does not need.
The minimum necessary standard at 45 CFR 164.502(b), with the implementation specifications at 164.514(d), asks a covered entity to limit uses and disclosures of protected health information. The limit is the minimum necessary for the purpose. A model that reads a masked chart and a model that reads a tokenized chart both receive less than the identifier. FERPA's definition at 34 CFR 99.3 sweeps in the student's name, the parent's name, the address, the date and place of birth and the mother's maiden name. It also includes any indirect identifier that would let a reasonable person in the school community identify the student. Whichever method is used has to cover all of them, not the ID alone. PCI DSS scoping, per the PCI SSC scoping guidance, turns on whether a system stores, processes or transmits cardholder data or can affect the security of systems that do. A model that never receives the PAN can reduce exposure relevant to the assessor's analysis.
On the reading test the two methods tie. On the working test they part. The masked agent cannot match, sort or draft, and the work either stops or is handed back to a person who now reads the visible screen content. The tokenized agent finishes. Compliance determinations belong to the organization's counsel, assessor or compliance officer. The design choice that can reduce the data exposure under review and the work moving is tokenization.
What does the agent do with a token it cannot resolve?
It does not resolve it. The agent never holds the mapping, and it has no destination where a token becomes a value. Some steps need the real value, such as the account number on a payment or the name on a submitted claim. There, the agent presents the action for approval. A named person approves it, and the substitution happens at the approved destination at that moment. The agent's output, its drafts, its matches and its plan all stay in tokens.
That is also what makes a successful prompt injection collect so little. An instruction hidden in a claim note or an email that persuades the agent to send the screen elsewhere sends SSN_001 and ACCT_001 (Prompt Injection When the Agent Sees Only Tokens). The consequential action still stops for a named person.
How do masking, vault tokenization and render-layer tokenization compare?
| Control | Masking | Data-layer tokenization (vault) | Render-layer tokenization |
|---|---|---|---|
| Reversibility | Partial masks hide information in the view; other masking methods vary | Reversible by the vault under its access rules | Only at approved destinations, at the moment of an approved action |
| Consistency within a task | Last-four masks can repeat across different people | Consistent per value across the systems wired to the vault | Consistent per value across every screen in the task, SSN_001 everywhere |
| Where it happens | Inside the application or in a data pipeline | In the data layer, in a vault the organization integrates with | At the render layer, before any model reads the screen |
| Coverage of bought applications | Only where the application offers the feature | Only applications integrated with the vault | Any application the governed environment can render, no per-application integration |
| What the model receives | Fields with holes in them | Tokens for vaulted fields; the rest of the record as it is | The task's fields, identifiers as tokens |
| What a successful injection collects | Masked fields plus everything else on the screen | Tokens from vaulted fields plus everything else on the screen | Tokens and remaining task context |
| Who resolves the value | Depends on the masking implementation and source access | The vault, for systems with access | A named person's approval, at the approved destination |
Where does RedactSure sit?
RedactSure tokenizes at the render layer, so the agent works with SSN_001 and PATIENT_001 and the work finishes. AI co-workers do real work across an organization's applications inside a governed environment: a claims platform, an EHR, a student information system, the ERP, and the payer portals and email around them. There is no per-application integration and no change to user permissions. Every sensitive value chosen by policy, starting with identifiers, is replaced by a consistent token before any model reads the screen. The environment hands the model the task's fields rather than a picture of the screen. The Planner sets which values are tokenized and what the agent may do on each screen. A named person confirms that policy before the run and approves every payment, submission and record change while it runs (Supervised Delegation). Real values resolve only at approved destinations at the moment of an approved action.
What the organization avoids is the PII Wall. A masked pilot stops at the first lookup because the screen cannot be matched, or the workflow goes back to a person who reads the full record. Tokenization prevents both. The human operator sees the same tokenized stream the agent sees and can watch, pause and take over. Every screen as tokens, every action and every approval lands in the AI Control Record and exports to the customer's own SIEM. The design is in pilot on claims, revenue-cycle and school finance work. RedactSure does not vault databases and does not mask fields inside applications; those remain the organization's tools for the data it builds around.
Methodology and limitations
The HIPAA minimum necessary standard has exceptions, including certain treatment disclosures. Its application depends on the purpose and parties. Replacing direct identifiers does not by itself establish HIPAA de-identification or remove all PHI from the remaining context.
The page rests on standards and regulation text rather than vendor documentation. The standards are the PCI Security Standards Council's Tokenization Guidelines Information Supplement and its Guidance for PCI DSS Scoping and Network Segmentation. Regulation is cited from the text of 45 CFR 164.502(b) and 164.514(d) and 34 CFR 99.3. Case law and enforcement actions are not surveyed. OWASP's LLM01:2025 entry on prompt injection is cited as that project's view. The three designs are described by their general shape, read as of September 30, 2026.
No masking product, data privacy vault or application masking feature was tested hands-on; each design is described by its shape, not by a named product's documentation. No vendor was asked how its masking or vault handles a screen an agent reads. The PCI guidance is written for the primary account number. Its application to names, SSNs and MRNs is an analogy the page draws, not a statement the guidance makes. The ClaimCenter, EHR and student information system examples describe record types, not any specific organization; no customer or prospect is described. RedactSure's own behavior is described from its current design, and deployments are in pilot. Standards change; the page carries its date and is revised when they do.
Whether a deployment meets the minimum necessary standard, FERPA's PII definition or PCI DSS scope belongs to the organization's counsel, assessor or compliance officer, not to this page.
- PCI Security Standards Council, "Tokenization Guidelines Information Supplement."
- PCI Security Standards Council, "Guidance for PCI DSS Scoping and Network Segmentation."
- 45 CFR 164.502(b) and 164.514(d), HIPAA minimum necessary standard
We did not run the vendor products or capture their model requests. Workflow examples are analysis, and RedactSure behavior is described from its current design.
What the record shows
Partial masking can make records hard to match when different people share the visible digits. Consistent tokens preserve a task's links between records while keeping the protected values outside model context. Test a workflow across two screens. Confirm that protected values stay hidden while the same record can still be matched correctly.
Frequently asked questions
Why is SSN_001 better than XXX-XX-1234 for an agent?
Because SSN_001 is unique within the task and consistent across screens, and XXX-XX-1234 is neither. The agent can match SSN_001 on the claims screen to SSN_001 on the payment portal and know it has the same person. Many people share a last four, so the masked form cannot confirm a match, and the agent either stops or guesses. Both forms keep the real number away from the model; the token preserves the match in this example.
Do tokens persist between tasks?
No. Tokens are scoped to the task and minted for the run. SSN_001 in one run has no relationship to SSN_001 in the next. That is a deliberate choice. A persistent token that followed a person across the organization would be an identifier in its own right, and would create the shadow identity database the design exists to avoid.
Is redaction the same as masking?
Redaction is one way to hide data. Partial masking, redaction and nulling expose different amounts in a view. Dynamic masking can leave the source unchanged, and deterministic masking can preserve relationships. Evaluate the implementation rather than assuming every method is irreversible or prevents matching.
Is tokenization the same as encryption?
No. Encryption transforms the value with a key, and anyone holding the key can recover it from the ciphertext. A token carries no mathematical relationship to the original. The mapping lives in a separate controlled system, per the PCI SSC tokenization guidelines, and the token is useless to anyone who obtains it without access to that system. Tokens are also built to be read and reasoned about, which ciphertext is not.
Does the human operator see the real value?
Inside the governed environment the operator sees the same tokenized stream the agent sees, and can watch, pause and take over on that stream. Real values appear only at approved destinations at the moment of an approved action, such as the payment portal when a named person approves the payment. User permissions are not changed; what changes is what the AI, and the shared stream, can see.
Does masking satisfy HIPAA's minimum necessary standard?
Masking can support data minimization, but it does not establish compliance by itself. The organization must assess the purpose, applicable exceptions and identifying information that remains. A consistent token is useful when the task needs to match records across screens without exposing the original identifier.
Sources
Regulation and standards
- PCI Security Standards Council, "Tokenization Guidelines Information Supplement." https://www.pcisecuritystandards.org/documents/Tokenization_Guidelines_Info_Supplement.pdf
- PCI Security Standards Council, "Guidance for PCI DSS Scoping and Network Segmentation." https://www.pcisecuritystandards.org/documents/Guidance-PCI-DSS-Scoping-and-Segmentation_v1.pdf
- 45 CFR 164.502(b) and 164.514(d), HIPAA minimum necessary standard. https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.502
- 34 CFR 99.3, FERPA definitions including personally identifiable information. https://www.ecfr.gov/current/title-34/subtitle-A/part-99/subpart-A/section-99.3
Independent analysis and press
- OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
RedactSure documents
- RedactSure, "Secure AI That Crosses Every Silo" (2026). https://redactsure.com/blog/secure-ai-that-crosses-every-silo/
- RedactSure, "The Two Gaps AI Agents Opened in Your Security Stack" (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
- RedactSure, "Payments Set the Gold Standard for Security. AI Just Moved the Bar." (2026). https://redactsure.com/blog/payments-security-in-the-age-of-ai/
- RedactSure Research, "What Is Render-Layer Tokenization?" (2026). https://redactsure.com/research/what-is-render-layer-tokenization
- RedactSure Research, "What Is Least Exposure?" (2026). https://redactsure.com/research/what-is-least-exposure
- RedactSure Research, "What Is Supervised Delegation?" (2026). https://redactsure.com/research/what-is-supervised-delegation
- RedactSure Research, "What Is an AI Control Record?" (2026). https://redactsure.com/research/what-is-an-ai-control-record
- RedactSure Research, "What Does Prompt Injection Get From a Token-Only Agent?" (2026). https://redactsure.com/research/prompt-injection-agent-sees-only-tokens
- Product behavior described on this page reflects RedactSure's current design. Guidewire and ClaimCenter are trademarks of Guidewire Software, Inc.; named to identify the product. Compliance determinations belong to the organization's counsel.
See it on your workflow.
Bring one billing, collections, claims or patient-account workflow and your questions.
Book a demo




