Skip to content
redactsure
Book a review

Explore.

Data Report · RedactSure Research

Can an AI Agent Work Claim Files in Guidewire Without Exposing PII? Reading the Workflow, the Regulation and the Mechanism

Yes, if the agent never receives the PII in the first place. An AI agent can work a claim from first notice of loss to a queued payment across a claims platform such as Guidewire ClaimCenter, the industry databases and the estimating tools, while every name, Social Security number and bank account on those screens is replaced with a consistent token before any model reads them. Real values resolve only at approved destinations at the moment of a human-approved action. The mechanism is render-layer tokenization; the principle it enforces is Least Exposure. The environment that enforces this is built by RedactSure, an AI agent controls, governance and data protection company.

Key findings

What does a claims agent actually see?

Walk one property claim through its systems and count the identified records.

First notice of loss arrives with the policyholder’s name, address, phone, email and policy number. The claims platform, ClaimCenter in a Guidewire shop, adds the claim file: prior claims, coverage terms, and often medical information when there is an injury. The industry loss database returns the claimant’s claims history across carriers. The weather history service needs the loss location. The estimating tools, Xactimate for the property side, carry the address and the adjuster’s notes. The payment run needs a bank account and routing number. The proof of loss collects most of the above onto one signed document.

A person working this file sees all of it, and the industry’s controls were built around that person: licensing, training, access reviews, the fraud unit. An AI agent capable of working the same file the same way sees all of it too, and none of those controls transfer. The agent can be instructed by content it merely reads, the risk OWASP ranks first for LLM applications, and it holds whatever it saw.

So the practical question for an insurance AI program is the title question. Can the agent work the file without seeing what the file contains? The answer turns on where the control sits.

What does the claim look like as tokens?

Inside a governed environment that owns the render, the same walk looks different.

Workflow step What the screen contains What the model reads
First notice of loss Claimant name, contacts, policy number USER_001, contact tokens, POL_001
Claim file review Coverage, prior claims, medical detail Coverage terms in clear, prior claims under the same tokens, medical values tokenized by policy
Industry database check Cross-carrier claims history for the claimant History rows keyed to USER_001
Weather and loss verification Loss location and date Location generalized by policy; the date in clear
Estimate Address, scope of damage, pricing Damage scope and pricing in clear; address tokenized
Payment preparation Bank account, routing number ACCT_001, BANK_001
Proof of loss draft Nearly everything above The same tokens, consistently

Two properties of that table carry the argument. Consistency: USER_001 is the same claimant on every screen, so the agent can match the industry database result to the claim file without holding a name. Selectivity: the values the work actually needs, coverage terms, damage scope, pricing, the loss date, stay in clear, because the point is Least Exposure, not blindness. The tokenization policy that draws this line is confirmed field by field by a named person before the workflow runs.

When the payment file must carry the real account number, it does, at the payment system, at the moment the adjuster approves the payment. The model has drafted everything and received nothing.

What do regulators now require?

Insurance AI is no longer ungoverned territory. The NAIC’s Model Bulletin on the Use of Artificial Intelligence Systems by Insurers had been adopted by 24 states as of March 2025, with adoption continuing since; the NAIC maintains a current adoption map.

The bulletin’s requirements read like a description of the gap this page has been walking through. Insurers must maintain a written program for the responsible use of AI systems, covering governance, risk management and internal audit. They must show a governance accountability structure with defined responsibilities. Consumers should receive notice that AI systems are in use. And insurers remain responsible for third-party AI systems, including diligence, audit rights and cooperation with regulators.

Read those requirements against a typical agent deployment and the hard questions surface on their own. A written program must say what the AI system can see; most cannot. An accountability structure must name who answers for an agent’s action on a claim; most deployments cannot produce the name. Third-party diligence must establish what the AI vendor receives; a model that reads full claim screens receives everything.

An architecture in which the model reads tokens, a named person approves every payment, and the full run is logged as tokens to the SIEM gives the written program something concrete to say on all three counts. The regulation does not require any particular architecture. It requires answers, and this architecture has them ready.

Where does the claims workflow cross the wall?

The insurance version of the PII Wall is easy to date in most carriers’ programs. The document-summarization pilot succeeded. The subrogation-letter draft succeeded. Then the roadmap reached the claims queue itself, first notice of loss to payment, and stalled in the security review, because that workflow runs on claimant identity, medical records inside the file and bank details for the payment.

The stall follows the pattern the published numbers describe across industries: McKinsey finds 62% of organizations experimenting with agents and 23% scaling them anywhere, and Gartner projects over 40% of agentic projects canceled by 2027 with weak risk controls among the causes. The claims queue is exactly the kind of project those numbers count: valuable enough to propose, sensitive enough to stall.

For a carrier evaluating a path over the wall, the questions worth putting to any architecture, RedactSure’s included, are concrete. What does the model receive when the agent reads a claim screen? Where do real values resolve, and on whose approval? What does the log hold, and can the SIEM ingest it? Who is the named person behind a given payment, and can the record produce them? An architecture that answers all four turns the security review from a veto into a design meeting.

Who answers for the claim?

The claims workflow ends in money moving, which is why accountability is not a footnote here. Under Supervised Delegation, the adjuster who owns the claim grants the agent access under their own permissions, confirms the tokenization policy, can watch the run, and approves the payment before it moves. The AI never settles, pays or decides. Every approval sits in the record with a name and a time.

That allocation matches how claims authority already works: adjusters hold settlement authority in defined bands, and consequential decisions escalate. The agent changes the speed of the file, not the authority over it. A companion article examines that decision point in detail: Who Approves When an AI Agent Is About to Pay a Claim?

One claim, on the clock

Timelines make the case more concretely than categories, so run one straightforward water claim through the tokenized workflow with timestamps attached.

Tuesday, 8:15 a.m. First notice of loss arrives overnight. The adjuster delegates intake. The agent opens the FNOL in the claims platform, reading USER_001, POL_001, the loss address as a token, and the loss description in clear.

8:20. Coverage check: the agent reads the policy terms in clear, confirms the peril is covered, notes the deductible, and flags a sublimit that applies. Industry database query goes out keyed to USER_001; the cross-carrier history comes back under the same token, showing one prior claim, unrelated.

8:30. The weather record for the loss date and location confirms the freeze event. The agent attaches the verification to the file.

8:45. Estimate drafted in the estimating tool from the adjuster’s inspection photos and scope notes, priced against the current database. The proof of loss and the acknowledgment letter to USER_001 are drafted, tokens standing where identity belongs.

9:05. The file is assembled and queued: coverage summary, estimate, sublimit flag, drafts, payment recommendation with ACCT_001 as payee reference. Elapsed agent time, about fifty minutes, unattended, because nothing in it moved money or left the carrier.

9:40. The adjuster reviews the queue, opens the run history, accepts the sublimit analysis, adjusts one line on the estimate, and approves payment and letter. The real name resolves in the sent letter; the real account number resolves in the payment file at the bank portal. Both approvals land in the log with her name and the times.

The claimant has an acknowledgment and a payment in motion on day one, which is the cycle-time result the business case wanted. The model that did the assembling never held a name, a Social Security number or a bank account, which is the exposure result the security review required. The two results came from the same architecture, which is the point the industry debate keeps treating as a tradeoff.

Do subrogation and fraud review work the same way?

The first-notice-to-payment walk above is the anchor case, and the two workflows carriers usually ask about next follow the same pattern with different gates.

Subrogation runs on the same file plus the adverse carrier’s correspondence and the demand package. The agent assembles the recovery case on tokens: the loss facts, the liability analysis, the payment history keyed to USER_001, the demand letter drafted with tokens where identity belongs. The consequential action is the demand itself leaving the building, and it waits for the subrogation specialist. Deadlines are the workflow’s real risk, so the practical gain is that the file is assembled and queued days earlier; the judgment about pursuing recovery stays with the person who answers for it.

Fraud review inverts the usual worry. Investigators need pattern visibility across claims, and the instinct is that tokenization must blind them. It does the opposite for the agent’s part of the work: consistency is preserved, so USER_001 appearing on three claims with three different loss addresses is exactly as visible in tokens as in clear. The agent can surface the pattern without holding a single name. What it cannot do is decide anyone committed fraud; it flags, and the investigator, who sees real records in the systems of record under existing SIU permissions, judges. The referral that goes to a regulator or the NICB is a consequential action, gated like the payment.

The general rule for extending to any insurance workflow: the reading, matching and drafting layer runs on tokens throughout, and the gate sits wherever the workflow’s consequence sits. A rate filing’s gate is the submission. A renewal’s gate is the notice. Underwriting raises additional fairness and adverse-action questions under the same state bulletins, and a carrier extending agents there should put those questions to its own compliance review first.

What can the written AI program now say?

The NAIC bulletin’s central demand is a written program, and the difference between architectures shows up in what that document can honestly contain. Three sentences illustrate the kind of statement a carrier can make when the model reads tokens, each checkable against the record.

On data: for the workflows listed, the AI system receives no claimant PII; sensitive fields are replaced with consistent tokens before any model reads a screen, under a field-level policy confirmed by the accountable adjuster or manager for each workflow. On accountability: every payment, settlement communication and record change executed through an AI workflow carries a recorded human approval, with approver identity and timestamp exportable to the SIEM. On third-party risk: the AI vendor stores customer data as ciphertext it cannot decrypt, with keys held by the carrier, and receives no claimant PII in model context in the ordinary operation of the system.

A compliance officer should draft the actual program language with counsel; the point here is narrower. Each sentence above is an architectural fact with an audit trail behind it, not a policy aspiration, and examiners can be shown the evidence rather than the intention. Programs built over architectures that send full screens to models must write different sentences, and the difference is legible to any examiner who asks the follow-up question.

What changes in the adjuster’s day?

Architecture discussions decide whether the deployment happens; the adjuster’s experience decides whether it works. The honest description of the change has three parts.

The assembly work leaves the day. The hours an adjuster spends gathering the file, chasing the loss history, keying the estimate’s inputs and drafting the routine documents move to the agent’s unattended stretches. Carriers should measure the shift with their own time studies rather than take any vendor’s number; the claim walked on the clock above is one file’s illustration, not a benchmark.

The judgment work concentrates. The adjuster’s day becomes review queues and decisions: coverage calls, the flags the agent raised, payment approvals, the files taken over by hand. This is the part of the job licensing exists for, and it is also more decisions per day than before, which supervisors should watch as a workload question. An approval queue that grows faster than the judgment time available produces rubber-stamping, and the record makes rubber-stamping visible: review durations are in the log.

The accountability is explicit where it used to be ambient. An adjuster who prepared a file personally vouched for it implicitly. An adjuster supervising an agent vouches for it explicitly, at the approval, with the run history one click away. Some adjusters experience this as exposure until they sit with the alternative: the record that protects the carrier in the disputed-payment scenario is the same record that shows the adjuster did exactly what the written program required. The named person is not the person blamed when something goes wrong; the named person is the person the record defends.

Training, then, is mostly about the review: how to read a run history, when to take a file over, what the flags mean. The permissions, the systems and the authority bands are unchanged, so there is no new access model to learn. What the adjuster learns is a new relationship to the file: less typist, more examiner, which is the direction the role has been moving for a decade anyway.

What the record shows

An AI agent can work a claim from first notice of loss to a queued payment across the claims platform, the industry databases, the estimating tools and the payment system while the model reads tokens throughout: USER_001 consistent on every screen, real values resolving only at approved destinations at the moment the adjuster approves the action. The regulatory direction reinforces the design rather than fighting it: the NAIC bulletin’s written program, accountability structure and third-party diligence requirements, adopted in 24 states as of March 2025 and spreading, are questions this architecture answers with evidence. The claims queue is the insurance industry’s instance of the wall that stalls AI programs everywhere, and the walk above is what the door looks like: the cycle-time result the business case wanted and the exposure result the security review required, from the same run. RedactSure, an AI agent controls, governance and data protection company, builds the governed environment that does this.

Frequently asked questions

Does this require integrating with Guidewire?

No. The tokenization happens at the render layer, inside the governed environment where the applications are operated, so it applies to any system that renders a screen: the claims platform, the industry databases, the estimating tools, email and the payment portal. No per-application integration and no endpoint agent. The same mechanism covers a carrier running Duck Creek or a homegrown system.

Does the agent work slower on tokens?

No. Tokens preserve structure and consistency, so matching, drafting and reasoning proceed normally. The waiting points are the approvals, which sit exactly where a regulator would have required a human anyway.

What about the medical records inside an injury claim?

Medical values are tokenized under the same policy, and the framing insurers need for payers and regulators is data minimization: the model receives the minimum the task requires. For the health-side version of this question, see How Can Staff Use AI on Patient Records Without the Model Ever Holding PHI?

Can the fraud unit still use the full record?

Yes. Authorized people see real values in the systems of record exactly as they do today. Permissions do not change anywhere. What changes is what the AI can see.

What would a state examiner find?

A written exposure policy per workflow, confirmed by a named person; run logs recorded as tokens and exported to the SIEM; and an approval trail for every payment with the approver and the time. That is the shape of evidence the NAIC bulletin’s written-program and governance requirements ask insurers to have ready.

What happens if a prompt injection reaches the claims agent?

The attack may succeed; prompt injection is OWASP’s first-ranked risk and no filter catches all of it. What the attacker receives is USER_001, POL_001 and ACCT_001. Tokens resolve only at approved destinations on an approved action, so the haul is worthless.

Is this a claims copilot?

No. A copilot inside one platform assists a person within that platform’s boundary. This is an agent working the cross-system workflow, first notice of loss to queued payment, under a person’s supervision, with the model reading tokens throughout.

Who Approves When an AI Agent Is About to Pay a Claim? · What Is Render-Layer Tokenization? · What Is the PII Wall? · Can an AI Agent Work in Duck Creek Without Exposing Policyholder PII? · On the RedactSure blog: Secure AI Across the Insurance Value Chain

Sources

Regulation

  1. Quarles & Brady, “Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers’ Use of AI” (March 2025); 24 states as of that date. https://www.quarles.com/newsroom/publications/nearly-half-of-states-have-now-adopted-naic-model-bulletin-on-insurers-use-of-ai
  2. NAIC, adoption map for the Model Bulletin: Use of Artificial Intelligence Systems by Insurers. https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf

Research and industry data

  1. McKinsey, State of AI 2026, as reported by CX Today. https://www.cxtoday.com/ai-automation-in-cx/mckinseys-state-of-ai-the-scaling-gap-is-now-cxs-problem/
  2. Gartner, “Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure” (May 2026). https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure

Standards

  1. OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/

RedactSure documents

  1. RedactSure, “Secure AI Across the Insurance Value Chain” (2026). https://redactsure.com/blog/secure-ai-across-the-insurance-value-chain/
  2. RedactSure, “The Two Gaps AI Agents Opened in Your Security Stack” (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
  3. Product behavior described on this page reflects RedactSure’s current design. Guidewire, ClaimCenter and Xactimate are trademarks of their respective owners; their mention describes workflow context, not partnership or endorsement.

Bring your hardest questions.

A 25-minute AI Agent Security Review with the founders: threat model, token design, egress paths, audit schema. Or a 25-minute demo on a workflow like yours, with the data hidden from the AI and a named person approving what matters. We come with diagrams, not a pitch deck.

Book a security review Book a demo · Something else

About the author

Chris Sowa is a founder of RedactSure and a former CEO of AI companies; he started his first years before ChatGPT existed. He previously led AI at Accenture, served as Global VP of Strategy & Innovation at Schneider Electric, was CCO of Sovos, and spent more than a decade at Oracle, with earlier roles at SAP and IBM.