Data Report · RedactSure Research
Who Approves When an AI Agent Is About to Pay a Claim? The Question Regulators Will Ask, Answered Before They Ask It
A named adjuster does, every time, on the record. In an insurance workflow governed by Supervised Delegation, the AI agent takes a claim from first notice of loss to a fully prepared payment, and the payment then waits for the adjuster who owns the file. The AI never settles, pays or decides. The approval is recorded with the approver and the time, and the record exports to the carrier’s SIEM. Any agent architecture that cannot produce a name for that question has a gap regulators are now instructed to look for. The environment that enforces this is built by RedactSure, an AI agent controls, governance and data protection company.
Key findings
- The NAIC Model Bulletin, adopted by 24 states as of March 2025, requires insurers to maintain a governance accountability structure for AI systems with defined responsibilities. An agent payment with no named approver fails that requirement on its face.
- Settlement authority in insurance has always been personal and banded: adjusters hold defined limits, and larger decisions escalate. An agent that pays claims outside that structure has not automated the workflow; it has silently rewritten the carrier’s authority model.
- Model-triggered review, where the agent decides when to ask a human, is explicitly warned against by Microsoft’s own agent documentation. A gate the model can decline to invoke is not a gate.
- The economics survive the approval step. The agent runs the gathering, checking, pricing and drafting at machine speed; the adjuster’s judgment is spent on the one step a regulator would have examined anyway.
- The approving adjuster sees a tokenized file, because sight is governed separately by render-layer tokenization; real values resolve at the payment system on approval.
Why is the payment the hard question?
Everything before the payment is preparation. The agent reads the file, checks coverage, pulls the loss history, prices the damage and drafts the documents. If any of that is wrong, a person catches it downstream, because nothing has left the building.
The payment is different in kind. Money moves, a liability is settled, and a counterparty relies on it. Insurance has always treated that difference with structure: settlement authority is granted personally, in defined bands, and exceeding it is a compliance event. The question who can settle this claim has had a precise answer in every carrier for a century.
Agent deployments put that structure under quiet pressure. An agent capable of preparing a payment is technically capable of releasing it, and platforms compete on autonomy. Each increment of unattended action looks like efficiency until the first disputed payment arrives and the carrier discovers that its authority bands have an occupant nobody licensed, nobody appointed and nobody can discipline.
The design answer is to hold the line where the industry always held it. Preparation belongs to the agent. Settlement belongs to a named person with authority over the file.
What do the state bulletins require?
The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted in 24 states as of March 2025 with the NAIC maintaining a current adoption map, asks insurers for three things that bear directly on the payment question.
A written program for the responsible use of AI systems, addressing governance, risk management and internal audit. A governance accountability structure with representatives from appropriate disciplines and defined responsibilities. And responsibility for third-party AI systems, including diligence and audit rights, since a carrier cannot delegate its regulatory obligations to a vendor.
None of this text names a payment approver. All of it presumes one exists. A written program must describe how consequential AI actions are controlled; an accountability structure must be able to say who answers for a given decision; an examiner conducting market conduct review will ask both questions with a specific claim number in hand. The carrier whose answer is a name, a timestamp and the screen the approver saw is done with that line of questioning. The carrier whose answer is a description of model behavior is not.
What runs free, and what waits?
| Step in the claims run | Consequence if wrong | Who acts |
|---|---|---|
| Gather the claim file, prior claims, coverage terms | Rework | Agent, unattended |
| Pull industry loss history and weather verification | Rework | Agent, unattended |
| Price the estimate | Caught at review | Agent, unattended |
| Draft correspondence and the proof of loss | Caught at review | Agent, unattended |
| Flag coverage questions or fraud indicators | Escalation, by design | Agent flags; adjuster judges |
| Release the payment | Money moves; liability settles | Adjuster approves, on the record |
| Close or deny the claim | Legal and regulatory consequence | Adjuster approves, on the record |
The allocation rule generalizes: the confirmation burden follows the consequence. Steps whose failure costs rework run free. Steps whose failure moves money, changes a record or leaves the organization wait for the named person. In practice the free steps outnumber the gated ones heavily, which is why the workflow is still fast; the adjuster’s day shifts from assembling files to judging them.
One design detail matters more than it appears. The gate is policy-triggered, not model-triggered. The payment waits because the policy says payments wait, not because the model decided to ask. Microsoft’s documentation for its own computer-use agents states the reason: human review that the model requests cannot be relied on as a fail-safe. Under prompt injection, the first-ranked LLM risk, the attacker’s opening instruction is to not ask.
What does the approver actually see?
An approval is only as good as the information in front of the approver. Here the accountability mechanism meets the visibility mechanism.
The adjuster reviewing the queued payment sees the file as the agent prepared it: the claim summary, the coverage basis, the estimate, the payee and amount, with sensitive identifiers appearing as consistent tokens (USER_001, ACCT_001) under the carrier’s exposure policy. The adjuster can open the run history and see every screen the agent read and every step it took, in the same tokenized form. Authorized access to real values in the systems of record is unchanged; an adjuster who needs the bank detail for a verification call retrieves it exactly as today, under the same permissions.
On approval, the payment file carries the real account number to the payment system, at that destination and nowhere else. The model that drafted everything never received it. The full run, the approval, the approver and the time land in the log as tokens, exportable to the SIEM. When the market conduct exam or the internal audit arrives, the answer to who approved this payment is retrieved, not reconstructed.
What goes wrong without the named approver?
Carriers do not need hypotheticals here; adjacent history is instructive. Authority failures in claims have always produced the same sequence: a payment that should not have gone out, an investigation into who authorized it, and a control finding when no clean answer exists. An unsupervised agent reproduces that sequence with the name permanently blank.
There is also a quieter failure. When the sanctioned path cannot handle payments, pressure builds toward the unsanctioned one: an agent wired into the payment workflow by a capable engineer, outside review, with accumulated credentials. That untethered pattern, and the July 2026 OpenAI sandbox escape that shows what unattached agents do with permissive environments, are examined in What Is a Tethered Agent?
What happens to settlement authority when the agent arrives?
Settlement authority is the insurance industry’s oldest accountability technology, and it is worth being concrete about what the agent does and does not change in it.
Before the agent, the structure looks like this: an adjuster holds authority to settle within a band, say to twenty-five thousand dollars; a supervisor holds the next band; committee or officer approval sits above that. Authority is granted in writing, reviewed periodically, and exceeding it is a compliance event regardless of whether the payment was correct. The structure exists because carriers learned, over a century, that the question who may commit the company’s money must never have a fuzzy answer.
After the agent, under Supervised Delegation, the bands are untouched. The agent prepares payments; it holds no authority at all, in any amount, which places it outside the structure rather than at the bottom of it. The queued payment routes to the adjuster whose authority covers it, exactly as a file prepared by a junior examiner routes today. A payment above the adjuster’s band escalates on the same rails it always did. The record gains one new artifact: the run history showing how the payment was prepared, attached to the approval showing who committed it.
Contrast the alternative honestly, because vendors do propose it: granting the agent its own authority band, small at first, for payments below some threshold. Every increment is individually defensible and the sequence ends somewhere no carrier’s compliance framework has been: an occupant of the authority structure that no regulator licensed, no employment agreement binds, and, as OWASP’s first-ranked risk documents, content on a screen can instruct. The conservative design is not a smaller band for the agent. It is no band, with preparation unlimited and commitment human.
What will the internal auditor ask?
Internal audit reaches agent payments with a standard toolkit, and the architecture can be scored against the five questions that toolkit will produce.
Who authorized this payment? The approval record: name, timestamp, the file as the approver saw it. Was the authorizer within authority? The payment routed under the existing bands; the record shows the band and the amount. What information supported the decision? The run history: every screen the agent read, as tokens, replayable. Could the agent have paid without approval? No; the gate is policy-triggered and architectural, not a model behavior, and the auditor can verify the property in the design rather than sampling for exceptions. Who set the agent up, and under what policy? The setup record: applications connected, credentials used, exposure policy confirmed, with the name of the person who did each.
Auditors will notice what is missing from those answers: any reliance on the model having behaved well. The control set is the same shape as the controls around a human payment clerk, deliberate grant, defined authority, supervisory review, complete records, which is what lets the audit close without a novel-technology finding. The deployments that generate findings are the ones where the answer to any of the five questions is a description of what the model usually does.
The disputed payment, replayed
The architecture earns its keep on the bad day, so construct one. Eight months from now, a claimant’s attorney disputes a settled claim: the payment was wrong, the process was automated, and discovery wants everything.
In the unsupervised deployment, the carrier’s counsel assembles what exists: system logs showing an agent executed the payment, model vendor documentation about how the agent generally behaves, and whatever prompt history survived retention policies. The question who decided this claim has no name in it. Counsel is now defending a process, and the process is a black box with the carrier’s name on it. Regulators reading the same file reach the same place, and the market conduct finding writes itself: consequential claim decisions executed without human authorization, contrary to the written program the state bulletin required.
In the supervised deployment, counsel pulls the record. The delegation: adjuster’s name, systems connected, exposure policy confirmed on a date. The run: every screen the agent read, as tokens, replayable in order. The decision: the payment approved by the named adjuster, timestamp attached, with the file as she saw it preserved, coverage summary, estimate, the sublimit flag she accepted. The dispute is now about whether a licensed adjuster’s documented judgment was correct, which is a dispute insurance has litigated for a century and knows how to win or settle on the merits.
The comparison is not about which carrier behaved better. The claims may be identical, the payments identical, the outcomes identical. The difference is entirely in what can be shown afterward, and that difference was decided months earlier, at architecture selection, by whether the payment gate put a person on the record. Carriers evaluating agent platforms can run this thought experiment against any vendor’s audit story before buying, and should.
What belongs on the approval screen?
If the whole design funnels judgment to one moment, the screen at that moment deserves design attention, and the requirements can be stated as what the approver must be able to answer without leaving the page.
What am I approving? The action, stated as an action: pay ACCT_001 the amount on claim CLM_001. Not a summary of the agent’s activity; the specific consequential step awaiting authority.
On what basis? The file as prepared: coverage basis, the estimate, the policy provisions applied, the sublimit or deductible arithmetic, with the figures in clear because figures are what the judgment runs on, and identity as tokens because identity is not what the judgment runs on.
What did the agent do to prepare this? The run history, one click deep: every screen read, every source consulted, every draft produced, in order, as tokens. The approver who wants to audit the weather verification or the loss-history pull can, in seconds. Most approvals will not open it; its availability is what makes the ones that do meaningful.
What was flagged? Anything the agent escalated, and anything policy requires surfacing: a prior claim pattern, a coverage question, a mismatch between estimate and photos. Flags the approver has to hunt for are flags that will be missed, and the log records which flags were on screen at approval.
What happens on my approval? Where the real values resolve and what leaves the building: the payment file to this bank portal, the letter to this address. The approver should never discover after the fact what their approval released.
The list is short because attention is the resource being budgeted. Every element that does not serve the decision dilutes the elements that do, and a cluttered approval screen is how conscientious reviewers become rubber stamps. Carriers evaluating platforms can put the five questions to any vendor’s demo and watch how many clicks each answer takes.
What the record shows
A named adjuster approves the payment, every time, on the record: that is the answer, and everything else in the design exists to make it true and provable. The agent prepares at machine speed; settlement authority stays personal, banded and unchanged, with the agent holding no band at all. The gate is policy-triggered because a gate the model can decline to invoke is not a gate, a point Microsoft’s own agent documentation concedes. The NAIC bulletin’s accountability requirements, now adopted across half the states, presume a name behind every consequential decision, and the disputed-payment scenario shows what the presence or absence of that name costs eight months later. The approver sees a tokenized file with the figures in clear, the run history a click away, and real values resolving only at the destination on their approval. The question in this article’s title will be asked by an examiner with a claim number in hand. The architecture decides today whether the answer is a name or an investigation. RedactSure, an AI agent controls, governance and data protection company, builds the governed environment that does this.
Frequently asked questions
Doesn’t the approval step give back the efficiency the agent created?
No. The hours in a claim are in the gathering, matching, pricing and drafting, which run unattended. The approval is minutes, and it replaces a review the adjuster was already obligated to perform. What disappears is the assembly work, not the judgment.
Can approval thresholds vary by claim size?
The carrier’s existing authority bands apply unchanged. A payment within the adjuster’s authority waits for that adjuster; one above it escalates exactly as it does today. The agent changes nothing about who may approve what.
What if the adjuster rubber-stamps?
Rubber-stamping is a supervision quality problem, and it is at least visible: the record shows who approved, when, and how long the review took. The alternative designs make the problem invisible by removing the approver entirely.
Does the model ever release a payment on its own?
No. The AI never settles, pays or decides. The gate is policy-triggered and applies every time, whether or not the model would have asked.
How is this different from the agent asking for confirmation?
An agent that asks is exercising discretion about whether to ask. A policy gate removes the discretion. Microsoft’s own guidance for computer-use agents warns against relying on model-requested review as a fail-safe; the difference is a gate versus a doorbell.
What does the examiner get?
The written exposure policy for the workflow, the setup record showing who connected which systems, tokenized run logs, and the approval trail with names and times, exportable from the SIEM. That evidence set maps directly onto the NAIC bulletin’s written-program and accountability requirements.
Who supervises agents on non-payment workflows, like subrogation letters?
The same model at lower intensity: the person who owns the work grants access and confirms the exposure policy, and the consequential gate sits wherever the workflow’s consequences sit, a filing deadline submission for subrogation, a report release for fraud review.
Related reading
Can an AI Agent Work Claim Files in Guidewire Without Exposing PII? · What Is Supervised Delegation? · What Is a Tethered Agent? · On the RedactSure blog: Secure AI Across the Insurance Value Chain
Sources
Regulation
- Quarles & Brady, “Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers’ Use of AI” (March 2025). https://www.quarles.com/newsroom/publications/nearly-half-of-states-have-now-adopted-naic-model-bulletin-on-insurers-use-of-ai
- NAIC, adoption map for the Model Bulletin: Use of Artificial Intelligence Systems by Insurers. https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf
Standards and vendor documentation
- Microsoft Learn, “Human supervision for computer use,” Microsoft Copilot Studio. https://learn.microsoft.com/en-us/microsoft-copilot-studio/human-supervision-computer-use
- OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
Incidents and reporting
- NPR, “OpenAI blamed a hacking event on its AI models gone rogue” (July 2026). https://www.npr.org/2026/07/23/g-s1-135085/openai-hacking-ai-models
RedactSure documents
- RedactSure, “Secure AI Across the Insurance Value Chain” (2026). https://redactsure.com/blog/secure-ai-across-the-insurance-value-chain/
- Product behavior described on this page reflects RedactSure’s current design.
Bring your hardest questions.
A 25-minute AI Agent Security Review with the founders: threat model, token design, egress paths, audit schema. Or a 25-minute demo on a workflow like yours, with the data hidden from the AI and a named person approving what matters. We come with diagrams, not a pitch deck.
Book a security review Book a demo · Something elseAbout the author
Chris Sowa is a founder of RedactSure and a former CEO of AI companies; he started his first years before ChatGPT existed. He previously led AI at Accenture, served as Global VP of Strategy & Innovation at Schneider Electric, was CCO of Sovos, and spent more than a decade at Oracle, with earlier roles at SAP and IBM.