Skip to content
redactsure
Book a review

Explore.

Explainer · RedactSure Research

What Should the AI See for This Piece of Work? The One Question That Turns AI Security From a Posture Into a Strategy

Exactly the data the task requires and nothing more, decided in advance by the person who owns the work. The question sounds small. It is the working form of an AI security strategy, because every current posture is an evasion of it: lock-everything-down never asks it, a Copilot that sees too little answers it by producing nothing valuable, and an agent wired into everything answers it with whatever is on the screen. An organization that can answer the question, workflow by workflow, field by field, in writing, has a strategy. One that cannot has a posture, and postures are what the cancellation statistics count. The environment that enforces this is built by RedactSure, an AI agent controls, governance and data protection company.

Key findings

Why is this the strategy question?

Enterprise AI strategy, as practiced, oscillates between two failure poles that the opening of every RedactSure strategy document names. Lock everything down: ban the tools, restrict the copilot to public documents, keep the logs quiet. The result is measured elsewhere in this series: 78% bring-your-own use, shadow AI in one of five breaches, value forgone and exposure retained. Or wire an agent into everything: give it the screens, the credentials and the queue, and collect the capability along with a model context holding every record the workflow touches, one successful prompt injection from becoming a disclosure.

Both poles share a root: nobody decided what the AI should see. The lock-down never reaches the question because nothing is allowed; the wire-in never reaches it because everything is. The question forces the third position, which is the only one that survives both the security review and the value review: for this piece of work, these fields in clear because the task runs on them, those fields as tokens because the task does not, decided by name, enforced at the screen.

Copilot alone is not AI transformation. An enterprise LLM contract alone is not AI governance. The enterprise must govern the workflow between them, and this question is what governing the workflow means in practice.

How is the question answered for one workflow?

The exercise is concrete enough to run in a single meeting per workflow, and the output fits on a page.

Name the task, not the tool: prepare denial appeals; reconcile purchase orders; draft attendance letters. Then walk the screens the task touches and sort every field into two columns. What does the reasoning actually run on? Amounts, dates, codes, statuses, histories, policy language: the structure of the work. That column stays in clear. What identifies a person or an account without feeding the reasoning? Names, government identifiers, account and card numbers, contact details, addresses: that column tokenizes, with consistency preserved so the agent can still match records it cannot read.

Three tests keep the sort honest. The task test: would the work product change if this field were a consistent token? For identifiers the answer is no in workflow after workflow, which is the empirical surprise this whole approach rests on. The regulator test: which column would the examiner, under the minimum necessary standard, FERPA’s conditions or PCI scope, expect this field in? The breach test: if the model’s context from this run leaked in full, which fields would trigger a notification? Fields failing the breach test and passing the task test have no business in clear, and almost none survive the pair.

Workflow Runs on (stays clear) Identifies (tokenizes)
Denial appeal CPT and diagnosis codes, dates of service, denial reason, payer policy Patient name, member ID, SSN
Claims intake to payment Coverage terms, loss details, estimate pricing Claimant identity, bank details, medical identifiers
Vendor payment run Amounts, PO numbers, terms, dates Vendor identity, account and routing numbers
Attendance follow-up Attendance pattern, dates, program rules Student name, family contacts
Dispute investigation Transaction facts, merchant descriptors, network rules PAN, cardholder identity

The person who signs the page is the person who owns the work, because they are the one who can answer the task test and the one who answers for the workflow under Supervised Delegation. The signature is the point: an exposure decision nobody owns is an exposure default nobody chose.

What makes the answer real?

A signed page changes nothing by itself; three mechanisms turn it into a control.

Enforcement at the render. The policy binds at the one point every application passes through on the way to being read: the screen, as it renders, inside a governed environment, before any model sees it. Enforced there, the policy holds for every application in the workflow without integration projects; enforced anywhere else, upstream in pipelines that miss most applications, or downstream in filters that inspect what the model already received, it leaks by construction. The mechanics are the subject of render-layer tokenization.

Verification per run. Every screen the agent read is logged as the model received it, tokens included, and the log exports to the SIEM. The question did the policy hold has a per-run answer, which makes the exposure policy one of the few AI governance documents whose implementation an auditor can check directly rather than attest around.

Review on change. Workflows change, fields get added, and the policy is a living document with an owner, revisited when the workflow is, with changes on the record. The staleness failure mode of governance documents is answered the same way access reviews answer it: by making review someone’s named job.

An organization can begin before any of this tooling exists, and should: put the question in the security review template today, require the two-column page for every proposed agent workflow, and watch which proposals cannot produce one. The stalled and the sound projects sort themselves immediately, which is why the failure analysis keeps arriving back at this question.

How does the working session actually run?

The one-meeting claim deserves an agenda, because the session succeeds or fails on its structure. The room needs three people: the workflow owner, who knows the task; someone who knows the screens, often the same person or a systems analyst; and whoever will sign off for security. An hour covers a workflow of ordinary complexity.

First fifteen minutes: name the task and walk it once, screen by screen, listing every field that appears. No judgments yet; the output is the inventory, and it is usually the first time anyone has written down what the workflow’s screens actually contain, which is itself a finding.

Next twenty minutes: run the sort. For each field, the task test first, would the work product change if this were a consistent token, answered by the owner, who is the only person in the building who actually knows. Disagreements get resolved by walking the specific step where the field is claimed to matter; most dissolve there, because “we need the name” usually means “the letter needs the name,” which resolution at the approved send handles without the model ever holding it.

Next fifteen minutes: apply the regulator and breach tests to the fields the task test left in clear, which is where the occasional real tension appears, a field the task plausibly uses that the breach test flags. Those go to the policy as tokenized with a documented exception path, or in clear with a documented justification; either way the decision is written, which is the entire point.

Final ten minutes: the owner reads the two columns back, confirms, and signs. The page gets a name, a date and a review trigger, when the workflow changes, and enters the same binder as the access reviews.

Two session failure modes are worth naming. A room without the workflow owner produces a guessed policy that the real owner will quietly route around, which is the shadow pattern reborn inside the governed path. And a session that tries to cover five workflows produces five vague pages instead of one usable one; the discipline is one task, one hour, one signed page.

What the record shows

What should the AI see for this piece of work is the question that separates an AI strategy from an AI posture. Its answer is a field-by-field exposure policy per workflow, owned and signed by the person accountable for the work: the structure the reasoning runs on stays in clear, the identifiers tokenize with consistency preserved, and the sort is disciplined by three tests any reviewer can apply. The answer becomes real through enforcement at the render, verification per run and review on change. Lock-everything-down never asks the question and collects shadow use; wire-in-everything answers it by accident and collects the breach surface; the strategy position answers it on purpose, in writing, and is the only one of the three that survives both the security review and the value review. The question costs nothing to start asking. Programs are sorted by whether they can answer it. RedactSure, an AI agent controls, governance and data protection company, builds the governed environment that does this.

Frequently asked questions

Who should ask this question inside an organization?

Security should ask it of every proposed agent workflow; workflow owners should answer it; executives should ask why any deployment lacks the page. It works as a review gate exactly because each role has a natural claim on one side of it.

Is the answer ever “the AI should see everything”?

For record-bearing enterprise workflows, a full-visibility answer that survives the three tests has not shown up. Where a field is load-bearing for the task it stays in clear; that is the policy working, not an exception to it.

How is this different from data classification programs we already run?

Classification labels data at rest by sensitivity. This question is task-relative and enforced at read time: the same field can tokenize in one workflow and stay clear in another, because the tasks differ. Classification informs the sort; it cannot replace the per-workflow decision.

What about AI uses that are not agents, like chat assistants?

The question applies to any system that puts enterprise content in front of a model; agents make it urgent because they read whole screens at scale. Asking it of the chat estate usually reveals the middle-row problem described in the ban analysis: the sanctioned tool is barred from the work that needed governing.

How long does answering take for one workflow?

One working session with the workflow owner and someone who knows the screens, in most cases. The two-column page is deliberately short; the discipline is in the three tests, not the length.

What is the fastest way to see the question’s value?

Ask it, today, of the agent deployment or proposal furthest along. Either the page appears, which is evidence of a governable program, or it cannot, which is the finding.

What Is Least Exposure? · What Is Render-Layer Tokenization? · Why Do Most Agentic AI Projects Fail to Reach Production in Regulated Industries? · On the RedactSure blog: Your AI Strategy Is Probably Wrong

Sources

Research and industry data

  1. Microsoft and LinkedIn, Work Trend Index. https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part
  2. IBM, Cost of a Data Breach Report 2025. https://www.ibm.com/reports/data-breach
  3. OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/

RedactSure documents

  1. RedactSure, “Your AI Strategy Is Probably Wrong” (2026). https://redactsure.com/blog/your-ai-strategy-is-probably-wrong/
  2. RedactSure, “The Two Gaps AI Agents Opened in Your Security Stack” (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
  3. Product behavior described on this page reflects RedactSure’s current design.

Bring your hardest questions.

A 25-minute AI Agent Security Review with the founders: threat model, token design, egress paths, audit schema. Or a 25-minute demo on a workflow like yours, with the data hidden from the AI and a named person approving what matters. We come with diagrams, not a pitch deck.

Book a security review Book a demo · Something else

About the author

Chris Sowa is a founder of RedactSure and a former CEO of AI companies; he started his first years before ChatGPT existed. He previously led AI at Accenture, served as Global VP of Strategy & Innovation at Schneider Electric, was CCO of Sovos, and spent more than a decade at Oracle, with earlier roles at SAP and IBM.