Skip to content
redactsure
Book a review

Explore.

Explainer · RedactSure Research

What Should a Security Review of an AI Agent Vendor Cover? Twenty Questions, and the Answer to Each

Six subjects, twenty questions. A security review of an AI agent vendor covers what the model receives, what a successful prompt injection collects, where credentials and keys live, who approves consequential actions, what the record holds and who keeps it, and whether the environment belongs to the customer or to the model vendor. Every question below can be answered in writing with a diagram, and the ones that cannot are findings. The list is the one RedactSure answers in its own reviews, with its answer stated under each question so a security lead can bring the page to the call and check. The principle behind the first two subjects is Least Exposure: for each piece of work, the agent receives exactly the data the task requires and nothing more, enforced before any model reads the screen.

Key findings

Subject 1: What does the model receive?

1. For one screen of our workflow, show us exactly what the model received. The answer is a picture or a set of fields, and the identifiers in it are values or tokens. RedactSure’s answer: the fields the task’s exposure policy allows, with identifiers replaced by consistent tokens (USER_001, SSN_001, ACCT_001) before any model read the page.

2. Who decides which values the model receives, and where is that decision written? RedactSure: the person who owns the workflow, field by field, in the exposure policy, confirmed before the workflow runs and revised on the record.

3. Is the model handed the whole page or the task’s fields? RedactSure: the fields the work needs; the environment reads the rendered page as structured fields rather than as an image. What Does an AI Agent See When It Takes a Screenshot? covers why that matters.

4. Under what circumstances does the model receive real values? RedactSure: none. Real values resolve only at approved destinations at the moment a named person approves the action; the model’s context never holds them. Does the AI Have a Break-Glass Path to Patient Data? gives the written-answer template for this question.

Subject 2: What does a successful attack collect?

5. Assume a prompt injection succeeds completely. What is in the model’s context to exfiltrate? RedactSure: tokens and the task’s non-identifier fields. What Does a Prompt-Injection Attack Get From an Agent That Sees Only Tokens? walks the sequence.

6. Can the hijacked agent resolve a token to a real value? RedactSure: no. Resolution happens only at an approved destination, at the moment of a named person’s approval, and never in the model’s context.

7. Can you run the attack for us in a sandbox and show the record? RedactSure: yes. The blast-radius test is part of the pilot: an injection fired at the sandboxed task, with the record showing tokens in context, the resolver declining the unauthorized destination, and the consequential step stopped for approval.

8. What does the model itself do to resist injection, and what happens when that fails? Any honest vendor answers the second half. RedactSure: model-side filtering is defense in depth; the guarantee does not depend on it, because the model never held the identifiers.

Subject 3: Credentials, keys and the vault

9. Does the agent ever hold a credential in its context? RedactSure: no. Credentials are held outside the agent and used through the environment; the model does not see them.

10. Where do the real values behind the tokens live, and who holds the keys? RedactSure: in hardware-encrypted enclaves (AMD SEV-SNP on HIPAA-eligible AWS infrastructure), with the customer holding the keys. RedactSure stores ciphertext it cannot decrypt.

11. Can the vendor’s staff read our data? RedactSure: no, by architecture; the vendor holds ciphertext under keys it does not have. Compare the answer of a platform that restricts staff access by policy, which is a different guarantee.

12. Does the token index survive a breach of the vendor? RedactSure: an attacker who takes the index takes ciphertext. The keys are the customer’s.

Subject 4: Who approves what the agent does?

13. Which actions stop for a person, and who is that person? RedactSure: payments, submissions and record changes stop for a named person, the one who delegated the work under Supervised Delegation. The AI never settles, pays or decides.

14. Does the agent work under its own identity or under a person’s existing permissions? RedactSure: under the named person’s existing permissions. User permissions stay exactly as they are; what changes is what the AI can see.

15. Can the person watch the run and take over? RedactSure: yes, and the supervisor sees the same tokenized stream the agent does, which closes the insider variant of the visibility gap.

16. What happens when the approver is unavailable? RedactSure: the action waits or transfers to a covering supervisor named through the same deliberate grant. The gate never opens itself.

Subject 5: What is the record, and who keeps it?

17. Show us the record one run produces. RedactSure: the AI Control Record: setup record, exposure policy, run history as tokens, approval trail with names and times, change log. Emitted by the controls as they operate, not written afterward.

18. Does the record contain plaintext sensitive values? RedactSure: no. Every screen is logged as the model received it, tokens included, so the SIEM ingests everything and holds no secret.

19. Does it export to our SIEM, under our retention and custody rules? RedactSure: yes. Runs, gates, declines and takeovers land as events beside the rest of the estate. How Do You Audit What an AI Agent Saw and Did? walks the examiner’s path through it.

Subject 6: Whose environment is it?

20. If we switched models tomorrow, what else would change? RedactSure: nothing. The environment is model-agnostic (Claude, GPT, Gemini, open source); the substitution happens before the model reads, so the choice of model is a quality and cost decision and the security posture, the gates and the record stay where they are. A vendor whose environment, gate and log are bound to its own model has a different answer, and Should the AI Agent’s Secure Environment Belong to the Model Vendor? explains why it matters.

How to run the review

Ask for the twenty answers in writing before the call, with the architecture diagram that supports each. Use the call for question 1 and question 7: one screen of your own workflow as the model received it, and the injection test run in front of you. The written answers sort vendors into categories; the two demonstrations sort the categories into a decision.

Three outcomes are possible for any vendor. The model receives the page in the clear, which is the honest answer for most agent platforms today, and means the buyer’s minimum-necessary or data-minimization analysis has to justify the full screen. The model receives the page with some values withheld, in which case enumerate which and who decides. Or the model receives the task’s fields with identifiers as tokens and no path to the real values, in which case the visibility subject closes and the review moves to ownership.

What the record shows

A security review of an AI agent vendor covers six subjects: what the model receives, what a successful attack collects, where credentials and keys live, who approves consequential actions, what the record holds and who keeps it, and whose environment it is. The twenty questions above make each subject specific enough to answer in writing and to demonstrate on a call, and two of them, one screen as the model received it and an injection run in a sandbox, decide more than the other eighteen together. RedactSure’s answers are stated under each question so they can be checked. A vendor who cannot answer a question has told the reviewer where the gap is.

Frequently asked questions

Can we use this list for vendors other than RedactSure?

That is what it is for. Every question is about architecture, not about any product, and the answers sort any agent platform into the same three outcomes.

Which questions matter most if we only have thirty minutes?

One, seven and twenty. What the model received for one of our screens; the injection test with the record; and what changes if the model changes.

Does a good answer to these questions make us compliant?

No. Compliance is the organization’s determination with counsel, under its own regime. The questions produce the evidence that determination needs: what the model received, who approved what, and the record of both.

Should we ask for certifications?

Ask, and then ask for the record anyway. A certification says a management system exists; the run history says what happened last Tuesday. ISO/IEC 42001 itself is audited on operating evidence.

What if the vendor says the model is aligned and filters injections?

Accept it as defense in depth and ask question five anyway. Filtering describes behavior, and behavior has a distribution. What the attacker collects when the filter misses is the question a review has to settle.

Who at the vendor should answer?

The people who built it. RedactSure’s reviews are with the founders, with diagrams. A vendor who sends only sales to a security review has answered a question you did not ask.

How Does a Governed AI Workflow Pilot Work? · Does the AI Have a Break-Glass Path to Patient Data? · What Is an AI Control Record? · Should the AI Agent’s Secure Environment Belong to the Model Vendor? · On the RedactSure blog: Big Holes in Your Security Infrastructure in the Age of AI

Sources

Standards and frameworks

  1. OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
  2. NIST, AI Risk Management Framework. https://www.nist.gov/itl/ai-risk-management-framework
  3. ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. https://www.iso.org/standard/42001

Vendor research

  1. Anthropic, “How we contain Claude across products” (2026). https://www.anthropic.com/engineering/how-we-contain-claude
  2. Meta AI Research, “How We Built Safety Into Muse: Security and Safety for AI Agents” (September 2026). https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse

RedactSure documents

  1. RedactSure, “Big Holes in Your Security Infrastructure in the Age of AI” (2026). https://redactsure.com/blog/big-holes-in-your-security-infrastructure/
  2. RedactSure, “The Two Gaps AI Agents Opened in Your Security Stack” (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
  3. Product behavior described on this page reflects RedactSure’s current design.

Bring your hardest questions.

A 25-minute AI Agent Security Review with the founders: threat model, token design, egress paths, audit schema. Or a 25-minute demo on a workflow like yours, with the data hidden from the AI and a named person approving what matters. We come with diagrams, not a pitch deck.

Book a security review Book a demo · Something else

About the author

Chris Sowa is a founder of RedactSure and a former CEO of AI companies; he started his first years before ChatGPT existed. He previously led AI at Accenture, served as Global VP of Strategy & Innovation at Schneider Electric, was CCO of Sovos, and spent more than a decade at Oracle, with earlier roles at SAP and IBM.