Explainer · RedactSure Research
What Does an AI Agent See When It Takes a Screenshot? Everything on the Page, Including What the Task Never Needed
Everything the page renders, whether or not the task needs it. A screenshot-based agent receives the whole picture: the record it was asked to work, every other value on the screen, text styled so a person cannot see it, and any instruction an attacker placed on the page. Least Exposure is a security principle for AI agents: for each piece of work, the agent receives exactly the data the task requires and nothing more, enforced before any model reads the screen. A screenshot is the opposite design. It hands the model the maximum and relies on the model to ignore the rest. This page walks what a screenshot carries, why the picture is the injection surface, what the 2026 agent platforms protect and what they read raw, and the two mechanisms that limit what a model is handed. The environment that enforces this is built by RedactSure, an AI agent controls, governance and data protection company.
Key findings
- A computer-use agent that works from screenshots reads the rendered page as an image. What enters its context is everything visible: the fields the task uses, the fields it does not, sidebars, notifications, and any text on the page regardless of who put it there.
- Prompt injection, which the OWASP Top 10 for LLM Applications 2025 ranks first, works through what the agent reads. The screenshot is the largest possible reading surface, so it is the largest possible injection surface.
- The agent platforms that shipped in 2026 protect what the user deposits. Meta’s Muse holds credentials outside the agent and swaps a surrogate for the real secret at the network boundary; 1Password for Claude fills passwords through a channel the model never sees. Neither touches what the agent reads on the page.
- Meta’s own safety write-up lists injection through data the agent observes on the web as a known risk and names human approval and deterministic boundaries as the mitigation. The boundaries limit what a persuaded agent can do afterward; they do not change what it read.
- Two mechanisms limit what a model is handed under Least Exposure. Render-layer tokenization governs which values it receives. Handing the agent the fields the work needs, rather than the picture, governs which parts of the screen it receives at all.
- A successful injection against a screenshot agent collects the context, identifiers included. Against an agent holding tokens and the task’s fields, it collects tokens.
What is in a screenshot that the task never asked for?
Take a claims screen. The task is to verify coverage for a wind and hail loss. The fields the work needs are the policy status, the deductible, the dwelling limit and the loss date. The screen also shows the policyholder’s name, address, date of birth, Social Security number, the mortgagee’s account and routing numbers, the adjuster’s queue of other claimants in a sidebar, and an internal note pasted into a free-text box. A person working the screen reads the four fields and ignores the rest. A screenshot agent reads all of it, because the image does not know which pixels the task cares about.
The same is true of every application an agent might operate. A patient account screen carries the diagnosis the task needs and the name, date of birth and insurance ID it does not. A student record carries the attendance pattern and the child’s name and family contacts. An email client carries the message the agent was asked to draft a reply to and the fifty other subject lines in the inbox. The picture is the whole window, and the whole window is what enters model context.
Three categories of unneeded content arrive with every screenshot. The first is other people’s data, on the same screen for the operator’s convenience. The second is the application’s own furniture: menus, banners, notifications, advertisements, which carry no task value and can carry anything else. The third is content nobody intended the agent to read, because it was never intended for a person either: text set to the background color, text positioned off the visible area but still rendered, instructions inside an image. A person never sees the third category. A model reading pixels reads all three the same way.
Why is the picture the injection surface?
An AI agent can be instructed by content it reads. That is the mechanism behind prompt injection, the risk OWASP ranks first for language-model applications, and it has no dependence on where the content came from. A hidden line on a web page, a comment field in a record, a sentence in an email: if it enters the model’s context, it can steer the model.
The research on model-layer defenses points the same way. Anthropic’s engineering account of how it contains its agents describes an internal red-team exercise in which a malicious prompt, delivered through a phished employee, got the coding agent to read cloud credentials and post them to an external endpoint in 24 of 25 attempts, and notes that model-layer defenses anchor on user intent, which is exactly what an injection borrows. Meta’s write-up on Muse safety names prompt injection via data the agent observes on the web as a known risk and places the defense below the model: approvals before actions that move data out, and deterministic boundaries that hold even if the agent is persuaded to behave badly. Both vendors, from different directions, reach the same position. Assume the agent can be talked into misbehaving; limit what it can do and what it holds when that happens.
A screenshot maximizes what it holds. Every one of the three unneeded categories above is a place an instruction can sit, and the agent that reads the whole picture reads every one of them. Reducing the reading surface is therefore a structural defense, not a probabilistic one: content the model is never handed cannot instruct it, whatever the detector missed.
What do the current agent platforms protect, and what do they read raw?
The consumer and developer agents of 2026 converged on one pattern for secrets, and it is worth stating what it covers and what it does not, as of September 2026 and from each vendor’s own page.
| Platform | What is protected | How | What the agent reads raw |
|---|---|---|---|
| Meta Muse (Secure VM) | Passwords, OAuth tokens, payment cards | Credentials held outside the agent’s runtime; the agent holds a surrogate token and the Sentinel swaps in the real value at the network boundary after the action is authorized; a one-time card number at checkout | The rendered page. Muse runs a full Chromium browser behind a virtualization layer, and everything the browser displays enters the agent’s context |
| 1Password for Claude | Passwords and one-time codes | Filled into the page through a channel 1Password manages, scoped to the task; the secret never enters the model’s context or the vendor’s systems | The rendered page and every value on it |
| Screenshot-based computer-use agents generally | Nothing by default | The agent receives an image of the screen and acts on it | Everything visible |
The pattern is a vault. It protects what the user hands over in advance: a password typed into the store, a card saved for checkout. The strength is real; the 1Password integration and Meta’s surrogate-token design both mean a perfect injection cannot extract a credential, because the agent never held one. The limit is equally real. A vault has no entry for the claimant’s Social Security number on a claims screen, because nobody deposited it; it was already there when the agent arrived. Enterprise records live in the second category. Employees do not deposit customer PII into a vault before opening the system of record. It is on every screen, and a screenshot agent reads it as it reads everything else.
Muse and the credential vaults are strong at what they set out to do. The question this page asks is about the data they were not designed to protect, and the answer is that the picture is read in full.
Two ways an agent can see a screen
An agent can be handed the picture, or it can be handed the fields.
The picture is the screenshot: a rendering of the whole window, read as an image, sometimes accompanied by an accessibility tree or the page’s structure. It requires no knowledge of the application and works on anything that draws to a screen, which is why it became the default for computer-use agents. It also carries everything, for the same reason.
The fields are the structured content of the page: this label, this value, this control, in this order. Handing an agent the fields means the environment decides which parts of the screen the model receives for this piece of work, and hands it those. The claims task above receives policy status, deductible, dwelling limit and loss date. The sidebar of other claimants, the banner, the hidden text and the free-text note outside the task’s fields are not in the model’s context, because the environment did not put them there.
Least Exposure names the principle behind the second design and applies it to the mechanism itself. For each piece of work, the agent receives exactly the data the task requires and nothing more. Two mechanisms enforce that at the screen. Render-layer tokenization governs which values the model receives: identifiers are replaced with consistent tokens (USER_001, SSN_001, ACCT_001) before any model reads the page, and real values resolve only at approved destinations at the moment a named person approves the action. Handing the agent the fields governs which parts of the screen it receives at all. The RedactSure environment does both: the agent operates the application through a structured reading of the rendered page rather than through a picture of it, which is more efficient than reading pixels and hands the model a smaller surface than the whole window, with the identifiers already tokenized when it arrives.
A structured reading is not, by itself, a smaller one. A raw accessibility tree or a full DOM extract carries the same content as the screenshot in a different form. What makes the surface smaller is the policy: the field-by-field decision, confirmed by the named person who owns the workflow under Supervised Delegation, about what this task is handed. That decision is the exposure policy, and it is the same document that governs tokenization.
What does a successful injection get in each design?
Assume the attack works completely in both cases. A page the agent reads carries a hidden instruction to send everything it knows to an outside address, and the model complies.
In the screenshot design, the agent’s context holds the whole picture: the record, the other claimants in the sidebar, the identifiers, the account numbers. The credential vault holds the password back. Everything else is in context in the clear, and the injected instruction moves it through whichever channel the agent is permitted to use. If the user has set that channel to always allow, no approval intervenes.
In the fields-and-tokens design, the agent’s context holds the four fields the task needed, with USER_001 and ACCT_001 where the identifiers were. The hidden instruction may never have reached the model at all, because it sat outside the task’s fields. If it did reach the model, through a free-text field the task legitimately reads, the agent complies and sends tokens, which resolve nowhere except at an approved destination at the moment a named person approves the action. The prompt-injection article walks the full sequence and the SIEM record it leaves.
The residual case is worth stating plainly. Reading fields narrows the surface; it does not remove it. A note box the workflow needs is a place an instruction can still sit, and non-identifier content the model legitimately reads, amounts, narrative text, contract terms, can be moved by the same route. What the second design guarantees is that identifiers cannot leak, because the model never held them, and that the consequential action stops for a person. What it does not guarantee is that the model was never instructed. No agent design guarantees that, and a vendor who claims otherwise is describing a detector’s success rate, not an architecture.
What the record shows
An AI agent that takes a screenshot sees everything the page renders: the fields its task needs, the fields it does not, other people’s records on the same screen, the application’s furniture, and any text an attacker placed there, visible or not. That reading surface is the injection surface OWASP ranks first, and the agent platforms of 2026, Meta’s Muse and the credential-vault integrations among them, protect only what the user deposits in advance, leaving the page itself read raw, as their own documentation describes. Least Exposure names the alternative: the agent receives exactly the data the task requires and nothing more, enforced before any model reads the screen. Two mechanisms carry it. Render-layer tokenization governs which values the model receives; handing the agent the fields the work needs, rather than the picture, governs which parts of the screen it receives at all. Under both, a successful injection collects tokens from a narrower surface, and the consequential action waits for a named person. Neither mechanism makes the model immune to instruction. They make the instruction find less, and the identifiers nothing. RedactSure, an AI agent controls, governance and data protection company, builds the governed environment that does this.
Frequently asked questions
Does reading fields instead of pixels stop prompt injection?
No. It narrows the surface. Content outside the task’s fields never reaches the model, so instructions hidden there cannot steer it, but a free-text field the task legitimately reads can still carry one. The protection that holds regardless is the token: what the persuaded agent sends is SSN_001, and the consequential action stops for a named person.
Is an agent that reads the accessibility tree or the DOM already doing this?
Partly. A structured reading is easier to filter than an image, but a full tree or DOM extract carries the same content as the screenshot in a different form. The surface becomes smaller when a policy decides which fields the task is handed, confirmed by the person who owns the workflow, and that policy is what Least Exposure asks for.
Doesn’t Meta’s Muse already hide data from the agent?
It hides credentials and payment cards, and does so well: the agent holds a surrogate and the real value is substituted at the network boundary after the action is authorized. The page the agent reads is not covered by that mechanism. Meta’s own safety write-up describes a full browser behind a virtualization layer and lists injection through observed web data as an open risk mitigated by approvals and boundaries.
Does the agent lose anything by not seeing the picture?
For the workflows measured so far, no. The work runs on the fields: amounts, dates, codes, statuses, the relationships between records. Layout is what a person needs to find those fields; the environment finds them for the agent. The steps that need a real value are the consequential actions, where the value resolves at the destination when a named person approves.
Does this apply to an agent we already run somewhere else?
No. Both mechanisms work because the agent does its work inside the RedactSure environment, where the render can be read as fields and changed before the model sees it. An agent running on another platform reads whatever its own screen shows. The workflow moves into the environment; the applications and the permissions stay exactly where they are.
What should a security team ask any agent vendor?
Show me what the model received for one screen of my workflow. Whether the answer is a picture or a set of fields, and whether the identifiers in it are values or tokens, settles the visibility question in one document. The audit article walks the rest of the record.
Related reading
What Is Least Exposure? · What Is Render-Layer Tokenization? · What Does a Prompt-Injection Attack Get From an Agent That Sees Only Tokens? · Enterprise Browser vs. Render-Layer Tokenization · Can a Credential Vault Protect the Data an AI Agent Reads? · Secure VM, Confidential VM, or Render Layer? · What Is the Enterprise Version of Meta Muse? · On the RedactSure blog: The Two Gaps AI Agents Opened in Your Security Stack
Sources
Standards
- OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
Vendor documentation
- Meta AI Research, “How We Built Safety Into Muse: Security and Safety for AI Agents” (September 2026). https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
- 1Password, “1Password and Anthropic Bring Secure Credential Access to Claude” (July 16, 2026). https://1password.com/press/2026/july/1password-for-claude
- Anthropic, “How we contain Claude across products” (2026). https://www.anthropic.com/engineering/how-we-contain-claude
RedactSure documents
- RedactSure, “The Two Gaps AI Agents Opened in Your Security Stack” (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
- Product behavior described on this page (structured reading of the rendered page, render-layer tokenization, task-level exposure policy, supervised approval) reflects RedactSure’s current design.
Bring your hardest questions.
A 25-minute AI Agent Security Review with the founders: threat model, token design, egress paths, audit schema. Or a 25-minute demo on a workflow like yours, with the data hidden from the AI and a named person approving what matters. We come with diagrams, not a pitch deck.
Book a security review Book a demo · Something elseAbout the author
Chris Sowa is a founder of RedactSure and a former CEO of AI companies; he started his first years before ChatGPT existed. He previously led AI at Accenture, served as Global VP of Strategy & Innovation at Schneider Electric, was CCO of Sovos, and spent more than a decade at Oracle, with earlier roles at SAP and IBM.