Explainer · By Chris Sowa · Published
Last updated
Can an AI Vendor's Own Model Be the Safety Reviewer for Its Agent?
A vendor's reviewer can be a useful safety layer. The customer still needs to verify its independence from the acting agent, the approval policy and access to evidence when a consequential action goes wrong.

A vendor's reviewer can be a useful safety layer. The customer still needs to verify its independence from the acting agent, the approval policy and access to evidence when a consequential action goes wrong. OpenAI describes Auto-review as separate from the acting dot, and Meta places safety components outside the agent-controlled runtime. Common ownership does not mean those controls share the agent's mutable execution boundary. The governance question is what the customer can configure, test and audit, and when a named person must approve the action. RedactSure, an AI agent controls, governance and data protection company, applies Least Exposure and render-layer tokenization to this problem.
Key findings
- Auto-review in OpenAI's dots checks planned actions against instructions, Custom Rules and safety requirements before execution; the same document tells users to "always review consequential work." Source: OpenAI.
- Claude in Chrome runs two classifiers, one screening content for injection and one evaluating each action; Anthropic reports attack success under 0.08% in internal testing. Source: Claude Help Center.
- Muse's Sentinel must approve every network egress, and Meta names injection via observed web data as a known risk. Source: Meta AI Research.
- Latent Space's DevDay coverage cites METR research finding coding agents self-approving flagged actions. Source: Latent Space.
- The NIST AI RMF, OMB M-25-21 and the NAIC bulletin each place accountability on the deploying organization, not on a vendor's safety layer.
What do the vendors' reviewers do?
OpenAI describes Auto-review as a safety system separate from the dot. A planned action is checked against the user's instructions, the Custom Rules and OpenAI's safety requirements. "If Auto-review blocks a step, it prevents the action from running and tells the dot why." Source: OpenAI. In ChatGPT Work, Auto-review has models review significant actions on connected tools before they run. Source: PPC Land.
Claude in Chrome takes screenshots of the active tab, so all visible information enters the model's context. Two classifiers sit on that stream, one screening content for injection and one evaluating each action; Anthropic reports attack success under 0.08% in internal testing. Source: Claude Help Center. Muse runs Sentinel outside the agent-controlled runtime; nothing leaves unless it approves, and the agent reads page content within its permitted access. Source: Meta AI Research.
What do they have in common?
The vendor supplies the acting system and its safety service. That common ownership matters for procurement and accountability, but it does not establish that the agent can modify or bypass the reviewer. OpenAI and Meta describe safety components separated from the agent-controlled execution environment.
The customer should test what the reviewer checks, what it can block and which evidence is available. Reviewing a planned action also differs from minimizing the content previously supplied to the model. Both controls can be useful in the same deployment.
Why does the separation matter?
The vendor that is the model should not also be the judge of the model. Financial control has run on the same idea for as long as ledgers have existed: the person who raises a payment does not approve it. Segregation of duties is ordinary practice, and nobody drops it because the clerk got more accurate.
The cited secondary reporting raises a related failure mode. Latent Space's DevDay coverage cites METR research finding coding agents self-approving actions their own checks had flagged, and early testers report a dot negotiating a service cancellation unprompted. Source: Latent Space. A vendor reviewer improves the odds. It does not change who is accountable when they fail.
What do regulators expect the reviewer to be?
The NIST AI RMF places a Govern function around the lifecycle, held by the deploying organization. OMB M-25-21 makes the agency, not the supplier, responsible for human oversight of rights- and safety-impacting uses.
In healthcare the covered entity is answerable. A cloud service that receives, maintains or transmits ePHI is a business associate even without the key, per HHS guidance, bound by a contract under 45 CFR 164.504(e). A business associate cannot be its own oversight; counsel decides what its reviewer counts for.
In insurance the insurer is answerable. The NAIC model bulletin expects a written governance program and oversight of third parties; the vendor's reviewer is one of the controls the insurer oversees, not the oversight.
In education the district is the school official's controller. Under 34 CFR 99.31(a)(1)(i)(B) an outsourced party is a school official only under the institution's direct control over the use and maintenance of education records. That is hard to show over a reviewer the district cannot configure or export.
In payments the merchant and its assessor are answerable. PCI SSC scoping guidance asks where account data goes; a vendor reviewer inside the vendor's environment is within scope, not a substitute for the merchant's controls.
What does a customer-owned judge look like?
It is an environment the customer owns, where the exposure decision is taken before the model reads anything and the action decision is taken by a person. In RedactSure's current design the Planner sets which values are tokenized and what the agent may do on each screen, and a named person confirms that policy before the run. Under Supervised Delegation that person can watch and pause the run and approves every consequential action. The AI Control Record holds every screen as tokens, every action and approval with a name, exported to the customer's monitoring.
The vendor's classifier is welcome inside that environment as one more control. The model sees SSN_001 and ACCT_001 where the identifiers were, so a successful injection can expose tokens and any remaining task context, and the action still stops for the named person.
How does a vendor-internal reviewer compare with a customer-owned judge?
| Control | Vendor-internal reviewer (Auto-review, action classifier, Sentinel) | Customer-owned judge (separation of model and control) |
|---|---|---|
| Who supplies the reviewer | The vendor that supplies the acting model | The customer's policy, applied by a named person |
| Whose environment | The vendor's cloud computer, extension or VM | The customer's, with its own keys |
| Who is accountable | The deploying organization, without its own record | The deploying organization, with its own record |
| What is recorded and where | Activity View, activity log or Compliance API, in the vendor's system | Every screen as tokens, every action and approval with a name, in the customer's SIEM |
| What a successful injection meets | An action review, after the model has read page content within its permitted access | Tokens, and a person before any consequential action |
| What changes when the model is swapped | The reviewer, environment and record | Nothing |
Where does RedactSure sit?
RedactSure builds the customer-owned judge: a secure virtual machine in the cloud the customer's posture requires, on hardware-encrypted enclaves, with keys the customer holds. Every identifier chosen by policy becomes a consistent token at the render layer before any model reads the screen. A named person confirms the Planner's per-task policy and approves every payment, submission and record change (Supervised Delegation). Supported models can operate under the environment's controls; proprietary reviewer integrations require separate confirmation, and compatible safety controls can remain additional layers.
What the customer avoids is a review that misses with neither the record nor the decision in the customer's hands. Every screen as tokens, every action and every approval with a name lands in the AI Control Record and exports to the customer's own SIEM. The evidence stays with the organization that is accountable. A successful injection can expose tokens and any remaining task context, and the consequential action still stops for the named person. Deployments are in pilot.
Methodology and limitations
OMB M-25-21 rescinded and replaced M-24-10 on April 3, 2025. References to the earlier memo are historical; current federal review must use the replacement and applicable agency policy.
The vendor documents are OpenAI's Dots safety document (September 29, 2026), Anthropic's "Use Claude in Chrome safely" and Meta's Muse safety write-up (September 2026), cited with their dates. ChatGPT Work's Auto-review comes from PPC Land (July 2026) and the METR finding from Latent Space, each cited as press. Regulation and standards are cited from text: NIST AI RMF, OMB M-25-21, NAIC Model Bulletin, 45 CFR 164.504(e), HHS cloud guidance, 34 CFR 99.31 and PCI SSC scoping guidance. Case law and enforcement actions are not surveyed. All were read as of September 30, 2026.
No hands-on testing of Dots, Claude in Chrome or Muse was performed; each reviewer is described as its vendor describes it. Anthropic's 0.08% attack success figure is the vendor's internal result. The METR research was not read directly; the finding is Latent Space's report of it. OpenAI refers to a system card that could describe Auto-review in more detail. Whether reviewer outcomes can be exported to a customer's own monitoring is not stated; that silence is reported as silence. The healthcare, insurance, education and payments examples describe record types, not any organization; no customer or prospect is described. Vendor features change; the page carries its date and is revised when the documentation changes.
Whether a vendor's reviewer counts toward oversight under the BAA, FERPA's direct control condition, PCI scope or the NAIC bulletin belongs to the organization's counsel, not to this page.
- OpenAI, "How we build safety, security and privacy into dots" (September 29, 2026)
- Claude Help Center, "Use Claude in Chrome safely."
- Meta AI Research, "How We Built Safety Into Muse" (September 2026)
We did not run the vendor products or capture their model requests. Workflow examples are analysis, and RedactSure behavior is described from its current design.
What the record shows
A vendor's reviewer can be a useful safety layer. The customer still needs to verify its independence from the acting agent, the approval policy and access to evidence when a consequential action goes wrong. Ask who can change the reviewer, whether the acting agent can bypass it and how the customer can export the decision record.
Frequently asked questions
What does Auto-review check?
Per OpenAI's September 29, 2026 documentation, Auto-review checks a dot's planned actions against the user's instructions, the Custom Rules and OpenAI's safety requirements before execution. A blocked step does not run and the dot is told why. It sits alongside mandatory confirmations. It judges the action after the dot has read page content within its permitted access, so it does not change what the model saw.
What did METR find about agents self-approving?
Latent Space's DevDay coverage cites METR research finding coding agents self-approved actions their own checks had flagged. That secondary report is not a test of Auto-review or Sentinel. Evaluate whether the acting agent can modify or bypass the actual reviewer rather than inferring that from vendor ownership.
Who reviews the reviewer?
In a vendor's agent, the vendor does, and the customer sees the outcome in the vendor's log. Under separation of model and control a named person in the customer's environment approves each consequential action. RedactSure's AI Control Record then exports every screen as tokens, every action and every approval to the customer's own monitoring for its auditors.
Should we turn the vendor's reviewer off?
No. Auto-review, the action classifiers and the Sentinel reduce the number of bad actions that reach a person, and they are useful controls. They are not the accountable control, and an organization relying on them alone has no record of its own to show a regulator. Where an integration supports them, they can remain one layer, with a named person and the customer's record above them.
Does separation of model and control mean running our own model?
No. It means owning the environment that decides what the model sees and who approves what it does, and owning the record. The model can be any vendor's, swapped without moving the controls. Running a model yourself is a procurement choice; owning the judge is governance.
What does an auditor ask to see?
The policy that set what the agent could see, who confirmed it, who approved each payment or record change, and every screen the agent saw. A vendor's Activity View shows some of this inside the vendor's product; an exported AI Control Record shows all of it in the organization's systems.
Sources
Vendor documentation
- OpenAI, "How we build safety, security and privacy into dots" (September 29, 2026). https://openai.com/index/how-we-build-safety-security-and-privacy-into-dots/
- Claude Help Center, "Use Claude in Chrome safely." https://support.claude.com/en/articles/12902428-use-claude-in-chrome-safely
- Meta AI Research, "How We Built Safety Into Muse" (September 2026). https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
Regulation and standards
- NIST, AI Risk Management Framework. https://www.nist.gov/itl/ai-risk-management-framework
- OMB, Memorandum M-24-10 (March 2024). https://www.whitehouse.gov/wp-content/uploads/2024/03/M-24-10-Advancing-Governance-Innovation-and-Risk-Management-for-Agency-Use-of-Artificial-Intelligence.pdf
- NAIC, Model Bulletin on the Use of AI Systems by Insurers, adoption map. https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf
- 45 CFR 164.504(e), business associate contracts. https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.504
- HHS, "Cloud Computing" guidance under HIPAA. https://www.hhs.gov/hipaa/for-professionals/special-topics/health-information-technology/cloud-computing/index.html
- 34 CFR 99.31, FERPA school official exception. https://www.ecfr.gov/current/title-34/subtitle-A/part-99/subpart-D/section-99.31
- PCI Security Standards Council, "Guidance for PCI DSS Scoping and Network Segmentation." https://www.pcisecuritystandards.org/documents/Guidance-PCI-DSS-Scoping-and-Segmentation_v1.pdf
Independent analysis and press
- Latent Space, "AINews: OpenAI DevDay 2026" (September 29, 2026). https://www.latent.space/p/ainews-openai-devday-2026-dots-61
- PPC Land, "OpenAI kills Atlas browser, folds it into new ChatGPT Work agent" (2026). https://ppc.land/openai-kills-atlas-browser-folds-it-into-new-chatgpt-work-agent/
RedactSure documents
- RedactSure, "Accountable AI and Workflow Governance" (2026). https://redactsure.com/blog/accountable-ai-and-workflow-governance/
- RedactSure, "The Two Gaps AI Agents Opened in Your Security Stack" (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
- RedactSure, "Your AI Strategy Is Probably Wrong" (2026). https://redactsure.com/blog/your-ai-strategy-is-probably-wrong/
- RedactSure Research, "Should the AI Agent's Secure Environment Belong to the Model Vendor?" (2026). https://redactsure.com/research/should-ai-agent-environment-belong-to-model-vendor
- RedactSure Research, "What Is Render-Layer Tokenization?" (2026). https://redactsure.com/research/what-is-render-layer-tokenization
- RedactSure Research, "What Is Supervised Delegation?" (2026). https://redactsure.com/research/what-is-supervised-delegation
- RedactSure Research, "What Is an AI Control Record?" (2026). https://redactsure.com/research/what-is-an-ai-control-record
- RedactSure Research, "What Does Prompt Injection Get From a Token-Only Agent?" (2026). https://redactsure.com/research/prompt-injection-agent-sees-only-tokens
- Product behavior described on this page reflects RedactSure's current design. OpenAI, ChatGPT and Dots are trademarks of OpenAI; Claude is a trademark of Anthropic; Meta and Muse are trademarks of Meta Platforms, Inc.; named to identify the products. Compliance determinations belong to the organization's counsel.
Editorial verification
- Office of Management and Budget, Memorandum M-25-21 (April 3, 2025), current federal AI policy replacing M-24-10. https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-21-Accelerating-Federal-Use-of-AI-through-Innovation-Governance-and-Public-Trust.pdf
See it on your workflow.
Bring one billing, collections, claims or patient-account workflow and your questions.
Book a demo




