Data Report · By Chris Sowa · Published
Last updated
Can an Enterprise Change AI Models Without Changing Its AI Controls?
Yes, when input policy, approvals and audit records are enforced independently of the selected model. A model change still needs capability, security and supplier review, even when those controls remain in place.

Yes, when input policy, approvals and audit records are enforced independently of the selected model. A model change still needs capability, security and supplier review, even when those controls remain in place. The September price table compares published token rates and one benchmark's costs. Neither is a forecast for a claims or billing workflow. Keeping controls outside the model can reduce migration work and make more models practical to evaluate, but it does not make their reliability, hosting terms or supply-chain risks identical. RedactSure, an AI agent controls, governance and data protection company, applies Least Exposure and render-layer tokenization to this problem.
Key findings
- Anthropic released Claude Opus 5.5 on September 22 at $4 and $20 per million tokens, and Sonnet 5.5 on September 28 at $2 and $10. Source: Anthropic.
- OpenAI released GPT-6.1 Sol on September 29 at $2 and $10, "a fifth of Astra's price." Source: OpenAI.
- Meta's Muse Spark 1.3 lists at $1.25 and $4.25 under the new Meta Enterprise Platform. Source: Runtime.
- On the Artificial Analysis Intelligence Index v4.3.2, cost per index task runs from $0.27 for DeepSeek V4.1 Flash to $7.60 for Claude Sonnet 5.5. Source: Artificial Analysis.
- Dots run on GPT-6 Astra only, Muse on Meta's model, Claude in Chrome on Claude; in each, controls and record are the model vendor's. Source: The Next Web.
What moved in September?
Three frontier releases in one week: Claude Opus 5.5 on September 22, 2026, Claude Sonnet 5.5 on September 28 and GPT-6.1 Sol on September 29. Sol lists at a fifth of GPT-6 Astra's price. Sources: Anthropic; OpenAI. The same day Meta placed the Muse API, at Muse Spark 1.3's $1.25 and $4.25, under an Enterprise Platform led by CJ Desai. Sources: MindStudio; Runtime.
Below the frontier, open-weights models from Z.ai, Moonshot and DeepSeek sit at a fraction of these rates. An enterprise that can swap models can treat every one as an option.
What does the same work cost on different models?
| Model | Vendor | Weights | Input ($/M) | Output ($/M) | Blended 3:1 ($/M) | Intelligence Index | Cost per index task |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | Closed | 4.00 | 20.00 | 8.00 | 58 | $5.98 |
| Claude Sonnet 5.5 | Anthropic | Closed | 2.00 | 10.00 | 4.00 | 56 | $7.60 |
| GPT-6 Astra | OpenAI | Closed | 10.00 | 50.00 | 20.00 | 53 | $3.26 |
| GPT-6.1 Sol | OpenAI | Closed | 2.00 | 10.00 | 4.00 | 52 | $0.72 |
| Muse Spark 1.3 | Meta | Closed | 1.25 | 4.25 | 2.00 | 48 | $1.60 |
| GLM-5.3 | Z.ai | Open | 1.40 | 4.40 | 2.15 | 45 | $2.01 |
| Kimi K3 | Moonshot | Open | 3.00 | 15.00 | 6.00 | 44 | $2.00 |
| Gemini 3.8 Flash | Closed | 0.75 | 3.75 | 1.50 | 41 | $1.24 | |
| DeepSeek V4.1 Flash | DeepSeek | Open | 0.30 | 1.20 | 0.53 | 39 | $0.27 |
| DeepSeek V4 Pro | DeepSeek | Open | 1.32 | 3.96 | 1.98 | 36 | $0.68 |
Prices as of September 30, 2026; the table is refreshed monthly. List prices, standard tier, USD per million tokens, from the pricing pages in Sources. Blended is (3 x input + output) / 4, a 3:1 ratio typical of agent work that reads more than it writes. DeepSeek figures are peak rates; Kimi K3 is Moonshot's rate; Gemini's runs through December 31, 2026. Index and cost per task are Artificial Analysis v4.3.2. Cost per task counts the tokens a model spends on the benchmark, which is why Sonnet 5.5 costs more per task than Opus 5.5 at half the price. Per-token price is not cost per workflow.
Why can a vendor agent not make this trade?
Because the model is not a component of the agent; it is the agent. Dots run on GPT-6 Astra, with Auto-review enforced by OpenAI and the record in Activity View. Sources: The Next Web; OpenAI. Muse runs Meta's model in Meta's VM with Meta's Sentinel and log. Source: Meta AI Research. Claude in Chrome runs Claude with Anthropic's classifiers. Source: Claude Help Center.
In each case the exposure policy, the reviewer and the record are properties of the model vendor's product. An enterprise that wants Sol's price for a task it runs in Muse cannot switch the model; it switches vendors, environments and logs. That is a migration, not procurement.
What has to stay fixed when the model changes?
Four things, none of which belong to the model. The exposure policy: which values on each screen are tokenized before any model reads it, and what the agent may do there. The named approver, who confirms that policy and approves every payment, submission and record change under Supervised Delegation. The AI Control Record, exported to the customer's monitoring. The keys, held by the customer, so real values never move when the model does.
When those four are the customer's, a model swap is a configuration change. The security review that examined the environment, the token policy and the approval flow can reuse the evidence for unchanged controls. The new model and provider still need review and regression testing.
What does the token layer have to do with price?
Tokenization can widen the set of models an organization is willing to evaluate by reducing the sensitive identifiers in their input. It does not make all providers or models equally safe: remaining task context, hosting terms, model behavior and supply-chain risks still differ.
The table's benchmark costs are useful for choosing candidates, not forecasting a claims, revenue-cycle or student-services task. Test candidates on representative work and include errors, retries, human review and hosting costs. The same input policy and approval record can be held constant during those tests.
Where does RedactSure sit?
RedactSure builds the environment in which the model is a swappable part: a secure virtual machine in the cloud the customer's posture requires, on hardware-encrypted enclaves, with keys the customer holds. Every sensitive value chosen by policy becomes a consistent token at the render layer before any model reads the screen. A named person confirms the Planner's per-task policy and approves every consequential action, on the AI Control Record. The environment is model-agnostic: Claude, GPT, Gemini or open-weights models, swappable without moving the controls.
What the enterprise avoids is riding one vendor's price curve and redoing the security review to leave it. Because every model receives tokens where the identifiers were, a cheaper or open-weights model can be evaluated under the same input policy. Its remaining risks and performance still need review. The policy, the approver, the record and the keys stay the customer's through every swap. Deployments are in pilot.
Methodology and limitations
Prices are taken from the pricing pages of OpenAI, Anthropic, Google and DeepSeek, read as of September 30, 2026. Muse Spark 1.3, GLM-5.3 and Kimi K3 prices come from MindStudio, VentureBeat and OpenRouter, cited as press. Scores and cost per task are the Artificial Analysis Intelligence Index v4.3.2 of the same date, cited as that analyst's measurement. Vendor design comes from OpenAI's Dots safety document (September 29, 2026), Meta's Muse safety write-up (September 2026) and Anthropic's Claude in Chrome help article, cited. No regulation is cited.
Prices are list prices at the standard tier as of the date and exclude discounts, caching and batch rates. DeepSeek figures are peak rates; Gemini's runs through December 31, 2026. Per-token price is not cost per workflow; the blended figure is a calculation, not a measured cost. No workflow was benchmarked on any model, and no vendor product was tested hands-on. The index is one benchmark, and its cost per task depends on the tokens each model spends. Whether Dots, Muse or Claude in Chrome will add model choice is not stated. The claims, revenue cycle and student services examples describe workflow types, not any specific organization; no customer or prospect is described. Prices change; the page carries its date and is revised when they do.
Whether any control meets an organization's requirements belongs to its counsel or compliance officer, and which model is capable enough to the workflow's owner, not to this page.
We did not run the vendor products or capture their model requests. Workflow examples are analysis, and RedactSure behavior is described from its current design.
What the record shows
Yes, when input policy, approvals and audit records are enforced independently of the selected model. A model change still needs capability, security and supplier review, even when those controls remain in place. Compare candidate models on the actual workflow, including errors, review effort and total cost. Keep the policy and audit format stable where possible.
Frequently asked questions
Does a model swap change the audit record?
Not in a customer-owned environment. The AI Control Record captures every screen as tokens, every action and approval with a name; which model proposed each action is one field, and a swap changes that field only. In a vendor's agent the record is the vendor's, so a new vendor means a new record.
How much do frontier models differ in price for the same work?
On the September 30, 2026 table, the four models scoring 52 or above list between $2 and $10 per million input tokens, a factor of five on blended price. Cost per index task runs from $0.72 for GPT-6.1 Sol to $7.60 for Sonnet 5.5. Price rank and capability rank differ.
Does the security review have to be redone?
Reuse evidence for controls that remain unchanged, but review the new model and provider. Hosting terms, residual data exposure, tool behavior, reliability and supply-chain risks can change even when tokenization and approval policies stay fixed. Test representative workflows before switching.
Is the cheapest model good enough for regulated work?
Capability is chosen per task; the exposure does not change with the model. A model that reads a remittance and posts lines needs less than one drafting a coverage position; the workflow owner sets the bar. In RedactSure's environment what the model receives and who approves the action are the same at every price.
Where do the price figures come from?
From each vendor's pricing page as of September 30, 2026: OpenAI, Anthropic, Google and DeepSeek directly, MindStudio for Muse Spark 1.3, VentureBeat for GLM-5.3 and OpenRouter for Kimi K3. Scores and cost per task are from the Artificial Analysis Intelligence Index v4.3.2, same date, linked in Sources.
What is a blended price?
A single per-million-token figure that weights input and output by an assumed ratio. The table uses 3:1, computed as (3 x input + output) / 4, because agent work on business applications reads far more than it writes. A drafting-heavy workflow would reorder the table. It is a comparison aid, not a forecast.
Sources
Vendor documentation
- OpenAI, API pricing. https://developers.openai.com/api/docs/pricing
- Anthropic, Claude pricing. https://platform.claude.com/docs/en/about-claude/pricing
- Google, Gemini API pricing. https://ai.google.dev/gemini-api/docs/pricing
- DeepSeek, API pricing. https://api-docs.deepseek.com/quick_start/pricing
- OpenAI, "How we build safety, security and privacy into dots" (2026). https://openai.com/index/how-we-build-safety-security-and-privacy-into-dots/
- Meta AI Research, "How We Built Safety Into Muse" (September 2026). https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
- Claude Help Center, "Use Claude in Chrome safely." https://support.claude.com/en/articles/12902428-use-claude-in-chrome-safely
Independent analysis and press
- Artificial Analysis, Intelligence Index v4.3.2 (September 30, 2026). https://artificialanalysis.ai/leaderboards/models
- MindStudio, "Muse Spark 1.3 pricing" (2026). https://www.mindstudio.ai/blog/muse-spark-1-3-pricing
- VentureBeat, "GLM-5.3 hits the API at $1.4/$4.4 per million tokens" (2026). https://venturebeat.com/technology/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens
- OpenRouter, Kimi K3 model page. https://openrouter.ai/moonshotai/kimi-k3
- Runtime, "Meta puts Muse models, agents and coding tools under a new enterprise platform" (2026). https://runtimewire.com/article/meta-enterprise-platform-muse-cj-desai
- The Next Web, "OpenAI launches dots, always-on AI agents with their own cloud computers" (2026). https://thenextweb.com/news/openai-dots-always-on-ai-agents-cloud-computers-devday
RedactSure documents
- RedactSure, "Your AI Strategy Is Probably Wrong" (2026). https://redactsure.com/blog/your-ai-strategy-is-probably-wrong/
- RedactSure, "The Two Gaps AI Agents Opened in Your Security Stack" (2026). https://redactsure.com/blog/two-gaps-ai-agents-opened-in-your-security-stack/
- RedactSure, "Accountable AI and Workflow Governance" (2026). https://redactsure.com/blog/accountable-ai-and-workflow-governance/
- RedactSure Research, "Should the AI Agent's Secure Environment Belong to the Model Vendor?" (2026). https://redactsure.com/research/should-ai-agent-environment-belong-to-model-vendor
- RedactSure Research, "What Is Render-Layer Tokenization?" (2026). https://redactsure.com/research/what-is-render-layer-tokenization
- RedactSure Research, "What Is Supervised Delegation?" (2026). https://redactsure.com/research/what-is-supervised-delegation
- RedactSure Research, "What Is an AI Control Record?" (2026). https://redactsure.com/research/what-is-an-ai-control-record
- RedactSure Research, "Can an AI Vendor's Own Model Be the Safety Reviewer for Its Agent?" (2026). https://redactsure.com/research/can-an-ai-vendors-own-model-be-the-safety-reviewer
- Product behavior described on this page reflects RedactSure's current design. OpenAI, GPT and Dots are trademarks of OpenAI; Claude is a trademark of Anthropic; Gemini is a trademark of Google; Meta and Muse are trademarks of Meta Platforms, Inc.; DeepSeek, GLM and Kimi are trademarks of their respective owners; named to identify the products.
See it on your workflow.
Bring one billing, collections, claims or patient-account workflow and your questions.
Book a demo



