Any organization that lets an AI system touch real data or real decisions needs an enforcement layer that sits outside the model, where the model cannot reach it or argue with it. The same holds for a mid-sized manufacturer using generative AI to draft procurement contracts and a healthcare provider running an agentic scheduling assistant, and it applies to frontier, generative, and agentic AI alike.
The first gap is the provider. Frontier developers are advancing capability faster than independent safety evidence is arriving. More than a dozen of them publish voluntary safety frameworks, and thresholds and evaluation rigor vary so widely that framework adoption alone tells a buyer little. Generative AI embedded in SaaS products and productivity tools adds dependencies that nobody scores, often with little visibility into which model and provider sit underneath. Agentic AI takes autonomous actions across multi-step workflows, sometimes for days, on systems that touch real data. EU AI Act enforcement against providers of general-purpose AI models, including fines of up to 3% of global annual turnover, began on August 2, 2026, which makes a provider’s regulatory posture a financial risk signal for any buyer’s vendor evaluation.
The second gap is architectural. Most organizations rely on the model itself to decide what it should and should not do, through system prompts, safety fine-tuning, and content filters that live in the same weights and context window the model uses for every other output. Those controls can be argued with, role-played around, encoded past, or overridden by a sufficiently creative prompt, because responsiveness to prompts is what makes the model useful. Prompt injection, jailbreaks, and tool-calling exploits are now a standing category of incident response, and attackers succeed by working inside a boundary that organizations assumed sat outside the model.
Putting enforcement inside the AI does not work. Every model, frontier, generative, or agentic, can find its way around controls embedded in its own context. The fix is the same regardless of AI type: an enforcement layer sitting outside the model, which the model cannot see, reach, or negotiate with. This is not optional infrastructure. It is the minimum viable governance for any organization deploying AI.
Both gaps come from treating governance as something the model participates in. The model provides the reasoning. The agent uses that reasoning to do work. The harness supplies the context, tools, and orchestration around the agent, along with the control points that shape how that work happens. Security boundaries have to be independent of the model’s decision-making, whether they sit at the agent level, the harness level, or in an external enforcement layer. A system prompt or safety fine-tuning instruction that the model is asked to follow cannot serve as one.
“
Frontier AI providers have made real progress on safety frameworks, but voluntary commitments with inconsistent thresholds are not a substitute for verifiable evidence. Enterprises need to treat frontier, generative, and agentic model providers as the critical, systemically important vendors they have become, not as black boxes
Why the model cannot govern itself
The intuitive response to AI risk is better instructions: stronger system prompts, more restrictive fine-tuning, better content filters. Those are courtesy controls: the model follows them through the same mechanism it uses to follow every other instruction. Deterministic controls at the agent or harness level work differently. They run independently of the model’s decision-making while staying close enough to the agent to understand its context and act as the work happens.
- Prompt injection: An attacker embeds instructions in content the model processes (a document, a webpage, a tool response) and overrides the system prompt’s constraints. The model cannot distinguish the attack from a legitimate instruction.
- Jailbreaks: Role-play scenarios, encoding, and multi-step conversational sequences exploit the model’s instruction-following to bypass safety fine-tuning.
- Tool-calling exploits: Generative and agentic systems that call tools and APIs or execute code create an attack surface that grows with capability. A compromised tool response or an ambiguous permission boundary can set off autonomous action that cascades across systems before a human reviewer sees it. A generative AI with file-access permissions is exposed to this as directly as a full agentic workflow.
- Multi-agent amplification: When models or agents call each other without a human in the loop, one successful injection propagates down the chain. Generative AI feeding outputs to downstream agentic workflows carries the same risk without any formal multi-agent orchestration, and the problem compounds at every system boundary the output crosses.
Recent incidents show the exposure. In July 2026, a security incident at a widely used model-evaluation platform demonstrated how fragile third-party access controls for AI testing remain. An autonomous coding agent operated unsupervised for approximately a week before its own developer identified and attributed the activity. A brief export-control suspension of a newly released frontier model in June 2026 showed how quickly regulatory and geopolitical action can disrupt access to a model that an organization has already built a process around.
“
AI cannot be responsible for enforcing the controls that govern its own behavior. As enterprises move toward autonomous agents, enforcement must remain architecturally independent of the model and informed by a continuously changing risk context. When the risk of an agent, model, or third-party dependency changes, what AI is permitted to do should change with it.
Independent controls and where they belong
Independent controls can sit at several layers, and most organizations will use all of them. Each must operate deterministically and outside the model’s reasoning, so the model cannot talk its way around it. Agent-level controls can enforce identity delegation and tool-call boundaries while keeping the context of the work in progress. Harness-level controls can enforce policy, manage orchestration boundaries, and log interactions without relying on the model to surface anything. External layers such as proxies and gateways provide independent boundaries where traffic naturally passes through them. The coverage gap is real: a gateway can act only on the interactions it mediates, and agents working across tools, identities, APIs, endpoints, browsers, and Model Context Protocol (MCP) connections will always have surfaces that no single gateway sees, so the layers have to be used together.
Independent controls enforce policy at the interaction level, in real time. The broader governance program covers AI model inventory, life-cycle management, provider risk assessment, continuous assurance, and reporting to leadership or a board, and it needs its own funding. Its shape scales with the organization, as the final section describes.
IDC recommends evaluating independent control architecture against five capability areas. Table 1 describes each area and why it has to be independent of the model to work as a security boundary. The five apply whether the implementation sits at the agent level, the harness level, or an external layer, and most deployments will combine all three. Treat them as a minimum: data sensitivity, regulatory environment, and deployment risk profile will call for additional controls, some of which the hub-and-spoke section below describes.
TABLE 1: Five Independent Control Capabilities: Applicable Across Frontier, Generative, and Agentic AI (Extensible by Organization)
| Layer | Function | Why It Must Be Independent of the Model |
| Identity and Access | Authenticates every caller (human, application, or agent) against the enterprise identity provider and delegates only the calling user’s actual entitlements downstream | A model holding a standing service identity with broad permissions can be jailbroken into using those permissions. Delegated identity bounds the blast radius of any exploit to whatever the calling user was already permitted to do. |
| Request Inspection | Screens every inbound prompt and tool-call parameter for policy violations, data-classification conflicts, and prompt-injection patterns before the model sees them | Prompt-injection defenses built into model weights or system prompts are arguable. A request inspector running outside the model cannot be negotiated with or role-played around. |
| Policy Enforcement | Evaluates and enforces business rules, regulatory constraints, and data-handling requirements as code, not as instructions that the model is asked to follow | Instructions inside the model’s context window compete with every other input for attention. Policy running as code in an external layer has no attention mechanism to exploit. |
| Response Mediation | Screens, redacts, or blocks the model’s outbound content and tool-call authorizations before they leave the enforcement layer, regardless of what the model produced | A model that generates a non-compliant response cannot self-censor reliably. External mediation acts on the output regardless of the model’s own judgment about its content. |
| Audit and Telemetry | Logs every mediated request and response immutably outside the model’s environment; hosts the kill switch at three grains: session, user/application, and full deployment | Audit logs that the model can reach can be suppressed, rewritten, or influenced. External, immutable logging is the only basis for forensic investigation, and the only kill switch the model cannot cooperate with or obstruct. |
Source: IDC, 2026
The hub-and-spoke MCP gateway pattern
For the highest-risk integrations, where traffic passes through a defined chokepoint, IDC recommends evaluating an MCP gateway in a hub-and-spoke topology as a strong technical substrate for external enforcement. A single hub centralizes policy, identity, and audit for the estate. Spokes deployed close to tools and data enforce that policy locally, so governance logic is not duplicated at every integration point. The pattern works best as one layer in a defense-in-depth architecture that also includes agent-level and harness-level controls, because no gateway mediates every interaction an agent has across its tools, APIs, browsers, and direct system connections.
Defining policy, identity, and audit once at the hub and enforcing them at every spoke avoids the redundancy and configuration drift that build up when each AI integration manages its own access controls. In high-volume or latency-sensitive environments, a well-designed hub-and-spoke enforcement layer can reduce total governance overhead compared with point-to-point control architectures, because inspection and policy evaluation happen at the spoke nearest the data source instead of through a central bottleneck. Where AI is embedded in operational workflows that cannot tolerate added latency, that property is a deployment prerequisite.
Two further considerations belong in any hub-and-spoke MCP evaluation. First, secure communications between AI systems and the hub, and between the hub and each spoke, are first-class architecture requirements that the network layer does not supply on its own. Most high-sensitivity deployments will need encrypted transport and mutual authentication between AI agents and the MCP gateway, integrity verification of messages between spokes and connected data repositories, and cross-agent session integrity validation, all as extensions to the five-layer stack. Second, data repository connections need their own enforcement: the spoke closest to a database, document store, or data lake should enforce access policy and log retrieval operations with the same rigor the identity and request layers apply to AI queries. The pattern does not yet cover skills, hooks, plugins, or direct API calls, so scope it to the high-risk chokepoints you can name today.
Specifying the kill switch
The audit and telemetry layer also hosts the kill switch, and it has to be specified deliberately. It should offer three suspension grains: a single session, a specific user or application, and the full deployment. Because it sits outside the model, it triggers instantly and needs no acknowledgment or cooperation from the model. Split authority over the enforcement layer across two administrative roles, one controlling policy and configuration and a separate one with sole authority to activate the kill switch. Neither role should hold both. Pair suspension with a defined rollback capability wherever the underlying system supports it, because halting future actions addresses only half the containment problem.
“
We shouldn’t rely on the model to enforce the controls that govern it, but that doesn’t mean every control needs to sit outside the agent. Deterministic controls at the agent and harness level can act with the context of the work, while external controls provide independent enforcement at the boundaries they mediate. Enterprises need both, informed by a continuous understanding of how the agent is built, what it can access and how it actually behaves.
Rating the providers behind your AI
An enforcement layer governs how models behave inside your organization. Governing the providers whose models you depend on is a separate and equally urgent job, and it falls on every organization using AI, including those without a formal third-party risk management (TPRM) program. A company buying a generative AI writing tool has an AI provider relationship, and so does a development team embedding a frontier model in a customer-facing product. In each case the provider sits inside operational processes as an unscored dependency, evaluated, if at all, against generic software-vendor criteria that miss safety-framework maturity and concentration risk.
Provider risk assessment for AI is a basic organizational responsibility that scales with the depth of the dependency. IDC recommends that any organization depending on an AI model provider, directly or through software it procures, assess that provider across five dimensions. Table 2 maps each dimension to its decision gate and explains why it belongs in TPRM.
TABLE 2: IDC Five-Dimension TPRM Scoring Framework for AI Model Providers (Frontier, Generative, and Agentic)
| Dimension | What It Measures | Decision Gate | Why It Belongs in TPRM |
| Safety-Framework Maturity | Whether the provider publishes a Frontier AI Safety Framework, how specific capability thresholds are, and whether commitments extend to deployment pauses | Onboarding | Framework adoption is a start, not proof. Threshold specificity and pause commitments distinguish genuine from performative frameworks. |
| Evaluation Transparency | Whether the provider permits independent red-teaming, publishes methodology and results rather than summary assurances | Onboarding / Contract | Summary assurances are unverifiable. Published methodology and results are auditable. Buyers cannot tell the difference without this dimension. |
| Incident Disclosure History | How quickly and completely the provider has disclosed past security incidents, model-weight exposures, or capability-tier reclassifications | Contract / Exit | A provider’s disclosure behavior under pressure is more informative about future reliability than its posture under normal conditions. |
| Regulatory Posture | EU AI Act Code of Practice adherence, engagement with other applicable frameworks, responsiveness to regulatory information requests | Contract | EU AI Act fines on general-purpose AI model providers of up to 3% of global turnover make regulatory posture a live financial risk, not a compliance checkbox. |
| Concentration and Dependency Exposure | How many of the enterprise’s other critical vendors embed the same provider’s models; whether infrastructure or capacity is shared across them | Exit / BCP | A single provider’s outage or incident can cascade simultaneously across vendors the enterprise never directly contracted with. |
Source: IDC, 2026
The score should drive three decisions: whether to accept a new AI provider dependency, what terms to require before the relationship proceeds, and when score deterioration justifies evaluating alternatives. Smaller organizations can make these calls informally, and larger ones may have defined governance gates. Score on trigger events as well as on an annual cycle, because AI providers change materially between assessments in ways conventional software vendors usually do not. Generative AI embedded in SaaS tools can swap its underlying model without telling the customer, so require contractual disclosure of model version changes in any AI tool agreement. Concentration risk deserves particular attention from organizations that do not think they have it: the same handful of underlying providers often powers frontier, generative, and agentic deployments across an entire software stack.
What every organization using AI must do
The enforcement layer and the provider scoring framework complement each other, and both apply even where no governance, risk, and compliance (GRC) team exists. The enforcement layer governs real-time model behavior; the scoring framework governs the provider relationship. These actions apply at any scale:
- Demand architectural separation from every AI vendor. A gateway or filter running as a plug-in inside the model vendor’s own console is still reachable by the model through shared memory or a persistent prompt. A terms-of-service promise that the model will not look is not evidence of separation either, so whether you are deploying one generative AI tool or a fleet of agentic workflows, ask for an architecture diagram showing no network path from the model to the policy engine, the audit store, or the enforcement configuration.
- Require identity parity across every AI system. Extend the conditional access, multifactor authentication (MFA), and risk-based policies already running for human identities to every AI model, agent, and tool, using on-behalf-of delegation so the model acts only within the calling user’s entitlements.
- Insist on model neutrality in the enforcement layer. You will run models from more than one provider, and an enforcement layer tied to a single vendor’s API becomes a second lock-in point.
- Specify the kill switch in every AI vendor contract. Require session, user-application, and full-deployment suspension grains, a defined rollback capability for actions already taken, and a documented split between whoever configures policy and whoever can pull the switch. A small organization using one generative AI tool needs to know it can turn that tool off instantly. An organization running agentic workflows needs the same capability.
- Score every AI model provider you depend on. Use the five-dimension framework whether or not you have a formal vendor risk program. Require software vendors to disclose which AI models sit inside the tools you have bought, with advance notice of model version changes. A generative AI model embedded in a productivity application you already license is a provider dependency even if nobody has mapped it.
- Plan for simultaneous provider failure. Build concentration-risk scenarios into business continuity planning that assume your primary and backup model providers fail at the same time. Documented correlated outages and shared infrastructure across labs make this a planning requirement.
- Size the governance program to your organization and build it around the enforcement layer. A large enterprise needs a formal AI governance program with dedicated ownership and board reporting, backed by continuous assurance. A mid-market company needs a named AI risk owner, a model inventory, and a quarterly review cadence. A smaller organization needs, at minimum, a list of every AI tool in use, with an approver and an off-switch for each.
What AI solution vendors must build
The enforcement obligation falls on the vendor side of the market as much as on buyers. Any company building, selling, or embedding an AI solution, whether a frontier model API, a generative AI feature in a SaaS product, or an agentic workflow platform, has to make external enforcement possible and, increasingly, provide it. Buyers who cannot enforce governance from outside the model are being failed by the products they buy. The build priorities:
- Treat AI provider scoring as a native capability in GRC platforms. Frontier, generative, and agentic AI providers each need a distinct scoring category; a custom field on a generic SaaS questionnaire cannot supply one. The same scoring logic serves a Fortune 500 company with a dedicated GRC team and a mid-market organization managing vendor risk through a single analyst; the interface and workflow complexity can scale.
- Offer fourth-party dependency discovery. Most enterprises do not know which foundation models sit inside the software they have already procured. A capability that traces embedded foundation models across a vendor portfolio and triggers scoring updates on version changes closes a gap that no generic TPRM questionnaire addresses today.
- Separate the model from the enforcement layer at the network level. Any vendor selling an AI solution needs network-layer separation between the model and the policy engine, audit store, and configuration, demonstrable in an architecture diagram. This is the minimum architecture for any AI product whose output matters, and it should ship as a standard feature. Buyers in every market segment are starting to ask for the diagram, and vendors without one will lose evaluations.
- Build chain-aware policy evaluation. Agentic and multi-agent workflows need policy engines that understand a session’s accumulated context and evaluate chains of related actions, beyond single requests in isolation. Generative AI with tool access creates shorter but still exploitable chains, and frontier models with external retrieval take on injection vectors at every hop.
- Package concentration-risk modeling and continuous provider monitoring. Annual scoring is too slow for providers that can change materially between assessment cycles. Rescore on safety-framework updates, model version changes, and material incidents, and offer concentration-risk scenarios as a standing business continuity service.
- Develop prebuilt EU AI Act Code of Practice scoring modules. Many buyers lack the in-house capacity to interpret the general-purpose AI (GPAI) Code of Practice. A prebuilt module that maps onto the regulatory posture dimension and updates as enforcement guidance evolves is a concrete differentiator now that enforcement has begun, on August 2, 2026.
- Integrate AI risk indicators into existing GRC and continuous controls monitoring (CCM) platforms. The OWASP Top 10 for Agentic Applications for 2026 covers agentic risk. Generative AI embedded in enterprise tools needs its own content-policy and output-integrity control category, and frontier AI providers need capability-threshold monitoring as a standing signal. Each belongs in continuous monitoring as a native control category. Most large enterprises deploy all three today and need a unified view across them.
“
A voluntary safety framework you cannot score is just a press release. And a system prompt you rely on as a security boundary is not a control; it is a courtesy. The principle is not that every control must sit outside the agent. It is that every control must operate independently of the model’s decision-making. Deterministic controls at the agent level, the harness level, and external enforcement boundaries each play a role. What no organization can afford is to have the model as both the worker and the referee of its own behavior.