Direct answer: treat the agent as an untrusted, nonhuman identity

Runtime agent access controls are the policies, technical restrictions, and monitoring applied while an AI agent operates rather than only before deployment. They determine which tools, data, models, credentials, and actions an agent may use at that moment, and under what conditions. The practical standard in 2026 is to give every agent a unique identity, issue only task-specific and short-lived permissions, evaluate every sensitive action, and record an audit trail. A prompt saying “do not delete production data” is not an access control; an API policy that denies the production account a delete operation is. Runtime controls should combine identity governance, policy enforcement, sandboxing, approval gates, telemetry, and rapid revocation because no single product handles all of these jobs well. This approach suits operational agents, coding assistants, research systems, and customer-service agents whose actions can affect real systems rather than merely generate text.

Also worth reading: What is an enterprise agent runtime security architecture and how should organizations implement it in 2026? · How Do Organizations Manage the Full Lifecycle of AI Agents in Production? · How Should Enterprises Design Runtime Governance for AI Agents in 2026?

The unit of protection must be the individual agent and its current session, not just the underlying foundation model. Two agents built from the same model can have different repositories, customer records, credentials, tools, and risk tolerances. Permissions should therefore be attached to an agent identity and purpose, then narrowed further by session, environment, action, data classification, and time. The broader market reflects this change: projects such as Faramesh describe open-source runtime enforcement, Prismor describes an open-source runtime control plane, and Station presents a sandboxed runtime for operational agents with MCP support. Commercial activity around agent runtime security, including Kontext Security’s reported $4 million round in 2026, also indicates that runtime governance has become a distinct product category rather than a generic feature of identity management.

How runtime controls work across the action path

A runtime policy sits between the agent and the resources it wants to reach. Before a tool call, the enforcement layer evaluates the caller’s identity, requested operation, target resource, current context, and applicable business rules. It can allow the request, rewrite it to remove unnecessary fields, require human approval, return a constrained result, or deny it outright. This decision should be repeated throughout a task because an agent may behave differently after reading new content or receiving tool output. For example, access to a public website might be permitted initially, while access to secrets discovered in that page remains prohibited. A static role assigned months earlier cannot express that distinction without becoming either too restrictive or dangerously broad.

Controls commonly operate at four levels: identity, authorization, execution environment, and observation. Identity assigns a verifiable principal such as agent://research/instance-47; authorization decides what that principal may do; the environment limits the blast radius of code or tool execution; and observation records prompts, decisions, tool calls, outputs, and administrative changes. Existing systems can contribute parts of this structure. Delinea supplies privileged-access controls, segregation-of-duties functions, and access reviews, while Okta and Microsoft provide identity and agent-platform capabilities. Agent runtimes such as Firecracker microVMs add process and workload isolation. These technologies solve related but different problems, so connecting them through one policy decision point is usually more reliable than expecting an LLM gateway to enforce every control by itself.

Control layerMain question answeredTypical mechanismCommon failure if omitted
Agent identityWhich agent is acting?Unique workload identity and short-lived tokenShared credentials make attribution impossible
AuthorizationWhat may it access or change?Role, attribute, purpose, and resource policiesExcess permission enables indirect harm
Tool and data filteringWhich content may be used?Schema validation, masking, allowlists, taint trackingSensitive data enters prompts or logs
Execution isolationWhat damage can its code cause?MicroVM, container, restricted user, egress filterOne workload compromises others
Human approvalWhich consequential action needs review?Risk-based approval with bounded scopeHumans approve too many routine actions
Monitoring and responseWhat happened, and can it stop?Immutable logs, anomaly alerts, kill switchIncidents cannot be reconstructed or contained
## A practical control model for production agents

Start by inventorying every autonomous or semi-autonomous component, including direct API clients, background workers, browser operators, coding tools, and agents that can call other agents. Classify actions by reversibility, data sensitivity, financial exposure, regulatory impact, and reach. A read-only search over public documents is materially different from issuing a payment, changing an IAM role, or sending customer email. Set explicit thresholds—for example, require approval for external messages above a defined recipient count, all production writes, all secret access, and any transaction above $500. These numbers should reflect the organization’s actual loss limits; universal thresholds create false precision and can still be too permissive for high-risk systems.

Issue credentials only after the runtime evaluates identity, environment, requested scope, and purpose. Prefer short-lived, audience-bound tokens over static API keys, and create separate read and write credentials where practical. Replace broad cloud-admin access with narrowly scoped roles that name permitted resources and operations. The runtime should also block direct egress except through approved gateways, because allowing an agent to reach an arbitrary endpoint can bypass a database policy. Microsoft Foundry’s agent-building capabilities and OpenAI’s agent-engineering guidance illustrate how model orchestration, tools, and operational context are becoming platform concerns, but platform convenience does not remove the need for organization-specific authorization decisions.

Design approvals as bounded decisions rather than vague “proceed” buttons. The approval record should show the exact command or API mutation, affected resource, predicted data exposure, estimated financial effect, and expiry time. A human should approve that version, and any material change should invalidate it. For low-risk, repeatable operations, a policy engine can allow the action after deterministic validation; for novel or high-impact actions, the system should pause. If approval is unavailable, a fail-closed rule is usually safer for production writes, though a temporary deny may be operationally unacceptable for some read-only workflows and therefore needs an explicit exception process.

Tool, network, and data controls that prevent indirect actions

Agents often cause harm indirectly by passing tainted information between tools. A web page can contain instructions that persuade an agent to disclose an environment variable, a support document can trigger unauthorized record changes, or one compromised tool can return a malicious command. Runtime controls should consequently track data provenance and restrict capabilities based on source, not just on the final destination. Secrets should never enter general conversation context merely because an agent reads a file. Production systems can use token brokers that return only the fields required for a task, while masking values that the model does not need to see.

Tool contracts should be narrow and schema-validated. Instead of exposing a shell capable of arbitrary commands, provide operations such as search_orders(order_id) or draft_refund(order_id, amount). Validate types, ranges, resource ownership, and state transitions on the server side. Network controls should resolve destinations against an allowlist and prevent access to cloud metadata endpoints, private administrative networks, and unapproved data-transfer services. Firecracker microVMs or comparable isolated runtimes can reduce kernel- and process-level exposure, but isolation cannot substitute for least privilege: a correctly confined agent can still perform a destructive action through an approved API.

Data controls should separate retrieval authority from action authority. An agent authorized to read a customer record does not automatically need permission to export it, email it, or use it to alter another customer’s record. Add field-level controls for regulated or personally identifiable information, retention limits for prompts and traces, and tenant identifiers in every storage and log key. The 2026 agent-security direction represented by products such as Backbase: Layer is consistent with this model, where access policies, policy rules, and audit logging operate across agent actions. The useful question is not whether a vendor calls a layer a “control plane,” but whether its policy can be tied to concrete identities, resources, and consequences.

Comparison of control approaches and alternatives

Organizations can implement runtime controls through a centralized policy plane, platform-native features, a sandboxed execution runtime, or a managed security product. These choices are not mutually exclusive. A smaller team may begin with platform-native identity and audit functions, then add a dedicated gateway when tool count and risk grow. Larger environments often use a central decision service with specialized executors, gateways, and secret brokers behind it. Open-source projects can provide flexibility and inspectable policy logic, while commercial services may shorten deployment time and supply support, anomaly detection, and integrations.

ApproachStrengthsLimitationsBest fit
Platform-native controlsSimple deployment; aligned with model, tools, and identity platformMay not protect actions taken through other tools; policy portability can be weakOne platform and moderate risk
API or agent gatewayCentral policy point; request inspection; easy revocationCan miss local code, browser, or non-API actionsTool-mediated enterprise agents
Sandboxed runtimeLimits code, filesystem, and network blast radiusRequires secure image and egress design; cannot stop every approved API actionCode execution and operational agents
Central runtime control planeConsistent policy, identity, approval, and audit across agentsIntegration cost; policy design and latency require careRegulated or multi-agent environments
Manual process and promptsFast to start; useful for prototypesInconsistent, hard to audit, and vulnerable to indirect prompt injectionExperiments only
A defense-in-depth architecture is usually stronger than selecting one row from this table. Identity-aware gateways can govern API calls, microVMs can contain untrusted execution, and a central control plane can coordinate decisions and evidence. The trade-off is complexity: every component that can execute an action must be instrumented, and every policy hop creates latency or failure modes. Before buying a product, test whether it covers direct tool calls, subprocesses, browser actions, retrieved instructions, agent-to-agent messages, and administrative overrides. Also verify that “allowed” decisions are enforced outside the model and that revocation takes effect within the organization’s required incident-response window.

Common mistakes and weak security patterns

The most common mistake is confusing a system prompt with a security boundary. Models may follow instructions inconsistently, and retrieved text can attempt to override their original task. Sensitive operations must therefore fail when the model asks for prohibited access; they must not depend on the model choosing compliance. Another error is giving all agents from one team a shared service account. This destroys attribution, prevents per-agent revocation, and makes least-privilege review impractical. Use unique workload identities and avoid long-lived secrets wherever the underlying platform supports federation or short-lived credentials.

Overusing human approval is a different failure. If agents require approval for every read, reviewers will approve mechanically or become a bottleneck. A better control uses deterministic thresholds based on action risk, amount, sensitivity, and scope. Teams also make the mistake of logging prompts without logging policy decisions, or logging both without preventing untrusted data from entering the telemetry pipeline. An audit record needs the agent identity, policy version, evaluated conditions, decision, resource, result, and correlation ID; it should not contain unnecessary secrets or full customer records.

Finally, test the control system rather than only the agent’s answer quality. Red-team direct privilege escalation, prompt injection, tool-output poisoning, secret exfiltration, confused-deputy requests, and agent-to-agent privilege propagation. A benchmark with 900 questions can measure model capability, but it cannot prove that a particular production token cannot delete a database. Measure policy coverage, unauthorized-action denial, revocation time, false-approval rates, logging completeness, and containment success. The July 2026 emphasis on runtimeWire-style activity and the wider move toward runtime security products reflects a useful correction, but naming a category does not guarantee an enforcement architecture.

When organizations should act and how quickly

Act before an agent can access production data or change production systems, not after an incident. This is especially important for agents with tool use, persistent memory, broad enterprise search access, code execution, browser control, financial authority, or permissions to contact customers. A useful trigger is the first time one agent can trigger a side effect, even if a human is nominally supervising it, because unattended or misrouted actions can scale quickly. Pilot teams should implement controls in parallel with a small prototype because retrofitting identity, logging, and approval workflows into an already-deployed agent is harder than designing them into the runtime.

For low-risk internal research, a limited implementation can begin within 1–2 weeks: identify agents, define allowed read-only tools, use isolated credentials, log every call, and prohibit network egress by default. Production deployment normally deserves a 4–12-week control cycle, depending on integrations and risk, with threat modeling, policy tests, owner approval, and rollback procedures. High-risk domains such as payments, healthcare, identity administration, utility control, or regulated record modification may require months of engineering and independent review. A reasonable go-live threshold is zero known paths to unlogged production writes, tested revocation of all active agent identities, and verified that high-impact operations cannot proceed without an enforceable decision.

Urgency should be risk-weighted rather than driven by hype. Agent runtimes based on Firecracker microVMs, MCP-based operational systems, and centralized enforcement projects show that execution and control are converging, but a new category can also produce overlapping claims. Demand a small proof of concept that includes an allowed action, a denied action, a prompt-injection attempt, a credential-exfiltration attempt, an approval failure, and immediate revocation. If the vendor cannot demonstrate those cases against your environment, the product is not ready for the claimed scope. This evidence-first sequence is faster and safer than waiting for a fully autonomous system to be “mature.”

Cost, pricing, and build-versus-buy decisions

There is no defensible single market price for runtime agent access controls because the cost depends on whether the organization uses existing platform features, adds a gateway and sandbox, or buys a specialized control plane. Open-source runtimes and enforcement projects may have no license fee, but they still carry infrastructure, integration, policy-engineering, monitoring, and security-review costs. A small proof of concept might use a few managed API and logging services, but its monthly cost cannot honestly be generalized to production. Production expenses commonly include identity federation, secret management, isolated compute, data stores, telemetry retention, policy evaluation, and staff time. The largest cost is often not the runtime license; it is assigning an owner to model, tool, and access risk across business and security teams.

Build-versus-buy should follow the need for differentiated policy and control of execution. Existing identity, cloud, and agent platforms may already cover workload identity, role permissions, deployment boundaries, and audit events for a narrow deployment. Buying a dedicated control plane can be economical when several agent frameworks and business units need one policy model, especially if it provides connectors, approvals, anomaly detection, and compliance evidence. Building internally makes sense when tool protocols, legacy systems, or regulatory constraints are unusual, but only if the team can maintain secure defaults, versioned policy, availability, incident response, and independent testing. Avoid purchasing primarily for a polished dashboard; test whether the product can intercept the actual action paths and deny them at the resource boundary.

A practical cost calculation should include the number of agents, active sessions, tool calls, protected resources, integrations, policy evaluations, retained events, and engineers required to operate the system. Include expected review time and incident costs, not just seats. A control that eliminates 99% of routine approval prompts may justify more expensive policy evaluation, while a product that generates thousands of alerts without actionable context may increase labor rather than reduce risk. Obtain current pricing and service terms directly from providers, because vendor packaging in this rapidly changing category can change quickly. The relevant 2026 investment signal—Kontext Security’s reported $4 million financing—supports market attention, not a specific return-on-investment claim.

A minimum viable governance policy

The minimum viable policy defines ownership, permitted actions, identity requirements, approval thresholds, isolation, logging, and response. Name a business owner for each agent and a security owner for its tools and data. Document the agent’s purpose, allowed resources, prohibited actions, model and prompt version, credential issuer, runtime image, network destinations, and escalation path. Require unique identities, short-lived credentials, server-side authorization, default-deny egress for code execution, and tamper-resistant records of every tool call and policy decision. State which actions require human approval and how approvals expire or change when the proposed action changes.

Set review intervals based on risk rather than habit. A low-risk internal read-only agent might be reviewed quarterly, while an agent with production write access should be reviewed before every material model, tool, prompt, or permission change and at least monthly thereafter. Revoke credentials immediately when an owner, purpose, environment, or incident status changes. Test restoration separately, because a system that can stop access but cannot safely resume may create an operational outage. The policy should also cover agents that create other agents, because delegated authority can otherwise become an undocumented privilege-escalation path.

Measure success with operational metrics: percentage of tool calls with a recorded policy decision, time to revoke an identity, number of standing credentials, production actions approved or denied, policy-evaluation latency, prompt-injection containment rate, and completeness of audit evidence. Targets should be explicit—for example, revoke a high-risk agent credential within 15 minutes, record 100% of production mutations, and block undeclared destinations. These are examples, not universal requirements. A defensible program reports exceptions, false positives, and unresolved risk instead of claiming that “all actions are secure,” because runtime controls reduce probability and blast radius but cannot eliminate model error, compromised dependencies, or human approval failures.