Direct Answer: Treat the Agent as an Adversarial System

Agentic AI security testing should evaluate the complete system that can perceive instructions, choose tools, retain state, delegate work, and take actions—not merely the underlying language model. A useful test program asks four linked questions: what actions is the agent permitted to take, what could cause it to take those actions incorrectly, can those failures be detected quickly, and can containment stop consequential damage? As of 29 September 2026, the subject has moved beyond static prompt review because autonomous tools, multi-agent coordination, and code-oriented security agents can produce effects through several API calls. Reported incidents and emerging products, including NVIDIA’s open agent safety platform, OX Security’s agentic pentesting offering, and OpenAI’s March 2026 Codex Security announcement, show how testing is becoming a continuous engineering discipline. No commercial scanner or simulated breach proves an agent is secure. The defensible goal is evidence that risks are identified, bounded, observable, and acceptable for the specific deployment rather than a claim that all attacks have been eliminated.

Also worth reading: What are the definitive AI agent security best practices 2026 for enterprise deployment? · What Are the Essential Enterprise MLOps Governance Controls Required for Agentic AI Deployment in 2026? · How Do Agentic Workflow Security Architectures Actually Work in 2026?

What Agentic AI Security Testing Actually Covers

An agentic system differs from a chatbot because it can perform multistep work with limited intervention. It may interpret a user request, retrieve documents, invoke a shell, edit a repository, call a payment service, or communicate with another agent. Testing must therefore cover the model, system prompt, retrieved content, tool schemas, credentials, memory, delegation paths, permission policies, and the environment in which actions occur. Prompt-injection tests are only one layer: indirect instructions hidden in web pages, files, tool results, or conversation history can redirect an otherwise compliant agent. Traditional application-security methods still apply, including authentication, authorization, input validation, dependency scanning, secret detection, and vulnerability management. Agent-specific tests add questions about action selection, tool-call integrity, sequence length, state corruption, retry behavior, and the ability to resist instructions supplied by untrusted data.

The test boundary must include every actor that can influence execution. This often includes developers who change prompts, operators who approve actions, data providers whose content is retrieved, external services that return tool results, and attackers who create adversarial accounts or files. The March 2026 launch of Codex Security, described as an application-security agent for identifying and fixing vulnerabilities, illustrates why code agents need controlled execution and verification rather than unrestricted trust. Testing should establish whether agent-generated patches compile, pass tests, preserve intended behavior, and avoid inserting backdoors. It should also determine whether the agent can be tricked into modifying security controls, disclosing secrets, or escalating privileges. ISO/IEC 42001:2023 can help organizations manage AI risk, but certification to that management-system standard does not prove that an individual agent resists prompt injection or unsafe tool use.

A Practical Test Program From Threat Model to Retest

Begin with an asset and action inventory. Record what the agent can read, modify, send, purchase, delete, deploy, or approve, then classify each capability by reversibility and business effect. A read-only support agent has a different risk profile from one that can issue refunds or deploy production code. For high-impact actions, define quantitative gates such as a maximum single-transfer value, a permitted list of repositories, or a requirement for two-person approval above a set threshold. Run the same abuse cases in unit, integration, adversarial, and production-like environments. As a practical baseline, organizations can require a zero-tolerance result for unauthorized production changes and cross-tenant access, while tracking less destructive anomalies by rate and severity. The numbers must reflect the deployment; blindly copying a universal threshold creates false confidence.

Use a documented corpus of attacks rather than one-off demonstrations. Include direct prompt injection, indirect injection in retrieved content, malicious tool output, poisoned memory, credential theft, encoded instructions, role confusion, excessive agency, delegated-task attacks, and attempts to bypass human approval. Repeat high-priority tests across multiple runs because agent behavior is stochastic; a single successful refusal is not enough. A reasonable initial gate for a high-risk workflow might be at least 1,000 adversarial executions per major release, with 100% blocking of tested critical escapes and no unresolved high-severity code vulnerabilities. Human reviewers should also score false refusals, task success, latency, and cost. The purpose is not maximum test volume; it is repeatable evidence tied to real permissions and credible attacker paths.

What to Measure and What to Record

Security metrics must describe system behavior rather than model impressions. Record attempted unauthorized actions, blocked actions, approvals requested, policy violations, sensitive-data exposures, privilege escalation, tool-call anomalies, and the time from malicious input to detection. For each attack, preserve the exact model and system-prompt versions, tool definitions, context contents, memory state, policy configuration, and external responses. Otherwise, a failed test cannot be reproduced and a successful attack cannot be generalized. Report both safety and operational performance: a test suite can achieve a 100% block rate by disabling the agent, but that is not a useful security result if the intended task cannot be completed.

Measure defense in depth. Use least-privilege credentials, short-lived tokens, network allowlists, separate data from instructions, approval gates, sandboxed execution, and logging independent of the agent’s own reporting. The agent should not be the only component capable of deciding whether an action was authorized. Where feasible, place policy enforcement in a deterministic gateway that validates tool arguments before execution. Compare intended and actual actions after each step, and alert when an agent requests resources inconsistent with the active task. For code agents, require static analysis, unit tests, dependency checks, and human review before merge. For agents with external side effects, require compensating controls such as transaction limits, reversible workflows, canary environments, and rollback plans.

Comparison of Testing Approaches and Alternatives

Organizations can combine several methods, but the methods answer different questions. Automated red-team scanners are fast and repeatable, while expert tabletop exercises reveal governance failures that scanners cannot detect. A model red team may also test the underlying model, but it cannot prove that a production tool configuration is safe. Conventional penetration testing remains necessary because agents call ordinary APIs and operate with ordinary software flaws, yet a conventional test may miss multistep prompt injection and delegated-agent attacks.

FeatureAutomated Agent Red TeamingManual Expert Penetration TestConventional Application PentestTabletop Simulation
Primary purposeRepeat attack prompts and tool sequencesDiscover chained failures in the deployed designFind exploitable software and configuration flawsExercise decisions, escalation, and incident response
Best useContinuous release regressionValidate high-risk workflowsTest APIs, identity, networks, and codeTrain responders and test human controls
Typical scaleHundreds to thousands of executionsDays to several weeksDays to several weeksHours to one day per scenario
Main limitationMay miss novel business-logic attacksExpensive and not fully repeatableOften misses agent-specific manipulationDoes not directly exploit the system
Evidence producedBlock rates, traces, regression logsReproducible attack chains and remediationSeverity-rated vulnerabilitiesDecision-time and response data
Recommended frequencyEvery model, prompt, tool, or policy releaseAt least annually and before major redesignPer release cycle or material changeQuarterly for critical systems
The strongest program joins all four approaches. Automated testing provides breadth, manual testing provides depth, conventional pentesting covers inherited infrastructure, and tabletop work tests whether people can stop or recover from failures. Management frameworks can structure ownership and evidence, but they should not be mistaken for adversarial validation. Likewise, using a more capable “red-team” model does not automatically create a secure test environment; the attacker model still needs tools, credentials, telemetry, and access comparable to those of the system under review.

Common Mistakes That Produce False Confidence

A frequent mistake is treating refusal to produce harmful text as proof of operational safety. A model may avoid answering a direct request while still complying with malicious instructions embedded in a document read through a retrieval tool. Another error is granting a test account realistic permissions during development and assuming staging behavior transfers to production. Secrets, internal data, third-party tools, and long-term memory can change attack results. Teams also err by evaluating only final answers rather than intermediate actions, allowing an agent to expose data or execute a destructive step and then fail to summarize the task correctly.

Benchmarking can also be misleading. A small collection of public jailbreak prompts is easy to overfit and may fail against adaptive attackers. Conversely, a scanner that reports thousands of paraphrased attacks may create activity without meaningful coverage. Deduplicate equivalent cases, assign severity, and retain only tests linked to a documented abuse path. Do not use a vendor risk score as the sole acceptance criterion; opaque scores make it difficult to know whether the tool tested tools, memory, retrieval, or only text generation. Ask whether the product supports reproducible traces, versioned configurations, private deployment, and exportable evidence compatible with the organization’s audit process.

Finally, do not confuse an agent’s self-reported confidence with verified safety. Agents can misunderstand context, hallucinate tool results, or describe an action that policy enforcement never actually completed. The test plan therefore needs an independent ground truth: API audit logs, repository diffs, database changes, identity-provider records, network telemetry, and signed approvals. “Human in the loop” is not itself a control unless the reviewer sees enough reliable information, has time to intervene, and retains authority to reject the action. This matters in high-risk domains where the report of the first reported agentic AI data breach to a Spanish regulator raised questions about developer testing and supervision.

When to Act, and How to Budget the Work

Testing should begin before an agent receives production credentials and continue whenever behavior changes. At minimum, rerun the suite after changes to the model, system prompt, tool schema, retrieval source, memory policy, identity permissions, or safety filter. Also trigger a full review after an incident, a new external data source, or a new capability that can alter state. Organizations should act immediately when an agent can transfer funds, change access, publish content, manipulate scientific processes, operate machinery, or modify production software. Lower-risk read-only assistants still need testing, but the release gate can be proportionate to their permissions and data exposure.

There is no dependable universal price because pricing depends on whether the organization buys software scanners, engages a specialist firm, builds internal red-team capacity, or runs a conventional pentest. Open-source command-line tools can reduce direct licensing cost, but training, compute, test data, and engineering time remain. Commercial platforms may charge per seat, test, agent, or workload, with enterprise contracts that dominate total cost. Specialist adversarial testing can range from several thousand dollars for a bounded assessment to tens of thousands or more for a multi-agent, regulated deployment; these are budgeting ranges, not vendor quotations. A small pilot can start with 20 to 30 high-value attack scenarios and 100 to 500 repeated executions, then expand if the system can cause material effects.

For cost control, prioritize actions rather than conversational polish. Spending 80% of effort on ambiguous wording while production credentials remain overly privileged reverses the proper risk ratio. Compute budgets should also cover repeated trials, because stochastic behavior requires more than one run, but repetition should target uncertain and high-impact paths. Track the average cost per completed test, number of manual review hours, and time to remediate a failed scenario. NVIDIA’s 2026 agent-safety initiative and industry moves toward open testing tools may improve access, but teams should still budget for integration, independent validation, and governance rather than assuming an open tool eliminates specialist work.

A Defensible Release Decision for 2026

A deployment decision should state exactly what evidence supports release. The record can identify the tested agent version, architecture, data sources, tools, identities, threat model, attack corpus, number of executions, severity thresholds, detected failures, residual risks, and accountable owners. Use a clear decision rule: block release for any reproducible path to cross-tenant access, unauthorized production modification, credential disclosure, or unreviewed high-impact action. Conditional approval may be reasonable when lower-severity issues are bounded by sandboxing, restricted networks, short-lived credentials, transaction limits, and rapid rollback. A general statement that the model is “AI secure” is not adequate evidence.

Regulation of agentic AI remained less mature by September 2026 than regulation of generative AI, and unresolved terminology can weaken governance. Legal obligations may depend on sector, jurisdiction, and the agent’s function, so security teams should work with counsel rather than wait for a single universal agent framework. Relevant controls can still be defined now through ISO/IEC 42001:2023-aligned management practices, established secure-development standards, privacy obligations, and existing software-security requirements. The aim is proportionate assurance, not regulatory theater. If the business cannot explain the agent’s authority, reproduce its behavior, detect misuse, and limit harm, it does not yet have enough evidence for a responsible production launch.