THE BUYER'S FIELD GUIDE

AI Agent Security & Threat Detection Tools

Products focused on adversarial risk to and from agents: prompt injection, tool abuse, agent-mediated data exfiltration, malicious MCP servers, plus adversarial testing, detection and response.

29 related offerings · 14 primary listings · 15 overlapping listings

Product research snapshot September 6, 2026 · Editorial guide September 7, 2026 · Published by DutyGraph

When to explore this layer

Use this category when agents consume untrusted content or can act through tools. Define the attack surface: incoming documents, retrieved pages, prompts, connectors and outputs. Ask which surfaces are inspected and whether the product blocks, alerts or only records a suspicious event.

ILLUSTRATIVE EVALUATION · NOT A CUSTOMER RESULT

Put a real task in the demonstration.

In a controlled fictional test, put an instruction to export supplier information inside a document that an assistant is meant to summarize. Observe whether the instruction is treated as document content, whether a prohibited tool call is attempted, and what an operator can investigate afterward. Use synthetic data and an isolated target.

Questions to bring to the demonstration

  1. Which prompt-injection and tool-abuse surfaces are tested, and which are excluded?
  2. Does the product prevent the action or notify someone after it occurred?
  3. How are benign unusual requests separated from attacks, and how can false positives be reviewed?
  4. What evidence is available for response without exposing unnecessary sensitive content?

Evidence to request

  • A scoped adversarial test report
  • A prevented or detected event with its action timeline
  • False-positive review and incident-response procedures

Record what was demonstrated, what was only described, and what remains unknown. Preserve the product version, environment and date beside each observation.

Use the editable Markdown worksheet →

Where this layer stops

No single demonstration establishes complete protection. Combine security testing with least privilege, data controls and task boundaries. A security alert does not replace a business decision about whether the agent should have been doing that work.

Connect it to the work

A documented task gives investigators an expected behavior to compare against: which information the agent should read, which artifact it should produce, and which actions lie outside its purpose.

Read our perspective on the demand side of agents →

29 offerings to investigate

Alphabetical, not ranked. Membership includes primary and secondary research categories. These products have different scopes; inspect the evidence profile before comparing capabilities.

Primary category

Amazon Bedrock Guardrails

Configurable safeguard service inside Amazon Bedrock that evaluates user inputs and model responses against content filters, denied topics, sensitive information filters and word filters, including a prompt attack category. It can be applied at inference or via a standalone API, but does not authorize agent tool calls.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me a Bedrock guardrail blocking a prompt attack and masking PII for a Bedrock Agent, then show the same guardrail invoked through ApplyGuardrail for a non-Bedrock model.

Read sources and limitations →

Also covers this layer

AppOmni Agent Inventory

Capability of AppOmni's SaaS security platform that surfaces AI agents running inside connected SaaS tenants - such as Salesforce Agentforce, ServiceNow Now Assist and Microsoft 365 Copilot - including agents enabled without security review, with their declared tools, identities, permissions and over-privilege findings. Discovery depends on AppOmni's SaaS API connections.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate listing every Agentforce and Now Assist agent in our tenants, flagging which ones were enabled without approval and which hold write or destructive permissions.

Read sources and limitations →

Also covers this layer

Astrix Agent Control Plane

Astrix discovers AI agents, MCP servers, service accounts, OAuth apps, API keys and other non-human identities across cloud, SaaS, CI/CD and vaults, maps each to a human owner in an identity graph, and applies agent policies plus onboarding and offboarding actions. Enforcement depth outside integrated platforms is unclear.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me the identity graph for one shadow agent — its NHIs, credentials, reachable resources and owner — then apply a policy that blocks it and offboard it.

Read sources and limitations →

Primary category

Azure AI Content Safety Prompt Shields

Microsoft API in Azure AI Content Safety that analyses user prompts and supplied documents for adversarial instructions, returning attack-detected flags for direct user prompt attacks and indirect attacks embedded in external content. It is a detection endpoint that applications must act on rather than an enforcement gateway.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me a shieldPrompt call flagging an indirect injection hidden in an uploaded document, and how the calling agent enforces a block based on that response.

Read sources and limitations →

Primary category

Check Point AI Agent Security (formerly Lakera Guard)

Runtime protection for AI agents that inspects prompts, reference material, tool responses and tool descriptions for injections and manipulation, applies tool allow and deny lists, and flags actions outside an agent's mandate. Also builds an inventory of agents and connected MCP servers across supported agent platforms.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me Check Point AI Guardrails detecting an injection hidden in a tool response and blocking the resulting tool call, plus the agent inventory entry for that agent's MCP servers.

Read sources and limitations →

Primary category

Cisco MCP Scanner

Apache-2.0 Python tool from Cisco's AI Defense group that scans MCP servers, their tools, prompts and resources plus server source code, combining YARA rules, LLM-based analysis and Cisco's hosted inspection API, and audits dependencies and bundled binaries. Full detection depth depends on optional third-party services.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Show me a scan of an untrusted MCP server package that flags a docstring-versus-implementation mismatch, and which findings required the Cisco AI Defense API.

Read sources and limitations →

Also covers this layer

CSA STAR for AI

Cloud Security Alliance assurance program extending its STAR registry to AI services, with a Level 1 self-assessment against the AI Controls Matrix questionnaire, an automated validation option, and a Level 2 tier referencing third-party certification. Registry-based transparency for AI providers rather than a regulatory conformity assessment.

service · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate a completed AI-CAIQ submission for an agent platform and what the Valid-AI-ted scoring adds over a plain Level 1 self-assessment.

Read sources and limitations →

Also covers this layer

Entro Security NHI & Agentic AI Platform

Entro inventories non-human identities, secrets and agentic AI deployments across cloud, code, CI/CD, on-prem and SaaS, links each agent to the NHIs, entitlements and secrets it uses and to a human owner, and monitors agent behaviour for anomalies through its NHIDR detection engine. Credential issuance is not part of the evidenced scope.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me an agent's NHI lineage — creator, secrets used, entitlements, resources touched — and a live NHIDR alert for anomalous agent behaviour.

Read sources and limitations →

Also covers this layer

garak

Apache-licensed LLM vulnerability scanner maintained by NVIDIA. It fires static, dynamic and adaptive probes at a model or dialog system to test for jailbreaks, prompt injection, toxicity, data leakage and misinformation, logs each generation and detector verdict, and outputs a report with failure rates and hit logs.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Show me a garak scan of my chatbot with the probe-by-probe failure rates and the hit log for successful jailbreaks.

Read sources and limitations →

Also covers this layer

Giskard Hub

French vendor pairing an open-source Python testing library with a delivered assessment service. Automated and expert-led testing probes conversational agents for prompt injection, data disclosure, sycophancy, hallucination and inappropriate refusals, returning a severity-ranked vulnerability report and a signed go/no-go deployment recommendation.

hybrid · Research snapshot 2026-09-06

Ask for a demonstration
Show me a full assessment report for my customer-facing agent, with vulnerabilities ranked by severity and the go/no-go deployment recommendation.

Read sources and limitations →

Primary category

Google Cloud Model Armor

Google Cloud service that screens LLM prompts and responses for prompt injection, jailbreaks, unsafe content and sensitive data, optionally returning sanitised text. Integrations extend screening to Google-managed MCP server traffic and the Gemini Enterprise agent platform, while the Agent Gateway integration is documented as preview.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me Model Armor floor settings screening traffic to a Google-managed MCP server, blocking an injected prompt, and clarify which agent integrations are GA versus preview.

Read sources and limitations →

Primary category

LlamaFirewall

Open-source guardrail framework from Meta that runs layered scanners around agent execution: PromptGuard 2 for jailbreak detection, Agent Alignment Checks that audit chain-of-thought for goal misalignment, and CodeShield static analysis of generated code. Meta describes the alignment auditor as still experimental.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Show me LlamaFirewall scanning an email agent's inputs with PromptGuard 2 and flagging an indirect injection through Agent Alignment Checks, including the experimental caveats.

Read sources and limitations →

Also covers this layer

Mindgard

UK vendor running automated red-team tests against AI models, applications and agents. It profiles the target, enumerates attack surface, executes techniques from a maintained attack library via CLI or SDK, and reports exploitable findings with remediation guidance. Documentation covers testing mechanics rather than audit, retention or access controls.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me an automated red-team run against my production agent, including which attack techniques were executed and the remediation guidance produced.

Read sources and limitations →

Also covers this layer

Netskope One AI Command Center

Module of Netskope One AI Security that discovers AI assets - corporate or personal, managed or shadow, cloud or on-premises - from the vendor's SSE/proxy vantage point and maps them to the identities, data stores and tools they connect to, adding risk correlation and response. Discovery leans on traffic and platform telemetry rather than code scanning.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me the asset-to-identity-to-data-store map for a shadow AI application discovered from our traffic, including any MCP servers it reaches.

Read sources and limitations →

Primary category

NeuralTrust (TrustGate)

Barcelona-based platform combining an agent gateway (TrustGate) with runtime protection over the models, tools, MCP servers and data agents touch. Documented gateway behaviour includes per-user and per-tool RBAC, end-user identity forwarding across hops and cryptographic audit trails, with SaaS, hybrid and air-gapped deployment options.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me TrustGate forwarding end-user identity through two agent hops while denying a payments API tool for that user, plus the cryptographic audit record of each tool call.

Read sources and limitations →

Primary category

Noma Security Platform

Platform that inventories agents, MCP servers, skills and models across endpoints, SaaS agent builders and homegrown AI stacks, maps each agent's permissions and data access, red teams them before production, and evaluates runtime actions to alert, block, mask data or route to a human. Claims rest on vendor pages.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me Noma discovering an unapproved MCP server on a developer laptop, mapping its blast radius, and then routing a risky agent action to a human for approval.

Read sources and limitations →

Primary category

NVIDIA NeMo Guardrails

Open-source Python toolkit that intercepts LLM application inputs, outputs and custom action calls, applying configurable rails written in YAML and Colang to block or modify content. The guardrail catalogue includes jailbreak detection, PII handling and agentic security checks; the repository labels the release beta and not production-recommended.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Show me NeMo Guardrails applying an execution rail around a tool-calling action plus jailbreak detection, and explain the beta production caveat in the repository.

Read sources and limitations →

Also covers this layer

Obsidian Security Shadow AI

Module of Obsidian's SaaS security platform that builds a continuously updated inventory of AI tools and agents by combining a managed browser extension, API integrations into SaaS tenants, and mapping of agent-to-MCP connections. Aimed at security teams; agent coverage depends on which SaaS tenants and endpoints are instrumented.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me an agent discovered only by your browser extension that never appeared in the SaaS platform's own API-reported agent list, with its creator, permissions, and MCP connections.

Read sources and limitations →

Also covers this layer

OpenSSF Model Signing (OMS) / model-transparency

OpenSSF-backed specification with an Apache-2.0 library and CLI that signs and verifies machine learning model artifacts using Sigstore, self-signed certificates, public keys or PKCS#11 devices, producing signature bundles that let consumers check model integrity and provenance before deployment or reuse.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate signing a multi-gigabyte model with Sigstore and verifying the signature in a deployment pipeline gate, including what the bundle attests to.

Read sources and limitations →

Also covers this layer

Operant Semantic Firewall

Inline enforcement layer that inspects prompts, plans, tool calls, generated commands and data payloads before execution and returns an allow, block or redact decision, with a companion MCP gateway applying least-privilege controls and trust zones. Evidence comes from vendor product pages rather than reference documentation.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me the Semantic Firewall blocking a shell command produced by an injected instruction and redacting a bulk data read, with the policy that produced each decision.

Read sources and limitations →

Primary category

Pillar Security

Israeli platform covering the AI agent lifecycle: cataloguing agents, models, prompts, MCP servers and coding agents through agentless integrations, then applying runtime guardrails that monitor prompts, tool calls and commands for prompt injection, tool poisoning and data exfiltration. Product claims come from vendor pages, not reference docs.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me Pillar validating that an agent's tool call matches its declared schema, flagging a permission-scope deviation, and blocking a poisoned instruction in an agent-to-agent handoff.

Read sources and limitations →

Primary category

Prisma AIRS AI Runtime Security

Palo Alto Networks security service that scans prompts and model responses via API or network enforcement to detect prompt injection, sensitive data leakage and malicious content, with agent-focused detections such as MCP threat detection, tool chaining attack analysis and privilege misuse. Delivered as a managed enterprise service.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me Prisma AIRS scanning an agent request through the AI Runtime API, flagging a tool chaining attack and a prompt injection, and the security profile that blocked it.

Read sources and limitations →

Also covers this layer

Promptfoo

Open-source evaluation and red-teaming tool that generates adversarial inputs from configurable plugins, runs them against an LLM application, and grades outputs with deterministic and model-graded metrics in CI. The paid enterprise editions add RBAC and team scoping; audit logging and retention are not documented.

hybrid · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate an end-to-end red-team scan of my deployed agent, including which plugins ran and how findings are scoped to a team.

Read sources and limitations →

Also covers this layer

PyRIT

MIT-licensed Python framework from Microsoft for probing generative AI systems for risk. It is aimed at security engineers running automated adversarial testing campaigns rather than at governance teams, and the repository provides no multi-user controls, evidence retention or reporting workflow of its own.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate an automated PyRIT attack run against my deployed model endpoint and show what artefacts the run leaves behind.

Read sources and limitations →

Primary category

Snyk Agent Scan (formerly MCP-Scan)

Apache-2.0 command line scanner that discovers locally installed agent components — harnesses, MCP servers, skills — and checks tools, prompts and resources for prompt injection, tool poisoning, cross-origin escalation and tool changes, with a proxy mode that inspects live MCP traffic. Some checks call Snyk's hosted API.

open_source · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate scanning our developers' MCP configurations and show what a tool-poisoning and rug-pull finding looks like, plus which checks require the hosted API.

Read sources and limitations →

Also covers this layer

SPLX AI Asset Management

AI-BOM and inventory module of the SPLX platform, now part of Zscaler. It connects to cloud platforms, code repositories and ML/AI platforms to detect LLMs in use, scan repositories to map agents, tools and MCP servers in AI workflows, and run risk assessments on discovered agents. Detection is scan-based, not user-traffic based.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me the agentic workflow map produced from scanning one of our repositories, including each agent, its tools, and the MCP servers it connects to.

Read sources and limitations →

Primary category

Straiker Defend AI

Runtime security product for AI agents that inspects prompts, reasoning steps and tool calls across coding assistants, productivity copilots and custom agents, blocking direct and indirect injection, destructive actions such as file deletion, and data exfiltration. Vendor pages also describe shutting down rogue agents and connections.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me Defend AI blocking an indirect injection delivered in an email to a productivity copilot and stopping a coding agent from deleting files, then show the rogue-agent shutdown action.

Read sources and limitations →

Also covers this layer

Valence SaaS and AI Discovery

Part of Valence's SaaS security platform: it inventories sanctioned and unsanctioned SaaS and AI applications, and continuously identifies OAuth tokens, API keys, connected apps and service accounts linking business SaaS tenants to third-party AI tools. Detection is API-based against connected SaaS tenants, so unconnected apps stay invisible.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Demonstrate how you surface an OAuth grant that connects an employee's unsanctioned AI tool to our Google Workspace tenant, including the scopes granted and the granting identity.

Read sources and limitations →

Primary category

Zenity AI Detection and Response (AIDR)

Runtime security layer for AI agents that analyses full execution sequences, including chained tool calls, retrievals and agent-to-agent handoffs, to detect direct and indirect prompt injection, unauthorised tool invocations and sensitive data leaving through agent activity, with inline blocking. Evidence is drawn from vendor platform pages.

commercial · Research snapshot 2026-09-06

Ask for a demonstration
Show me AIDR detecting a slow-building indirect injection across several turns, blocking the unauthorised tool call it triggers, and the agent-to-agent handoff record.

Read sources and limitations →

About this guide

The evaluation questions and fictional scenario are DutyGraph's editorial guidance. Product listings use the supplied source-linked research snapshot. We have not independently tested these offerings. Listing is not an endorsement, certification or working integration.

Read the directory methodology · Suggest a correction