Agent security & threat detection

LlamaFirewall

Open-source guardrail framework from Meta that runs layered scanners around agent execution: PromptGuard 2 for jailbreak detection, Agent Alignment Checks that audit chain-of-thought for goal misalignment, and CodeShield static analysis of generated code. Meta describes the alignment auditor as still experimental.

open_source · generally available · Research snapshot 2026-09-06

Visit the official product source ↗

Where it fits

Agent security & threat detection · Evaluation & testing

Useful conversation with: AI engineer, Application security engineer, AI safety researcher.

Ask for a demonstration

Show me LlamaFirewall scanning an email agent's inputs with PromptGuard 2 and flagging an indirect injection through Agent Alignment Checks, including the experimental caveats.

Capabilities and evidence

Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.

Documented by provider

LlamaFirewall provides PromptGuard 2 as a jailbreak detector, Agent Alignment Checks as a chain-of-thought auditor for prompt injection and goal misalignment, and CodeShield as an online static analysis engine for insecure generated code.

Limit: Published as research and open source; production support commitments are not stated.

Source s2

Documented by provider

The framework is designed for integration into existing AI agents and LLM applications, with scanners configured per role and invoked on input at runtime.

Limit: The documentation page fetched shows prompt-guard scanning only and does not demonstrate tool-call or output enforcement.

Source s1

Documented by provider

Meta states Agent Alignment Checks show stronger efficacy against indirect injections than prior approaches but remain experimental.

Limit: No accuracy figures for enterprise workloads are provided.

Source s2

Limitations to discuss

Sources

  1. LlamaFirewall · Meta · official docs
    Access date reported by researcher: 2026-09-06
  2. LlamaFirewall: An open source guardrail system for building secure AI agents · Meta · official release
    Access date reported by researcher: 2026-09-06

Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.

Suggest a correction