Agent security & threat detection
LlamaFirewall
Open-source guardrail framework from Meta that runs layered scanners around agent execution: PromptGuard 2 for jailbreak detection, Agent Alignment Checks that audit chain-of-thought for goal misalignment, and CodeShield static analysis of generated code. Meta describes the alignment auditor as still experimental.
open_source · generally available · Research snapshot 2026-09-06
Visit the official product source ↗Where it fits
Agent security & threat detection · Evaluation & testing
Useful conversation with: AI engineer, Application security engineer, AI safety researcher.
Ask for a demonstration
Show me LlamaFirewall scanning an email agent's inputs with PromptGuard 2 and flagging an indirect injection through Agent Alignment Checks, including the experimental caveats.
Capabilities and evidence
Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.
Documented by provider
LlamaFirewall provides PromptGuard 2 as a jailbreak detector, Agent Alignment Checks as a chain-of-thought auditor for prompt injection and goal misalignment, and CodeShield as an online static analysis engine for insecure generated code.
Limit: Published as research and open source; production support commitments are not stated.
Source s2
Documented by provider
The framework is designed for integration into existing AI agents and LLM applications, with scanners configured per role and invoked on input at runtime.
Limit: The documentation page fetched shows prompt-guard scanning only and does not demonstrate tool-call or output enforcement.
Source s1
Documented by provider
Meta states Agent Alignment Checks show stronger efficacy against indirect injections than prior approaches but remain experimental.
Limit: No accuracy figures for enterprise workloads are provided.
Source s2
Limitations to discuss
- No managed service or dashboard
- Alignment checking described as experimental
Sources
- LlamaFirewall · Meta · official docs
Access date reported by researcher: 2026-09-06 - LlamaFirewall: An open source guardrail system for building secure AI agents · Meta · official release
Access date reported by researcher: 2026-09-06
Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.
Suggest a correction