Evaluation & testing
Promptfoo
Open-source evaluation and red-teaming tool that generates adversarial inputs from configurable plugins, runs them against an LLM application, and grades outputs with deterministic and model-graded metrics in CI. The paid enterprise editions add RBAC and team scoping; audit logging and retention are not documented.
hybrid · generally available · Research snapshot 2026-09-06
Visit the official product source ↗Where it fits
Evaluation & testing · Agent security & threat detection
Useful conversation with: AI security engineer, AI engineering lead.
Ask for a demonstration
Demonstrate an end-to-end red-team scan of my deployed agent, including which plugins ran and how findings are scoped to a team.
Capabilities and evidence
Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.
Documented by provider
Promptfoo documents generating a wide range of adversarial inputs via plugins covering prompt injection, jailbreaking, PII leakage from RAG context and tool-based vulnerabilities such as unauthorized data access and privilege escalation, then grading outputs with deterministic and model-graded metrics.
Limit: The guide does not quantify detection coverage or false-positive rates, and does not document capture of LLM spans or tool-call traces.
Source s1
Documented by provider
Promptfoo Enterprise is a hosted SaaS edition and Promptfoo Enterprise On-Prem is self-hosted on any major cloud with a dedicated runner inside the customer network; both list RBAC and team management.
Limit: Audit logs, SSO, retention settings and compliance certifications are not stated on the enterprise page.
Source s2
Documented by provider
The red-teaming guide describes running scans continuously in CI/CD and feeding user-reported issues back into the red-team configuration on a review cadence.
Limit: No documented mechanism that blocks a release automatically; gating depends on the customer's pipeline.
Source s1
Limitations to discuss
- No documented audit logs, retention configuration, PII redaction or compliance certifications.
- Findings are test results rather than governance evidence records; no immutable report store is documented.
Sources
- LLM red teaming guide (open source) · Promptfoo · official docs
Access date reported by researcher: 2026-09-06 - Promptfoo Enterprise - Secure LLM Application Testing · Promptfoo · official docs
Access date reported by researcher: 2026-09-06
Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.
Suggest a correction