Evaluation & testing
Inspect AI
MIT-licensed evaluation framework from the UK AI Security Institute for running model and agent evaluations. It executes dataset samples through solvers, supports tool calling, multi-turn dialog, multi-agent primitives and sandboxed execution in Docker or Kubernetes, and writes evaluation logs for analysis. No hosted service or access control.
open_source · generally available · Research snapshot 2026-09-06
Visit the official product source ↗Where it fits
Evaluation & testing
Useful conversation with: AI safety researcher, Model evaluation lead.
Ask for a demonstration
Demonstrate running one of the pre-built agentic evaluations against my model in a sandboxed environment and show the resulting evaluation log.
Capabilities and evidence
Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.
Documented by provider
Inspect supports evaluations measuring coding, agentic tasks, reasoning, knowledge, behaviour and multi-modal understanding, with datasets of labelled samples loaded from Hugging Face, CSV, JSON or memory.
Limit: Does not document human feedback collection, red-team probe libraries or production monitoring.
Source s1
Documented by provider
Documented features include flexible tool calling with custom and MCP tools, built-in bash, python, web browsing and computer tools, multi-agent primitives, external agents such as Claude Code and Codex CLI, and sandboxed execution of untrusted model code in Docker, Kubernetes, Modal, Proxmox or Vagrant.
Limit: Tool-call approval is referenced in navigation but no policy-gating feature is described on the page fetched.
Source s1
Documented by provider
The repository states Inspect was created by the UK AI Security Institute, is MIT licensed, and ships over 200 pre-built evaluations runnable against any model.
Limit: No enterprise tier, hosted deployment or support commitment is documented.
Source s2
Limitations to discuss
- Research and assurance tooling: no RBAC, audit logging, retention control or immutable evidence store.
- Evaluation logs are local artefacts; nothing documented about tamper-evidence or long-term evidence retention.
Sources
- Inspect AI · UK AI Security Institute · official docs
Access date reported by researcher: 2026-09-06 - Inspect: A framework for large language model evaluations · GitHub / UK AI Security Institute · official repository
Access date reported by researcher: 2026-09-06
Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.
Suggest a correction