Evaluation & testing

Inspect AI

MIT-licensed evaluation framework from the UK AI Security Institute for running model and agent evaluations. It executes dataset samples through solvers, supports tool calling, multi-turn dialog, multi-agent primitives and sandboxed execution in Docker or Kubernetes, and writes evaluation logs for analysis. No hosted service or access control.

open_source · generally available · Research snapshot 2026-09-06

Visit the official product source ↗

Where it fits

Evaluation & testing

Useful conversation with: AI safety researcher, Model evaluation lead.

Ask for a demonstration

Demonstrate running one of the pre-built agentic evaluations against my model in a sandboxed environment and show the resulting evaluation log.

Capabilities and evidence

Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.

Documented by provider

Inspect supports evaluations measuring coding, agentic tasks, reasoning, knowledge, behaviour and multi-modal understanding, with datasets of labelled samples loaded from Hugging Face, CSV, JSON or memory.

Limit: Does not document human feedback collection, red-team probe libraries or production monitoring.

Source s1

Documented by provider

Documented features include flexible tool calling with custom and MCP tools, built-in bash, python, web browsing and computer tools, multi-agent primitives, external agents such as Claude Code and Codex CLI, and sandboxed execution of untrusted model code in Docker, Kubernetes, Modal, Proxmox or Vagrant.

Limit: Tool-call approval is referenced in navigation but no policy-gating feature is described on the page fetched.

Source s1

Documented by provider

The repository states Inspect was created by the UK AI Security Institute, is MIT licensed, and ships over 200 pre-built evaluations runnable against any model.

Limit: No enterprise tier, hosted deployment or support commitment is documented.

Source s2

Limitations to discuss

Sources

  1. Inspect AI · UK AI Security Institute · official docs
    Access date reported by researcher: 2026-09-06
  2. Inspect: A framework for large language model evaluations · GitHub / UK AI Security Institute · official repository
    Access date reported by researcher: 2026-09-06

Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.

Suggest a correction