Inspect: An open-source framework for large language model evaluations The UK AI Security Institute and Meridian Labs released Inspect, an open-source framework for large language model evaluations that ships with over 200 pre-built evaluations and supports more than 20 model providers. Inspect is installable from PyPI via `pip install inspect-ai` and offers composable datasets, solvers, scorers, agents, and tools, plus sandboxing through Docker, Kubernetes, Modal, Proxmox, and Vagrant. The framework also supports running external agents such as Claude Code, Codex CLI, and Gemini CLI, and includes a web-based Inspect View tool and a VS Code extension. Inspect An open-source framework for large language model evaluations Welcome Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute https://aisi.gov.uk and Meridian Labs https://meridianlabs.ai . Inspect can be used for a broad range of evaluations that measure coding, agentic tasks, reasoning, knowledge, behavior, and multi-modal understanding. Core features of Inspect include: - Composable building blocks—datasets, agents, tools, and scorers—that make evaluations easy to write and reuse. - A collection of over 200 pre-built evaluations ready to run on any model. - Extensive tooling, including a web-based Inspect View tool for monitoring and visualizing evaluations and a VS Code Extension that assists with authoring and debugging. - Flexible support for tool calling—custom and MCP tools, as well as built-in bash, python, text editing, web search, web browsing, and computer tools. - Support for agent evaluations, including flexible built-in agents, multi-agent primitives, and the ability to run arbitrary external agents like Claude Code, Codex CLI, and Gemini CLI. - A sandboxing system that supports running untrusted model code in Docker, Kubernetes, Modal, Proxmox, Vagrant, and other systems via an extension API. We’ll walk through two short “Hello, Inspect” examples below. Read on to learn the basics, then read the documentation on Datasets ./datasets.html , Solvers ./solvers.html , Scorers ./scorers.html , Tools ./tools.html , and Agents ./agents.html to learn how to create more advanced evaluations. If you are primarily interested in running evaluations rather than developing new ones, see the Evals ./evals/index.html listing where you’ll find implementations for over 200 popular benchmarks. Getting Started To get started using Inspect: 1. Install Inspect from PyPI with: pip install inspect-ai 2. If you are using VS Code, install the Inspect VS Code Extension ./vscode.html not required but highly recommended . To develop and run evaluations, you’ll also need access to a model, which typically requires installation of a Python package as well as ensuring that the appropriate API key is available in the environment. For example: pip install openai export OPENAI API KEY=your-openai-api-key inspect eval simpleqa.py --model openai/gpt-4o pip install anthropic export ANTHROPIC API KEY=your-anthropic-api-key inspect eval simpleqa.py --model anthropic/claude-sonnet-4-0 pip install google-genai export GOOGLE API KEY=your-google-api-key inspect eval simpleqa.py --model google/gemini-2.5-pro pip install torch transformers export HF TOKEN=your-hf-token inspect eval simpleqa.py --model hf/meta-llama/Llama-2-7b-chat-hf Inspect has built-in support for over 20 model providers as well as support for local inference with HuggingFace, vLLM, and SGLang. See the documentation on Model Providers ./providers.html for details on all supported providers. If you use a coding agent alongside Inspect, the inspect-skills https://github.com/meridianlabs-ai/inspect-skills install plugin provides skills that teach it to monitor running evals, read logs efficiently, and analyze results. Hello, Inspect An Inspect evaluation is a Task ./tasks.html that brings together three things: 1. Dataset ./datasets.html that provides labelled samples—typically a table with input and target columns, where input is the prompt and target is the ideal answer or grading guidance. 2. Solver ./solvers.html that produces an answer for each sample. This can be as simple as a single generate ./reference/inspect ai.solver.html generate call to the model, or as sophisticated as a full agent that uses tools over many turns. 3. Scorer ./scorers.html that evaluates the output—using text comparisons, model grading, or other custom schemes. Let’s look at two short examples: a question-answering benchmark and a capture the flag challenge. Benchmark: SimpleQA This task evaluates a model on SimpleQA https://openai.com/index/introducing-simpleqa/ , a benchmark of short, fact-seeking questions click on the numbers at right for further explanation : simpleqa.py python from inspect ai import Task, task from inspect ai.dataset import FieldSpec, hf dataset from inspect ai.scorer import model graded qa from inspect ai.solver import generate 1@task def simpleqa : return Task 2 dataset=hf dataset "codelion/SimpleQA-Verified", split="train", 3 sample fields=FieldSpec input="problem", target="answer", , , 4 solver=generate , 5 scorer=model graded qa , - 1 - The @task decorator registers the function with Inspect so that inspect eval can discover and run it by name. - 2 - hf dataset ./reference/inspect ai.dataset.html hf dataset loads samples directly from Hugging Face. Inspect also reads CSV, JSON, and in-memory datasets. - 3 - FieldSpec ./reference/inspect ai.dataset.html fieldspec declaratively maps the dataset’s problem and answer columns onto the sample’s input and target —no custom conversion function required. - 4 - The generate ./reference/inspect ai.solver.html generate solver simply sends each input to the model and collects its response. - 5 - Because the answers are free-form text, model graded qa ./reference/inspect ai.scorer.html model graded qa uses a model to grade each response against the target . Run it from the command line with inspect eval , choosing a model with --model : inspect eval simpleqa.py --model openai/gpt-5 Use inspect view to view the results: inspect view Agent: CTF Challenge Agent evaluations require the model take actions rather than just answer a question. Here’s a Capture the Flag CTF task where the react ./agents.html agent explores a sandboxed system using bash ./reference/inspect ai.tool.html bash and todo write ./reference/inspect ai.tool.html todo write tools to find a hidden flag: ctf.py python from inspect ai import Task, task from inspect ai.agent import react from inspect ai.dataset import json dataset from inspect ai.scorer import includes from inspect ai.tool import bash, todo write @task def ctf : return Task dataset=json dataset "challenges.json" , 1 solver=react prompt= "You are a Capture the Flag player. Explore the system and find the flag." , tools= bash , todo write , attempts=3, , 2 scorer=includes , 3 sandbox="docker", - 1 - react ./reference/inspect ai.agent.html react is a built-in agent that runs a reason-act-observe loop, giving the model the supplied tools until it submits an answer here allowing up to 3 attempts . - 2 - The includes ./reference/inspect ai.scorer.html includes scorer passes if the target flag appears in the agent’s submitted answer. - 3 - sandbox="docker" provides the isolated Docker container used by bash ./reference/inspect ai.tool.html bash configured by a Dockerfile or compose.yaml alongside the task . Use inspect view to view the results and look more carefully at individual transcripts: See the Tutorial ./tutorial.html to explore more in-depth examples that demonstrate additional Inspect features and techniques. Python API Above we demonstrated using inspect eval from CLI to run evaluations—you can perform all of the same operations from directly within Python using the eval ./reference/inspect ai.html eval function. For example: python from inspect ai import eval from simpleqa import simpleqa eval simpleqa , model="openai/gpt-5" LLM Assistance As you learn and use Inspect we recommend you provide an LLM with the documentation required for it to assist. There are two versions of LLM friendly markdown documentation available: - llms.txt llms.txt : Documentation index, articles fetched as required ~2k tokens . - llms-guide.txt llms-guide.txt : Full contents of all documentation ~185k tokens . There is also a Copy Page button at the top of every page that provides a markdown version of the page. Learning More To learn more about using Inspect see the following documentation sections: - Tutorial ./tutorial.html includes several annotated examples demonstrating various features an capabilities. - Components ./tasks.html are the building blocks of an evaluation: tasks, datasets, solvers, and scorers. - Models ./models.html covers specifying models and providers, along with caching, multimodal input, reasoning, batch mode, and concurrency. - Agents ./agents.html combine planning, memory, and tool use for longer-horizon tasks, including the built-in ReAct agent, multi-agent architectures, and bridges to external frameworks. - Tools ./tools.html extend models with custom and built-in tools, MCP integrations, sandboxing, and tool-call approval. - Running ./running.html covers running larger eval sets, with error handling, limits, parallelism, and early stopping. - Analysis ./analysis.html explains how to read eval logs, extract data frames, and scan transcripts for issues. - Extensions ./extensions.html shows how to extend Inspect with new model APIs, components, sandboxes, approvers, hooks, and filesystems. You may also want to explore the Evals ./evals/index.html listing of ready-to-run benchmark implementations, the Extensions ./extensions/index.html gallery of community packages, and the Reference ./reference/index.html for the complete Python and CLI API. Citation @software{UK AI Security Institute Inspect AI Framework 2024, author = {AI Security Institute, UK}, title = {Inspect {AI:} {Framework} for {Large} {Language} {Model} {Evaluations}}, date = {2024-05}, url = {https://github.com/UKGovernmentBEIS/inspect ai}, langid = {en} } Inspect AI: Framework for Large Language Model Evaluations . Released May. https://github.com/UKGovernmentBEIS/inspect ai https://github.com/UKGovernmentBEIS/inspect ai .