{"slug": "inspect-an-open-source-framework-for-large-language-model-evaluations", "title": "Inspect: An open-source framework for large language model evaluations", "summary": "The UK AI Security Institute and Meridian Labs released Inspect, an open-source framework for large language model evaluations that ships with over 200 pre-built evaluations and supports more than 20 model providers. Inspect is installable from PyPI via `pip install inspect-ai` and offers composable datasets, solvers, scorers, agents, and tools, plus sandboxing through Docker, Kubernetes, Modal, Proxmox, and Vagrant. The framework also supports running external agents such as Claude Code, Codex CLI, and Gemini CLI, and includes a web-based Inspect View tool and a VS Code extension.", "body_md": "# Inspect\n\nAn open-source framework for large language model evaluations\n\n## Welcome\n\nInspect is a framework for frontier AI evaluations developed by the [UK AI Security Institute](https://aisi.gov.uk) and [Meridian Labs](https://meridianlabs.ai). Inspect can be used for a broad range of evaluations that measure coding, agentic tasks, reasoning, knowledge, behavior, and multi-modal understanding. Core features of Inspect include:\n\n- Composable building blocks—datasets, agents, tools, and scorers—that make evaluations easy to write and reuse.\n- A collection of over 200 pre-built evaluations ready to run on any model.\n- Extensive tooling, including a web-based Inspect View tool for monitoring and visualizing evaluations and a VS Code Extension that assists with authoring and debugging.\n- Flexible support for tool calling—custom and MCP tools, as well as built-in bash, python, text editing, web search, web browsing, and computer tools.\n- Support for agent evaluations, including flexible built-in agents, multi-agent primitives, and the ability to run arbitrary external agents like Claude Code, Codex CLI, and Gemini CLI.\n- A sandboxing system that supports running untrusted model code in Docker, Kubernetes, Modal, Proxmox, Vagrant, and other systems via an extension API.\n\nWe’ll walk through two short “Hello, Inspect” examples below. Read on to learn the basics, then read the documentation on [Datasets](./datasets.html), [Solvers](./solvers.html), [Scorers](./scorers.html), [Tools](./tools.html), and [Agents](./agents.html) to learn how to create more advanced evaluations.\n\nIf you are primarily interested in running evaluations rather than developing new ones, see the [Evals](./evals/index.html) listing where you’ll find implementations for over 200 popular benchmarks.\n\n## Getting Started\n\nTo get started using Inspect:\n\n1. Install Inspect from PyPI with: \n\n```\npip install inspect-ai\n```\n\n2. If you are using VS Code, install the [Inspect VS Code Extension](./vscode.html) (not required but highly recommended).\n\nTo develop and run evaluations, you’ll also need access to a model, which typically requires installation of a Python package as well as ensuring that the appropriate API key is available in the environment. For example:\n\n```\npip install openai\nexport OPENAI_API_KEY=your-openai-api-key\ninspect eval simpleqa.py --model openai/gpt-4o\npip install anthropic\nexport ANTHROPIC_API_KEY=your-anthropic-api-key\ninspect eval simpleqa.py --model anthropic/claude-sonnet-4-0\npip install google-genai\nexport GOOGLE_API_KEY=your-google-api-key\ninspect eval simpleqa.py --model google/gemini-2.5-pro\npip install torch transformers\nexport HF_TOKEN=your-hf-token\ninspect eval simpleqa.py --model hf/meta-llama/Llama-2-7b-chat-hf\n```\n\nInspect has built-in support for over 20 model providers as well as support for local inference with HuggingFace, vLLM, and SGLang. See the documentation on [Model Providers](./providers.html) for details on all supported providers.\n\nIf you use a coding agent alongside Inspect, the [inspect-skills](https://github.com/meridianlabs-ai/inspect-skills#install) plugin provides skills that teach it to monitor running evals, read logs efficiently, and analyze results.\n\n## Hello, Inspect\n\nAn Inspect evaluation is a [Task](./tasks.html) that brings together three things:\n\n1. [Dataset](./datasets.html) that provides labelled samples—typically a table with`input` and`target` columns, where`input` is the prompt and`target` is the ideal answer or grading guidance.\n2. [Solver](./solvers.html) that produces an answer for each sample. This can be as simple as a single[generate()](./reference/inspect_ai.solver.html#generate) call to the model, or as sophisticated as a full agent that uses tools over many turns.\n3. [Scorer](./scorers.html) that evaluates the output—using text comparisons, model grading, or other custom schemes.\n\nLet’s look at two short examples: a question-answering benchmark and a capture the flag challenge.\n\n### Benchmark: SimpleQA\n\nThis task evaluates a model on [SimpleQA](https://openai.com/index/introducing-simpleqa/), a benchmark of short, fact-seeking questions(click on the numbers at right for further explanation):\n\n```\nsimpleqa.py\npython\nfrom inspect_ai import Task, task\nfrom inspect_ai.dataset import FieldSpec, hf_dataset\nfrom inspect_ai.scorer import model_graded_qa\nfrom inspect_ai.solver import generate\n\n1@task\ndef simpleqa():\n    return Task(\n2        dataset=hf_dataset(\n            \"codelion/SimpleQA-Verified\",\n            split=\"train\",\n3            sample_fields=FieldSpec(\n                input=\"problem\",\n                target=\"answer\",\n            ),\n        ),\n4        solver=generate(),\n5        scorer=model_graded_qa(),\n    )\n```\n\n- 1\n- \nThe `@task` decorator registers the function with Inspect so that`inspect eval` can discover and run it by name.\n- 2\n- \n[hf_dataset()](./reference/inspect_ai.dataset.html#hf_dataset) loads samples directly from Hugging Face. Inspect also reads CSV, JSON, and in-memory datasets.\n- 3\n- \n[FieldSpec](./reference/inspect_ai.dataset.html#fieldspec) declaratively maps the dataset’s`problem` and`answer` columns onto the sample’s`input` and`target` —no custom conversion function required.\n- 4\n- \nThe [generate()](./reference/inspect_ai.solver.html#generate) solver simply sends each`input` to the model and collects its response.\n- 5\n- \nBecause the answers are free-form text, [model_graded_qa()](./reference/inspect_ai.scorer.html#model_graded_qa) uses a model to grade each response against the`target` .\n\nRun it from the command line with `inspect eval`, choosing a model with `--model`:\n\n```\ninspect eval simpleqa.py --model openai/gpt-5\n```\n\nUse `inspect view` to view the results:\n\n```\ninspect view\n```\n\n### Agent: CTF Challenge\n\nAgent evaluations require the model take actions rather than just answer a question. Here’s a Capture the Flag (CTF) task where the [react()](./agents.html) agent explores a sandboxed system using [bash()](./reference/inspect_ai.tool.html#bash) and [todo_write()](./reference/inspect_ai.tool.html#todo_write) tools to find a hidden flag:\n\n```\nctf.py\npython\nfrom inspect_ai import Task, task\nfrom inspect_ai.agent import react\nfrom inspect_ai.dataset import json_dataset\nfrom inspect_ai.scorer import includes\nfrom inspect_ai.tool import bash, todo_write\n\n@task\ndef ctf():\n    return Task(\n        dataset=json_dataset(\"challenges.json\"),\n1        solver=react(\n            prompt=(\n                \"You are a Capture the Flag player.\n                Explore the system and find the flag.\"\n            ),\n            tools=[bash(), todo_write()],\n            attempts=3,\n        ),\n2        scorer=includes(),\n3        sandbox=\"docker\",\n    )\n```\n\n- 1\n- \n[react()](./reference/inspect_ai.agent.html#react) is a built-in agent that runs a reason-act-observe loop, giving the model the supplied`tools` until it submits an answer (here allowing up to 3 attempts).\n- 2\n- \nThe [includes()](./reference/inspect_ai.scorer.html#includes) scorer passes if the target flag appears in the agent’s submitted answer.\n- 3\n- \n`sandbox=\"docker\"` provides the isolated Docker container used by[bash()](./reference/inspect_ai.tool.html#bash) (configured by a`Dockerfile` or`compose.yaml` alongside the task).\n\nUse `inspect view` to view the results and look more carefully at individual transcripts:\n\nSee the [Tutorial](./tutorial.html) to explore more in-depth examples that demonstrate additional Inspect features and techniques.\n\n## Python API\n\nAbove we demonstrated using `inspect eval` from CLI to run evaluations—you can perform all of the same operations from directly within Python using the [eval()](./reference/inspect_ai.html#eval) function. For example:\n\n``` python\nfrom inspect_ai import eval\nfrom simpleqa import simpleqa\n\neval(simpleqa(), model=\"openai/gpt-5\")\n```\n\n## LLM Assistance\n\nAs you learn and use Inspect we recommend you provide an LLM with the documentation required for it to assist. There are two versions of LLM friendly markdown documentation available:\n\n- [llms.txt](llms.txt) : Documentation index, articles fetched as required (~2k tokens).\n- [llms-guide.txt](llms-guide.txt) : Full contents of all documentation (~185k tokens).\n\nThere is also a **Copy Page** button at the top of every page that provides a markdown version of the page.\n\n## Learning More\n\nTo learn more about using Inspect see the following documentation sections:\n\n- [Tutorial](./tutorial.html) includes several annotated examples demonstrating various features an capabilities.\n- [Components](./tasks.html) are the building blocks of an evaluation: tasks, datasets, solvers, and scorers.\n- [Models](./models.html) covers specifying models and providers, along with caching, multimodal input, reasoning, batch mode, and concurrency.\n- [Agents](./agents.html) combine planning, memory, and tool use for longer-horizon tasks, including the built-in ReAct agent, multi-agent architectures, and bridges to external frameworks.\n- [Tools](./tools.html) extend models with custom and built-in tools, MCP integrations, sandboxing, and tool-call approval.\n- [Running](./running.html) covers running larger eval sets, with error handling, limits, parallelism, and early stopping.\n- [Analysis](./analysis.html) explains how to read eval logs, extract data frames, and scan transcripts for issues.\n- [Extensions](./extensions.html) shows how to extend Inspect with new model APIs, components, sandboxes, approvers, hooks, and filesystems.\n\nYou may also want to explore the [Evals](./evals/index.html) listing of ready-to-run benchmark implementations, the [Extensions](./extensions/index.html) gallery of community packages, and the [Reference](./reference/index.html) for the complete Python and CLI API.\n\n## Citation\n\n```\n@software{UK_AI_Security_Institute_Inspect_AI_Framework_2024,\n  author = {AI Security Institute, UK},\n  title = {Inspect {AI:} {Framework} for {Large} {Language} {Model}\n    {Evaluations}},\n  date = {2024-05},\n  url = {https://github.com/UKGovernmentBEIS/inspect_ai},\n  langid = {en}\n}\n```\n\n*Inspect AI: Framework for Large Language Model Evaluations*. Released May.\n\n[https://github.com/UKGovernmentBEIS/inspect_ai](https://github.com/UKGovernmentBEIS/inspect_ai).", "url": "https://wpnews.pro/news/inspect-an-open-source-framework-for-large-language-model-evaluations", "canonical_source": "https://inspect.aisi.org.uk/", "published_at": "2026-09-29 23:40:04+00:00", "updated_at": "2026-09-29 23:47:39.560538+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "developer-tools", "ai-agents"], "entities": ["Inspect", "UK AI Security Institute", "Meridian Labs", "Claude Code", "Codex CLI", "Gemini CLI", "PyPI", "VS Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/inspect-an-open-source-framework-for-large-language-model-evaluations", "markdown": "https://wpnews.pro/news/inspect-an-open-source-framework-for-large-language-model-evaluations.md", "text": "https://wpnews.pro/news/inspect-an-open-source-framework-for-large-language-model-evaluations.txt", "jsonld": "https://wpnews.pro/news/inspect-an-open-source-framework-for-large-language-model-evaluations.jsonld"}}