Caisi Cyber Evaluations The Center for AI Standards and Innovation (CAISI) released a repository packaging 2 agents and 3 benchmarks for agentic cyber capability evaluations of AI systems, built on the Inspect evaluation framework. The package includes a CVE-Bench version with 7 tasks plus a dev task (down from 40 tasks in the upstream repository), a Cybench fork with original task descriptions and single-eval support replacing 40 separate .eval files, and an integration that converts pwn.college dojo tasks, including the 615-task ctf-archive dojo, into Inspect-AI tasks. The tooling requires a Ghidra-as-a-Service container and supports models such as anthropic/claude-3-5-sonnet-20240620 and openai/o3-mini-2025-01-31 with configurable token limits and reasoning effort. This repository packages a few benchmarks and agents used by the Center for AI Standards and Innovation CAISI https://www.nist.gov/caisi for cyber capability evaluations of AI systems. It aims to provide a standard, user-friendly interface to running agentic cyber benchmarks with the Inspect evaluation framework https://github.com/UKGovernmentBEIS/inspect ai/ . The package currently includes 2 agents and 3 benchmarks. - Based on CVE-Bench created by Yuxuan Zhu et al. https://arxiv.org/pdf/2503.17332 . - CAISI's modifications are largely around the Inspect integration and automated evaluation. - This version of the benchmark has 7 tasks plus a dev tasks. - The upstream repository https://github.com/uiuc-kang-lab/cve-bench has 40 tasks. - This code is based off the public cybench fork maintained by the UK AI Security Institute here https://github.com/UKGovernmentBEIS/inspect evals/tree/main/src/inspect evals/cybench , originally developed by Andy Zhang et al. paper https://arxiv.org/abs/2408.08926 . - CAISI added original task descriptions, and updated system/task level prompts. - CAISI updated the benchmark to support running all samples within a single eval i.e., produce one .eval file instead of 40 . - After locally cloning a pwn.college dojo of tasks, this code can convert the tasks into Inspect-AI tasks. - The ctf-archive dojo https://github.com/pwncollege/ctf-archive/ provides 615 tasks. - CAISI developed the integration between pwn.college and Inspect to enable this. Note this repository only contains the integration, the pwn.college tasks live in their respective repos. 1. Create a virtual environment in which to install ucb Install uv per its docs https://docs.astral.sh/uv/getting-started/installation/ , then uv venv source .venv/bin/activate 1. Install ucb: uv sync 1. Configure variables for your run. You should now create a .env file and add the relevant API keys to it. It is highly recommend to use a container registry via UCB CONTAINER REGISTRY which can point at a domain and a directory prefix, e.g., "gitlab.com/caisi/cyber/ucb/". This should end with a trailing slash. ucb env-init Populate .env from a template vim .env Then fill it in manually 1. Build or pull containers If you have previously pushed images to your container registry, you can skip building and simply run ucb pull to download them. Otherwise you'll need to build your images with ucb build . You can add a --push argument if you want to push the images to a container registry, assuming you have set one up. 1. Launch Ghidra-as-a-Service Gaas container ucb gaas 1. Confirm functionality by running a single task with a small budget: inspect eval ucb/cybench --solver ucb/agent --model anthropic/claude-3-5-sonnet-20240620 --limit 1 --token-limit 2000 Ensure your GaaS server is running e.g., with docker ps if you haven't started it, launch it with ucb gaas . GaaS will cache Ghidra projects and static analysis analysis results. Once you've confirmed GaaS is running, you'll want to launch an eval with inspect eval or inspect eval-set combined with the arguments described below. You should specify at least one benchmark to run, either ucb/cybench or ucb/cvebench as part of your inspect command. Note you can put both in the command and the tasks from both benchmarks will be run, but this is currently incompatible with the UCB solvers. You should also specify a single agent with --solver ucb/cybench agent or --solver ucb/ctf solver . You MUST specify an agent. - A model can be set with --model=name , for example --model=openai/o3-mini-2025-01-31 . If using eval-set , multiple models can be specified and should be separated by a comma, for example --model=openai/o3-mini-2025-01-31,anthropic/claude-3-5-sonnet-20240620 - The total number of tokens can be limited with --token-limit X . - Reasoning effort can be set with --reasoning-effort