Build an agent harness from scratch to pass Ian Goodfellow's intelligence test A developer tutorial published in 2026 walks through building a modern agent harness from scratch in Python, using only the standard library plus the OpenAI client, and tests the finished agent against a challenge Ian Goodfellow said in a 2019 interview would convince him "real AI" had been achieved. The guide follows the seven-part harness taxonomy from Barbaste et al., 2026, which studied eleven production coding harnesses including Claude Code, Codex CLI, Gemini CLI and Pi across the loop, LLM integration, tools, context management, safety controls, orchestration and extension surfaces. The core harness is a loop that builds context, calls a large language model, runs tools, and appends results back into context. Building an agent harness from scratch to pass Ian Goodfellow's intelligence test The topic of agent harness building has exploded in popularity in 2026, both as an area of development and of research. Almost all my colleagues and peers are thinking about harnesses - either directly, by building agentic products and services, or indirectly, by tuning and experimenting with the coding agents they use every day, like Claude Code or Codex. In this article, I want to walk through building a modern agent harness from scratch, starting with the simplest possible loop. Along the way, I'll share research and opinions I've come across about different approaches to building harnesses. By the end, you'll understand exactly what goes into a modern agent harness and have the skills to build your own. To test ours, I'll give the finished agent a challenge that Ian Goodfellow said in a 2019 interview https://www.youtube.com/watch?v=Z6rxFNMGdn0&t=3765s would convince him we've achieved "real AI". A coding harness is a harness built to write software, but it seems many general-purpose agent harnesses are now converging on being coding harnesses, so I'll use the terms interchangeably. What is an agent harness? The harness is everything in an AI Agent https://notesbylex.com/ai-agent that isn't the model. The simplest possible harness is a loop that: - builds context. - calls a large language model LLM . - runs some tools or finishes if the task is done . - adds the result back into context Willison, 2025 willisonAgentMayFinally2025 . Here's what that looks like in a few lines of Python: while True: reply = model context if not reply.tool call: break result = run tool reply.tool call context.append result Of course, in the real world, there's more to think about. The agent needs safety checks and sandboxing. Users need a way to extend their harness with extra skills and tools. There's typically an interface, either in the terminal or browser. Some agents can orchestrate sub-agents, and so on. Even so, the core of a harness is pretty straightforward. Barbaste et al., 2026 https://notesbylex.com/harness-engineering-anatomy-architecture-and-evolution-of-coding-agents studied eleven production coding harnesses, including Claude Code, Codex CLI, Gemini CLI and Pi, and compared them across seven aspects of harness design Barbaste et al., 2026 barbasteHarnessEngineeringAnatomy2026 : 1. the loop 2. the LLM integration 3. tools 4. context management 5. safety controls 6. orchestration running several agents or tasks together 7. extension surfaces places users can plug in their own tools and skills I'll use those as our guide, building one piece at a time. The paper also includes a minimal agent, which inspired this post. First, the imports. I'll stick to the standard library, apart from the OpenAI client, which saves writing our own API calls. python import html import inspect import json import pathlib import subprocess import sys import textwrap from functools import partial from itertools import islice from pprint import pprint from tempfile import TemporaryDirectory from typing import Any, Callable, Literal, NotRequired, Protocol, TypedDict, get type hints import openai We've already seen a basic loop, so we'll come back to it at the end, when we put it all together. First on our list is the model. Model Modern agentic AI is only possible thanks to the incredible magic of language models. Since the capabilities of new LLMs are constantly on an upward trajectory and the space is very competitive, today's best model may not be next month's, so it's useful to be as model-agnostic as possible, making it easy to switch. So, it's typical to hide the vendor's client behind a wrapper, and just plug a generic model into the mix. I've also added a cost field to track what the task has cost so far. class ToolCall TypedDict : id: str name: str arguments: str class ToolSpec TypedDict : name: str description: str parameters: dict str, Any class TextMessage TypedDict : role: Literal "system", "user" content: str class ToolResult TypedDict : role: Literal "tool" tool call id: str content: str class ModelReply TypedDict : role: Literal "assistant" content: str tool calls: NotRequired list ToolCall output items: NotRequired list dict str, Any Message = TextMessage | ToolResult | ModelReply class Model Protocol : cost: float def complete self, messages: list Message , tools: list ToolSpec - ModelReply: ... Your wrapper will likely also need to handle reasoning LLM Reasoning https://notesbylex.com/llm-reasoning , which most modern LLMs support, along with streaming responses, errors and so forth. Here's a basic model wrapper for GPT-6 Luna, using the Responses API. The output items field keeps the API's reasoning, text and function call items together for the next turn OpenAI openaiResponsesFunctionCalling . The rest of the harness passes this provider state through untouched: class ResponsesModel: """OpenAI Responses adapter. Swap this for another model provider.""" def init self, client, name: str = "gpt-6-luna", input price: float = 0.10, output price: float = 0.50, web search: bool = False, : self.client = client self.name = name self.input price = input price / 1e6 self.output price = output price / 1e6 self.web search = web search self.cost = 0.0 def complete self, messages: list Message , tools: list ToolSpec - ModelReply: api input = for message in messages: if message "role" == "tool": api input.append {"type": "function call output", "call id": message "tool call id" , "output": message "content" } elif message "role" == "assistant" and "output items" in message: api input.extend message "output items" else: api input.append {"role": message "role" , "content": message "content" } response = self.client.responses.create model=self.name, input=api input, tools= {"type": "function", tool, "strict": False} for tool in tools + {"type": "web search"} if self.web search and tools else , reasoning={"effort": "low"}, if response.status = "completed": raise RuntimeError f"model response was {response.status}" usage = response.usage if usage: self.cost += usage.input tokens self.input price + usage.output tokens self.output price reply: ModelReply = { "role": "assistant", "content": response.output text, "output items": item.model dump exclude none=True for item in response.output , } calls = item for item in response.output if item.type == "function call" if calls: reply "tool calls" = {"id": call.call id, "name": call.name, "arguments": call.arguments} for call in calls return reply The tool schemas below have optional arguments, so this adapter sets strict=False . Otherwise, the Responses API tries to make them strict OpenAI openaiResponsesFunctionCalling . And give that a little test run: client = openai.OpenAI model = ResponsesModel client reply = model.complete {"role": "user", "content": "What is 1+1?"} , tools= print reply "content" 2 That's the model done - now for the tools. Tools In a lot of ways, tools are the most interesting part of the harness. They let the model act on the world and get feedback from it. Tool Use https://notesbylex.com/tool-use in LLMs dates back to 2022-2023, with papers like: - TALM: Tool Augmented Language Models Parisi et al., 2022 parisiTALMToolAugmented2022 - PAL: Program-aided Language Models Gao et al., 2022 gaoPALProgramaidedLanguage2022 - Toolformer Schick et al., 2023 schickToolformerLanguageModels2023 Each showed how tools can extend what LLMs can do. Later in 2023, OpenAI introduced function calling, which gave us a schema for structured tool definitions and a way of returning tool outputs to the model OpenAI, 2023 openaiFunctionCallingOther2023 . Other vendors soon adopted the idea. Tool use has seen the most debate and change in harness design in recent months. The community seems to be heading toward a consensus: as LLMs get more capable, the harness should get simpler, with fewer tools. Garry Tan describes Thin Harnesses https://notesbylex.com/thin-harnesses as systems that run the loop, handle files, manage context and enforce safety, while skills and deterministic tools do the rest Tan, 2026 tanThinHarnessFat2026 . At one extreme, mini-SWE-agent https://github.com/SWE-agent/mini-swe-agent uses only bash. I'll follow Pi https://pi.dev/ , a minimal open-source harness, and implement just four tools: read, write, edit and bash. I'll call them read file , write file , search replace and bash . Tool use has two parts: the code that runs each tool, and registering the tools with the model. I'll start with the former. It's good practice to truncate tool outputs, so a runaway log doesn't exhaust the context budget. The loop will run every tool result through this: php def truncate text text: str, limit: int = 25 000 - str: return text if len text <= limit else text :limit + "\n... truncated " print truncate text "This is some long winded text...", limit=15 This is some lo ... truncated Now the bash tool. I'll write a simple tool that delegates to subprocess.run , which means it can run anything on your machine. We'll come back to that in safety controls. I've also included a timeout so a hung command can't stall the agent: python def tool bash cmd: str, timeout: int = 120, , base dir: pathlib.Path = pathlib.Path "." - str: """Run a shell command from the working directory. Timeout is in seconds max 3600 .""" if not 1 <= timeout <= 3600: raise ValueError "timeout must be between 1 and 3600 seconds" result = subprocess.run cmd, shell=True, cwd=base dir, capture output=True, text=True, timeout=timeout output = f"exit={result.returncode}" if result.stdout: output.append "stdout:\n" + textwrap.indent result.stdout.rstrip "\n" , " " if result.stderr: output.append "stderr:\n" + textwrap.indent result.stderr.rstrip "\n" , " " return "\n".join output The examples use a small workspace folder, agent-harness-working : BASE DIR = pathlib.Path "agent-harness-working" print tool bash "ls", base dir=BASE DIR exit=0 stdout: AGENTS.md data Before the file tools, here's a helper that rejects paths outside the working directory. It doesn't restrict bash. php def workspace path path: str, base dir: pathlib.Path - pathlib.Path: root = base dir.resolve target = root / path .resolve if not target.is relative to root : raise ValueError "path is outside the working directory" return target The read tool returns numbered lines. You can choose how many lines to skip and how many to return: python def tool read file path: str, offset: int = 0, limit: int = 2000, , base dir: pathlib.Path = pathlib.Path "." - str: """Read numbered lines from a file in the working directory.""" if offset < 0 or limit < 1: raise ValueError "offset must be non-negative and limit must be positive" with workspace path path, base dir .open as file: lines = islice file, offset, offset + limit return "\n".join f"{i:4}: " + line.rstrip "\r\n" for i, line in enumerate lines, offset + 1 print tool read file "AGENTS.md", limit=5, base dir=BASE DIR 1: AGENTS file for Lex's Simple Agent Harness 2: 3: This is the working directory for the harness built in the post "An Agent Harness in one blog post" on notesbylex.com. The harness loads this file into context at the start of every session. 4: 5: Environment The write tool saves text to a file, creating any missing folders: python def tool write file path: str, content: str, , base dir: pathlib.Path = pathlib.Path "." - str: """Write a file in the working directory.""" target = workspace path path, base dir target.parent.mkdir parents=True, exist ok=True target.write text content return f"wrote {len content.encode } bytes" Finally, the edit tool replaces an exact string, but only when it appears once: python def tool search replace path: str, search: str, replace: str, , base dir: pathlib.Path = pathlib.Path "." - str: """Replace one exact match in a file in the working directory.""" if not search: raise ValueError "search string must not be empty" p = workspace path path, base dir text = p.read text count = text.count search if count = 1: raise ValueError f"search string occurs {count}x, expected exactly once" p.write text text.replace search, replace, 1 return "OK" Let's try the write and edit tools together, in a temporary folder that cleans up after itself: with TemporaryDirectory dir=BASE DIR as tmp: demo dir = pathlib.Path tmp tool write file "demo.txt", "Hello, world \n", base dir=demo dir print tool search replace "demo.txt", "world", "agent", base dir=demo dir print tool read file "demo.txt", base dir=demo dir OK 1: Hello, agent Now register the four tools by name: TOOLS: dict str, Callable ..., str = { "bash": tool bash, "read file": tool read file, "write file": tool write file, "search replace": tool search replace, } Then describe them in the JSON schema format that tool-calling APIs expect. Rather than writing each schema by hand, I'll build it from the function's arguments: php def schema name: str, f - ToolSpec: params = {key: param for key, param in inspect.signature f .parameters.items if key = "base dir"} types = get type hints f props = {key: {"type": "integer" if types.get key is int else "string"} for key in params} required = key for key, param in params.items if param.default is inspect.Parameter.empty return { "name": name, "description": inspect.getdoc f or "", "parameters": {"type": "object", "properties": props, "required": required, "additionalProperties": False}, } pprint schema "read file", tool read file , sort dicts=False {'name': 'read file', 'description': 'Read numbered lines from a file in the working directory.', 'parameters': {'type': 'object', 'properties': {'path': {'type': 'string'}, 'offset': {'type': 'integer'}, 'limit': {'type': 'integer'}}, 'required': 'path' , 'additionalProperties': False}} Finally, let's check that OpenAI can call a tool. I'll give it only bash and ask it to run pwd once: messages: list Message = {"role": "user", "content": "Call the bash tool once with pwd , then report the result."} reply = model.complete messages, tools= schema "bash", tool bash call = reply "tool calls" 0 result = TOOLS call "name" json.loads call "arguments" , base dir=BASE DIR print result messages.extend reply, {"role": "tool", "tool call id": call "id" , "content": result} print model.complete messages, tools= "content" exit=0 stdout: /Users/lex/code/private-notes/public/code/agent-harness-working pwd returned /Users/lex/code/private-notes/public/code/agent-harness-working . Context & memory Even with a giant modern context window of 1M+ tokens, a long task will eventually overflow it, so we need a strategy for Context Management https://notesbylex.com/context-management . The typical approaches are to truncate the old context or to have an LLM summarise the conversation so far. Fan et al., 2026 https://notesbylex.com/an-empirical-study-of-harness-design-for-coding-agents tested a few methods and found the right one depends largely on the model: context management mattered more as the context window shrank, mostly by preventing overflow failures Fan et al., 2026 fanEmpiricalStudyHarness2026 . Unsurprisingly, the more context the model has, the less the approach matters. We'll use the common approach of asking the model to summarise. To decide when to compact, we'll estimate about four characters per token, a rough rule of thumb. When we do, the system prompt, the task and the recent messages stay, and everything in the middle gets swapped for a summary. We'll wrap the summary in a