{"slug": "harness-as-a-language", "title": "Harness as a Language", "summary": "Researchers at MIT CSAIL led by Zhening Li, Omar Khattab and Armando Solar-Lezama built JAZ, an agent framework whose only primitive is an LLM-backed `invoke` function that lets the model write arbitrary executable code, including recursive `invoke` calls, with all inputs and interaction history exposed as variables in the code environment. Using prompting alone and no external memory or file-system systems, JAZ beat Letta (MemGPT) by 8% at half the cost on the recall-heavy portion of StuLife, where recall exceeds the context window, and outperformed ACE by 4% at lower cost on AppWorld for continual self-improvement. The authors frame `invoke` as a language primitive whose implementation an LLM supplies at runtime on each call, generalizing existing code-mode agent loops and adding built-in hooks for constraints and monitoring instead of dedicated memory or self-improvement subsystems.", "body_md": "# Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity\n\nZhening Li, Omar Khattab, Armando Solar-Lezama and colleagues at MIT CSAIL build JAZ, an agent framework whose only primitive is an LLM-backed `invoke` function, and test whether it can replace dedicated memory and self-improvement systems.\n\n## Ask this paper\n\nSingle primitive. `invoke` lets the LLM write arbitrary executable code, including recursive calls to `invoke`, and every input and all interaction history are variables in the code environment.\n\nLanguage view. The authors treat `invoke` as a function whose implementation an LLM supplies at runtime on each call, generalizing existing code-mode agent loops.\n\nHooks, not subsystems. Programmers add constraints and monitoring through built-in hooks instead of attaching memory stores, file systems or custom tools.\n\nMemory result. With prompting only, JAZ beats Letta (MemGPT) by 8% at half the cost on the recall-heavy part of StuLife, where recall exceeds the context window.\n\nSelf-improvement result. On AppWorld it beats ACE by 4% at lower cost.\n\n## Abstract\n\nModern language-model agents are built around the \\textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the agent loop itself, such as memory systems and self-improving systems. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints and monitoring. Generalizing existing code-mode agent loops, \\texttt{invoke} is the simplest loop that satisfies two defining properties: (1) the LLM can write arbitrary executable code that can include recursive \\texttt{invoke}; (2) everything visible to the LLM --- all inputs to \\texttt{invoke} as well as its interaction history with the code environment --- are variables in the code environment. We motivate our design from first principles, viewing \\texttt{invoke} as a language primitive representing a function whose implementation is provided at runtime by an LLM every time it is called. To validate the design of our core \\texttt{invoke} primitive, we evaluate \\texttt{invoke} --- with only prompting, no manually designed tools, harness, or external systems (e.g., memory or the file system) --- on workflows traditionally implemented through specialized external harnesses. On long-horizon workflows requiring recall beyond the context window, JAZ invoke outperforms Letta (MemGPT) by 8\\% at half its cost on the recall-heavy portion of StuLife. On continual self-improvement, JAZ invoke outperforms ACE by 4\\% at a lower cost on AppWorld.", "url": "https://wpnews.pro/news/harness-as-a-language", "canonical_source": "https://academy.dair.ai/papers/harness-as-a-language-a-minimalist-agent-framework-with-maximal-expressivity-2609.26891", "published_at": "2026-09-28 01:06:34+00:00", "updated_at": "2026-09-28 01:30:58.420858+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["MIT CSAIL", "JAZ", "Zhening Li", "Omar Khattab", "Armando Solar-Lezama", "Letta", "MemGPT", "ACE"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/harness-as-a-language", "markdown": "https://wpnews.pro/news/harness-as-a-language.md", "text": "https://wpnews.pro/news/harness-as-a-language.txt", "jsonld": "https://wpnews.pro/news/harness-as-a-language.jsonld"}}