Long-horizon agents at half the cost Unreal Labs' Unreal Agent harness now matches Codex's score on the SWE-Marathon benchmark at 48% lower inference cost, $445.75 versus $850.52, after adding a context compaction mechanism. Both harnesses ran GPT-6.1 Sol with xhigh reasoning across 20 tasks × 8 attempts, each scoring 43.125% ±3.92 pp pooled pass@1 (69/160 passes). Unreal Labs also shipped a terminal interface installable via Homebrew and added support for Claude models through Anthropic's native Messages API. Long-horizon agents at half the cost Long-running agents are still too expensive. With the recent release of GPT ultrafast mode, we all noticed just how quickly agents can burn through token limits. An agent’s ability to tackle long-running tasks depends on several pieces working together: a capable language model, a harness, and an execution environment. Today, we’re sharing an improvement to our harness: context compaction, now available in Unreal Agent. With the new compaction mechanism, Unreal Agent achieves the same score as Codex at 48% lower cost on SWE-Marathon https://www.swe-marathon.org/ . Both agents used GPT-6.1 Sol with xhigh reasoning. Compaction Architecture compaction-architecture When designing the compaction mechanism, we wanted to preserve Unreal Agent’s non-blocking async architecture https://unreallabs.ai/blog/unreal-agent/ motivation-and-architecture and avoid a “stop-the-world” pause waiting for all tool calls to finish before performing compaction. Here’s how Unreal Agent’s compaction works: When Unreal Agent’s context size reaches a configured threshold, it submits a request to summarise the middle portion of the session context, keeping the original system instructions and a few recent turns intact. Unreal Agent doesn’t interrupt any tool calls that are running when compaction begins. Their initial tool call blocks are appended after the summarised message in the compacted context. Once those tool calls complete, their results enter the compacted context. Benchmark details benchmark-details Both harnesses use GPT-6.1 Sol, xhigh , evaluated on 20 tasks × 8 attempts . Scores are pooled pass@1. We follow the scoring methodology described in the SWE-Marathon paper https://arxiv.org/html/2606.07682v1 S3 . | Harness | Passes | Score ±1 SE | Total inference cost USD | Expected repeat-run cost USD , 95% CI | Harbor | |---|---|---|---|---|---| | Unreal Agent | 69/160 | 43.125% ±3.92 pp | $445.75 | $419.82–$472.65 | Harbor job https://hub.harborframework.com/jobs/18eab680-0191-4368-963d-0f2c82e61a13 | | Codex | 69/160 | 43.125% ±3.92 pp | $850.52 | $781.60–$920.40 | Harbor job https://hub.harborframework.com/jobs/1f7a26f1-af67-4df6-958f-1de559e3593d | TUI tui We’re also shipping a terminal interface TUI for Unreal Agent — an easy way to try it out without diving deep into the SDK. This TUI is our love letter to Borland Turbo Vision — the OG user interface many of our team members learned to program with. Install it with Homebrew, then launch the TUI: % brew install unreallabsai/tap/unreal-agent % unreal-agent Anthropic support anthropic-support Unreal Agent now supports Claude models through Anthropic’s native Messages API. Compared with OpenAI’s Responses API, it imposes stricter requirements on the ordering of messages and content blocks. In our harness, this requires custom handling of asynchronous tool results. Unlike Responses, Messages API expects tool results immediately after their corresponding tool calls. Results from tool calls that finish after the next model turn are delivered as synthetic user-content blocks.