{"slug": "long-horizon-agents-at-half-the-cost", "title": "Long-horizon agents at half the cost", "summary": "Unreal Labs' Unreal Agent harness now matches Codex's score on the SWE-Marathon benchmark at 48% lower inference cost, $445.75 versus $850.52, after adding a context compaction mechanism. Both harnesses ran GPT-6.1 Sol with xhigh reasoning across 20 tasks × 8 attempts, each scoring 43.125% ±3.92 pp pooled pass@1 (69/160 passes). Unreal Labs also shipped a terminal interface installable via Homebrew and added support for Claude models through Anthropic's native Messages API.", "body_md": "# Long-horizon agents at half the cost\n\nLong-running agents are still too expensive. With the recent release of GPT ultrafast mode, we all noticed just how quickly agents can burn through token limits.\n\nAn agent’s ability to tackle long-running tasks depends on several pieces working together: a capable language model, a harness, and an execution environment. Today, we’re sharing an improvement to our harness: context compaction, now available in Unreal Agent.\n\nWith the new compaction mechanism, Unreal Agent achieves the same score as Codex at 48% lower cost on [SWE-Marathon](https://www.swe-marathon.org/). Both agents used GPT-6.1 Sol with xhigh reasoning.\n\n## [Compaction Architecture](#compaction-architecture)\n\nWhen designing the compaction mechanism, we wanted to preserve Unreal Agent’s [non-blocking async architecture](https://unreallabs.ai/blog/unreal-agent/#motivation-and-architecture) and avoid a “stop-the-world” pause waiting for all tool calls to finish before performing compaction.\n\nHere’s how Unreal Agent’s compaction works:\n\nWhen Unreal Agent’s context size reaches a configured threshold, it submits a request to summarise the middle portion of the session context, keeping the original system instructions and a few recent turns intact.\n\nUnreal Agent doesn’t interrupt any tool calls that are running when compaction begins. Their initial tool call blocks are appended after the summarised message in the compacted context. Once those tool calls complete, their results enter the compacted context.\n\n## [Benchmark details](#benchmark-details)\n\nBoth harnesses use **GPT-6.1 Sol, xhigh**, evaluated on **20 tasks × 8 attempts**. Scores are pooled pass@1.\n\nWe follow the scoring methodology described in the [SWE-Marathon paper](https://arxiv.org/html/2606.07682v1#S3).\n\n| Harness | Passes | Score ±1 SE | Total inference cost (USD) | Expected repeat-run cost (USD), 95% CI | Harbor | \n|---|---|---|---|---|---|\n| Unreal Agent | 69/160 | **43.125% ±3.92 pp** | **$445.75** | $419.82–$472.65 | [Harbor job](https://hub.harborframework.com/jobs/18eab680-0191-4368-963d-0f2c82e61a13) | \n| Codex | 69/160 | **43.125% ±3.92 pp** | **$850.52** | $781.60–$920.40 | [Harbor job](https://hub.harborframework.com/jobs/1f7a26f1-af67-4df6-958f-1de559e3593d) | \n\n## [TUI](#tui)\n\nWe’re also shipping a terminal interface (TUI) for Unreal Agent — an easy way to try it out without diving deep into the SDK. This TUI is our love letter to Borland Turbo Vision — the OG user interface many of our team members learned to program with.\n\nInstall it with Homebrew, then launch the TUI:\n\n```\n% brew install unreallabsai/tap/unreal-agent\n% unreal-agent\n```\n\n## [Anthropic support](#anthropic-support)\n\nUnreal Agent now supports Claude models through Anthropic’s native Messages API. Compared with OpenAI’s Responses API, it imposes stricter requirements on the ordering of messages and content blocks. In our harness, this requires custom handling of asynchronous tool results.\n\nUnlike Responses, Messages API expects tool results immediately after their corresponding tool calls. Results from tool calls that finish after the next model turn are delivered as synthetic user-content blocks.", "url": "https://wpnews.pro/news/long-horizon-agents-at-half-the-cost", "canonical_source": "https://unreallabs.ai/blog/long-horizon-agents-at-half-the-cost/", "published_at": "2026-10-06 17:30:45+00:00", "updated_at": "2026-10-06 17:50:27.725334+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "ai-infrastructure", "developer-tools"], "entities": ["Unreal Labs", "Unreal Agent", "Codex", "GPT-6.1 Sol", "SWE-Marathon", "Anthropic", "Claude", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/long-horizon-agents-at-half-the-cost", "markdown": "https://wpnews.pro/news/long-horizon-agents-at-half-the-cost.md", "text": "https://wpnews.pro/news/long-horizon-agents-at-half-the-cost.txt", "jsonld": "https://wpnews.pro/news/long-horizon-agents-at-half-the-cost.jsonld"}}