Holo4: How H Company Built a Single Open-Weight Agent That Clicks, Codes, and Calls APIs H Company released Holo4 on September 28, 2026, a series of open-weight agentic models in 27B dense and 35B-A3B Mixture of Experts configurations that interact with software through GUIs, code execution, MCP servers, and REST APIs using a shared action space. The models are trained on roughly 10,000 tasks generated by an "Agentic Task Factory" pipeline and a three-stage process ending in merged domain-specific LoRA experts, scoring 85.2% on OSWorld and 85.1% on AndroidWorld for the 27B model. On OSWorld 2.0 the 27B model reaches 61.7% at an estimated $1.22 per task, versus 81.8% for Claude Opus 5.5. Most agentic AI models are specialists. A GUI-focused model can click through a browser but is lost when the task requires an API call. A tool-calling model can chain function calls but has no idea what to do when it needs to interact with a desktop application that has no API. Real work rarely respects these boundaries — a single business task might require reading a PDF, filling out a web form, querying a database via MCP, and writing a script to process the result. H Company's Holo4 https://hcompany.ai/newsroom/holo4 , released on September 28, 2026, is a direct attempt to close that gap. It is a series of open-weight agentic models — available in 27B dense and 35B-A3B Mixture of Experts configurations — that interact with software through any available interface: GUIs, code execution, MCP servers, and REST APIs. The same model, called the same way, regardless of platform. The fragmentation in agentic AI is not accidental. Training a model to reliably click through a GUI requires a very different kind of data than training it to chain API calls. Most labs have solved this by building separate models for separate contexts, or by routing between specialists at inference time. Holo4 takes a different approach: train a single model on all interface types simultaneously, using a shared action space. The model decides at each step which interface to use — screenshot-based GUI interaction, code execution in a sandbox, MCP tool calls, or direct API requests — based on what the task requires. H Company reports that this unified approach outperforms their previous models on every benchmark they tested, including OSWorld https://os-world.github.io/ 85.2% for the 27B model , AndroidWorld 85.1% , and AutomationBench 45.4% . The most technically interesting part of Holo4 is not the model architecture — both sizes use Qwen3.8 as the base — but the training data pipeline H Company calls the Agentic Task Factory https://hcompany.ai/newsroom/holo4 . The factory is a set of agentic pipelines that build interactive environments and verifiable tasks from documentation alone: product manuals, help center articles, screenshots of real websites, and open-source software. It has produced approximately 10,000 tasks across three categories: web app tasks 40% , MCP server tasks 30% , and desktop/OS tasks 30% . What makes this pipeline notable is its quality gate. A task is only kept if: Failed attempts loop back into an audit and hardening process rather than being discarded. This is a meaningful departure from the common practice of filtering trajectories by outcome alone — the factory is designed to produce tasks that are genuinely hard and genuinely solvable, not just tasks that happened to be completed. Holo4's training pipeline has three stages. First, supervised fine-tuning on 127B tokens, roughly 75% of which are successful agentic trajectories from the Task Factory, with the remainder covering multimodal reasoning, GUI grounding, and text-only tool use. Second, asynchronous online reinforcement learning on long-horizon tasks. Rather than training a single RL policy, H Company trains two specialized LoRA experts: one for desktop and web environments, one for terminal, MCP, and API environments. This separation allows each expert to develop strong priors for its domain without the two domains interfering with each other during RL. Third, both experts are merged back into the fine-tuned base model with equal weight. The result is a single model that carries the learned behaviors of both specialists. The HuggingFace model card https://huggingface.co/blog/Hcompany/holo4 notes that the training token distribution reflects this multi-domain focus: desktop 45% , web 14% , MCP/API 12% , mobile 3% , and other 26% . On OSWorld 2.0 — a benchmark focused on long, multi-step desktop workflows — Holo4 27B scores 61.7% at an estimated cost of $1.22 per task. For comparison, Claude Opus 5.5 scores 81.8% and GPT-6 Astra scores 73.5%, but both are closed models with significantly higher inference costs. The Holo4 35B-A3B MoE model scores 30.9% on the same benchmark at $0.61 per task, making it the more cost-efficient option for high-volume deployments. On AutomationBench, which tests business automation workflows, Holo4 27B scores 45.4% — competitive with GPT-5.6 Sol 45.8% and Kimi K3 46.7% , and above Claude Opus 5 50.3% is the frontier here . The 35B-A3B model scores 34.5% at $0.02 per task, which is a meaningful cost advantage for production use cases. H Company has published all trajectories behind their benchmark scores publicly, which is unusual in the agentic model space and allows independent verification of the reported numbers. H Company rebuilt their execution harness alongside the model — the loop that executes actions and manages context over hundreds of steps. The two largest changes were giving the agent a reliable memory capable of tracking hundreds of steps, and providing it with a shell on the desktop machine itself. Long-horizon agentic tasks fail in ways that are not always the model's fault. A model that loses track of what it has already done, or cannot recover from a failed action, will fail tasks it could theoretically complete. The harness improvements are likely responsible for a meaningful portion of the OSWorld 2.0 gains. Both Holo4 sizes are available on the H Models API and as open weights on HuggingFace in BF16, FP8, NVFP4, and 4-bit GGUF formats. The 27B model is released under CC BY-NC 4.0 research use . The 35B-A3B model is under Apache 2.0, making it suitable for commercial deployment. API pricing is $0.40 per million output tokens for the 27B and $0.30 for the 35B-A3B. H Company also released Holotron4 Nano, applying the same post-training stack to NVIDIA's Nemotron 3 Nano Omni model. On OSWorld CUA GUI, Holotron4 Nano improves over the base model by 55.3 percentage points — suggesting the training recipe transfers well to smaller base models. For teams building agentic workflows, Holo4 offers concrete advantages over current alternatives. The unified interface model eliminates the need to route between specialists, simplifying harness design and reducing routing errors. The Apache 2.0 license on the 35B-A3B model makes it viable for production deployment, and the published trajectories provide a starting point for domain-specific fine-tuning. The gap to frontier closed models on long-horizon benchmarks 61.7% vs. 81.8% on OSWorld 2.0 is real. For tasks requiring sustained multi-step reasoning over hundreds of actions, closed frontier models still have an edge. But for well-defined, verifiable business automation tasks, Holo4's cost-performance profile is competitive. The Agentic Task Factory is the most transferable idea here. Generating verifiable tasks from documentation — with quality gates that test both solvability and verifier precision — is applicable to any domain where you need agentic training data at scale. It is a more principled approach than filtering model rollouts by outcome alone.