Holo4: The Open-Weight Agent That Clicks, Codes, and Calls Tools Paris-based lab H Company released Holo4, an open-weight agent model that unifies GUI clicking, code execution, and MCP/API tool calls in a single agent loop instead of chaining separate specialized models. Two main sizes ship on Hugging Face: a 27B dense model built on Qwen3.8-27B and a 35B mixture-of-experts model with 3B active parameters built on Qwen3.6, both released under license alongside a Nemotron Nano variant trained with the same recipe. Benchmarks show the 27B dense model trailing frontier models like Claude Opus and GPT-5.5 on OSWorld but at a fraction of the per-task cost since it runs locally, and the dense 27B model clearly outperforms the 35B mixture-of-experts model on vision-heavy benchmarks. Holo4: The Open-Weight Agent That Clicks, Codes, and Calls Tools H Company's Holo4 is an open-weight agent model that chooses between GUI clicks, code execution, and MCP tool calls. Here's how it works. What is Holo4? Holo4 is an open-weight agent model from the Paris-based lab H Company, built to handle three different ways of getting work done on a computer: clicking through graphical interfaces, writing and running code, and calling tools through MCP servers or APIs. Instead of using separate models for visual tasks and API tasks, Holo4 is trained to decide which approach fits a given step and switch between them inside a single agent loop. TL;DR - Holo4 unifies three action types GUI clicking, code execution, and MCP/API tool calls in one model instead of requiring separate specialized models chained together. - Two main sizes ship on Hugging Face : a 27B dense model built on Qwen3.8-27B, and a 35B mixture-of-experts model 3B active built on Qwen3.6, both released under license alongside a Nemotron Nano variant trained with the same recipe. - The model only decides, it never acts : a separate harness H Company’s own “H agents” SDK, or any compatible agent framework executes the clicks, runs the code, and calls the tools, then feeds results back to the model. - Benchmarks show the 27B dense model trailing frontier models like Claude Opus and GPT-5.5 on OSWorld , but at a fraction of the per-task cost since it runs locally instead of through expensive vision-heavy API calls. - The dense 27B model clearly outperforms the mixture-of-experts 35B model on vision-heavy benchmarks, suggesting more active parameters matter a lot when a model has to interpret screenshots. - Training relies on an “agentic task factory” : automatically checkable tasks across real websites, open-source software, and desktop environments, some of which can be solved either visually or via tools, with reinforcement learning rewarding whichever path actually works. - A hands-on demo showed the same task running two ways : with MCP access turned off, the model navigated a web UI step by step and finished with a code step; with MCP turned on, it skipped the GUI entirely and called a tool directly, cutting the number of steps substantially. Plans first. Then code. Remy writes the spec, manages the build, and ships the app. Why does an agent need to choose between clicking, coding, and calling tools? Most real-world automation tasks don’t stay on one side of the line between “visual” and “API-driven” work. A typical job might involve pulling a report out of a legacy application with no API, cleaning the resulting data with a script, and then posting it to Slack. That’s a GUI step, a code step, and a tool-call step, back to back. Up to now, teams have generally built or picked models for one side of that divide. A GUI agent trained on screenshots can act like a human using a browser: it works on anything a person can see, but it’s slow, fragile when buttons move, and expensive in token terms since every step burns context on an image. A tool-calling model is fast and precise when a clean API or MCP server exists, but it’s stuck the moment there isn’t one. The common workaround has been to chain models: one model handles screenshots, another handles orchestration and tool calls, and a harness stitches the outputs together. Holo4’s pitch is that a single model can learn when to use which lane, cutting out the handoff between separate systems. How does the Holo4 agent loop actually work? The basic structure is the standard agent loop: a task comes in, the model looks at the current state, picks one action, and hands it off. The harness, meaning the surrounding code, is what actually executes that action, whether that’s clicking a button, running a script, or calling an MCP tool. The result a new screenshot, a return value, an API response gets captured and fed back into the model along with the memory of prior steps. This repeats until the task is done. The important distinction is that the model itself never touches anything directly. It only decides. All the doing happens in the harness, which is one reason benchmark scores for agent models can vary a lot depending on which harness they’re run through. Holo4 was built and optimized around three specific action types treated as equally valid, first-class options rather than one being the main mode and the others being add-ons: - GUI actions : the model reads a screenshot and issues clicks or typed input, working anywhere a human could, but slower and more brittle to interface changes. - Code execution : the model writes a script that runs in a sandbox or shell and returns an output, fast and precise but only usable when the task can actually be scripted. - Tool calls : the model calls an MCP server or API function and gets structured data back, generally the most reliable option when a suitable tool exists. Earlier Holo models from H Company relied primarily on function calling. The shift in Holo4 is training a single model to move fluidly across all three lanes in the same task, picking the cheapest or most reliable option available at each step rather than defaulting to one method. What models did H Company actually release? H Company shipped two main models plus a variant: - Holo4-27B , a dense model built on Qwen3.8-27B. - Holo4 35B-A3B , a mixture-of-experts model with 3 billion active parameters, built on Qwen3.6. - Holo4 Nemotron Nano , the same training recipe applied to Nvidia’s Nemotron 3 Nano Omni model. Everyone else built a construction worker. We built the contractor. One file at a time. UI, API, database, deploy. Weights for all of them are published on Hugging Face the 27B model’s repository shows a Qwen3.5-architecture based checkpoint released under a CC-BY-NC-4.0 license , with multiple quantization options available and an H Company-hosted API for anyone who doesn’t want to self-host. The full bfloat16 version of the 27B model runs around 54 GB, meaning it needs a serious GPU a workstation-class card with large VRAM to run locally, and anyone deploying it should budget extra headroom in the KV cache since screenshots consume it fast. How well does Holo4 actually perform? On benchmarks like OSWorld, the 27B dense model lands behind current frontier multimodal models from major labs, but not by a dramatic margin, and it reportedly edges past older frontier checkpoints like GPT-5.5 on some measures. The more notable gap shows up on cost: frontier models charging per API call with image inputs can run several times more expensive per task than running Holo4 locally or through H Company’s API. The more striking internal comparison is between H Company’s own two main releases. The 35B mixture-of-experts model, despite having a larger total parameter count, scores meaningfully lower than the 27B dense model on vision-heavy benchmarks like OSWorld, in some cases less than half the score. That gap narrows on less visual tasks. The likely explanation is that interpreting screenshots benefits more from raw active parameter count 27B active vs. 3B active than from a larger but sparser mixture-of-experts setup. How was Holo4 trained to choose between clicking, coding, and calling tools? H Company built what they call an “agentic task factory”: a set of pipelines that generate interactive environments and tasks with automatically checkable outcomes. These include screenshots of real websites, open-source software, specific MCP servers, and full desktop environments. Some of these environments are deliberately hybrid, meaning the same underlying task say, finding information on a website can be solved either by navigating the UI visually or by calling a tool or writing code against it. Reinforcement learning sits on top of this setup, rewarding whichever approach actually solves the task more efficiently in a given context. Over enough examples, the model learns patterns like “clicking through several menus is slower than calling the API that’s sitting right there” and the reverse, when a screenshot captures details no API exposes. The training pipeline itself starts with supervised fine-tuning on roughly 127 billion tokens, followed by training two separate reinforcement learning experts that are later merged back into a single model. H Company also rebuilt its own agent harness around this training approach, using failure analysis to find weak points, with reliable long-horizon memory tracking and relating hundreds of steps flagged as one of the bigger improvements that came out of that process. The same training recipe was applied across the Qwen3.8 dense model, the Qwen3.6 mixture-of-experts model, and the Nemotron Nano variant, suggesting H Company is investing in a reusable post-training pipeline rather than a one-off model. With Qwen4 already confirmed to include a 27B variant, a future Holo release built on that base seems like a natural next step. Is Holo4 worth using as your main agent model? Remy doesn't write the code. It manages the agents who do. Remy runs the project. The specialists do the work. You work with the PM, not the implementers. It depends on the setup. H Company runs its own benchmarks with Holo4 handling both planning and acting end to end, and the hands-on demo in the source material showed it working that way: completing a task purely through GUI navigation when no MCP tool was available, then completing the identical task in far fewer steps once an MCP tool was switched on, skipping the UI entirely. But Holo4 doesn’t have to be the only model in the loop. H Company’s own harness and SDK, called H agents, is built to let a larger reasoning model handle high-level planning while delegating screen-heavy subtasks to Holo4, including integration with tools like Claude Code over MCP. For teams that already have a strong orchestration model and just need something to reliably handle GUI automation, running Holo4 as a specialized sub-agent for the “computer use” portion of a workflow looks like a reasonable middle ground. Frequently Asked Questions What makes Holo4 different from other computer-use agents? Most computer-use or GUI agent models are trained primarily on screenshots and clicking. Holo4 is trained to treat GUI actions, code execution, and MCP/API tool calls as equally valid options and to choose between them within the same task, rather than relying on a separate model or harness for each. What hardware do you need to run Holo4 locally? The full bfloat16 version of the 27B model is around 54 GB, which requires a workstation-class GPU with substantial VRAM. Quantized versions are also available on Hugging Face for lower-resource setups, though screenshot-heavy tasks still consume KV cache quickly regardless of quantization. Is the mixture-of-experts version better than the dense 27B model? Not for vision-heavy tasks. The 35B mixture-of-experts model 3B active parameters scores noticeably lower than the 27B dense model on benchmarks like OSWorld, suggesting that processing screenshots benefits more from higher active parameter count than from a sparser, larger total model. Does Holo4 need MCP tools to work? No. Holo4 can complete tasks purely through GUI navigation when no tools are available, falling back to screenshots and clicks. But when MCP tools or APIs exist for a task, it will generally use them instead, since they’re faster and more reliable than clicking through an interface. Can Holo4 be combined with other AI models? Yes. H Company’s own H agents SDK supports using a larger model for high-level planning and orchestration while handing off GUI-heavy subtasks to Holo4, including integration with tools like Claude Code through MCP.