Open-Weight Computer Use AI Matches Proprietary Models at 1/100th the Cost


Imagine an AI that can watch your screen, move your mouse, type in fields, and complete multi-step tasks across any software -- without being built into that software, without sending your data to a cloud API, and without paying dollars per task.
Until this week, that was the promise of proprietary computer-use agents from OpenAI, Anthropic, and Google. The trade-off was clear: access to frontier-level computer-use AI meant accepting closed APIs, per-task pricing that runs into dollars, and a black box where you cannot inspect or modify how the model works.
H Company just made that trade-off obsolete.
Holo4: Open weights, proprietary-level performance
On September 28, H Company released Holo4, a family of open-weight models built specifically for computer-use agents. Available in two sizes -- a 27B dense model and a 35B-A3B Mixture-of-Experts variant -- Holo4 matches the performance of proprietary frontier models on standard computer-use benchmarks while costing a fraction as much per task.
The headline number: Holo4-27B scores 85.2% on OSWorld, the industry standard benchmark for computer-use AI, at a cost of just $0.08 per task. For comparison, Anthropic's Fable 5 scores 86.0% and OpenAI's GPT-5.5 scores 78.7% -- both at significantly higher inference costs. On OSWorld 2.0, which measures long, multi-step computer workflows, Holo4-27B reaches 61.7% at $1.22 per task, while Anthropic's Opus 5.5 hits 81.8% but costs $8.48 per task.
The cost advantage is even more dramatic with the smaller MoE variant: Holo4-35B-A3B scores 80.8% on OSWorld at just $0.05 per task.
What Holo4 actually does
Holo4 is not just another language model with some vision capabilities tacked on. It is purpose-built for the full spectrum of computer interaction:
- GUI control: It watches screenshots and decides where to click, what to type, and which elements to interact with -- the same way a human operates a desktop, web app, or Android interface.
- Code execution: It can write and run its own code, inspect the output, and adjust course.
- API and MCP calls: It calls Model Context Protocol tools and REST APIs directly, choosing the best interface for each task.
The same model works across desktop, web, and mobile without needing a different configuration. It handles up to 262,144 tokens of context, enough to maintain awareness across hundreds of interaction steps.
The architecture behind the numbers
Holo4-27B is built on the Qwen3.8 dense architecture and was trained through a three-stage pipeline: supervised fine-tuning on 127 billion tokens of agentic trajectories (roughly 45% desktop interactions, 14% web, 12% MCP and API, and 3% mobile), followed by asynchronous online reinforcement learning that produced two specialized LoRA experts -- one focused on desktop and web, the other on terminal, MCP, and API interactions. These experts were merged back into the base model with equal weight.
The company also rebuilt its agentic harness -- the loop that executes the model's decisions and manages context across hundreds of steps. Engineers analyzed why agents failed on each task and redesigned the harness to include a reliable long-term memory and a shell running directly on the desktop machine. The result is a system that does not just see the screen better but remembers what it learned across multiple interactions.
Open weights, but not fully open
The Holo4-27B model weights are published on Hugging Face under the CC BY-NC 4.0 (non-commercial) license, available in BF16, FP8, NVFP4, and 4-bit GGUF formats. The larger Holo4-35B-A3B variant is available under Apache 2.0 through H Company's API at $2.00 per million output tokens.
H Company has also open-sourced every trajectory behind its public benchmark scores, replayable through a dedicated viewer and downloadable as a Hugging Face dataset -- a level of transparency that proprietary labs have not matched.
One notable limitation: 480 of the 600 tasks in the AutomationBench public v1.0.6 set fall within the data split from which the company collected training data, potentially inflating those scores. On the 120 held-out tasks, Holo4-27B scores 49.3% compared to 45.4% on the full set -- still strong, but the gap is worth noting.
The Holotron4 Nano bonus
Alongside Holo4, H Company released Holotron4 Nano, an updated version of its compact agent model. As a member of the NVIDIA Nemotron Coalition, the company applied its post-training stack to Nemotron 3 Nano Omni, achieving dramatic gains: OSWorld scores jumped from 21.0% to 76.3%, and AutomationBench from 19.4% to 35.6%. These gains demonstrate that H Company's training recipe transfers across foundation models and is not dependent on model size.
What this means for the agent economy
The AI agent market is on track to become a trillion-dollar industry, but so far most agentic AI has been concentrated in the hands of a few companies with expensive proprietary models. Holo4 changes that calculation. An open-weight model that matches frontier performance at 1-2% of the cost means that startups, academic researchers, and enterprises in cost-sensitive markets can now experiment with computer-use agents without committing to expensive API contracts.
The question now is whether the proprietary labs respond by lowering prices or by building moats elsewhere -- in tooling, in integrations, or in the quality of the agentic harness rather than the model itself. H Company's bet is that the model is no longer the bottleneck. The benchmark numbers suggest they may be right.