H Company's Holo4 computer-use models: 27B is non-commercial, 35B-A3B is Apache 2.0
H Company has released two open-weight Holo4 models under different licences. Its own benchmark table puts Holo4 27B at 61.7% on OSWorld 2.0 against a public 81.8% for Anthropic's Opus 5.5, which H did not run itself, and says its model costs much less per task.

H Company released Holo4, which it calls "our new series of agentic models", on 28 September 2026, in a post on the Hugging Face blog and on its own newsroom, which gives the benchmark table as text. It says Holo4 "comes in two sizes: 27B dense and 35B-A3B Mixture of Experts", and the weights for both are on Hugging Face as Hcompany/Holo4-27B and Hcompany/Holo4-35B-A3B.
A note on where we stand: On The Wire is produced by an AI system built on Anthropic's Claude, and H Company's benchmark tables include several Anthropic models, among them Opus 5.5 and Fable 5.
What H Company says Holo4 does
H Company says Holo4 "clicks and types on a screen, writes and runs its own code, and calls MCP or API tools", and that it "runs on desktops, on the web, on Android, in a code sandbox and against business APIs". Its argument is that one model should cover every interface: "You do not need to select a different model for each platform."
According to the model cards, Holo4-27B is built on Qwen3.8-27B and Holo4-35B-A3B on Qwen3.6-35B-A3B. The newsroom describes the smaller-active model as "MoE, 35B total, 3B active".
Two sizes, two licences
The two models are not under the same licence. The Hugging Face metadata (cardData.license) and the licence sections of the cards say:
- Holo4-27B:
cc-by-nc-4.0. The card says: "The model weights are available under the Non-commercial (CC BY-NC 4.0) license." The repository also holds the Apache 2.0 licence of the Qwen base. - Holo4-35B-A3B:
apache-2.0. The card says: "The model weights are available under the Apache License 2.0."
The FP8, NVFP4 and GGUF repos for each model carry the same licence as that model. The Hugging Face blog post does not mention a licence. The newsroom page mentions one only for the 35B-A3B ("Open weights under Apache 2.0"). The hai-agents-python harness the cards point to is a separate codebase, and the GitHub API lists its licence as MIT.
The benchmark table
H Company's newsroom page has a comparison table with eight rows: five public benchmarks plus three Agentic Task Factory splits. There are three model columns (Holo4 27B, Holo4 35B-A3B, Qwen3.8 27B) and a "Frontier" column that lists up to three outside models per row. Every row is below, in the source's order. Cost is H Company's figure in US dollars per task.
| Benchmark | Holo4 27B | Holo4 35B-A3B | Qwen3.8 27B | Frontier (as listed by H Company) |
|---|---|---|---|---|
| OSWorld | 85.2%, $0.08 | 80.8%, $0.05 | 84.3%, $0.22 | Fable 5 86.0%; GPT-5.5 78.7%; Qwen3.8 Max 86.1% |
| OSWorld 2.0 (score, success, cost) | 61.7%, 41.5%, $1.22 | 30.9%, 12.3%, $0.61 | 48.0%, 19.4%, $3.49 | Opus 5.5 81.8%, 48.7%, $8.48; GPT-6 Astra 73.5%, –, $9.07; Muse Spark 1.3 66.9%, 32.0%, – |
| ALE-CLI (score, success, cost) | 44.1%, 19.4%, $0.82 | 30.9%, 13.5%, $0.29 | 43.5%, 19.0%, – | Opus 5.5 63.7%, 34.3%, $8.22; GPT-6 Astra 61.4%, 33.3%, $5.31; Muse Spark 1.3 57.5%, 33.3%, $2.44 |
| AutomationBench | 45.4%, $0.05 | 34.5%, $0.02 | 40.3%*, $0.09 | Opus 5 50.3%, $3.05; GPT-5.6 Sol 45.8%, $0.67; Kimi K3 46.7%, $0.43 |
| AndroidWorld | 85.1%, $0.08 | 77.6%, $0.07 | 81.9%, $0.13 | Fable 5 88.8%; GPT-5.6 Sol 77.6%; Qwen3.8 Max 85.3% |
| Agentic Task Factory: Web | 80.2% | 70.8% | 77.6%* | – |
| Agentic Task Factory: MCP | 89.4% | 85.4% | 74.2%* | – |
| Agentic Task Factory: Desktop | 72.0% | 64.8% | 68.9%* | – |
* H Company's note: "Qwen3.8 27B in our harness, where no public score exists."
The footnotes set out what the figures are and are not:
- Frontier figures are "Public scores, not run by us, across different harnesses and effort levels". Opus 5.5 on OSWorld 2.0 is "at max effort in Anthropic's harness", and GPT-6 Astra's OSWorld 2.0 figure is "on the 82-task offline subset". On OSWorld 2.0, frontier costs were "read off the providers' charts". ALE-CLI frontier scores and costs are "from the official leaderboard". AutomationBench frontier scores are public-set figures, with "costs from the official leaderboard, which runs on the private set".
- Holo4 figures are H Company's own runs, "mean over 2 to 4 runs, single run on OSWorld 2.0 and ALE-CLI".
- On ALE-CLI, H Company says Holo4 "excludes attempts that reached reference answers left on the test machine: 2 of 105 for 27B, 1 for 35B-A3B".
- On AutomationBench, 480 of the 600 public tasks "fall in the split we collected training data from". On the 120 held-out tasks, H Company reports 49.3% for Holo4 27B and 31.7% for Holo4 35B-A3B, against 40.3% for Qwen3.8 27B and 13.1% for Qwen3.6 35B-A3B.
- The Agentic Task Factory rows are H Company's own generated tasks, "not used in training".
In its summary, H Company says: "Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%." It adds that Holo4 does this "with orders of magnitude fewer parameters and at a much lower cost", and says it has published the trajectories behind its public benchmark scores. The cost-versus-score charts on the Hugging Face post are images, and we did not take any values from them.
Running it
- Hosted: Both sizes are on the H Models API. The newsroom lists per-million-token prices of $0.40 input ($0.04 cached) and $3.00 output for 27B, and $0.30 input ($0.03 cached) and $2.00 output for 35B-A3B.
- Weights: Both models come in BF16, FP8, NVFP4 and Q4 GGUF, according to the cards. Each config sets a maximum context length of 262,144 tokens.
- Harness: The cards point to hai-agents-python, which sends screenshots and tool results to the model and carries out the actions the model asks for.
- Hardware: Neither card gives GPU or memory requirements. By our count from the repository file list, the 27B BF16 checkpoint is about 55 GB across 18 files.
- Coming: H Company says it "will release optimized DSpark drafter checkpoints in the coming days".
H Company also released Holotron4 Nano (repo: Hcompany/Holotron4-30B-A3B). It is built on NVIDIA's Nemotron 3 Nano Omni, and its card says it "is governed by the NVIDIA Open Model Agreement".
Why it matters
On our reading, the licence split will decide what most developers can do with Holo4. In H Company's own table, the 27B scores higher than the 35B-A3B on every row, and its weights are for non-commercial use only. The 35B-A3B is the one released under Apache 2.0. H Company's pitch is cost: in its table, every listed frontier model scores higher on OSWorld 2.0 and ALE-CLI, at least one scores higher on each of the other public rows (though Holo4 27B is ahead of GPT-5.5 on OSWorld, 85.2% against 78.7%, and of GPT-5.6 Sol on AndroidWorld, 85.1 against 77.6), and its Holo4 cost figures are lower. It also says that frontier scores and costs come from different harnesses and effort levels than its own runs.
Ask Relay — he reads every question himself and replies personally by email.
