H Company’s new models can move between screens, code and business tools. The practical limits emerge in full-task success, licensing and the software that executes their actions.
Holo4 can request clicks, write code and call tools through one model, allowing an agent to work across software with different interfaces. H Company released it on September 28 in 27B dense and 35B-A3B variants. Its published results support supervised automation experiments; they do not establish dependable, unattended completion of arbitrary office work. H Company’s release
The useful question is where this flexibility saves work. A business process might begin in an old desktop application, continue in a spreadsheet and end in a service with an API. Holo4 gives builders a model that can attempt those transitions. Whether the finished record is correct still depends on the surrounding software and the acceptance checks.
A model needs something to execute its actions
Downloading Holo4 does not give the weights direct control of a computer. The model receives observations, such as screenshots and tool results, and proposes actions. An execution layer performs them and returns the resulting state. The 27B model card describes that loop explicitly.
This distinction matters when evaluating a deployment. An agent that can identify a button still needs a driver that clicks the intended window, handles the application’s response and reports what changed. EyesTech’s computer-use explainer walks through these observation, decision, execution and verification layers.
Consider an invoice-reconciliation pilot. The agent could retrieve an invoice from a browser portal, use code to compare its fields with an exported ledger, then prepare a correction through an application tool. That is a proposed use case, not a Holo4 result measured by EyesTech.
The decisive check would be whether the invoice number, currency, amount and destination account survive every transition. A successful API response cannot prove that the agent chose the right invoice. A populated form cannot prove that the application saved it.
What the demonstrations establish
H Company presents FreeCAD modeling and a Godot game-building example. In its release, the Eiffel Tower run uses 84 calls and 1.3 million tokens; the Pac-Man-style game uses 68 calls and 2.4 million tokens. Those are vendor demonstrations of sustained software interaction, rather than independent measurements of everyday reliability. Published demonstrations
They provide a useful starting point for inspection: did the agent build a valid artifact, use code effectively and check its output? They leave other questions open, including how often the same task fails across fresh runs and whether the result satisfies every requirement.
For a CAD task, a rendered object is only one acceptance condition. The saved file also needs to reopen, preserve dimensions and contain the required geometry. For a game, launching a window establishes less than an extended run that exercises scoring, collisions and resets. These are proposed verification criteria; this article does not claim to have tested the released artifacts.
The 61.7% score is not a completion rate
Holo4’s benchmark figures need their labels intact. The older OSWorld result and the OSWorld 2.0 result measure different evaluations. Within OSWorld 2.0, average partial score and full-task success are also separate measures.
| Vendor-reported measure | ;text-align: right;”>Holo4 27B ;text-align: right;”>Holo4 35B-A3B
|---|
| OSWorld score | ;text-align: right;”>85.2% ;text-align: right;”>80.8%
| OSWorld 2.0 average partial score | ;text-align: right;”>61.7% ;text-align: right;”>30.9%
| OSWorld 2.0 full-task success | ;text-align: right;”>41.5% ;text-align: right;”>12.3%
| OSWorld 2.0 estimated inference cost per attempted task | ;text-align: right;”>$1.22 ;text-align: right;”>$0.61
Source: H Company’s evaluation table. OSWorld 2.0 results are from a single run in H’s harness. Cross-provider rows use differing harnesses, effort settings and task subsets.

Partial credit is useful for research because it reveals progress within a difficult task. A production workflow may have a stricter requirement: the correct record must be saved, with no duplicate and no unrelated changes. A partly completed workflow can still require full human recovery.

H also discloses that 480 of its 600 public AutomationBench tasks belong to a split used to collect training data. It reports 120 held-out tasks separately; private-set evaluation is pending. The main public-set score is not wholly unseen-task evidence. Evaluation notes

For reproduction, pin the benchmark release as well as the model. The OSWorld-V2 repository already recommends a newer v2.1 release, illustrating why a benchmark name alone is insufficient to identify a test.
“Open” means different things for the two models
The deployment choice has a licensing constraint. Holo4-27B’s weights are marked CC BY-NC 4.0, while Holo4-35B-A3B’s weights are marked Apache 2.0. The family therefore cannot be treated as one uniformly permissive release. 27B model card, 35B-A3B model card
For a commercial self-hosting project, the Apache-licensed model is the straightforward candidate to evaluate under its terms. The non-commercial checkpoint requires a separate assessment of permitted use; hosted access is a different arrangement with its own service terms.
This creates a practical tension: the more permissive checkpoint reports substantially lower full-task success on OSWorld 2.0. A team needs to test the model it can actually deploy, rather than transferring the stronger model’s results to its chosen configuration.
The “A3B” designation also describes active parameters, not the entire model’s storage requirement. H identifies the MoE model as 35B total with approximately 3B active. Using fewer parameters for a token does not make all the remaining weights disappear from a deployment. Model specifications
As a rough weight-only calculation, 27 billion parameters at two bytes each occupy about 54 GB; 35 billion occupy about 70 GB. At an idealized four bits per parameter, those figures become 13.5 GB and 17.5 GB. Actual allocations include quantization metadata, vision components, context caches and runtime buffers. These calculations are not measured VRAM requirements or speed estimates.
The public traces make failures inspectable
H publishes a trajectory dataset containing 7,366 runs, with reasoning, actions, tool results and screenshots. Its index includes task scores, success labels, duration and step counts. The dataset also documents redactions and omitted tasks.
That makes a more useful review possible than watching a selected demo. Choose failures resembling the intended workflow and inspect where execution first diverged: wrong record selection, lost state, an unsuccessful action, or an omitted final check.
A trace can support a specific conclusion about a particular run. It cannot establish how the model will behave on a company’s application, permissions and data. The next step is a controlled evaluation in that environment.
Test a whole workflow before scaling it
A useful pilot has a small set of task families and explicit final-state checks. For invoice reconciliation, the following would be a reasonable starting design:
| Test case | What the evaluator should check |
|---|---|
| Similar invoice numbers | Correct source invoice and destination record |
| Delayed save response | Confirmed saved state without a duplicate submission |
| Spreadsheet discrepancy | Preserved source values and a clearly identified difference |
| Session interruption | Explicit recovery or a clear request for intervention |
| Successful routine case | Correct output with no unrelated changes |
Use fresh task instances and repeat runs. Record exact model and quantization, runtime version, elapsed time, inference charges, interventions and recovery work. Define a consequential error separately from an ordinary failure to finish.
The economic measure should be cost per accepted outcome. Divide inference, runtime, review and recovery costs by the number of outputs that pass the checks. A cheap attempt that generates a repair task can cost more than a slower attempt that finishes correctly.
The initial deployment should also give the agent only the applications and tools required for that pilot. Put any approval boundary in the executor, where an action is performed. That is an engineering recommendation for the proposed workflow, not a verified Holo4 security property.
Where Holo4 is worth evaluating
Holo4 is a credible candidate for experiments that require movement between a screen, code and structured tools, especially where an existing process lacks a complete integration. Its released weights and traces let builders inspect more of the system than a hosted demonstration alone.
The current evidence supports a narrower conclusion about dependable automation. The stronger model’s reported full-task success remains well below universal completion, the two checkpoints have different licenses, and the execution environment contributes to the result.
Independent repeated tests on unfamiliar workflows would strengthen the case. So would lower intervention rates and verified final outputs under a deployment’s actual constraints. Until then, Holo4 belongs in a measured pilot with a clear definition of finished work.
