GUI automation developers
Build agents for interfaces that do not offer a reliable API, including browser, desktop, mobile, and legacy enterprise software.
Independent tool overview
Holo3 is H Company's vision-language model family for agents that interpret screenshots, locate interface elements, and operate web, desktop, and mobile software.
Visit the official Holo 3 site ↗
Overview
Holo3 is a computer-use model family for developers building agents that navigate graphical interfaces. It reads screenshots, reasons about the current state, and returns structured actions or screen coordinates for an external agent harness to execute.
The original Holo3 launch introduced a 35B model with 3B active parameters and a larger 122B model with 10B active parameters. H Company has since released Holo3.1, which adds smaller 0.8B, 4B, and 9B options, broader mobile and cross-harness support, function calling, and quantized checkpoints for local inference. The Holo3 page remains useful as the foundation, but new deployments should evaluate the current Holo3.1 family.
Holo is a model and agent component, not permission to run unattended on sensitive systems. Teams still need a secure execution environment, explicit tool definitions, approval gates, credential controls, logging, rollback plans, and human review for irreversible actions.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Build agents for interfaces that do not offer a reliable API, including browser, desktop, mobile, and legacy enterprise software.
Exercise user journeys visually and return evidence from a real interface rather than stopping at code-level tests.
Navigate applications and collect information from screens when conventional scraping or direct integrations are unavailable.
Self-host open Holo3.1 checkpoints when screenshots and application content must stay on controlled hardware.
Use Holo for UI grounding inside an existing orchestration layer while retaining ownership of permissions, tools, policy, and monitoring.
Capabilities
Consumes conversation context and screenshots, then returns a structured note, thought, and tool call for an external browser or desktop harness.
Can run in a single-turn grounding mode that maps a text description of an interface element to normalized screen coordinates.
The current Holo3.1 family is trained and evaluated across multiple interface environments rather than browser navigation alone.
The 35B mixture-of-experts model activates 3B parameters and is published under Apache 2.0 for self-hosting and commercial development.
The larger API model activates 10B of 122B total parameters for more demanding multi-step navigation, but its published license is research-only and non-commercial.
Newer 0.8B, 4B, 9B, and 35B-A3B checkpoints let teams trade accuracy, latency, memory, and local hardware requirements.
Holo3.1 includes FP8, Q4 GGUF, and NVFP4 options, with specific formats and sizes varying by checkpoint.
Hosted Holo models can be called through a familiar chat-completions pattern with text and image inputs.
H Company's open client exposes CLI, MCP, ACP, and A2A surfaces so another agent can delegate visual desktop tasks.
Process
Step 1
Use the Holo Models API or self-hosted checkpoint when you own the agent loop; use H Company's managed Computer-use Agents when you do not want to build the runtime and environment layer.
Step 2
Run the agent in a dedicated browser profile, virtual machine, test account, or sandbox with the least privilege needed for the task.
Step 3
Expose only approved actions such as click, type, scroll, wait, and screenshot, and validate every model-generated argument before execution.
Step 4
Require a person to approve purchases, messages, submissions, account changes, file deletion, credential use, and any action that is difficult to reverse.
Step 5
Measure task success, wrong-action rate, recovery behavior, latency, token cost, and performance after interface changes on your own workflows.
Step 6
Keep screenshots, actions, approvals, errors, and final outcomes so failures can be audited and the agent can be stopped quickly.
Cost
H Company offers rate-limited access to Holo3.1-35B-A3B at no charge, paid API usage for higher limits, and downloadable Apache 2.0 weights for self-hosting. The hosted 35B model costs $0.25 per million input tokens and $1.80 per million output tokens; Holo3-122B-A10B costs $0.40 and $3.00 respectively. Self-hosting avoids per-token model fees but adds substantial hardware and operations costs.
Free, rate limited
No-card hosted access for evaluation.
$0.25 input / $1.80 output per 1M tokens
Pay-as-you-go hosted inference with higher rate limits.
$0.40 input / $3.00 output per 1M tokens
Flagship hosted model for complex navigation.
No model usage fee
Run an open checkpoint on your own compatible hardware.
Check current platform pricing
H Company supplies the agent runtime and managed execution environment.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Agents
Google's specialized computer-use model is an alternative for developers already building on Gemini and Google Cloud tooling.
Explore Gemini Computer Use →Agents
A more packaged browser experience for users who want Claude to work directly in Chrome rather than build a model-level agent loop.
Explore Claude in Chrome →Agents
A broader managed general-purpose agent for users who prioritize completed tasks over self-hosting a computer-use model.
Explore Manus →Questions
Holo3 is H Company's family of vision-language models trained for UI grounding and computer-use agents across web, desktop, and mobile interfaces.
The Holo3 family remains available, but Holo3.1 is the current generation for new evaluations. It adds smaller models, local quantizations, mobile improvements, and broader agent-harness support.
The 35B-A3B model is openly available under Apache 2.0. The flagship 122B-A10B model is not equivalent: H Company lists it as research-only and non-commercial.
Holo3.1-35B-A3B is $0.25 per million input tokens and $1.80 per million output tokens on the paid tier, with a free 10-RPM tier. Holo3-122B-A10B is $0.40 input and $3.00 output per million tokens.
Yes. Holo3.1 publishes open 35B-A3B checkpoints including Q4 GGUF, FP8, and NVFP4 options. Smaller 0.8B, 4B, and 9B family members are also available, with support varying by checkpoint.
No. The model interprets screenshots and proposes actions or coordinates. A surrounding agent harness must capture the screen, execute tools, manage permissions, and decide when a human must approve.
Not by default. Use isolation, least privilege, action validation, spending and messaging approvals, detailed logs, and a reliable stop mechanism before giving any computer-use model autonomy.
Bottom line
Holo3 is attractive for teams that need a specialized, economical computer-use model and want control over the agent harness or deployment location. Holo3.1 is the better starting point today, but the real production decision depends less on a benchmark headline than on end-to-end task success, hardware, action safety, and the licensing difference between the open 35B and restricted 122B models.
Visit Holo 3 website ↗
MolmoWeb - Ai2's open-source web browsing agent that navigates sites using screenshots

Perplexity Personal Computer - Perplexity's local orchestrator working across your files, native apps, and the web

Dynamic Workers - Cloudflare's isolate-based sandbox for running AI agent code at scale

Workspace Agents - OpenAI's new Codex-powered team agents for shared multi-step workflows in ChatGPT and Slack

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.