The Rundown AI homepage

Independent tool overview

Matrices at a glance

Matrices is now an infrastructure company building realistic simulated computer environments for training and evaluating multimodal LLM agents. The original self-filling research spreadsheet described in older tool listings is no longer the product offered at Matrices.app.

Visit the official Matrices site ↗
Matrices product preview
Current product
LLM agent training environments
Earlier product
AI research spreadsheet
Primary users
AI labs and agent teams
Agent type
Multimodal computer-use agents
Access
Contact the company
Public price
Not published

Overview

What Matrices is

Matrices builds what it describes as training environments for multimodal, computer-using agents. Instead of allowing a model to practice on live accounts and create real side effects, the company recreates websites, workflows, and surrounding context so an agent can attempt a task repeatedly and receive feedback inside a controlled simulation.

This is a substantial pivot from the earlier Matrices research tool, which was presented as an AI-powered spreadsheet that could gather and organize information. The current offering is aimed at frontier AI labs and agent-development teams, not analysts looking for a no-code spreadsheet or a public self-serve data-analysis app.

Use cases

Who Matrices is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Frontier model labs

Create large sets of interactive computer-use tasks for reinforcement learning, post-training, or evaluation.

Computer-use agent teams

Test agents against realistic websites and multi-step workflows without touching real user or financial accounts.

Agent reliability research

Measure success, failure, recovery, and behavior over repeatable tasks with controlled starting states.

Training-data operations

Produce trajectories and feedback from agents completing realistic interface tasks at scale.

Custom environment programs

Commission simulations for valuable domains that cannot safely or cheaply be exercised thousands of times in production.

Capabilities

Core Matrices features

1

Simulated computer tasks

Recreates the interfaces and state needed for an agent to perform realistic digital work without the real-world consequences.

2

Multimodal interaction

Targets agents that must perceive screens and act through user interfaces rather than answer a single text prompt.

3

Repeatable starting states

Lets teams rerun tasks from controlled conditions for training comparisons and regression evaluation.

4

Feedback-ready tasks

Environments can define whether an attempt succeeded so learning systems can associate actions with outcomes.

5

Longer computer workflows

The company's roadmap focuses on increasingly complex tasks that span multiple steps, sites, and decisions.

6

Distributed agent runner

Matrices' engineering materials describe infrastructure for running many agent attempts with higher throughput and lower latency.

7

Task factory

Internal tooling is designed to help analyze, create, and operate a growing library of agent challenges.

8

Internet cloning

The team describes its broader technical goal as recreating the parts of the internet and workflows agents need to practice on.

9

Future level editor

Matrices has said it plans to open a level editor so external contributors can create simulated tasks, but public availability and terms are not yet published.

Process

How the Matrices workflow works

  1. Step 1

    Choose the target work

    Identify a high-value computer task, its real systems, success criteria, and the side effects that make live practice unsafe.

  2. Step 2

    Specify the environment

    Define the interfaces, data, identities, dependencies, and hidden state the agent could plausibly encounter.

  3. Step 3

    Build the simulation

    Recreate the necessary websites and workflow behavior with enough depth that the agent can explore beyond one scripted path.

  4. Step 4

    Create tasks and graders

    Establish initial states, user instructions, acceptable outcomes, failure conditions, and machine-verifiable feedback where possible.

  5. Step 5

    Run agent trajectories

    Execute repeated attempts across models or checkpoints and collect actions, observations, errors, and results.

  6. Step 6

    Train and regression-test

    Use the trajectories and rewards for post-training, then keep stable evaluation sets to detect overfitting and new failures.

Cost

Matrices pricing and free plan

Matrices does not publish self-serve plans or unit pricing. The current site routes prospective customers to direct contact, indicating a custom commercial engagement rather than a subscription an individual developer can buy online.

Training environment engagement

Custom

Commercial access for AI labs and agent teams is arranged directly with Matrices.

  • Contact required
  • Scope depends on target domains and task volume
  • No public rate card

Custom task and simulation work

Custom

Environment creation is likely scoped around the websites, task complexity, graders, and scale the customer needs.

  • Bespoke environments
  • Training and evaluation use cases
  • Terms not publicly listed

Public level editor

Not yet announced

Matrices has discussed opening environment creation to outside contributors, but access, compensation, and pricing are not public.

  • Roadmap item
  • No launch date published
  • Do not assume current availability

Pricing checked . Check current pricing at the source ↗

Assessment

Matrices strengths and limitations

Where it stands out

  • Addresses a real bottleneck for agents that need interactive experience rather than static examples.
  • Simulations reduce the risk of repeated training attempts changing real accounts or sending real transactions.
  • Controlled starting states make model and checkpoint comparisons more meaningful.
  • Realistic interface tasks expose planning, perception, action, and recovery failures that text benchmarks miss.
  • Custom environments can target proprietary or economically important workflows.
  • A distributed runner can generate substantially more trajectories than manual evaluation.
  • Task-level feedback can support both reinforcement learning and repeatable evaluation.
  • The company is focused specifically on computer-use environment engineering rather than a general agent wrapper.

What to consider

  • The old spreadsheet and research description is obsolete; Matrices is no longer a general data-analysis tool for end users.
  • There is no public self-serve demo, API documentation, plan comparison, or price list for the current offering.
  • A simulated website may omit edge cases, organizational context, latency, fraud, or human behavior that appears in production.
  • Agents can overfit to environment artifacts or graders and appear capable without generalizing to the real system.
  • A verifiable task outcome does not necessarily capture safety, policy compliance, efficiency, or user satisfaction.
  • Building high-fidelity environments is expensive and requires ongoing updates as source interfaces and workflows change.
  • Training tasks need careful licensing, privacy, security, and intellectual-property review.
  • Synthetic identities and data still require safeguards to prevent accidental use of real sensitive information.
  • Computer-use agents can discover unintended paths through a simulation, so the environment needs adversarial testing.
  • Commercial claims and internal task counts are company-reported and not a substitute for customer-specific validation.
  • The platform is aimed at well-resourced model teams, not individual researchers seeking an open benchmark package.
  • A training environment provider is only one layer; customers still need models, inference, orchestration, reward design, and evaluation analysis.
  • The announced public level editor is a future plan, not a feature buyers should assume is available now.
  • Custom engagements can create vendor dependency if task specifications, graders, and environments are not portable.

Compare

Matrices alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Agents

Gemini Computer Use

Use Gemini Computer Use when you need a model designed to operate interfaces rather than infrastructure for training such a model.

Explore Gemini Computer Use

Agents

Manus Cloud Computer

Use Manus Cloud Computer when the goal is deploying an always-on agent in a hosted environment instead of building RL simulations.

Explore Manus Cloud Computer

Agents

Perplexity Computer

Use Perplexity Computer for a finished multi-model agent experience rather than custom agent-training environments.

Explore Perplexity Computer

Agents

My Computer

Use Manus My Computer when you want an agent to work on a local desktop with user oversight.

Explore My Computer

Questions

Matrices FAQs

What is Matrices now?

Matrices now builds simulated training environments for multimodal LLM agents that perform realistic computer-use tasks. It sells to AI labs and agent-development teams rather than general spreadsheet users.

Is Matrices still an AI spreadsheet?

No. Older listings described a self-filling research spreadsheet, but the current Matrices.app site and company materials are about training environments for computer-use agents.

Why train agents in a simulation?

An agent may need thousands of attempts to improve. A simulation lets it practice workflows involving accounts, websites, files, or transactions without repeatedly changing a real system or affecting real people.

Can I sign up for Matrices online?

The current public site does not offer self-serve registration for the training product. Prospective customers are directed to contact the team.

How much does Matrices cost?

Pricing is not publicly listed. Expect a custom enterprise or research engagement based on environment complexity, task volume, integrations, and operational scale.

Does Matrices provide the agent model?

Matrices describes itself as the environment layer for training computer-use agents. Customers should confirm model hosting, inference, runners, rewards, and training responsibilities in the engagement scope.

Are simulated task results reliable?

They can be useful for controlled comparison, but only if the simulation and grader reflect the real workflow. Teams should hold out tasks, test for environment shortcuts, and validate improvements in a safely limited real setting.

Is the Matrices level editor available?

Matrices has said it plans to open a level editor to outside contributors. The company has not publicly documented general availability, pricing, or contributor terms, so it should be treated as a roadmap item.

Bottom line

Our Matrices verdict

Matrices is relevant to a narrow but important buyer: AI labs that need realistic, repeatable computer-use experience for agent training. Its pivot makes the old spreadsheet framing misleading. The current product looks like custom infrastructure, so buyers should evaluate environment fidelity, grader quality, portability, security, and evidence of real-world transfer before committing.

Visit Matrices website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.