The Rundown AI homepage

Independent tool overview

Petri at a glance

Petri is an open-source alignment-auditing framework that has an auditor agent create multi-turn test scenarios for a target model, then uses a judge model to score the transcripts for concerning behavior.

Visit the official Petri site ↗
Petri product preview
Current steward
Meridian Labs
Original developer
Anthropic Fellows researchers
License
MIT
Framework
Inspect AI and Python

Overview

What Petri is

Petri, short for Parallel Exploration Tool for Risky Interactions, helps AI safety researchers test concrete hypotheses about model behavior at scale. A researcher supplies seed instructions; an auditor model builds and drives realistic scenarios with simulated users and tools; the target model responds; and a judge scores the transcript across behavioral dimensions with evidence-linked explanations.

Anthropic created Petri and transferred stewardship to the nonprofit Meridian Labs in May 2026. The current Inspect Petri 3 architecture separates auditor and target components, supports custom agents and rollback branches, and integrates with Petri Dish for testing real agent scaffolds and Petri Bloom for deeper evaluation of a specific behavior.

Use cases

Who Petri is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Exploratory alignment audits

Probe models across many multi-turn scenarios to surface deception, sycophancy, self-preservation, reward hacking, and other behaviors for review.

Model comparisons

Run consistent seed sets and judging dimensions against several targets, then compare transcripts and dimension deltas.

Custom agent evaluation

Adapt the target interface or use Dish to test real scaffolds such as Claude Code, Codex CLI, and Gemini CLI.

Capabilities

Core Petri features

1

Three-model audit loop

Separates the auditor that designs the test, the target being evaluated, and the judge that scores the completed transcript.

2

Built-in seed library

Ships with more than 170 scenario seeds that can be filtered by tags or replaced with custom Markdown, JSON, CSV, or inline instructions.

3

Behavioral dimensions

Includes dozens of scored dimensions with 1–10 rubrics, written justifications, and references to supporting messages.

4

Rollback and branching

Lets the auditor return to an earlier trajectory point and try a different approach while preserving branches for evaluation.

5

Inspect result viewer

Stores logs and provides transcript views, scores, annotations, auditor reasoning, tools, and branch navigation.

6

Dish and Bloom extensions

Dish evaluates real deployment scaffolds, while Bloom generates targeted suites for repeatedly measuring one chosen behavior.

Process

How the Petri workflow works

  1. Step 1

    Define the question

    Choose a behavior or scenario worth testing and translate it into a specific seed instruction and scoring dimensions.

  2. Step 2

    Assign model roles

    Select auditor, target, and judge models, then set turn limits, tools, realism filters, and other audit controls.

  3. Step 3

    Run a small pilot

    Start with a handful of seeds to check scenario validity, provider policies, token use, judge behavior, and log quality.

  4. Step 4

    Scale and inspect

    Run the validated suite, sort scores, read high-signal transcripts and branches, and compare results across targets or versions.

  5. Step 5

    Confirm findings

    Reproduce important cases with additional seeds, human review, alternative judges, and complementary evaluation methods before drawing conclusions.

Cost

Petri pricing and free plan

Petri is free, MIT-licensed software. Running it is not free: each audit can invoke separate auditor, target, judge, and optional realism models across many turns and seeds, so users pay the connected model providers and any compute or storage costs.

Inspect Petri

Free and open source

Install the Python package and run it on infrastructure you control.

  • MIT license
  • Install from PyPI
  • Model credentials required

Model usage

Provider charges vary

Pay for the auditor, target, judge, and optional utility-model calls used by each run.

  • Full default audit has 170+ seeds
  • Up to 30 turns per seed by default
  • Frontier-model judging can create substantial API spend
  • Start with a sample limit

Pricing checked . Check current pricing at the source ↗

Assessment

Petri strengths and limitations

Where it stands out

  • Turns a natural-language safety hypothesis into parallel, multi-turn behavioral tests
  • Preserves transcripts, scores, justifications, tool activity, and alternate trajectory branches for inspection
  • Separates the auditor and target so researchers can customize either side independently
  • Supports broad exploration with Petri and focused repeated measurement with the Bloom extension

What to consider

  • Automated scenarios can be unrealistic, and a target may change its behavior when it recognizes that it is being evaluated
  • LLM judges are imperfect and reductive; a numerical score should guide human review rather than become a standalone safety verdict
  • Large audits can consume significant time and API budget because several models run across many turns and seeds
  • Audits intentionally explore concerning content, so researchers must follow provider policies and handle logs and results responsibly

Compare

Petri alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Coding

Claude Security

Choose Claude Security when the target is source-code vulnerability discovery rather than language-model behavioral alignment.

Explore Claude Security

Agents

Codex Security

Choose Codex Security for agentic repository scanning and patch validation rather than multi-turn model-behavior research.

Explore Codex Security

Questions

Petri FAQs

What does Petri test?

Petri tests model behavior in generated multi-turn situations, including hypotheses about deception, sycophancy, harmful cooperation, self-preservation, power seeking, reward hacking, evaluation awareness, and custom behaviors.

Is Petri still an Anthropic project?

Anthropic created Petri and continues to support and use it, but current development is stewarded by the independent nonprofit Meridian Labs.

Is Petri free?

The MIT-licensed software is free. Model API calls and any compute or storage used to run audits are separate costs.

How does a Petri audit work?

An auditor model turns a seed into a scenario and interacts with the target over multiple turns. A judge then scores the transcript against selected dimensions and cites relevant messages.

Can Petri prove that a model is safe?

No. It can surface and measure behavior under tested conditions, but results depend on seeds, scaffolds, models, judges, and realism. Important conclusions require reproduction, human review, and complementary methods.

Bottom line

Our Petri verdict

Petri is a strong open tool for researchers who need to explore model behavior beyond static prompt-response benchmarks. Its value lies in generating inspectable hypotheses and transcripts quickly, not in producing a definitive safety grade; careful experiment design, cost controls, multiple judges, and expert human review remain essential.

Visit Petri website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.