A real build, replayed below

Rehearse your AI agent on a copy of your real app.

Rehearsal compiles your application into a sandboxed practice world with fictional data, runs your agent through real jobs, and scores every outcome against the application's own database.

rehearsal / applications / gitea / build
  1. 1. Read the source
    waiting
  2. 2. Plan the runtime
    waiting
  3. 3. Start the sandbox
    waiting
  4. 4. Find the API
    waiting
  5. 5. Seed fictional data
    waiting
  6. 6. Choose agent tools
    waiting
  7. 7. Write scenarios
    waiting
  8. 8. Validate and repair
    waiting
  9. 9. Publish
    waiting
Loading the recording

Recorded, not simulated: the World Compiler turning Gitea into a practice world.

Powered by

Everything a rehearsal needs, built from your app

Mocks agree with whatever the agent does, and production is too risky to practise in. Rehearsal builds the middle ground automatically.

world build · giteapublished
  1. Read the source
  2. Plan the runtime
  3. Start the sandbox
  4. Find the API
  5. Seed fictional data
  6. Choose agent tools
  7. Write jobs
  8. Validate
  9. Publish
› published · world v3

The World Compiler

An agent reads your source, starts your real images in an isolated sandbox, seeds fictional data, chooses the tools each role may use, and writes jobs it proves can be passed and failed.

Watch a build
Refund order 4182, please.
Looking up order 4182…
request changed · v2
Actually, make it store credit instead.
$40 store credit added, no refund.

Customers who change their minds

Simulated customers revise requests mid-task. Agents that finish the old request get caught.

episode · tool calls
POST/refundstimeout · injected
GET/refunds?order=41821 found
already refunded, no retry

Failures on purpose

Requests time out after the change went through. Blind retries show up as duplicates.

check · read-only · ep_0193
SELECT count(*) FROM refunds
WHERE order_id = 4182;
→ count = 1, read after the episode
  • Goal reachedpass
  • No duplicate refundpass
  • Claim matches recordspass

Verified in your database, not by an LLM

Read-only checks against the application's own records decide every outcome, so a confident agent can't talk its way to a pass.

agent instructionsv1 → v2
Retry any refund call that fails.
Before retrying, check whether the refund already exists.
Re-read the latest customer message before acting.
Proposed from failed training jobs
Held-out jobs
91%
verified success, jobs neither version has seen

Improve the agent, then compare

The improvement agent proposes changes from failed training jobs. Compare versions on jobs neither version has seen.

trace · rehearse-agent4.8 s
episode · changed-mind
llm · deepseek-v4.1-flash
tool · create_refund
tool · list_refunds
verifier · 3 passed

Every step traced

Each build and rehearsal is a Neatlogs trace: model calls, tool calls, and checks.

Coming soonTrain your own model on your practice worlds.Every episode already ends with a score verified in your database: the reward reinforcement learning needs.

A readiness verdict, with evidence

For every agent on every application: ready to pilot, or not yet. Each measure links to the episodes behind it. Switch agents, or pick a measure to see its evidence.

readiness · support-agent on Gitea · 48 held-out episodes
Verdict
Not yet

Duplicate refunds and false “done” claims stop it short of a pilot.

71%
verified success

A pilot needs 90% verified success, no rule violations and no duplicate side effects.

Evidenceep_0391 · retried POST /refunds after a timeout; two refunds exist.

Pricing

Start free and keep the free plan. Pay when you rehearse more often. The open-source core stays free to self-host with your own model key.

Free

$0 forever

  • 1 application, built into a practice world once
  • 40 episodes a month, run in quiet hours
  • 1 sandbox at a time

Pro

Recommended

$29 / month

  • 3 applications, 2 world builds a month
  • 750 episodes a month, run at once
  • 2 sandboxes at a time
  • Top-ups when you need more

Team

$149 / month

  • 20 applications, 8 world builds a month
  • 3,000 episodes a month
  • 4 sandboxes at a time
  • Top-ups when you need more

Enterprise

Custom from $3,000 a month

  • Runner inside your own cloud account
  • Single sign-on, audit export, retention controls
  • Private held-out evaluation of vendor agents

Yearly plans get two months free. Need more in a month? Add 1,000 episodes for $30, or a world build for $15. Nothing renews by itself, and every job has a spend cap.

FAQ

How Rehearsal treats your data, what it runs, and what it costs. Anything else: ask us.

Find out what breaks before your customers do.

Design-partner pilot: four weeks, free for the first five teams shipping an agent into their own product.

Fictional data only. No production access. Isolated sandboxes.