# Show HN: Understudy: Scenario Testing for AI Agents

> Source: <https://github.com/gojiplus/understudy>
> Published: 2026-08-28 03:51:56+00:00

Understudy is a scenario-driven testing framework for AI agents that simulates realistic multi-turn users, runs those scenes against an agent through a simple app adapter, records a structured execution trace of messages, tool calls, and handoffs, and then evaluates behavior with deterministic checks, optional LLM judges, and run reports.

Testing with understudy is **4 steps**:

**Wrap your agent**— Adapt your agent (ADK, LangGraph, HTTP) to understudy's interface** Mock your tools**— Register handlers that return test data instead of calling real services** Write scenes**— YAML files defining what the simulated user wants and what you expect** Run and assert**— Execute simulations, check traces, generate reports

The key insight: **assert against the trace, not the prose**. Don't check what the agent said—check what it did (tool calls).

Simulate multi-turn conversations with personas to test dialogue agents.

- Use case: Customer service bots, assistants, chatbots
- Assert on tool calls:
`trace.called("tool_name")`

Evaluate autonomous agents executing multi-step tasks.

- Use case: Code agents, research agents, task automation
- Assert on actions:
`trace.performed("action")`

See [examples/README.md](/gojiplus/understudy/blob/main/examples/README.md) for complete examples of both paradigms.

**See real examples:**

[Example scene](https://github.com/gojiplus/understudy/blob/main/examples/scenes/return_eligible_backpack.yaml)— YAML defining a test scenario[ADK test file](https://github.com/gojiplus/understudy/blob/main/examples/adk/test_adk_returns.py)— pytest assertions against traces[LangGraph test file](https://github.com/gojiplus/understudy/blob/main/examples/langgraph/test_langgraph_returns.py)— same tests, different framework[Agentic test file](https://github.com/gojiplus/understudy/blob/main/examples/agentic/test_agentic.py)— agentic flow evaluation[Agentic scene](https://github.com/gojiplus/understudy/blob/main/examples/agentic_scenes/code_review_task.yaml)— task-based scenario[Example report](https://htmlpreview.github.io/?https://github.com/gojiplus/understudy/blob/main/examples/langgraph/report/index.html)— HTML report with metrics and transcripts

```
pip install understudy[all]
python
from understudy.adk import ADKApp
from my_agent import agent

app = ADKApp(agent=agent)
```

Your agent has tools that call external services. Mock them for testing:

``` python
from understudy.mocks import MockToolkit

mocks = MockToolkit()

@mocks.handle("lookup_order")
def lookup_order(order_id: str) -> dict:
    return {"order_id": order_id, "items": [...], "status": "delivered"}

@mocks.handle("create_return")
def create_return(order_id: str, item_sku: str, reason: str) -> dict:
    return {"return_id": "RET-001", "status": "created"}
```

Create `scenes/return_backpack.yaml`

:

```
id: return_eligible_backpack
description: Customer wants to return a backpack

starting_prompt: "I'd like to return an item please."
conversation_plan: |
  Goal: Return the hiking backpack from order ORD-10031.
  - Provide order ID when asked
  - Return reason: too small

persona: cooperative
max_turns: 15

expectations:
  required_tools:
    - lookup_order
    - create_return
  forbidden_tools:
    - issue_refund
python
from understudy import Scene, run

scene = Scene.from_file("scenes/return_backpack.yaml")
trace = run(app, scene, mocks=mocks)

assert trace.called("lookup_order")
assert trace.called("create_return")
assert not trace.called("issue_refund")
```

Or with pytest (define `app`

and `mocks`

fixtures in conftest.py):

```
pytest test_returns.py -v
```

Run multiple scenes with multiple simulations per scene:

``` python
from understudy import Suite, RunStorage

suite = Suite.from_directory("scenes/")
storage = RunStorage()

# Run each scene 3 times and tag for comparison
results = suite.run(
    app,
    mocks=mocks,
    storage=storage,
    n_sims=3,
    tags={"version": "v1"},
)
print(f"{results.pass_count}/{len(results.results)} passed")
```

Understudy separates simulation (generating traces) from evaluation (checking traces). Use together or separately:

```
understudy run \
  --app mymodule:agent_app \
  --scene ./scenes/ \
  --n-sims 3 \
  --junit results.xml
```

Generate traces only:

```
understudy simulate \
  --app mymodule:agent_app \
  --scenes ./scenes/ \
  --output ./traces/ \
  --n-sims 3
```

Evaluate existing traces:

```
understudy evaluate \
  --traces ./traces/ \
  --output ./results/ \
  --junit results.xml
```

Python API:

``` python
from understudy import simulate_batch, evaluate_batch

# Generate traces
traces = simulate_batch(
    app=agent_app,
    scenes="./scenes/",
    n_sims=3,
    output="./traces/",
)

# Evaluate later
results = evaluate_batch(
    traces="./traces/",
    output="./results/",
)
# Run simulations
understudy run --app mymodule:app --scene ./scenes/
understudy simulate --app mymodule:app --scenes ./scenes/
understudy evaluate --traces ./traces/

# View results
understudy list
understudy show <run_id>
understudy summary

# Compare runs by tag
understudy compare --tag version --before v1 --after v2

# Generate reports
understudy report -o report.html
understudy compare --tag version --before v1 --after v2 --html comparison.html

# Interactive browser
understudy serve --port 8080

# HTTP simulator server (for browser/UI testing)
understudy serve-api --port 8000

# Cleanup
understudy delete <run_id>
understudy clear
```

For qualities that can't be checked deterministically:

``` python
from understudy.judges import Judge

empathy_judge = Judge(
    rubric="The agent acknowledged frustration and was empathetic while enforcing policy.",
    samples=5,
)

result = empathy_judge.evaluate(trace)
assert result.score == 1
```

Built-in rubrics:

```
from understudy.judges import (
    TOOL_USAGE_CORRECTNESS,
    POLICY_COMPLIANCE,
    TONE_EMPATHY,
    ADVERSARIAL_ROBUSTNESS,
    TASK_COMPLETION,
)
```

The `understudy summary`

command shows:

**Pass rate**— percentage of scenes that passed all expectations** Avg turns**— average conversation length** Tool usage**— distribution of tool calls across runs** Agents**— which agents were invoked

The HTML report (`understudy report`

) includes:

- All metrics above
- Full conversation transcripts
- Tool call details with arguments
- Expectation check results
- Judge evaluation results (when used)

See the [full documentation](https://gojiplus.github.io/understudy) for:

[Installation guide](https://gojiplus.github.io/understudy/installation.html)[Writing scenes](https://gojiplus.github.io/understudy/tutorial/scenes.html)[ADK integration](https://gojiplus.github.io/understudy/adk-integration.html)[LangGraph integration](https://gojiplus.github.io/understudy/langgraph-integration.html)[HTTP client for deployed agents](https://gojiplus.github.io/understudy/tutorial/http.html)[API reference](https://gojiplus.github.io/understudy/api/index.html)

MIT
