{"status": "confirmed"} cannot tell you whether the agent handles a double-booking, and it has no memory. The agent books an appointment, the caller changes their mind, the agent reschedules, and the mock confirms both because it never knew about the first. Every stateful bug survives a mocked suite.
Fork the world, keep the agent
Gradient’s answer is to run the production topology against a forked copy of a seeded dataset. Same prompts, same graph, samehttp_req rows, same model bindings, the same deployed handlers. Nothing is stubbed. What changes is the world the tools read and write: each redteam conversation gets its own writable fork of the tenant dataset, and your tool code is pointed at that fork.
Personas come from the data
The dataset is not only the world, it is the cast. A redteam instance picks a row and impersonates it: this patient, with this chart, this insurance, this appointment history. There is no persona file to keep in sync with the fixtures, because the persona is a fixture. That makes coverage aSELECT. Loop the patients table and each instance becomes a different caller; run 10 to 20 in parallel, each on its own fork, each impersonating someone real enough to have a history. Disposition layers on top of identity: the same row can be run cooperative, confused, or evasive.
State survives the whole conversation
Because each conversation has a private, writable database, tool calls persist.book_appointment inserts a row, reschedule_appointment sees it, cancel_appointment sees the reschedule. The agent is talking to a world that remembers what it just did, the only way to test an ordinary request like “book a visit, change your mind, move it to Thursday, then cancel.”
The rules that grade this run read both sides: the transcript (did it ask for new insurance?) and the fork’s final state (is the row new_patient_office rather than follow_up?). See rubrics.
One handler, no mock branch
Your custom tool resolves which database it is talking to from authenticated redteam context, so the same code serves production traffic and redteam traffic:db helper uses it automatically. This is the payoff of forking over stubbing: there is no second implementation to drift, because the code under test is the code that ships.
Tools that leave the dataset keep a mock
A fork can contain your dataset, not the outside world. Tools that reach past it (charging a card, sending an SMS) still need an explicitmock handler.
What a run produces
Each conversation produces a trace (every turn, every swap, every tool call with its arguments and response) plus the final state of its fork. Both feed evaluation.Datasets
The seeded world a redteam forks, and the rows a persona comes from.
Custom tools
Write one handler that serves production and redteam alike.