If AI Runs the Test, Who Tests the AI? — Auditing Jev Gameplay Runs
How to distinguish gameplay defects, agent decision errors, and test harness failures by recording and reviewing the full Jev gameplay run.
- unity
- ai
- qa
Imagine opening the dashboard after an overnight autoplay run and seeing one failure. The character is below the floor, but the last log says is_alive: true. Is that a physics bug in the game? Did Jev choose the wrong action? Or did the test harness pass stale state or judge the run against the wrong expectation? The visible result may look the same, but each explanation points to a different fix.
As automation grows, another question comes with it: who tests the AI that is testing the game, and what evidence do they use?
TypeSafe's Agent Trace Observability example reviews a customer-support agent run through its instructions, each conversation turn, tool calls with their arguments and results, and the final response. Customer feedback is included when available. The review can lead to an automatic close, a not-a-bug decision, human review, priority review, an issue, or an on-call page. The point is not to transplant customer support into a game. It is to look at the context around a run instead of judging it by its final answer alone.
In a game, inspect the gap between choice and execution
Conversation turns in a game run become the steps where the agent reads state and chooses an action. The trace should keep the run ID and test instruction (including its version), the state shown to Jev, the actions allowed at that point, Jev's choice and confidence, whether Unity accepted or rejected it, and the actual state after execution — all in order. A final failed tells us the outcome, but not where the run went off track.
With that chain, similar-looking failures split into three kinds. If Unity executes an allowed action but the game violates an invariant, the defect may be in gameplay. If Jev repeatedly picks a poor option from the actions it was allowed to take, inspect its decision or the state representation. If the state sent to Jev was stale, the legal-action list disagreed with Unity's execution rules, or the harness checked the result at the wrong time, the test setup may be at fault. Instead of examining one screen, connect what it saw → what it chose → what Unity did → what actually happened to narrow the cause.
A low-confidence choice, for example, can pause automatic progress and go to a person for review. But as TypeSafe's confidence documentation explains, confidence is a signal computed from the probability distribution over choices. It is neither permission nor a guarantee that the choice is correct. Code should therefore block an action outside the granted permissions, or a destructive action, before execution regardless of how high the confidence is.
if (!allowedAtThisStep(choice) || isDestructive(choice))
{
RecordDecision("blocked", choice, confidence);
StopRun();
return;
}
This trace is not decoration for watching AI. It helps decide whether a failure belongs in the game code, the agent's decision, or the test harness. Healthy runs can close, unexpected outcomes can be reviewed, and reproducible defects can become issues. A permission breach should be escalated separately from an ordinary quality finding. Sometimes the reason to stop is clearer than the reason to continue.
What Rekon preserves, and what it does not do yet
Rekon does not make those audit judgments. Today, it preserves video, logs, and game state when game code explicitly requests a capture. It does not automatically collect Jev decisions, check permissions, or stop execution. Each game must define which state and actions to record, then implement safety boundaries and review paths in its harness. The guide to connecting capture to Jev autoplay assumes the same boundary.
Automation's maturity may not be measured only by how many times it plays overnight. A result becomes accountable when we can compare what the agent saw, what it was allowed to do, and what it did. Distinguishing a broken game from a poor decision or a faulty test is still a human responsibility.