When AI Finds a Bug, What Should a Human Check First? — Triaging Automated Play Reports with Jev

Separate severity, reproducibility, and evidence quality when triaging Jev automated-play reports, while keeping scoring policy and human review explicit.

  • unity
  • ai
  • qa

The morning after an automated play run, we can end up facing a pile of failed reports before we look at the failed game itself. Hundreds of runs may have finished overnight, but the titles look alike: one blocks progress, one stalls briefly, and another leaves only a line in the log. They cannot all be read in the same order. Deciding what to inspect first has not kept pace with the speed of execution.

What we need is less a score that reduces a bug to one number than a classification that shows why a report should be read first. TypeSafe's Score documentation shows how to rate one criterion on ordered levels, then combine separate Scores with weights in application code. Applied to automated-play reports, severity and report quality should first be treated as different questions.

Separate the questions before combining them

Severity is the impact on play or the product. A failure that blocks progress or corrupts saved data should not have the same priority as a brief visual glitch at the edge of the screen. Reproducibility asks whether the issue occurs again under the same conditions. A stall after the same action every time calls for a different investigation from an issue seen once after many runs.

Evidence quality asks whether there are enough clues to support those judgments. Do the automated-play state, selected action, result, and logs line up in time? Can someone follow the report back to the point of failure? If sparse evidence also lowers the severity score, the most dangerous but least understood issues can sink to the bottom of the queue.

Suppose the screen froze during a boss fight, but the final state and logs are missing. The useful summary is closer to “high severity, reproducibility unknown, insufficient evidence.” Poor evidence is not another name for low severity. Keeping the three judgments separate lets someone explain whether to rerun the scenario, collect better logs, or send it to a person now.

Scores can sort the queue; the team owns the policy

Combining the three criteria with chosen weights can help sort a long queue. But Score does not decide those weights or the threshold that makes something P0. The code and team that understand the game's progression, player impact, and release risk have to own that policy. A score is a way to express the policy in operation; it should not become the source of the policy.

Borrowing the shape of Agent Trace observability, we can retain how each judgment and route followed from its inputs instead of saving only the final score. A high-scoring report might go to a priority queue, while an issue marked severe but uncertain in reproducibility or evidence goes to human review. When a person receives it, the reasoning should arrive with the number.

A Rekon evidence bundle could be input to that triage flow. Rekon does not currently provide Jev automated-play scoring or automatic report triage. The earlier Jev and capture article covered what to preserve as evidence during automated play; this one asks how a person might read that evidence afterward. Capture preserves material for investigation, but it does not take responsibility for deciding the importance of the question or its answer.

The more automation sorts reports, the more we need a way to retrace why a report landed where it did. Expressing quality as a number seems possible; making sure the number does not hide what it cannot know still belongs to the team.