University of Liverpool · Agent monitoring

PrefixGuardOnline Failure Warning and
Trace-Grounded Diagnosis for LLM Agents

Warn before an agent fails.
Trace what happened. Check what should have happened.

Xinmiao HuangJinwei HuQisong HeRajarshi RoyChangshun WuYi Dong†Xiaowei Huang†

University of Liverpool · † Corresponding authors

24 / 24

Benchmark–horizon settings with higher mean AUPRC than every evaluated baseline

90.9%

Paired GRU AUPRC retained by compact DFAs, on average

21 of 265

Benchmark-successful retail runs with verified specification violations

The idea

A warning is the beginning of a diagnosis.

A terminal verdict tells us whether an agent succeeded. PrefixGuard learns from those verdicts to estimate risk during execution, then connects the warning to an inspectable path through the agent’s actions and feedback.

A failure warning arrives at step 13 before termination at step 20; paired executions diverge after a shared diagnosis and data refill.
Warn early, then inspect the execution. In this example, PrefixGuard alerts at step 13 before failure at step 20. Replaying the learned events connects the warning to the actions that separate a successful repair from an unresolved one.
Read the abstract

Large language model (LLM) agents execute multi-step tasks, but terminal verdicts arrive too late for intervention. Agent traces mix messages, tool calls, and feedback, making it difficult to learn failure signals from terminal outcomes. Even if a model learns to predict failure from execution prefixes, its predictions alone cannot show which actions, feedback, or unmet conditions need inspection. Existing work studies online failure prediction and diagnosis from completed trajectories, yet few methods connect timely warning to a compact symbolic model of execution behavior. We introduce PrefixGuard, a neuro-symbolic method that uses learned events to connect online failure warning with finite-state execution diagnosis. A gated recurrent unit (GRU) predicts horizon-specific failure risk from these events. A deterministic finite automaton (DFA) organizes their histories into compact risk-labeled paths for replay and rule-monitor composition. Across all 24 benchmark–horizon settings, PrefixGuard exceeds the strongest evaluated baseline in mean test area under the precision–recall curve (AUPRC). Average gains range from 12.0 to 21.9 percentage points. At a nominal 10% FAR budget, averaged test recall spans 56.1%–93.0% across benchmarks. On one frozen τ²-Bench DFA, model checking five specifications identifies violations in 21 of 265 benchmark-successful runs.

01 / Method

One event representation.
Two connected views of execution.

The same learned events support a recurrent predictor for online warning and a finite-state model for replay and diagnosis.

PrefixGuard architecture: StepView and TF-IDF feed a shared event projection; soft events drive GRU warnings while hard events drive DFA replay. Bottom annotations show training, DFA induction, and calibration.
Figure 2 · PrefixGuard architecture. A shared projection supplies soft events for GRU warnings and hard events for DFA replay. DFA paths are inspected alongside recorded actions and feedback. Bottom annotations summarize training, induction from prefix labels, and calibration. Click the figure to view it at full resolution.

02 / Online warning

Earlier warnings across four agent benchmarks.

Across WebArena, τ²-Bench, SkillsBench, and TerminalBench, PrefixGuard exceeds PPM-LSTM and three log-anomaly baselines in all 24 benchmark–horizon settings. Horizon-averaged gains over the strongest baseline range from 12.0–21.9 percentage points in AUPRC.

Four benchmark columns compare test AUPRC, failed-run recall, and first-alert lead across warning horizons for PrefixGuard and the baselines.
Choose how far ahead to predict. Larger horizons increase warning lead but do not uniformly improve recall. H = ∞ predicts eventual failure. Thresholds use a nominal 10% calibration false-alarm budget; realized test FAR can exceed that budget.

A common operating view: H = 3

Full held-out test results · Mean ± population SD over three seeds

PrefixGuard GRU at a three-step prediction horizon
BenchmarkTest AUPRCFailed-run recallRealized test FARLead fraction
WebArena0.836 ± 0.0010.514 ± 0.0260.074 ± 0.0060.035 ± 0.003
τ²-Bench0.699 ± 0.0140.997 ± 0.0020.123 ± 0.0030.165 ± 0.006
SkillsBench0.591 ± 0.0080.741 ± 0.0800.090 ± 0.0120.059 ± 0.006
TerminalBench0.556 ± 0.0050.964 ± 0.0010.095 ± 0.0040.224 ± 0.009

The horizon is a prediction target, not a guaranteed warning lead. AUPRC ranks prefixes; recall and FAR measure alerted trajectories. See Appendix C.1 in the paper for the complete horizon sweep.

Evaluation protocol and what the ablations show

Each benchmark uses a 70/10/10/10 trajectory split for training, calibration, validation, and test. Validation AUPRC selects checkpoints. Successful calibration trajectories set the split-conformal alert threshold at a nominal 10% FAR budget. The collected traces cover 17 agent implementations and 39 backbone models.

The GRU on full StepView features without the event projection scores 0.4–2.4 AUPRC points higher than the event-based GRU. The projection supplies discrete histories for the DFA at a small predictive cost. A zero-shot DeepSeek-V4-Pro judge is a separate H = 3 reference evaluated on 200 sampled test prefixes per benchmark.

03 / Execution diagnosis

From a risk score to a replayable path.

Across benchmark–horizon settings, the induced DFAs retain 90.9% of their paired GRU’s AUPRC on average, with mean sizes ranging from 5.0 to 32.3 states. Aligning state paths with recorded actions and feedback exposes where executions diverge.

Explore DFA induction step by step

An interactive illustration with toy data. Use the playback controls to step through the animation.

Open the animation in a separate tab ↗

Successful and failed telecom executions replayed through the same DFA: the failed path reaches alert state s6, while the successful path reaches low-risk s8.
Same task, different execution paths. On τ²-Bench (H = 3, seed 13), the failed run reaches alert state s₆ at step 13 with risk 0.965 and returns there after further repairs. The successful run reaches low-risk s₈ while restoring MMS.

Successful execution

Check the network mode, then restore MMS.

The agent checks APN and network settings, selects 4G/5G, and restores service. The state path is read alongside these concrete actions.

Failed execution

Repeat repairs while the prerequisite stays unresolved.

The agent leaves the network mode unchanged, repeats repairs, and hands off with MMS unresolved. The replay identifies the divergence for inspection.

Inspect finite-state fidelity across benchmarks
DFA and GRU AUPRC, number of DFA states, and alert agreement across four benchmarks and their warning horizons.
DFA alert recall is 71–98% against GRU alerts. Agreement varies by benchmark: precision is 64–88% on τ²-Bench, SkillsBench, and TerminalBench, and 37–76% on WebArena. Compact states preserve much of the ranking but can group prefixes that lie on different sides of the GRU’s threshold.

04 / Task-grounded model checking

A successful outcome can still violate a requirement.

Compose the frozen DFA with monitors for five tool-policy specifications. The product state tracks both the learned execution state and whether a requirement has been satisfied or violated, without retraining the DFA.

21 / 265

Benchmark-successful runs with verified violations

7.92% of successful retail executions. The same audit flags 19 of 94 failed runs, across a population of 359 executions, 114 tasks, and four agents.

Verified violations by specification family
RequirementSuccessful runs (n = 265)Failed runs (n = 94)
Order status permits the operation1211
Replacement lists match and modifications change the item34
Requested quantities do not exceed stock01
Replacement is an available same-product variant11
Refund destination follows the payment policy63
Any verified violation2119

A run may violate multiple specifications. Unknown verdicts are excluded from violation counts; 6 successful runs have an unknown overall verdict. All flagged violations are checked against the original trajectories because the DFA can recombine transitions into unobserved paths.

Retail task 27 / Order matters

The customer prefers an exchange if both requests cannot be fulfilled.

Failed run

Return first → status changes → exchange blocked

Successful run

Exchange first → then handle the return

Both runs violate the order-status specification, but only the second passes the task’s database check. Both violations share DFA state s₂; the requirement monitor distinguishes their histories. Task context explains why the blocked precondition matters.

This study evaluates warning and trace-grounded diagnosis. Generalization to unseen tasks and agents, and whether acting on these diagnoses improves outcomes, remain open questions.

Resources

Read, reproduce, and cite.

BibTeX

@article{huang2026prefixguard,
  title={PrefixGuard: Online Failure Warning and Trace-Grounded Diagnosis for LLM Agents},
  author={Huang, Xinmiao and Hu, Jinwei and He, Qisong and Roy, Rajarshi and
          Wu, Changshun and Dong, Yi and Huang, Xiaowei},
  journal={arXiv preprint arXiv:2605.06455},
  year={2026},
  doi={10.48550/arXiv.2605.06455},
  url={https://arxiv.org/abs/2605.06455}
}

Correspondence: Xiaowei Huang · Yi Dong