University of Liverpool · Agent monitoring
PrefixGuardOnline Failure Warning and
Trace-Grounded Diagnosis for LLM Agents
Warn before an agent fails.
Trace what happened. Check what should have happened.
University of Liverpool · † Corresponding authors
Benchmark–horizon settings with higher mean AUPRC than every evaluated baseline
Paired GRU AUPRC retained by compact DFAs, on average
Benchmark-successful retail runs with verified specification violations
The idea
A warning is the beginning of a diagnosis.
A terminal verdict tells us whether an agent succeeded. PrefixGuard learns from those verdicts to estimate risk during execution, then connects the warning to an inspectable path through the agent’s actions and feedback.
Read the abstract
Large language model (LLM) agents execute multi-step tasks, but terminal verdicts arrive too late for intervention. Agent traces mix messages, tool calls, and feedback, making it difficult to learn failure signals from terminal outcomes. Even if a model learns to predict failure from execution prefixes, its predictions alone cannot show which actions, feedback, or unmet conditions need inspection. Existing work studies online failure prediction and diagnosis from completed trajectories, yet few methods connect timely warning to a compact symbolic model of execution behavior. We introduce PrefixGuard, a neuro-symbolic method that uses learned events to connect online failure warning with finite-state execution diagnosis. A gated recurrent unit (GRU) predicts horizon-specific failure risk from these events. A deterministic finite automaton (DFA) organizes their histories into compact risk-labeled paths for replay and rule-monitor composition. Across all 24 benchmark–horizon settings, PrefixGuard exceeds the strongest evaluated baseline in mean test area under the precision–recall curve (AUPRC). Average gains range from 12.0 to 21.9 percentage points. At a nominal 10% FAR budget, averaged test recall spans 56.1%–93.0% across benchmarks. On one frozen τ²-Bench DFA, model checking five specifications identifies violations in 21 of 265 benchmark-successful runs.
01 / Method
One event representation.
Two connected views of execution.
The same learned events support a recurrent predictor for online warning and a finite-state model for replay and diagnosis.
02 / Online warning
Earlier warnings across four agent benchmarks.
Across WebArena, τ²-Bench, SkillsBench, and TerminalBench, PrefixGuard exceeds PPM-LSTM and three log-anomaly baselines in all 24 benchmark–horizon settings. Horizon-averaged gains over the strongest baseline range from 12.0–21.9 percentage points in AUPRC.
A common operating view: H = 3
Full held-out test results · Mean ± population SD over three seeds
| Benchmark | Test AUPRC | Failed-run recall | Realized test FAR | Lead fraction |
|---|---|---|---|---|
| WebArena | 0.836 ± 0.001 | 0.514 ± 0.026 | 0.074 ± 0.006 | 0.035 ± 0.003 |
| τ²-Bench | 0.699 ± 0.014 | 0.997 ± 0.002 | 0.123 ± 0.003 | 0.165 ± 0.006 |
| SkillsBench | 0.591 ± 0.008 | 0.741 ± 0.080 | 0.090 ± 0.012 | 0.059 ± 0.006 |
| TerminalBench | 0.556 ± 0.005 | 0.964 ± 0.001 | 0.095 ± 0.004 | 0.224 ± 0.009 |
The horizon is a prediction target, not a guaranteed warning lead. AUPRC ranks prefixes; recall and FAR measure alerted trajectories. See Appendix C.1 in the paper for the complete horizon sweep.
Evaluation protocol and what the ablations show
Each benchmark uses a 70/10/10/10 trajectory split for training, calibration, validation, and test. Validation AUPRC selects checkpoints. Successful calibration trajectories set the split-conformal alert threshold at a nominal 10% FAR budget. The collected traces cover 17 agent implementations and 39 backbone models.
The GRU on full StepView features without the event projection scores 0.4–2.4 AUPRC points higher than the event-based GRU. The projection supplies discrete histories for the DFA at a small predictive cost. A zero-shot DeepSeek-V4-Pro judge is a separate H = 3 reference evaluated on 200 sampled test prefixes per benchmark.
03 / Execution diagnosis
From a risk score to a replayable path.
Across benchmark–horizon settings, the induced DFAs retain 90.9% of their paired GRU’s AUPRC on average, with mean sizes ranging from 5.0 to 32.3 states. Aligning state paths with recorded actions and feedback exposes where executions diverge.
Explore DFA induction step by step
An interactive illustration with toy data. Use the playback controls to step through the animation.
Successful execution
Check the network mode, then restore MMS.
The agent checks APN and network settings, selects 4G/5G, and restores service. The state path is read alongside these concrete actions.
Failed execution
Repeat repairs while the prerequisite stays unresolved.
The agent leaves the network mode unchanged, repeats repairs, and hands off with MMS unresolved. The replay identifies the divergence for inspection.
Inspect finite-state fidelity across benchmarks

04 / Task-grounded model checking
A successful outcome can still violate a requirement.
Compose the frozen DFA with monitors for five tool-policy specifications. The product state tracks both the learned execution state and whether a requirement has been satisfied or violated, without retraining the DFA.
Benchmark-successful runs with verified violations
7.92% of successful retail executions. The same audit flags 19 of 94 failed runs, across a population of 359 executions, 114 tasks, and four agents.
| Requirement | Successful runs (n = 265) | Failed runs (n = 94) |
|---|---|---|
| Order status permits the operation | 12 | 11 |
| Replacement lists match and modifications change the item | 3 | 4 |
| Requested quantities do not exceed stock | 0 | 1 |
| Replacement is an available same-product variant | 1 | 1 |
| Refund destination follows the payment policy | 6 | 3 |
| Any verified violation | 21 | 19 |
A run may violate multiple specifications. Unknown verdicts are excluded from violation counts; 6 successful runs have an unknown overall verdict. All flagged violations are checked against the original trajectories because the DFA can recombine transitions into unobserved paths.
Retail task 27 / Order matters
The customer prefers an exchange if both requests cannot be fulfilled.
Failed run
Return first → status changes → exchange blocked
Successful run
Exchange first → then handle the return
Both runs violate the order-status specification, but only the second passes the task’s database check. Both violations share DFA state s₂; the requirement monitor distinguishes their histories. Task context explains why the blocked precondition matters.
This study evaluates warning and trace-grounded diagnosis. Generalization to unseen tasks and agents, and whether acting on these diagnoses improves outcomes, remain open questions.
Resources
Read, reproduce, and cite.
BibTeX
@article{huang2026prefixguard,
title={PrefixGuard: Online Failure Warning and Trace-Grounded Diagnosis for LLM Agents},
author={Huang, Xinmiao and Hu, Jinwei and He, Qisong and Roy, Rajarshi and
Wu, Changshun and Dong, Yi and Huang, Xiaowei},
journal={arXiv preprint arXiv:2605.06455},
year={2026},
doi={10.48550/arXiv.2605.06455},
url={https://arxiv.org/abs/2605.06455}
}
Correspondence: Xiaowei Huang · Yi Dong