Briefing edition:
Research · · reported
ParanoiaEval preprint proposes benchmark for unnecessary defensive work in coding agents
An arXiv preprint by Hanjun Luo and co-authors introduces ParanoiaEval, described by its authors as a benchmark for unified evaluation of risk-treatment capabilities in coding agents, built on the Avoidance-Transfer-Mitigation-Acceptance framework with 200 evidence-controlled repository-level task pairs. The authors report that across 8 representative models, unnecessary risk treatment occurred in 11.2%-58.7% of runs despite explicit evidence, and that stronger task capability did not ensure more appropriate risk treatment.
5.0/10 significance · AI confidence estimate 62%
What changed
Researchers Hanjun Luo, Xiucheng Zhang, Zhuoning Xu, Zhimu Huang, Yingbin Jin, Xinfeng Li and Hanan Salam posted an arXiv preprint describing ParanoiaEval, which they describe as a benchmark for unified evaluation of risk-treatment capabilities in coding agents.
Why it matters
If the authors' reported pattern holds, teams deploying coding agents may need to evaluate defensive over-treatment as a capability dimension separate from task performance, though the preprint's metrics, judge calibration and results have not been independently replicated.
What remains uncertain
Still to verify for this briefing: when this specific development occurred; technical specifications; performance claims; independent corroboration.
What to watch
Watch for the full paper, peer review or independent replication of the reported violation rates and the human-calibrated agentic judge.
Sources
arxiv.org ↗
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
Why this ranks here
Relevant to this audience because coding agents are widely used by builders, and the preprint frames unnecessary defensive work as a measurable, separate capability dimension rather than a task-accuracy issue. The principal limitation is that this is an author-supplied abstract of a preprint: the full paper was not read, peer-review status and replication are unestablished, and the repository timestamps are metadata dates, not verified announcement times.
- impact
- 5/10
- reach
- 5/10
- novelty
- 6/10
- institutional
- 3/10
- evidence
- 5/10
- potential
- 5/10
Story development
First recorded development in this briefing.