All publications

TOSEM 2026

LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead

Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo

ACM Transactions on Software Engineering and Methodology (TOSEM), 2026

Pushing LLMs into the software development lifecycle produced far more artifacts — patches, tests, reviews, summaries — than anyone can evaluate by hand, and the field responded by having LLMs grade each other. That move is convenient and mostly unexamined.

We survey the SE studies that adopt the LLM-as-a-judge paradigm, analyze the limitations they share, and identify the research gaps that keep these judges from being dependable: unstable agreement with developers, prompt and position sensitivity, and evaluation setups that cannot distinguish a good judge from a lucky one. The paper closes with a roadmap for building judges whose verdicts a practitioner could reasonably act on.

An earlier version of this work circulated as From Code to Courtroom: LLMs as the New Software Judges (arXiv:2503.02246).