LLM-as-a-Judge & reliable automated evaluation
How LLMs are used to judge software artifacts — and what it would take for their verdicts to be trustworthy.
Ph.D. Candidate in Computer Science
I make large language models and AI agents trustworthy enough to rely on — by evaluating them rigorously and testing them systematically.
About
I am a PhD candidate at the School of Computing and Information Systems, Singapore Management University, advised by Prof. David Lo as part of the SOAR group. Before SMU, I received an MSc in Software Engineering (Distinction) and a BSc in Computer Science (First Class Honours) from University College London.
Benchmarks leak, automated judges disagree with developers, and learned agents fail in ways that ordinary tests never reach. I work on both halves of that problem: building evaluations — such as LLM-as-a-judge and BigCodeBench — whose numbers can be trusted, and building testing techniques that find the failures of reinforcement-learning and LLM-based agents before users do.
News
Our paper Learning from the Test: Self-Referential Differential Testing for Deep RL Agents will appear at ISSTA 2026.
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead is published in TOSEM.
Two co-authored papers at ASE 2026 (reference-free evaluation for binary reverse engineering; AgentChaos) and one at FSE 2026.
Invited talk at Korea University, Seoul: From Code to Courtroom: LLMs as the New Software Judges.
Invited talk at Osaka University: When LLMs Judge Software: Toward Trustworthy Automated Evaluation.
Received the SMU PhD Research Excellence Award and was named to the SCIS Dean’s List.
Research
Five threads, one goal: AI systems whose behaviour we can measure, test, and trust.
How LLMs are used to judge software artifacts — and what it would take for their verdicts to be trustworthy.
Curiosity-driven and self-referential differential testing for deep-RL and sequential decision-making systems; runtime safety shields, safety-violation search, and backdoors in offline RL data.
A literature review and vision for LLM-based multi-agent systems in SE, and chaos engineering for agent systems.
BigCodeBench, whose public leaderboard covers 160+ models; a critical assessment of SE benchmarks; reference-free evaluation; and data leakage across 83 SE benchmarks.
Pre-trained models for Stack Overflow posts — tag recommendation and a systematic study of post representations.
Publications
I am always happy to talk about agent evaluation, testing AI systems, benchmark design, or collaborations.