Ph.D. Candidate in Computer Science

Junda He

I make large language models and AI agents trustworthy enough to rely on — by evaluating them rigorously and testing them systematically.

  • SCIS, Singapore Management University
  • Advised by Prof. David Lo
  • Singapore
Junda He
LLM-as-a-Judge Testing AI agents

About

Hello, I'm Junda.

I am a PhD candidate at the School of Computing and Information Systems, Singapore Management University, advised by Prof. David Lo as part of the SOAR group. Before SMU, I received an MSc in Software Engineering (Distinction) and a BSc in Computer Science (First Class Honours) from University College London.

Benchmarks leak, automated judges disagree with developers, and learned agents fail in ways that ordinary tests never reach. I work on both halves of that problem: building evaluations — such as LLM-as-a-judge and BigCodeBench — whose numbers can be trusted, and building testing techniques that find the failures of reinforcement-learning and LLM-based agents before users do.

28publications
7first-authored
12journal articles
160+models on the BigCodeBench leaderboard

News

Recent updates

  1. 2026

    Our paper Learning from the Test: Self-Referential Differential Testing for Deep RL Agents will appear at ISSTA 2026.

  2. 2026

    LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead is published in TOSEM.

  3. 2026

    Two co-authored papers at ASE 2026 (reference-free evaluation for binary reverse engineering; AgentChaos) and one at FSE 2026.

  4. 2025.11

    Invited talk at Korea University, Seoul: From Code to Courtroom: LLMs as the New Software Judges.

  5. 2025.10

    Invited talk at Osaka University: When LLMs Judge Software: Toward Trustworthy Automated Evaluation.

  6. 2025

    Received the SMU PhD Research Excellence Award and was named to the SCIS Dean’s List.

Research

What I work on

Five threads, one goal: AI systems whose behaviour we can measure, test, and trust.

LLM-as-a-Judge & reliable automated evaluation

How LLMs are used to judge software artifacts — and what it would take for their verdicts to be trustworthy.

Representation learning for developer content

Pre-trained models for Stack Overflow posts — tag recommendation and a systematic study of post representations.

Publications

Selected papers

All 28 papers

Let's talk.

I am always happy to talk about agent evaluation, testing AI systems, benchmark design, or collaborations.

jundahe.2022@phdcs.smu.edu.sg