Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Junda He
Ph.D. candidate at Singapore Management University. I make LLMs and AI agents trustworthy enough to rely on — by evaluating them rigorously and testing them systematically.
Posts
portfolio
publications
Can Identifier Splitting Improve Open-Vocabulary Language Model of Code?
Published in IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), Early Research Achievements Track, 2022
Answer Summarization for Technical Queries: Benchmark and New Approach
Published in IEEE/ACM International Conference on Automated Software Engineering (ASE), 2022
Natural Attack for Pre-trained Models of Code
Published in IEEE/ACM International Conference on Software Engineering (ICSE), 2022
PTM4Tag: Sharpening Tag Recommendation of Stack Overflow Posts with Pre-trained Models
Published in IEEE/ACM 30th International Conference on Program Comprehension (ICPC), 2022
A triplet architecture that models a Stack Overflow post’s title, description, and code with independent pre-trained models to recommend tags.
CCBERT: Self-Supervised Code Change Representation Learning
Published in IEEE International Conference on Software Maintenance and Evolution (ICSME), 2023
Generation-Based Code Review Automation: How Far Are We?
Published in IEEE/ACM International Conference on Program Comprehension (ICPC), 2023
Representation Learning for Stack Overflow Posts: How Far Are We?
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2024
Curiosity-Driven Testing for Sequential Decision-Making Process
Published in IEEE/ACM 46th International Conference on Software Engineering (ICSE), 2024
CureFuzz: curiosity-driven black-box fuzzing that finds failure-triggering states in sequential decision-making systems without needing access to the policy.
Greening Large Language Models of Code
Published in IEEE/ACM International Conference on Software Engineering (ICSE), Software Engineering in Society Track, 2024
Baffle: Hiding Backdoors in Offline Reinforcement Learning Datasets
Published in IEEE Symposium on Security and Privacy (S&P), 2024
Stealthy Backdoor Attack for Code Models
Published in IEEE Transactions on Software Engineering (TSE), 2024
Leveraging Large Language Model for Automatic Patch Correctness Assessment
Published in IEEE Transactions on Software Engineering (TSE), 2024
PTM4Tag+: Tag Recommendation of Stack Overflow Posts with Pre-trained Models
Published in Empirical Software Engineering (EMSE), 30(28), 2025
An extended study of tag recommendation for Stack Overflow posts using a triplet of independent pre-trained models over a post’s title, description, and code.
LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks
Published in arXiv preprint, 2025
A large-scale measurement of how much of 83 SE benchmarks already appears in LLM pre-training data — and a leakage-filtered version of those benchmarks.
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Published in International Conference on Learning Representations (ICLR), 2025
A code generation benchmark built around realistic tasks that require composing many library function calls under complex instructions — much harder to game than short, self-contained problems.
LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision, and the Road Ahead
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 34(5), Article 124 — 2030 Special Issue, 2025
A literature review of LLM-based multi-agent systems in software engineering, with two case studies showing what current frameworks can and cannot do.
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2025
Finding Safety Violations of AI-Enabled Control Systems Through the Lens of Synthesized Proxy Programs
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2025
Compiling Code LLMs into Lightweight Executables
Published in Proceedings of the ACM on Software Engineering (FSE), 2026
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2026
LLMs are increasingly used to grade the artifacts other LLMs produce. We review how SE research actually uses this paradigm, where it breaks, and what a trustworthy version would require.
AgentChaos: Chaos Engineering for Agent Systems via Programmatic Fault Injection
Published in IEEE/ACM International Conference on Automated Software Engineering (ASE), 2026
Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering
Published in IEEE/ACM International Conference on Automated Software Engineering (ASE), 2026
Guidelines for Empirical Studies in Software Engineering involving Large Language Models
Published in Empirical Software Engineering (EMSE), 2026
Artificial Intelligence Support for Software Architecture Practice: A Systematic Review and Future Directions
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2026
Identifying and Mitigating API Misuse in Large Language Models
Published in IEEE Transactions on Software Engineering (TSE), 2026
Synthesizing Efficient and Permissive Programmatic Runtime Shields for Neural Policies
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2026
AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
Published in arXiv preprint arXiv:2604.02665, 2026
Learning from the Test: Self-Referential Differential Testing for Deep RL Agents
Published in ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), 2026