Previously, I was a PostDoc at UC Berkeley, working with Dawn Song. I received my PhD from ETH Zurich, advised by Martin Vechev. I finished my undergraduate at Zhejiang University.
I am looking for motivated PhD students, Postdocs, visiting students, and research assistants to join my group at NTU starting in Fall 2027. If you are interested or see a potential fit, I encourage you to apply by filling out this form.
Research
My research spans AI, programming languages, and security. I seek to answer a key question: How does AI reshape programming?
Security is a direct motivation. AI is becoming ever more capable of discovering and exploiting vulnerabilities, yet it keeps introducing them into the code it generates. My work has demonstrated these risks at scale and is widely used by the frontier AI industry (CyberGym, ExploitGym, BaxBench).
Guarantees by construction are my response. AI can absorb much of the effort traditionally required by secure- or correct-by-construction techniques, such as adopting safer programming languages, making them practical at unprecedented scale. I therefore build these techniques and their guarantees into how AI generates code, via typing, compilation, proofs, and training.
Honors and Awards
Oral Paper at ICLR 2026
Two Spotlight Papers at ICML 2025
ETH Medal for Oustanding Doctoral Thesis, 2024
ACM CCS Distinguished Paper Award, 2023
NeurIPS Top Reviewer, 2023
Publications
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Niels Mündler-Sasahara*, Hristo Venev*, Dawn Song, Martin Vechev, Jingxuan He
ICML 2026 Workshop on Deep Learning for Code (DL4C), 2026
papercode
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
Zhun Wang, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, Eric Wallace, Elie Bursztein, Luca Invernizzi, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, Dawn Song
arXiv preprint, 2026
Covered by OpenAI-Huggingface incident and adopted by frontier AI industry (e.g., OpenAI, Anthropic, Z.ai) papercode
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, Dawn Song
International Conference on Learning Representations (ICLR), 2026
(Oral) Adopted by frontier AI industry (e.g., OpenAI, Anthropic, Z.ai, DeepSeek, Microsoft) papercodewebsitedataset
Verina: Benchmarking Verifiable Code Generation
Zhe Ye, Zhengxu Yan, Jingxuan He, Timothe Kasriel, Kaiyu Yang, Dawn Song
International Conference on Learning Representations (ICLR), 2026
Adopted by startups working on AI for formal verification (e.g., Harmonic, Logical Intelligence, Axiom) papercodewebsitedataset
Type-Constrained Code Generation with Language Models
Niels Mündler*, Jingxuan He*, Hao Wang, Koushik Sen, Dawn Song, Martin Vechev
ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2025
papercode
BaxBench: Can LLMs Generate Secure and Correct Backends?
Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanović, Jingxuan He, Martin Vechev
International Conference on Machine Learning (ICML), 2025
(Spotlight) papercodewebsitedataset
Large Language Models for Code: Security Hardening and Adversarial Testing
Jingxuan He, Martin Vechev
ACM Conference on Computer and Communications Security (CCS), 2023
(Distinguished Paper) papercodeslides
Recent Preprints
Vero: Can AI Agents Build Formally Verified Software Repositories?
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Niels Mündler-Sasahara*, Hristo Venev*, Dawn Song, Martin Vechev, Jingxuan He
ICML 2026 Workshop on Deep Learning for Code (DL4C), 2026
papercode
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
Zhun Wang, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, Eric Wallace, Elie Bursztein, Luca Invernizzi, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, Dawn Song
arXiv preprint, 2026
Covered by OpenAI-Huggingface incident and adopted by frontier AI industry (e.g., OpenAI, Anthropic, Z.ai) papercode
Conference Publications
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities
Tianneng Shi*, Robin Rheem*, Dongwei Jiang, Mona Wang, Francisco De La Riega, Zhun Wang, Jingzhi Jiang, Alexander Cheung, Sean Tai, Jonah Cha, Jianhong Tu, Gabriel Han, Chenguang Wang, Jingxuan He, Wenbo Guo, Dawn Song
International Conference on Machine Learning (ICML), 2026
papercodewebsite
Hongwei Li*, Zhun Wang*, Qinrun Dai, Yuzhou Nie, Jinjun Peng, Ruitong Liu, Jingyang Zhang, Kaijie Zhu, Jingxuan He, Lun Wang, Yangruibo Ding, Yueqi Chen, Wenbo Guo, Dawn Song
International Conference on Machine Learning (ICML), 2026
papercodewebsite
SecCodePRM: A Process Reward Model for Code Security
Weichen Yu, Ravi Mangal, Yinyi Luo, Kai Hu, Jingxuan He, Corina S. Pasareanu, Matt Fredrikson
International Conference on Machine Learning (ICML), 2026
papercode
Position: Agent Security Needs Redefinition through a Holistic Framework
Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Chenguang Wang, Dawn Song
International Conference on Machine Learning (ICML) Position Paper Track, 2026
paper
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, Dawn Song
International Conference on Learning Representations (ICLR), 2026
(Oral) Adopted by frontier AI industry (e.g., OpenAI, Anthropic, Z.ai, DeepSeek, Microsoft) papercodewebsitedataset
Verina: Benchmarking Verifiable Code Generation
Zhe Ye, Zhengxu Yan, Jingxuan He, Timothe Kasriel, Kaiyu Yang, Dawn Song
International Conference on Learning Representations (ICLR), 2026
Adopted by startups working on AI for formal verification (e.g., Harmonic, Logical Intelligence, Axiom) papercodewebsitedataset
Type-Constrained Code Generation with Language Models
Niels Mündler*, Jingxuan He*, Hao Wang, Koushik Sen, Dawn Song, Martin Vechev
ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2025
papercode
BaxBench: Can LLMs Generate Secure and Correct Backends?
Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanović, Jingxuan He, Martin Vechev
International Conference on Machine Learning (ICML), 2025
(Spotlight) papercodewebsitedataset
Formal Mathematical Reasoning: A New Frontier in AI
Kaiyu Yang, Gabriel Poesia, Jingxuan He, Wenda Li, Kristin Lauter, Swarat Chaudhuri, Dawn Song
International Conference on Machine Learning (ICML), 2025
(Spotlight) Communications of the ACM, 2026
ICML paperCACM paper
Mind the Gap: A Practical Attack on GGUF Quantization
Kazuki Egashira, Robin Staab, Mark Vero, Jingxuan He, Martin Vechev
International Conference on Machine Learning (ICML), 2025
ICLR 2025 Workshop on Building Trust in Language Models and Applications, 2025
(Oral) papercode
Black-Box Adversarial Attacks on LLM-Based Code Completion
Slobodan Jenko*, Niels Mündler*, Jingxuan He, Mark Vero, Martin Vechev
International Conference on Machine Learning (ICML), 2025
papercode
Exploiting LLM Quantization
Kazuki Egashira, Mark Vero, Robin Staab, Jingxuan He, Martin Vechev
Neural Information Processing Systems (NeurIPS), 2024
ICML 2024 Workshop on the Next Generation of AI Safety, 2024
(Oral) papercodewebsite
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
Niels Mündler, Mark Niklas Müller, Jingxuan He, Martin Vechev
Neural Information Processing Systems (NeurIPS), 2024
papercodeposter
Instruction Tuning for Secure Code Generation
Jingxuan He*, Mark Vero*, Gabriela Krasnopolska, Martin Vechev
International Conference on Machine Learning (ICML), 2024
papercode
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
Niels Mündler, Jingxuan He, Slobodan Jenko, Martin Vechev
International Conference on Learning Representations (ICLR), 2024
papercodewebsite
Large Language Models for Code: Security Hardening and Adversarial Testing
Jingxuan He, Martin Vechev
ACM Conference on Computer and Communications Security (CCS), 2023
(Distinguished Paper) papercodeslides
On Distribution Shift in Learning-based Bug Detectors
Jingxuan He, Luca Beurer-Kellner, Martin Vechev
International Conference on Machine Learning (ICML), 2022
papercode
Learning to Explore Paths for Symbolic Execution
Jingxuan He, Gishor Sivanrupan, Petar Tsankov, Martin Vechev
ACM Conference on Computer and Communications Security (CCS), 2021
papercodeslides
TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer
Berkay Berabi, Jingxuan He, Veselin Raychev, Martin Vechev
International Conference on Machine Learning (ICML), 2021
papercodetalkslides
Learning to Find Naming Issues with Big Code and Small Supervision
Jingxuan He, Cheng-Chun Lee, Veselin Raychev, Martin Vechev
ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2021
papertalkslides
Learning Fast and Precise Numerical Analysis
Jingxuan He, Gagandeep Singh, Markus Püschel, Martin Vechev
ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2020
papercodetalkslides
Learning to Fuzz from Symbolic Execution with Application to Smart Contracts
Jingxuan He, Mislav Balunović, Nodar Ambroladze, Petar Tsankov, Martin Vechev
ACM Conference on Computer and Communications Security (CCS), 2019
papercodetalkslides
DeBin: Predicting Debug Information in Stripped Binaries
Jingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev, Martin Vechev
ACM Conference on Computer and Communications Security (CCS), 2018
papercodetalkslides
Workshop Papers and Other Preprints
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
Hao Wang, Niels Mündler, Mark Vero, Jingxuan He, Dawn Song, Martin Vechev
arXiv preprint, 2026
papercode
A Framework for Formalizing LLM Agent Security
Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Neil Gong, Chenguang Wang, Dawn Song
ICLR 2026 Workshop on Agents in the Wild (AIWILD), 2026
paper
Progent: Securing AI Agents with Privilege Control
Tianneng Shi, Jingxuan He, Zhun Wang, Hongwei Li, Linyu Wu, Wenbo Guo, Dawn Song
arXiv preprint, 2025
papercode
Reasoning Models Can Be Effective Without Thinking
Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, Matei Zaharia
arXiv preprint, 2025
paper