I am a second year PhD student studying at the Department of Computer Science and Technology, Tsinghua University, advised by Prof. Lifeng Sun. I received my bachelor’s degree from Tsinghua University in June 2024. My research interests primarily lie in the field of multimodal reasoning and self-evolving agents.
🔥 News
- 2026.06: 🎉 One paper accepted by TMLR 2026.
- 2026.02: 🎉 One paper accepted by CVPR 2026.
- 2026.01: 🎉 Two papers accepted by ICLR 2026.
📖 Educations
2024.09 - Present, Ph.D. @ Department of Computer Science and Technology, Tsinghua University.
2020.09 - 2024.06, B.Eng @ Department of Computer Science and Technology, Tsinghua University.
💻 Internships
2026.06 - 2026.08, Research Intern @ Yuanbao Group, Tencent
2024.11 - 2026.05, Research Intern @ DAMO Academy, Alibaba
📝 Publications
${*}$ Equal contribution, ${\dagger}$ Corresponding author

Scalable Behaviour Cloning on Browser Using via Skill Distillation
Kaisen Yang$^{*}$, Zheng Jiang$^{*}$, Yuzhao Peng$^{*}$, Houde Qian$^{*}$, Boshi Zhang$^{*}$, Youjie Zheng, Shijin Hong, Qingle Liu, Ruoyu Han, Bohan Lyu, Bingxiang He, Eren Cai, Calvin Xiao, Qinhuai Na$^{\dagger}$
- BrowserBC distills human browser interaction traces into reusable natural-language skills organized as a skill graph, providing retrievable and composable procedural priors for efficient browser-agent execution.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images
Zheng Jiang, Yiming Chen, Nan He, Jiahui Chen, Chaoyang Li, Houde Qian, Lifeng Sun$^{\dagger}$
- TTSP is a test-time scaling framework that resolves the grounding paradox in tool-augmented visual reasoning by scaling perception through parallel exploration, reliability filtering, and iterative knowledge refinement.

SubFLOT: Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport
Zheng Jiang, Nan He, Yiming Chen, Lifeng Sun$^{\dagger}$
- SubFLOT is a server-side personalized federated pruning framework that leverages optimal transport and adaptive regularization to address system and statistical heterogeneity without accessing local data.

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
Zheng Jiang$^{*}$, Heng Guo$^{*}$, Chengyu Fang$^{*}$, Changchen Xiao, Xinyang Hu, Lifeng Sun$^{\dagger}$, Minfeng Xu$^{\dagger}$
- MedVR is the first end-to-end reinforcement learning framework that seamlessly integrates visual and textual reasoning for medical VLMs, obviating the need for costly intermediate supervision.

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
Che Liu$^{*}$, Zheng Jiang$^{*}$, Chengyu Fang$^{*}$, Heng Guo, Yanjie Zhou, Jiaqi Qu, Le Lu, Minfeng Xu$^{\dagger}$
- A unified visual encoder without any modality-specific customization for various medical visual modalities in 2D and 3D.

Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
Chengyu Fang$^{*}$, Heng Guo$^{*}$, Zheng Jiang, Chunming He, Xiu Li$^{\dagger}$, Minfeng Xu$^{\dagger}$
- Photon is a variable-length 3D medical VQA framework with instruction-conditioned token scheduling and surrogate gradients, achieving adaptive acceleration and state-of-the-art performance.
🎖 Honors and Awards
- 2026.06: Tsinghua University Merit Student
- 2026.06: Outstanding Student Leader of Tsinghua University
- 2024.06: Outstanding Graduates of Department of Computer Science and Technology, Tsinghua University
- 2021-2023: Academic Excellence Scholarship, Tsinghua University
💬 Invited Talks
- 2026.04, “MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning”, AI TIME.