I am a second year PhD student studying at the Department of Computer Science and Technology, Tsinghua University, advised by Prof. Lifeng Sun. I received my bachelor’s degree from Tsinghua University in June 2024. My research interests primarily lie in the field of multimodal reasoning and self-evolving agents.

🔥 News

  • 2026.06:  🎉 One paper accepted by TMLR 2026.
  • 2026.02:  🎉 One paper accepted by CVPR 2026.
  • 2026.01:  🎉 Two papers accepted by ICLR 2026.

📖 Educations

  • 2024.09 - Present, Ph.D. @ Department of Computer Science and Technology, Tsinghua University.
  • 2020.09 - 2024.06, B.Eng @ Department of Computer Science and Technology, Tsinghua University.

💻 Internships

  • 2026.06 - 2026.08, Research Intern @ Yuanbao Group, Tencent
  • 2024.11 - 2026.05, Research Intern @ DAMO Academy, Alibaba

📝 Publications

${*}$ Equal contribution, ${\dagger}$ Corresponding author

Arxiv 2026
sym

Scalable Behaviour Cloning on Browser Using via Skill Distillation

Kaisen Yang$^{*}$, Zheng Jiang$^{*}$, Yuzhao Peng$^{*}$, Houde Qian$^{*}$, Boshi Zhang$^{*}$, Youjie Zheng, Shijin Hong, Qingle Liu, Ruoyu Han, Bohan Lyu, Bingxiang He, Eren Cai, Calvin Xiao, Qinhuai Na$^{\dagger}$

  • BrowserBC distills human browser interaction traces into reusable natural-language skills organized as a skill graph, providing retrievable and composable procedural priors for efficient browser-agent execution.
Arxiv 2026
sym

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

Zheng Jiang, Yiming Chen, Nan He, Jiahui Chen, Chaoyang Li, Houde Qian, Lifeng Sun$^{\dagger}$

  • TTSP is a test-time scaling framework that resolves the grounding paradox in tool-augmented visual reasoning by scaling perception through parallel exploration, reliability filtering, and iterative knowledge refinement.
CVPR 2026
sym

SubFLOT: Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport

Zheng Jiang, Nan He, Yiming Chen, Lifeng Sun$^{\dagger}$

  • SubFLOT is a server-side personalized federated pruning framework that leverages optimal transport and adaptive regularization to address system and statistical heterogeneity without accessing local data.
ICLR 2026
sym

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning

Zheng Jiang$^{*}$, Heng Guo$^{*}$, Chengyu Fang$^{*}$, Changchen Xiao, Xinyang Hu, Lifeng Sun$^{\dagger}$, Minfeng Xu$^{\dagger}$

  • MedVR is the first end-to-end reinforcement learning framework that seamlessly integrates visual and textual reasoning for medical VLMs, obviating the need for costly intermediate supervision.
TMLR 2026
sym

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

Che Liu$^{*}$, Zheng Jiang$^{*}$, Chengyu Fang$^{*}$, Heng Guo, Yanjie Zhou, Jiaqi Qu, Le Lu, Minfeng Xu$^{\dagger}$

  • A unified visual encoder without any modality-specific customization for various medical visual modalities in 2D and 3D.
ICLR 2026
sym

Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models

Chengyu Fang$^{*}$, Heng Guo$^{*}$, Zheng Jiang, Chunming He, Xiu Li$^{\dagger}$, Minfeng Xu$^{\dagger}$

  • Photon is a variable-length 3D medical VQA framework with instruction-conditioned token scheduling and surrogate gradients, achieving adaptive acceleration and state-of-the-art performance.

🎖 Honors and Awards

  • 2026.06: Tsinghua University Merit Student
  • 2026.06: Outstanding Student Leader of Tsinghua University
  • 2024.06: Outstanding Graduates of Department of Computer Science and Technology, Tsinghua University
  • 2021-2023: Academic Excellence Scholarship, Tsinghua University

💬 Invited Talks

  • 2026.04, “MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning”, AI TIME.