Currently, I am an
Associate Researcher at the
Suzhou Institute for Advanced Research, University of Science and Technology of China (USTC),
and a member of the
MIRACLE Lab,
led by Prof.
S. Kevin Zhou
(IEEE Fellow). I have held this position since August 2026.
Prior to this appointment, I was a
Postdoctoral Researcher at USTC
from 2024 to July 2026, working with Prof. S. Kevin Zhou and Prof.
Houqiang Li
(IEEE Fellow).
I received my Ph.D. degree from the Department of Electronic Engineering and Information Science
at USTC in 2024, under the supervision of Prof.
Yongdong Zhang
(IEEE Fellow) and Prof.
Zhendong Mao
.
From 2018 to 2020, I studied in the Department of Automation at USTC, advised by Prof.
Shuang Cong
.
I received my B.Eng. degree from the School of Internet of Things Engineering
at Jiangnan University in 2018.
My research interests broadly lie in the areas of Multimodal Artificial Intelligence and Deep Learning (e.g., vision-language alignment, cross-modal retrieval, report generation, retrieval augmented generation, hallucination evaluation, etc.). I am recently interested in multimodal large language models in medical scenarios, including explainable disease diagnosis, LLM-based clinical decision making, medical text processing, pathology, MRI image processing, etc.
(1) My research in vision-language alignment and retrieval focuses on learning fine-grained, reliable, and interpretable cross-modal correspondences between visual content and natural language. The studied scenarios cover both natural images and medical images, ranging from general image-text matching, cross-modal retrieval, and semantic alignment to medical report generation, pathology image analysis, and retrieval-augmented medical decision support.
Recently, we have been completing a review: Composed Multi-modal Retrieval: A Survey of Approaches and Applications. For more details, please see
CMR page.
This repo is used for recording and tracking recent Composed Multi-modal Retrieval (CMR) works, including Composed Image Retrieval (CIR), Composed Video Retrieval (CVR), Composed Person Retrieval (CPR), etc.
The survey can be found
here.
(2) Explainable artificial intelligence is an important research direction, and the Concept Bottleneck Model (CBM) is a current promising research paradigm.
CBMs typically involve a layer preceding the final fully connected classifier, where each neuron corresponds to a concept that can be interpreted by humans. CBMs also show advantages in improving accuracy through human intervention during testing.
We are maintaining a GitHub repository
CBM page, aiming to keep pace with its rapidly evolving. The survey can be found
Concept Bottleneck Models for Explainable Decision Making: A Survey of Progress, Taxonomy and Future Directions.
(3) I am also actively exploring
LLM- and Agent-based medical intelligence, with a particular focus on adapting large language models to downstream clinical applications. Representative directions include clinical decision support, auxiliary diagnosis, medical report generation, ICD coding, medical text processing, hallucination evaluation, and retrieval-augmented clinical reasoning. More recently, I have been interested in advanced agentic systems that integrate domain-specific skills, self-learned knowledge, tool use, retrieval modules, and multi-agent collaboration, aiming to build more reliable, interpretable, and clinically useful AI systems for real-world medical scenarios.
🔥 We are recruiting!
We are always looking for highly motivated undergraduate and graduate students,
particularly those interested in pursuing a Ph.D., to join us in exploring
research topics in medical imaging and medical artificial intelligence.
If you are interested in working with us, please feel free to contact me via email.
For outstanding students interested in pursuing careers in industry, I am also happy to recommend
strong candidates for internship opportunities at leading technology companies,
such as Huawei, ByteDance, Alibaba, Ant Group, and Tencent.