Report-guided
Uses radiology reports to identify clinically important image regions during masked image modeling.
Research & Publications
MICCAI 2023
Yutong Xie, Lin Gu, Tatsuya Harada, Jianpeng Zhang, Yong Xia, Qi Wu
2023
Yutong Xie, Lin Gu, Tatsuya Harada, Jianpeng Zhang, Yong Xia, Qi Wu
2023

MedIM introduces a novel approach for self-supervised medical image representation learning by incorporating clinical knowledge from radiology reports into the masking process used during masked image modeling.
Instead of randomly masking image regions, MedIM identifies clinically meaningful areas using report-derived knowledge and prioritizes them during pretraining. This encourages the model to learn richer semantic representations aligned with diagnostic reasoning.
The framework demonstrates strong transferability across classification and segmentation tasks, improving representation quality while reducing dependence on large labeled medical datasets.
“MedIM bridges the gap between medical images and clinical language by allowing radiology reports to guide self-supervised representation learning.”
Editorial research summary
OrthoAI research content adaptation
Uses radiology reports to identify clinically important image regions during masked image modeling.
Validated across medical image classification and segmentation benchmarks.
Produces transferable representations for downstream medical AI applications.
MedIM extracts clinically relevant concepts from radiology reports and uses them to guide masking decisions during self-supervised pretraining.
This strategy encourages the model to reconstruct diagnostically meaningful image regions rather than relying solely on random masking.

The framework leverages complementary information from radiology reports and medical images to learn semantically rich feature representations.
Experimental results show improved reconstruction quality and stronger downstream performance compared with conventional masking approaches.

MedIM enhances self-supervised medical image learning by integrating textual clinical knowledge into image pretraining.

MICCAI 2021
Yutong Xie, Jianpeng Zhang, Yong Xia, Qi Wu

CVPR 2024
Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia