Research & Publications

MICCAI 2023

MedIM: Radiology report-guided masking for medical foundation models

Yutong Xie, Lin Gu, Tatsuya Harada, Jianpeng Zhang, Yong Xia, Qi Wu

2023

Medical image representation learning guided by radiology reports

Publication summary

MedIM introduces a novel approach for self-supervised medical image representation learning by incorporating clinical knowledge from radiology reports into the masking process used during masked image modeling.

Instead of randomly masking image regions, MedIM identifies clinically meaningful areas using report-derived knowledge and prioritizes them during pretraining. This encourages the model to learn richer semantic representations aligned with diagnostic reasoning.

The framework demonstrates strong transferability across classification and segmentation tasks, improving representation quality while reducing dependence on large labeled medical datasets.

“MedIM bridges the gap between medical images and clinical language by allowing radiology reports to guide self-supervised representation learning.”

Editorial research summary

OrthoAI research content adaptation

Key findings

Report-guided

Uses radiology reports to identify clinically important image regions during masked image modeling.

Multi-task

Validated across medical image classification and segmentation benchmarks.

Foundation Model

Produces transferable representations for downstream medical AI applications.

Method

Knowledge word-driven masking

MedIM extracts clinically relevant concepts from radiology reports and uses them to guide masking decisions during self-supervised pretraining.

This strategy encourages the model to reconstruct diagnostically meaningful image regions rather than relying solely on random masking.

Knowledge word-driven masking using radiology reports

Report-image aligned representation learning

The framework leverages complementary information from radiology reports and medical images to learn semantically rich feature representations.

Experimental results show improved reconstruction quality and stronger downstream performance compared with conventional masking approaches.

Radiology report-guided representation learning pipeline

Key contributions

MedIM enhances self-supervised medical image learning by integrating textual clinical knowledge into image pretraining.

  • Introduces radiology report-guided masking for medical image pretraining.
  • Aligns image reconstruction objectives with clinically relevant semantics.
  • Improves representation quality beyond random masking strategies.
  • Supports downstream classification and segmentation tasks.
  • Demonstrates strong transfer learning potential for medical foundation models.
  • Bridges medical vision and clinical language understanding.

Similar publications

MICCAI 2021

UniMiSS: Universal Medical Self-Supervised Learning via Breaking Dimensionality Barrier

Yutong Xie, Jianpeng Zhang, Yong Xia, Qi Wu

view

CVPR 2024

Continual Self-supervised Learning: Towards Universal Multi-modal Medical Data Representation Learning

Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia

view