113,970
Joint self-supervised pretraining performed using 5,022 CT volumes and 108,948 X-ray images.
Research & Publications
MICCAI 2021
Yutong Xie, Jianpeng Zhang, Yong Xia, Qi Wu
2021
Yutong Xie, Jianpeng Zhang, Yong Xia, Qi Wu
2021

UniMiSS addresses a fundamental limitation of medical self-supervised learning: the scarcity of large-scale 3D datasets. The framework leverages abundant 2D medical images, such as chest X-rays, to complement limited 3D CT volumes during pretraining.
To bridge the dimensionality gap between 2D and 3D data, the authors introduce a dimension-free Medical Transformer (MiT) with a Switchable Patch Embedding (SPE) module that dynamically adapts to either 2D or 3D inputs.
Through joint self-supervised learning on 5,022 CT volumes and 108,948 X-ray images, UniMiSS learns transferable representations that improve performance across segmentation and classification tasks in both 2D and 3D medical imaging.
“UniMiSS demonstrates that medical foundation representations can be learned across imaging dimensions, allowing 2D and 3D modalities to reinforce each other during self-supervised training.”
Editorial research summary
OrthoAI research content adaptation
Joint self-supervised pretraining performed using 5,022 CT volumes and 108,948 X-ray images.
Classification improvement over DINO on the RICORD COVID-19 screening benchmark with full labels.
Best BCV online benchmark Dice score achieved using the UniMiSS ensemble configuration.
UniMiSS introduces the Medical Transformer (MiT), a pyramid U-shaped Transformer architecture capable of processing both 2D and 3D medical images.
The Switchable Patch Embedding module dynamically chooses either 2D or 3D patch embedding based on the incoming data, enabling a unified representation space.

The framework employs a student-teacher self-distillation strategy and alternates training between 2D and 3D datasets.
A volume-slice consistency objective aligns volumetric CT representations with their corresponding 2D slices, strengthening cross-dimensional representation learning.

UniMiSS establishes a universal medical self-supervised learning framework capable of bridging dimensionality barriers in medical imaging.

CVPR 2021 · IEEE/CVF Conference on Computer Vision and Pattern Recognition
Jianpeng Zhang, Yutong Xie, Yong Xia, Chunhua Shen

CVPR 2024 · IEEE/CVF Conference on Computer Vision and Pattern Recognition
Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia