Paper Lineage
Esc
MethodOct 2020arXiv 2010.00747cs.CV

Contrastive Learning of Medical Visual Representations from Paired Images and Text

Yuhao Zhang, Hang Jiang, Yasuhide Miura and 2 others

, X-rays) is core to medical image understanding but its progress has been held back by the scarcity of human annotations. Existing work commonly relies on fine-tuning weights transferred from ImageNet pretraining, which is suboptimal due to drastically different image characteristics, or rule-based label extraction from the textual report data paired with medical images, which is inaccurate and hard to generalize.

From the abstract

Built on

9 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (9)

  • SimCLR2020 · cited 9×, 3 in Method
    “Note that unlike previous work which use a contrastive loss between inputs of the same modality Chen et al. 2020a; He et al. 2020, our image-to-text contrastive loss is asymmetric for each input modality.”
    From this paper · §Methods
  • MoCo2019 · cited 5×, 1 in Method
    “Note that unlike previous work which use a contrastive loss between inputs of the same modality Chen et al. 2020a; He et al. 2020, our image-to-text contrastive loss is asymmetric for each input modality.”
    From this paper · §Methods
  • ResNet2015 · cited 1×, 1 in Method
    “For the image encoder fvf_{v}, we use the ResNet50 architecture He et al. 2016 for all experiments, as it is the architecture of choice for much medical imaging work and is shown to achieve competitive performance.”
    From this paper · §Methods
  • CPC2018 · cited 1×, 1 in Method
    “This loss takes the same form as the InfoNCE loss Oord et al. 2018, and minimizing it leads to encoders that maximally preserve the mutual information between the true pairs under the representation functions.”
    From this paper · §Methods
Show 5 more
  • BERT2018 · cited 1×, 1 in Method
    “For the text encoder fuf_{u}, we use a BERT encoder Devlin et al. 2019 followed by a max-pooling layer over all output vectors.”
    From this paper · §Methods
  • CPC v22019 · cited 2×
    “Our work is inspired by the recent line of work on image view-based contrastive learning Hénaff et al. 2020; Chen et al. 2020a; He et al. 2020; Grill et al. 2020; Sowrirajan et al. 2021; Azizi et al. 2021, but fundamenta…”
    From this paper · §Related Work
  • LXMERT2019 · cited 2×
    “Contrastive-Binary-Loss: This baseline differs from ConVIRT by contrasting the paired image and text representations with a binary classification head, as is widely done in visual-linguistic pretraining work Tan and Bans…”
    From this paper · §Experiments
  • VL-BERT2019 · cited 2×
    “Contrastive-Binary-Loss: This baseline differs from ConVIRT by contrasting the paired image and text representations with a binary classification head, as is widely done in visual-linguistic pretraining work Tan and Bans…”
    From this paper · §Experiments
  • BYOL2020 · cited 2×
    “Our work is inspired by the recent line of work on image view-based contrastive learning Hénaff et al. 2020; Chen et al. 2020a; He et al. 2020; Grill et al. 2020; Sowrirajan et al. 2021; Azizi et al. 2021, but fundamenta…”
    From this paper · §Related Work

Led to

  • CLIP2021 · cited 5×, 4 in Method
    “Zhang et al. 2020, Gomez et al. 2017, Joulin et al. 2016, and Desai & Johnson 2020 all introduce methods which learn visual representations from text paired with images but describe their approaches as unsupervised, self…”
    From CLIP · §Approach
  • LiT2021 · cited 2×, 1 in Method
    “Another approach, which we adopt in this work, is to learn an alignment between image and text embedding spaces devise; vse; karpathy; vse-pooling; virtex; convirt.”
    From LiT · §Related work
Abstract

Learning visual representations of medical images (e.g., X-rays) is core to medical image understanding but its progress has been held back by the scarcity of human annotations. Existing work commonly relies on fine-tuning weights transferred from ImageNet pretraining, which is suboptimal due to drastically different image characteristics, or rule-based label extraction from the textual report data paired with medical images, which is inaccurate and hard to generalize. Meanwhile, several recent studies show exciting results from unsupervised contrastive learning from natural images, but we find these methods help little on medical images because of their high inter-class similarity. We propose ConVIRT, an alternative unsupervised strategy to learn medical visual representations by exploiting naturally occurring paired descriptive text. Our new method of pretraining medical image encoders with the paired text data via a bidirectional contrastive objective between the two modalities is domain-agnostic, and requires no additional expert input. We test ConVIRT by transferring our pretrained weights to 4 medical image classification tasks and 2 zero-shot retrieval tasks, and show that it leads to image representations that considerably outperform strong baselines in most settings. Notably, in all 4 classification tasks, our method requires only 10\% as much labeled training data as an ImageNet initialized counterpart to achieve better or comparable performance, demonstrating superior data efficiency.