Paper Lineage
Esc
MethodMay 2018arXiv 1805.01978cs.CV

Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination

Zhirong Wu, Yuanjun Xiong, Stella Yu, Dahua Lin

Neural net classifiers trained on data with annotated class labels can also capture apparent visual similarity among categories without being directed to do so. We study whether this observation can be extended beyond the conventional domain of supervised learning: Can we learn a good feature representation that captures apparent similarity among instances, instead of classes, by merely asking the feature to be discriminative of individual instances?

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • Knowledge Distillation2015 · cited 2×, 1 in Method
    “Fig. 1 shows that an image from class leopard is rated much higher by class jaguar rather than by class bookcase hinton2015distilling.”
    From this paper · §Introduction
  • VGG2014 · cited 2×
    “As the network architecture has a big impact on the performance, we consider a few typical architectures: AlexNet krizhevsky2012imagenet, VGG16 simonyan2014very, ResNet-18, and ResNet-50 he2015deep.”
    From this paper · §Experiments
  • ResNet2015 · cited 2×
    “As the network architecture has a big impact on the performance, we consider a few typical architectures: AlexNet krizhevsky2012imagenet, VGG16 simonyan2014very, ResNet-18, and ResNet-50 he2015deep.”
    From this paper · §Experiments

Led to

  • Invariant & spreading instance features2019 · cited 20×, 2 in Method
    “To improve the inferior efficiency, Wu et al. cvpr18nce propose to set up a memory bank to store the instance features 𝐟if_{i} calculated in the previous step.”
    From Invariant & spreading instance features · §Proposed Method
  • MoCo2019 · cited 23×, 11 in Method
    “In this paper, we follow a simple instance discrimination task Wu2018a; Ye2019; Bachman2019: a query matches a key if they are encoded views (e.g., different crops) of the same image.”
    From MoCo · §Introduction
  • SimCLR2020 · cited 6×, 3 in Method
    “To keep it simple, we do not train the model with a memory bank (Wu et al. 2018; He et al. 2019).”
    From SimCLR · §Method
  • CLIP2021 · cited 1×, 1 in Method
    “The learnable temperature parameter τ\tau was initialized to the equivalent of 0.07 from (Wu et al. 2018) and clipped to prevent scaling the logits by more than 100 which we found necessary to prevent training instabilit…”
    From CLIP · §Approach
  • BEiT2021 · cited 3×, 1 in Method
    “Our augmentation policy includes random resized cropping, horizontal flipping, color jittering [43].”
    From BEiT · §Methods
  • MAE2021 · cited 2×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From MAE · §Related Work
  • BEiT-32022 · cited 1×, 1 in Method
    “We use the same image augmentation as in BEiT [3], including random resized cropping, horizontal flipping, and color jittering [56].”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

Neural net classifiers trained on data with annotated class labels can also capture apparent visual similarity among categories without being directed to do so. We study whether this observation can be extended beyond the conventional domain of supervised learning: Can we learn a good feature representation that captures apparent similarity among instances, instead of classes, by merely asking the feature to be discriminative of individual instances? We formulate this intuition as a non-parametric classification problem at the instance-level, and use noise-contrastive estimation to tackle the computational challenges imposed by the large number of instance classes. Our experimental results demonstrate that, under unsupervised learning settings, our method surpasses the state-of-the-art on ImageNet classification by a large margin. Our method is also remarkable for consistently improving test performance with more training data and better network architectures. By fine-tuning the learned feature, we further obtain competitive results for semi-supervised learning and object detection tasks. Our non-parametric model is highly compact: With 128 features per image, our method requires only 600MB storage for a million images, enabling fast nearest neighbour retrieval at the run time.