Paper Lineage
Esc
MethodNov 2021arXiv 2111.11432cs.CV

Florence: A New Foundation Model for Computer Vision

Lu Yuan, Dongdong Chen, Yi-Ling Chen and 20 others

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications.

From the abstract

Built on

21 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (21)

  • CLIP2021 · cited 16×, 3 in Method
    “In addition, we follow the sampling strategy introduced in (Radford et al. 2021; Ramesh et al. 2021) with the goal of achieving improved balance, informativeness, and learnability of the sampled dataset.”
    From this paper · §Approach
  • METER2021 · cited 5×, 3 in Method
    “We use METER (Dou et al. 2021) adapter to expand to fine-grained vision-language representation.”
    From this paper · §Approach
  • Video Swin2021 · cited 4×, 3 in Method
    “Our Video CoSwin adapter can borrow the image encoder from CoSwin for the video domain with minimum changes, similar to prior work (Liu et al. 2021b).”
    From this paper · §Approach
  • Dynamic Head2021 · cited 5×, 2 in Method
    “For this goal, we add an adaptor Dynamic Head (Dai et al. 2021a) (or Dynamic DETR (Dai et al. 2021b)), a unified attention mechanism for the detection head, to the pretrained image encoder (i.e. , CoSwin).”
    From this paper · §Approach
Show 17 more
  • Swin2021 · cited 3×, 2 in Method
    “For the image encoder, we chose hierarchical Vision Transformers (e.g. , Swin (Liu et al. 2021a), CvT (Wu et al. 2021), Vision Longformer (Zhang et al. 2021a), Focal Transformer (Yang et al. 2021), and CSwin (Dong et al.…”
    From this paper · §Introduction
  • ALIGN2021 · cited 8×, 1 in Method
    “To improve data quality, we performed rigorous data filtering, similar to ALIGN (Jia et al. 2021), including a simple hash-based near-duplicate image removal, small-size image removal, image-text relevance, etc.”
    From this paper · §Approach
  • VQA v22016 · cited 2×, 1 in Method
    “To evaluate the performance, we fine-tune the pre-trained model on the challenging VQA (Goyal et al. 2017) task, which is to answer a question based on the image context.”
    From this paper · §Experiments
  • SimVLM2021 · cited 2×, 1 in Method
    “Compared with SimVLM (Wang et al. 2021), which uses 1.81.8B image-text pairs, we only use 900900M data to pre-train the image encoder and 2020M for VLP, but achieve better results.”
    From this paper · §Experiments
  • Transformer2017 · cited 1×, 1 in Method
    “Our Florence pretrained model uses a two-tower architecture: a 12-layer transformer (Vaswani et al. 2017) as language encoder, similar to CLIP (Radford et al. 2021), and a hierarchical Vision Transformer as the image enc…”
    From this paper · §Approach
  • Bottom-Up Top-Down attention2017 · cited 1×, 1 in Method
    “Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”
    From this paper · §Approach
  • RoBERTa2019 · cited 1×, 1 in Method
    “In the Florence V+L adaptation model, we replace the image encoder of METER (Dou et al. 2021) with Florence pretrained model CoSwin, and use a pretrained Roberta (Liu et al. 2019) as the language encoder, shown in Figure…”
    From this paper · §Approach
  • LVIS2019 · cited 1×, 1 in Method
    “We merge several well-known object detection datasets, including COCO (Lin et al. 2015), LVIS (Gupta et al. 2019), OpenImages (Krasin et al. 2016), Object365 (Shao et al. 2019).”
    From this paper · §Approach
  • UNITER2019 · cited 1×, 1 in Method
    “Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”
    From this paper · §Approach
  • Oscar2020 · cited 1×, 1 in Method
    “Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”
    From this paper · §Approach
  • VinVL2021 · cited 1×, 1 in Method
    “Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”
    From this paper · §Approach
  • ViLT2021 · cited 1×, 1 in Method
    “Recently, there is an increasing trend (Huang et al. 2021; Xue et al. 2021; Wang et al. 2021; Kim et al. 2021; Dou et al. 2021) of end-to-end approaches to reduce dependency on the object bounding box, which instead cons…”
    From this paper · §Approach
  • DALL·E2021 · cited 1×, 1 in Method
    “In addition, we follow the sampling strategy introduced in (Radford et al. 2021; Ramesh et al. 2021) with the goal of achieving improved balance, informativeness, and learnability of the sampled dataset.”
    From this paper · §Approach
  • Noisy Student2019 · cited 3×
    “Linear probe as another main metric for evaluating representation quality has been used in most recent studies, including self-supervised learning (Chen et al. 2020b; Chen et al. 2020c), self-training with noisy student…”
    From this paper · §Experiments
  • Visual Genome2016 · cited 2×
    “We evaluate fine-tuning on three popular object detection datasets: COCO (Lin et al. 2015), Object365 (Shao et al. 2019), and Visual Genome (Krishna et al. 2016).”
    From this paper · §Experiments
  • “We evaluate our Florence model on the ImageNet-1K dataset and 11 downstream datasets from the well-studied evaluation suit introduced by (Kornblith et al. 2019).”
    From this paper · §Experiments
  • ViT2020 · cited 2×
    “While inheriting performance benefits of the transformer self-attention operations (Dosovitskiy et al. 2021b), these hierarchical architectures model the scale invariance nature of images and have linear computational co…”
    From this paper · §Introduction

Led to

  • VLMo2021 · cited 2×
    “Our large-size model even outperforms SimVLM-Huge [46] and Florence-Huge [48] by a large margin, which consists of more parameters and are also trained on larger-scale image-text pairs.”
    From VLMo · §Experiments
  • UniCL2022 · cited 2×
    “Finally, we scaled up UniCL to billions of image-text-label data in Florence yuan2021florence and demonstrated its superiority over CLIP radford2021learning and ALIGN jia2021scaling across dozens of benchmarks.”
    From UniCL · §Introduction
  • CoCa2022 · cited 6×
    “However, these models rely heavily on image annotations as labeled vectors and do not bake in knowledge of free-form human natural language, hindering their application to downstream tasks that involving both vision and…”
    From CoCa · §Introduction
  • BEiT-32022 · cited 3×, 2 in Method
    “In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications. While existing vision foundation models such as CLIP, ALIGN, and Wu Dao 2.0 focus mainly on mapping images and textual representations to a cross-modal shared representation, we introduce a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine (object), from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth). By incorporating universal visual-language representations from Web-scale image-text data, our Florence model can be easily adapted for various computer vision tasks, such as classification, retrieval, object detection, VQA, image caption, video retrieval and action recognition. Moreover, Florence demonstrates outstanding performance in many types of transfer learning: fully sampled fine-tuning, linear probing, few-shot transfer and zero-shot transfer for novel images and objects. All of these properties are critical for our vision foundation model to serve general purpose vision tasks. Florence achieves new state-of-the-art results in majority of 44 representative benchmarks, e.g., ImageNet-1K zero-shot classification with top-1 accuracy of 83.74 and the top-5 accuracy of 97.18, 62.4 mAP on COCO fine tuning, 80.36 on VQA, and 87.8 on Kinetics-600.