Paper Lineage
Esc
MethodNov 2022arXiv 2211.07636cs.CV

EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Yuxin Fang, Wen Wang, Binhui Xie and 6 others

We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches.

From the abstract

Built on

16 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (16)

  • BEiT-32022 · cited 11×
    “However, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision bao2021beit; beitv2; beit3, and (ii) good performances could be also achieved via…”
    From this paper · §Introduction
  • BEiT2021 · cited 10×
    “Recently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”
    From this paper · §Introduction
  • BEiT v22022 · cited 10×
    “However, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision bao2021beit; beitv2; beit3, and (ii) good performances could be also achieved via…”
    From this paper · §Introduction
  • ViT2020 · cited 7×
    “However, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”
    From this paper · §Introduction
Show 12 more
  • Swin V22021 · cited 7×
    “However, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”
    From this paper · §Introduction
  • LVIS2019 · cited 6×
    “Using 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K deng2009imagenet (89.7% top-1 accuracy…”
    From this paper · §Introduction
  • CLIP2021 · cited 5×
    “Recently, there are a few trials leveraging the semantic information from image-image or image-text contrastive learning caron2021emerging; chen2021mocov3; clip for MIM pre-training zhou2021ibot; wei2022mvp; hou2022milan…”
    From this paper · §Introduction
  • Scaling ViTs (ViT-G)2021 · cited 5×
    “However, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”
    From this paper · §Introduction
  • MS COCO2014 · cited 4×
    “Using 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K deng2009imagenet (89.7% top-1 accuracy…”
    From this paper · §Introduction
  • MAE2021 · cited 4×
    “Recently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”
    From this paper · §Introduction
  • ResNet2015 · cited 2×
    “Notably, the ImageNet-1K validation zero-shot top-1 accuracy is 78.2% without using any of its training set labels, matching the original ResNet-101 resnet.”
    From this paper · §Fly EVA to the Moon
  • ImageNetV22019 · cited 2×
    “We also evaluate the robustness & generalization capability of EVA along with our training settings & hyper-parameters using ImageNet-V2 matched frequency (IN-V2) inv2, ImageNet-ReaL (IN-ReaL) inreal, ImageNet-Adversaria…”
    From this paper · §Fly EVA to the Moon
  • T52019 · cited 2×
    “Scaling up pre-trained language models (PLMs) liu2019roberta; gpt3; t5 has revolutionized natural language processing (NLP) in the past few years.”
    From this paper · §Introduction
  • GPT-32020 · cited 2×
    “Scaling up pre-trained language models (PLMs) liu2019roberta; gpt3; t5 has revolutionized natural language processing (NLP) in the past few years.”
    From this paper · §Introduction
  • LAION-400M2021 · cited 2×
    “Meanwhile, these CLIP features is also widely used in other state-of-the-art representation learning & pre-training works such as the BEiT family beitv2; beit3, AI generated content dalle2; imagen; stablediffusion and la…”
    From this paper · §Fly EVA to the Moon
  • Emergent abilities2022 · cited 2×
    “With further scaling on compute, data, and model sizes, PLMs have led to not only continuous performance improvements t5; kaplan2020scalingLM; rae2021gopher, but also a surprising emergence of in-context learning capabil…”
    From this paper · §Introduction

Led to

  • BLIP-22023 · cited 1×, 1 in Method
    “For the frozen image encoder, we explore two state-of-the-art pre-trained vision transformer models: (1) ViT-L/14 from CLIP (Radford et al. 2021) and (2) ViT-g/14 from EVA-CLIP (Fang et al. 2022).”
    From BLIP-2 · §Method
Abstract

We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches. Via this pretext task, we can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks, such as image recognition, video action recognition, object detection, instance segmentation and semantic segmentation without heavy supervised training. Moreover, we observe quantitative changes in scaling EVA result in qualitative changes in transfer learning performance that are not present in other models. For instance, EVA takes a great leap in the challenging large vocabulary instance segmentation task: our model achieves almost the same state-of-the-art performance on LVISv1.0 dataset with over a thousand categories and COCO dataset with only eighty categories. Beyond a pure vision encoder, EVA can also serve as a vision-centric, multi-modal pivot to connect images and text. We find initializing the vision tower of a giant CLIP from EVA can greatly stabilize the training and outperform the training from scratch counterpart with much fewer samples and less compute, providing a new direction for scaling up and accelerating the costly training of multi-modal foundation models. To facilitate future research, we release all the code and models at https://github.com/baaivision/EVA.