EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches.
Also cited · not yet reviewed (16)
- BEiT-32022 · cited 11דHowever, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision bao2021beit; beitv2; beit3, and (ii) good performances could be also achieved via…”From this paper · §Introduction
- BEiT2021 · cited 10דRecently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”From this paper · §Introduction
- BEiT v22022 · cited 10דHowever, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision bao2021beit; beitv2; beit3, and (ii) good performances could be also achieved via…”From this paper · §Introduction
- ViT2020 · cited 7דHowever, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”From this paper · §Introduction
Show 12 more
- Swin V22021 · cited 7דHowever, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”From this paper · §Introduction
- LVIS2019 · cited 6דUsing 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K deng2009imagenet (89.7% top-1 accuracy…”From this paper · §Introduction
- CLIP2021 · cited 5דRecently, there are a few trials leveraging the semantic information from image-image or image-text contrastive learning caron2021emerging; chen2021mocov3; clip for MIM pre-training zhou2021ibot; wei2022mvp; hou2022milan…”From this paper · §Introduction
- Scaling ViTs (ViT-G)2021 · cited 5דHowever, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”From this paper · §Introduction
- MS COCO2014 · cited 4דUsing 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K deng2009imagenet (89.7% top-1 accuracy…”From this paper · §Introduction
- MAE2021 · cited 4דRecently, masked image modeling (MIM) bao2021beit; xie2021simmim; he2021masked has boomed as a viable approach for vision model pre-training and scaling.”From this paper · §Introduction
- ResNet2015 · cited 2דNotably, the ImageNet-1K validation zero-shot top-1 accuracy is 78.2% without using any of its training set labels, matching the original ResNet-101 resnet.”From this paper · §Fly EVA to the Moon
- ImageNetV22019 · cited 2דWe also evaluate the robustness & generalization capability of EVA along with our training settings & hyper-parameters using ImageNet-V2 matched frequency (IN-V2) inv2, ImageNet-ReaL (IN-ReaL) inreal, ImageNet-Adversaria…”From this paper · §Fly EVA to the Moon
- T52019 · cited 2דScaling up pre-trained language models (PLMs) liu2019roberta; gpt3; t5 has revolutionized natural language processing (NLP) in the past few years.”From this paper · §Introduction
- GPT-32020 · cited 2דScaling up pre-trained language models (PLMs) liu2019roberta; gpt3; t5 has revolutionized natural language processing (NLP) in the past few years.”From this paper · §Introduction
- LAION-400M2021 · cited 2דMeanwhile, these CLIP features is also widely used in other state-of-the-art representation learning & pre-training works such as the BEiT family beitv2; beit3, AI generated content dalle2; imagen; stablediffusion and la…”From this paper · §Fly EVA to the Moon
- Emergent abilities2022 · cited 2דWith further scaling on compute, data, and model sizes, PLMs have led to not only continuous performance improvements t5; kaplan2020scalingLM; rae2021gopher, but also a surprising emergence of in-context learning capabil…”From this paper · §Introduction
Led to
- BLIP-22023 · cited 1×, 1 in Method“For the frozen image encoder, we explore two state-of-the-art pre-trained vision transformer models: (1) ViT-L/14 from CLIP (Radford et al. 2021) and (2) ViT-g/14 from EVA-CLIP (Fang et al. 2022).”From BLIP-2 · §Method
Abstract
We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches. Via this pretext task, we can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks, such as image recognition, video action recognition, object detection, instance segmentation and semantic segmentation without heavy supervised training. Moreover, we observe quantitative changes in scaling EVA result in qualitative changes in transfer learning performance that are not present in other models. For instance, EVA takes a great leap in the challenging large vocabulary instance segmentation task: our model achieves almost the same state-of-the-art performance on LVISv1.0 dataset with over a thousand categories and COCO dataset with only eighty categories. Beyond a pure vision encoder, EVA can also serve as a vision-centric, multi-modal pivot to connect images and text. We find initializing the vision tower of a giant CLIP from EVA can greatly stabilize the training and outperform the training from scratch counterpart with much fewer samples and less compute, providing a new direction for scaling up and accelerating the costly training of multi-modal foundation models. To facilitate future research, we release all the code and models at https://github.com/baaivision/EVA.