Paper Lineage
Esc
MethodJun 2021arXiv 2106.04803cs.CV

CoAtNet: Marrying Convolution and Attention for All Data Sizes

Zihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing Tan

Transformers have attracted increasing interests in computer vision, but they still fall behind state-of-the-art convolutional networks. In this work, we show that while Transformers tend to have larger model capacity, their generalization can be worse than convolutional networks due to the lack of the right inductive bias.

From the abstract

Built on

7 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (7)

  • Swin2021 · cited 6×, 2 in Method
    “Enforce local attention, which restricts the global receptive field 𝒢G in attention to a local field ℒL just like in convolution [22, 21].”
    From this paper · §Model
  • MobileNetV22018 · cited 3×, 1 in Method
    “Traditionally, regular convolutions, such as ResNet blocks [3], are popular in large-scale ConvNets; in contrast, depthwise convolutions [28] are popular in mobile platforms due to its lower computational cost and smalle…”
    From this paper · §Related Work
  • T52019 · cited 3×, 1 in Method
    “Interestingly, while the idea seems overly simplified, the pre-normalization version yprey^{pre} corresponds to a particular variant of relative self-attention [30, 31].”
    From this paper · §Model
  • ViT2020 · cited 4×
    “In comparison, after the success of ViT and ResNet-ViT [13], another popular line of research starts with a Transformer backbone and tries to incorporate explicit convolution or some desirable properties of convolution i…”
    From this paper · §Related Work
Show 3 more
  • EfficientNet2019 · cited 3×
    “Recent works show that an improved inverted residual bottlenecks (MBConv [27, 35]), which is built upon depthwise convolutions, can achieve both high accuracy and better efficiency [5, 19].”
    From this paper · §Related Work
  • ResNet2015 · cited 2×
    “Traditionally, regular convolutions, such as ResNet blocks [3], are popular in large-scale ConvNets; in contrast, depthwise convolutions [28] are popular in mobile platforms due to its lower computational cost and smalle…”
    From this paper · §Related Work
  • Scaling ViTs (ViT-G)2021 · cited 2×
    “Finally, when JFT-3B is used for pre-training, CoAtNet exhibits better efficiency compared to ViT, and pushes the ImageNet-1K top-1 accuracy to 90.88% while using 1.5x less computation of the prior art set by ViT-G/14 [2…”
    From this paper · §Introduction

Led to

  • SimVLM2021 · cited 5×
    “Following Dai et al. 2021, we experiment with using either the first 2/3/4 ResNet Conv blocks, and empirically observe that the 3 conv block setup works best.”
    From SimVLM · §Experiments
  • LiT2021 · cited 2×
    “With the pre-trained model ViT-g/14 vitg, LiT achieves 85.2% zero-shot transfer accuracy on ImageNet, halving the gap between previous best zero-shot transfer results clip; align and supervised fine-tuning results vitg;…”
    From LiT · §Introduction
  • Swin V22021 · cited 4×
    “While it has long been recognized that larger vision models usually perform better on vision tasks simonyan2014vgg; he2015resnet, the absolute model size was just able to reach about 1-2 billion parameters very recently…”
    From Swin V2 · §Introduction
Abstract

Transformers have attracted increasing interests in computer vision, but they still fall behind state-of-the-art convolutional networks. In this work, we show that while Transformers tend to have larger model capacity, their generalization can be worse than convolutional networks due to the lack of the right inductive bias. To effectively combine the strengths from both architectures, we present CoAtNets(pronounced "coat" nets), a family of hybrid models built from two key insights: (1) depthwise Convolution and self-Attention can be naturally unified via simple relative attention; (2) vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency. Experiments show that our CoAtNets achieve state-of-the-art performance under different resource constraints across various datasets: Without extra data, CoAtNet achieves 86.0% ImageNet top-1 accuracy; When pre-trained with 13M images from ImageNet-21K, our CoAtNet achieves 88.56% top-1 accuracy, matching ViT-huge pre-trained with 300M images from JFT-300M while using 23x less data; Notably, when we further scale up CoAtNet with JFT-3B, it achieves 90.88% top-1 accuracy on ImageNet, establishing a new state-of-the-art result.