Paper Lineage
Esc
MethodOct 2021arXiv 2110.04627cs.CV

Vector-quantized Image Modeling with Improved VQGAN

Jiahui Yu, Xin Li, Jing Yu Koh and 7 others

Pretraining language models with next-token prediction on massive text corpora has delivered phenomenal zero-shot, few-shot, transfer learning and multi-tasking capabilities on both generative and discriminative language tasks. Motivated by this success, we explore a Vector-quantized Image Modeling (VIM) approach that involves pretraining a Transformer to predict rasterized image tokens autoregressively.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • VQ-VAE2017 · cited 3×
    “Unlike many autogressive methods which generate sequence directly in pixel space, VQVAE (van den Oord et al. 2017; Razavi et al. 2019) decomposes the image generation process into two stages: the first stage trains a vec…”
    From this paper · §Related Work
  • DALL·E2021 · cited 3×
    “DALL-E (Ramesh et al. 2021) improves token prediction in second stage by using Transformers (Vaswani et al. 2017), resulting in a strong text-to-image synthesis model.”
    From this paper · §Related Work
  • Transformer2017 · cited 2×
    “Driven by the effectiveness of VQVAE and progress in sequence modeling (Vaswani et al. 2017; Devlin et al. 2019), many approaches follow the two-stage paradigm.”
    From this paper · §Related Work
  • BERT2018 · cited 2×
    “Driven by the effectiveness of VQVAE and progress in sequence modeling (Vaswani et al. 2017; Devlin et al. 2019), many approaches follow the two-stage paradigm.”
    From this paper · §Related Work
Show 1 more
  • BYOL2020 · cited 2×
    “In computer vision, in contrast, most recent unsupervised or self-supervised learning research focuses on applying different random augmentations to images, with the pretraining objective to distinguish image instances (…”
    From this paper · §Introduction

Led to

  • BEiT v22022 · cited 4×, 2 in Method
    “where j∈{1,2,⋯,K}j\in\{1,2,\cdots,K\} and ℓ2\ell_{2} normalization is used for codebook lookup (Yu et al. 2021).”
    From BEiT v2 · §Methodology
Abstract

Pretraining language models with next-token prediction on massive text corpora has delivered phenomenal zero-shot, few-shot, transfer learning and multi-tasking capabilities on both generative and discriminative language tasks. Motivated by this success, we explore a Vector-quantized Image Modeling (VIM) approach that involves pretraining a Transformer to predict rasterized image tokens autoregressively. The discrete image tokens are encoded from a learned Vision-Transformer-based VQGAN (ViT-VQGAN). We first propose multiple improvements over vanilla VQGAN from architecture to codebook learning, yielding better efficiency and reconstruction fidelity. The improved ViT-VQGAN further improves vector-quantized image modeling tasks, including unconditional, class-conditioned image generation and unsupervised representation learning. When trained on ImageNet at \(256\times256\) resolution, we achieve Inception Score (IS) of 175.1 and Fr'echet Inception Distance (FID) of 4.17, a dramatic improvement over the vanilla VQGAN, which obtains 70.6 and 17.04 for IS and FID, respectively. Based on ViT-VQGAN and unsupervised pretraining, we further evaluate the pretrained Transformer by averaging intermediate features, similar to Image GPT (iGPT). This ImageNet-pretrained VIM-L significantly beats iGPT-L on linear-probe accuracy from 60.3% to 73.2% for a similar model size. VIM-L also outperforms iGPT-XL which is trained with extra web image data and larger model size.