Paper Lineage
Esc
MethodFeb 2021arXiv 2102.12092cs.CV

Zero-Shot Text-to-Image Generation

Aditya Ramesh, Mikhail Pavlov, Gabriel Goh and 5 others

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • VQ-VAE2017 · cited 2×, 2 in Method
    “We address these issues by using a two-stage training procedure, similar to (Oord et al. 2017; Razavi et al. 2019):”
    From this paper · §Method
  • Transformer2017 · cited 2×, 1 in Method
    “Our goal is to train a transformer (Vaswani et al. 2017) to autoregressively model the text and image tokens as a single stream of data.”
    From this paper · §Method
  • CLIP2021 · cited 2×, 1 in Method
    “Similar to Razavi et al. 2019, we rerank the samples drawn from the transformer using a pretrained contrastive model (Radford et al. 2021).”
    From this paper · §Method
  • MS COCO2014 · cited 1×, 1 in Method
    “Our preliminary experiments for models up to 1.21.2 billion parameters were carried out on Conceptual Captions, a dataset of 3.3 million text-image pairs that was developed as an extension to MS-COCO (Lin et al. 2014).”
    From this paper · §Method
Show 2 more
  • YFCC100M2015 · cited 1×, 1 in Method
    “This dataset does not include MS-COCO, but does include Conceptual Captions and a filtered subset of YFCC100M (Thomee et al. 2016).”
    From this paper · §Method
  • JFT-300M (unreasonable effectiveness)2017 · cited 1×, 1 in Method
    “To scale up to 1212-billion parameters, we created a dataset of a similar scale to JFT-300M (Sun et al. 2017) by collecting 250 million text-images pairs from the internet.”
    From this paper · §Method

Led to

  • BEiT2021 · cited 8×, 6 in Method
    “Following [34], we use the image tokenizer learned by discrete variational autoencoder (dVAE).”
    From BEiT · §Methods
  • Codex2021 · cited 2×
    “Further, just as text-conditional generative models in other modalities (Ramesh et al. 2021) have difficulty with binding attributes to objects, Codex can make mistakes binding operations to variables, especially when th…”
    From Codex · §Limitations
  • SimVLM2021 · cited 2×
    “Motivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”
    From SimVLM · §Related Work
  • ViT-VQGAN2021 · cited 3×
    “DALL-E (Ramesh et al. 2021) improves token prediction in second stage by using Transformers (Vaswani et al. 2017), resulting in a strong text-to-image synthesis model.”
    From ViT-VQGAN · §Related Work
  • LAION-400M2021 · cited 3×
    “Multi-modal language-vision models demonstrated recently strong transfer capability to novel datasets in absense of per-sample labels [1, 2, 3].”
    From LAION-400M · §Introduction
  • METER2021 · cited 1×, 1 in Method
    “Specifically, we first use the VQ-VAE van2017neural model in DALL-E ramesh2021zero to tokenize each image into a sequence of discrete tokens.”
    From METER · §The Meter Framework
  • MAE2021 · cited 3×
    “Specifically for this variant, we use the DALLE pre-trained dVAE Ramesh2021 as the tokenizer, following Bao2021.”
    From MAE · §ImageNet Experiments
  • Florence2021 · cited 1×, 1 in Method
    “In addition, we follow the sampling strategy introduced in (Radford et al. 2021; Ramesh et al. 2021) with the goal of achieving improved balance, informativeness, and learnability of the sampled dataset.”
    From Florence · §Approach
  • BEiT v22022 · cited 2×
    “DALL-E (Ramesh et al. 2021) uses the Gumbel-softmax relaxation for quantization instead of the nearest neighbor lookup in VQ-VAE.”
    From BEiT v2 · §Related Work
Abstract

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.