Zero-Shot Text-to-Image Generation
Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training.
Also cited · not yet reviewed (6)
- VQ-VAE2017 · cited 2×, 2 in Method“We address these issues by using a two-stage training procedure, similar to (Oord et al. 2017; Razavi et al. 2019):”From this paper · §Method
- Transformer2017 · cited 2×, 1 in Method“Our goal is to train a transformer (Vaswani et al. 2017) to autoregressively model the text and image tokens as a single stream of data.”From this paper · §Method
- CLIP2021 · cited 2×, 1 in Method“Similar to Razavi et al. 2019, we rerank the samples drawn from the transformer using a pretrained contrastive model (Radford et al. 2021).”From this paper · §Method
- MS COCO2014 · cited 1×, 1 in Method“Our preliminary experiments for models up to 1.21.2 billion parameters were carried out on Conceptual Captions, a dataset of 3.3 million text-image pairs that was developed as an extension to MS-COCO (Lin et al. 2014).”From this paper · §Method
Show 2 more
- YFCC100M2015 · cited 1×, 1 in Method“This dataset does not include MS-COCO, but does include Conceptual Captions and a filtered subset of YFCC100M (Thomee et al. 2016).”From this paper · §Method
- JFT-300M (unreasonable effectiveness)2017 · cited 1×, 1 in Method“To scale up to 1212-billion parameters, we created a dataset of a similar scale to JFT-300M (Sun et al. 2017) by collecting 250 million text-images pairs from the internet.”From this paper · §Method
Led to
- BEiT2021 · cited 8×, 6 in Method“Following [34], we use the image tokenizer learned by discrete variational autoencoder (dVAE).”From BEiT · §Methods
- Codex2021 · cited 2דFurther, just as text-conditional generative models in other modalities (Ramesh et al. 2021) have difficulty with binding attributes to objects, Codex can make mistakes binding operations to variables, especially when th…”From Codex · §Limitations
- SimVLM2021 · cited 2דMotivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”From SimVLM · §Related Work
- ViT-VQGAN2021 · cited 3דDALL-E (Ramesh et al. 2021) improves token prediction in second stage by using Transformers (Vaswani et al. 2017), resulting in a strong text-to-image synthesis model.”From ViT-VQGAN · §Related Work
- LAION-400M2021 · cited 3דMulti-modal language-vision models demonstrated recently strong transfer capability to novel datasets in absense of per-sample labels [1, 2, 3].”From LAION-400M · §Introduction
- METER2021 · cited 1×, 1 in Method“Specifically, we first use the VQ-VAE van2017neural model in DALL-E ramesh2021zero to tokenize each image into a sequence of discrete tokens.”From METER · §The Meter Framework
- MAE2021 · cited 3דSpecifically for this variant, we use the DALLE pre-trained dVAE Ramesh2021 as the tokenizer, following Bao2021.”From MAE · §ImageNet Experiments
- Florence2021 · cited 1×, 1 in Method“In addition, we follow the sampling strategy introduced in (Radford et al. 2021; Ramesh et al. 2021) with the goal of achieving improved balance, informativeness, and learnability of the sampled dataset.”From Florence · §Approach
- BEiT v22022 · cited 2דDALL-E (Ramesh et al. 2021) uses the Gumbel-softmax relaxation for quantization instead of the nearest neighbor lookup in VQ-VAE.”From BEiT v2 · §Related Work
Abstract
Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.