Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich descriptive explanations for visual content that go semantically far beyond the captions available in current datasets.
Also cited · not yet reviewed (1)
- Karpathy visual-semantic alignment2014 · cited 4דRecently, several approaches based on RNNs emerged, generating captions via a learned joint image-text embedding [13, 11, 36, 21].”From this paper · §Related Work
Led to
- UniLM2019 · cited 1×, 1 in Method“UniLM is initialized by BERTLARGE, and then pre-trained using English Wikipedia11 1 Wikipedia version: enwiki-20181101. and BookCorpus [53], which have been processed in the same way as [9].”From UniLM · §Unified Language Model Pre-training
- RoBERTa2019 · cited 2×, 2 in Method“BookCorpus Zhu et al. 2015 plus English Wikipedia.”From RoBERTa · §Experimental Setup
- ViLBERT2019 · cited 2דThis pretrain-then-transfer learning approach to vision-and-language tasks follows naturally from its widespread use in both computer vision and natural language processing where it has become the de facto standard due t…”From ViLBERT · §Introduction
- VL-BERT2019 · cited 3דTo better exploit the generic representation, we pre-train VL-BERT at both large visual-linguistic corpus and text-only datasets11 1 Here we exploit the Conceptual Captions dataset (Sharma et al. 2018) as the visual-ling…”From VL-BERT · §Introduction
- Megatron-LM2019 · cited 1×, 1 in Method“For BERT models we include BooksCorpus (Zhu et al. 2015) in the training dataset, however, this dataset is excluded for GPT-2 trainings as it overlaps with LAMBADA task.”From Megatron-LM · §Setup
- RLHF for LMs (Ziegler)2019 · cited 2×, 1 in Method“For stylistic continuation tasks we perform supervised fine-tuning of the language model to the BookCorpus dataset of Zhu et al. 2015 prior to RL fine-tuning; we train from scratch on WebText, supervised fine-tune on Boo…”From RLHF for LMs (Ziegler) · §Methods
- Kaplan scaling laws2020 · cited 1×, 1 in Method“We reserve 6.6×1086.6\times 10^{8} of these tokens for use as a test set, and we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-avail…”From Kaplan scaling laws · §Background and Methods
- The Pile2020 · cited 3דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From The Pile · §Related Work
- Fairseq MoE LMs2021 · cited 1×, 1 in Method“BookCorpus (Zhu et al. 2019) consists of more than 10K unpublished books (4GB);”From Fairseq MoE LMs · §Experimental Setup
- OPT2022 · cited 1×, 1 in Method“We included the BookCorpus Zhu et al. 2015 and Stories Trinh and Le 2018 subsets of the RoBERTa corpus and utilized an updated version of CCNews, containing news stories crawled through September 28, 2021.”From OPT · §Method
- BEiT-32022 · cited 1×, 1 in Method“For monomodal data, we use 1414M images from ImageNet-21K and 160160GB text corpora [4] from English Wikipedia, BookCorpus [64], OpenWebText22 2 http://skylion007.github.io/OpenWebTextCorpus, CC-News [33], and Stories [5…”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract
Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich descriptive explanations for visual content that go semantically far beyond the captions available in current datasets. To align movies and books we exploit a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book. We propose a context-aware CNN to combine information from multiple sources. We demonstrate good quantitative performance for movie/book alignment and show several qualitative examples that showcase the diversity of tasks our model can be used for.