Concepts
25 themes, each grouping the ideas that grew up inside it. Every concept is credited to the paper that introduced it in this dataset.
- Attention & Transformers9 conceptsAttention mechanisms and the Transformer family, plus their positional and feed-forward refinements.
- Benchmarks, evaluation & robustness8 conceptsHow progress is measured, and how models fail under distribution shift.
- Contrastive image–text models8 conceptsDual encoders that align images and text in one embedding space, enabling zero-shot transfer.
- Convolutional networks6 conceptsThe CNN era: deeper, wider and more modular convolutional backbones for vision.
- Datasets & data curation4 conceptsHow training and evaluation data were collected, cleaned and scaled.
- Detector-based vision-language pre-training9 conceptsBERT-style vision-language models built on object-detector region features (2019–2021).
- Discrete & generative image models3 conceptsImage tokenisers and models that generate images.
- Distributed training & systems5 conceptsParallelism and systems tricks for training models too big for one accelerator.
- Efficient & searched architectures5 conceptsMobile-friendly building blocks, neural architecture search and principled model scaling.
- End-to-end vision-language pre-training11 conceptsVision-language models trained from pixels, without a separate object detector.
- Image captioning11 conceptsGenerating natural-language descriptions of images, and how to evaluate them.
- Instruction tuning & alignment8 conceptsTeaching LMs to follow instructions and human preferences.
- Language model pre-training13 conceptsFrom word vectors to BERT-style and seq2seq pre-training objectives for text.
- Large language models & scaling11 conceptsVery large autoregressive LMs, what scale buys you, and how to spend compute.
- Object detection & grounding6 conceptsFinding objects and linking words or phrases to image regions.
- Parameter-efficient adaptation3 conceptsAdapt a big frozen model by training only a small number of new parameters.
- Self-supervised visual learning10 conceptsLearning visual features without labels: contrastive, mutual-information and masked-prediction methods.
- Sequence models & translation6 conceptsRecurrent, convolutional and encoder–decoder models that map one sequence to another.
- Sparse & mixture-of-experts models2 conceptsConditional computation: activate only part of a huge model for each input.
- Training & optimization7 conceptsNormalisation, initialisation, regularisation, distillation and optimiser fixes that make training work.
- Transfer learning & weak supervision10 conceptsPre-train on huge, cheap or noisy labels, then transfer to downstream tasks.
- Unified multitask models4 conceptsOne model, one format, many tasks.
- Vision Transformers7 conceptsTransformers applied to images and video, and how to train and scale them.
- Vision-language models on frozen LLMs7 conceptsConnect a pretrained vision encoder to a frozen large language model through a small trainable bridge.
- Visual question answering & reasoning7 conceptsAnswering questions and reasoning about images in natural language.