Paper Lineage
Esc
architectureIntroduced by Set Transformer · 2018

Inducing-point attention (learned latents)

A few learned vectors attend to a large set and summarise it: the idea behind later resamplers and query transformers.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2014

    Let a decoder look back at the most relevant input positions instead of one fixed vector.

    Cites · not yet reviewedcited 5× · §Model Architecture
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From Transformer · §Model Architecture
  2. 2017

    Every token attends to every other token, weighted by learned query–key similarity.

    Cites · not yet reviewedcited 3× · §Introduction
    “Transformer, (Vaswani et al. 2017)).”
    From Set Transformer · §Introduction
  3. 2018

    A few learned vectors attend to a large set and summarise it: the idea behind later resamplers and query transformers.

Papers using this