Paper Lineage
Esc
MethodOct 2018arXiv 1810.00825cs.LG

Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks

Juho Lee, Yoonho Lee, Jungtaek Kim and 3 others

Many machine learning tasks such as multiple instance learning, 3D shape recognition, and few-shot image classification are defined on sets of instances. Since solutions to such problems do not depend on the order of elements of the set, models used to address them should be permutation invariant.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • Transformer2017 · cited 3×, 1 in Method
    “Transformer, (Vaswani et al. 2017)).”
    From this paper · §Introduction

Led to

  • Perceiver2021 · cited 2×, 1 in Method
    “Attention is a permutation-invariant operation, and this property is preserved by the Perceiver and related models (Lee et al. 2019).”
    From Perceiver · §Methods
  • Scaling ViTs (ViT-G)2021 · cited 1×, 1 in Method
    “In particular, we evaluate global average pooling (GAP) and multihead attention pooling (MAP) lee2019set to aggregate representation from all patch tokens.”
    From Scaling ViTs (ViT-G) · §Method details
  • CoCa2022 · cited 2×, 2 in Method
    “As discussed in the previous section, CoCa adopts task-specific attentional pooling [42] (pooler for brevity) to customize visual representations for different types downstream tasks while sharing the backbone encoder.”
    From CoCa · §Approach
Abstract

Many machine learning tasks such as multiple instance learning, 3D shape recognition, and few-shot image classification are defined on sets of instances. Since solutions to such problems do not depend on the order of elements of the set, models used to address them should be permutation invariant. We present an attention-based neural network module, the Set Transformer, specifically designed to model interactions among elements in the input set. The model consists of an encoder and a decoder, both of which rely on attention mechanisms. In an effort to reduce computational complexity, we introduce an attention scheme inspired by inducing point methods from sparse Gaussian process literature. It reduces the computation time of self-attention from quadratic to linear in the number of elements in the set. We show that our model is theoretically attractive and we evaluate it on a range of tasks, demonstrating the state-of-the-art performance compared to recent methods for set-structured data.