architectureIntroduced by Video Swin · 2021
Local video Transformers
Extend windowed attention across space and time for video recognition.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Video Swin cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Attention inside local windows that shift between layers, with multi-scale feature maps.
Train ViTs on ImageNet alone by adding a token that learns from a CNN teacher.
Split an image into patches and feed them to a plain Transformer as tokens.
Papers using this
No other paper in this dataset is tagged with it yet.