architectureIntroduced by BEiT-3 · 2022
Multiway Transformer (image as a foreign language)
Treat images as another language and pre-train one multiway model with masked 'language' modelling on all.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2021
Shared attention with per-modality feed-forward experts, usable as a dual or fusion encoder.
Cites · not yet reviewedcited 5× · §BEiT-3: A General-Purpose Multimodal Foundation Model“We use Multiway Transformers [53] as the backbone model to encode different modalities.”
From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model - 2022
Treat images as another language and pre-train one multiway model with masked 'language' modelling on all.
Papers using this
No other paper in this dataset is tagged with it yet.