architectureIntroduced by VLMo · 2021
Mixture of modality experts
Shared attention with per-modality feed-forward experts, usable as a dual or fusion encoder.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that VLMo cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Align unimodal embeddings contrastively before fusing them with cross-attention.
Learn from soft targets produced by a moving-average model, tolerating noisy web pairs.
Feed raw image patches straight into the VL Transformer: no CNN, no detector.