training techniqueIntroduced by ALBEF · 2021
Momentum distillation
Learn from soft targets produced by a moving-average model, tolerating noisy web pairs.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2019
A queue of negatives encoded by a slowly-updated momentum encoder.
Cites · not yet reviewedcited 3× · §ALBEF Pre-training“Inspired by MoCo [24], we maintain two queues to store the most recent MM image-text representations from the momentum unimodal encoders.”
From ALBEF · §ALBEF Pre-training - 2021
Learn from soft targets produced by a moving-average model, tolerating noisy web pairs.
Also draws on: Knowledge distillation (Knowledge Distillation)
Papers using this
No other paper in this dataset is tagged with it yet.