architectureIntroduced by ALBERT · 2019
Cross-layer parameter sharing
Share weights across layers and factorise embeddings to shrink BERT.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that ALBERT cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Hide some tokens and predict them from context on both sides.
Train BERT longer, on more data, with bigger batches and dynamic masking.
Autoregressive training over all factorisation orders, capturing bidirectional context without masks.