Paper Lineage
Esc
training techniqueIntroduced by RoBERTa · 2019

Robustly optimised BERT recipe

Train BERT longer, on more data, with bigger batches and dynamic masking.

Drafted by AI · not yet reviewed

How this idea evolved

Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that RoBERTa cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.

Papers using this