objectiveIntroduced by VL-T5 · 2021
All VL tasks as text generation
Answer every vision-language task by generating its label as text.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that VL-T5 cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Cast every task as text in, text out: one model, one loss.
Cast ten NLP tasks as QA over a context so one model handles them all.
Papers using this
- 2021SimVLM