evaluationIntroduced by MMLU · 2020
Broad multitask knowledge benchmarks
Test models across dozens to hundreds of academic and reasoning tasks at once.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that MMLU cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Answer a question by selecting a span from a passage.
Papers using this
- 2021T0
- 2022Chinchilla
- 2022PaLM
- 2022BIG-bench
- 2022Emergent abilities
- 2022U-PaLM
- 2022Flan-T5 / Flan-PaLM