Paper Lineage
Esc
evaluationIntroduced by MMLU · 2020

Broad multitask knowledge benchmarks

Test models across dozens to hundreds of academic and reasoning tasks at once.

Drafted by AI · not yet reviewed

How this idea evolved

Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that MMLU cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.

Papers using this