Benchmarks, evaluation & robustness
How progress is measured, and how models fail under distribution shift.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Extractive reading comprehension Answer a question by selecting a span from a passage. | SQuAD | 2016–2021 | 16 |
| Retriever–reader open-domain QA Retrieve documents, then read them to extract an answer. | DrQA | 2017–2018 | 1 |
| ImageNet accuracy predicts transfer Better ImageNet models are, almost linearly, better feature extractors elsewhere. | Do better ImageNet models transfer bette | 2018 | 0 |
| Texture vs shape bias ImageNet CNNs recognise textures more than shapes, unlike humans. | ImageNet-trained CNNs are biased towards | 2018–2021 | 1 |
| Corruption robustness Measure accuracy under noise, blur, weather and other common corruptions. | ImageNet-C | 2019 | 1 |
| Natural distribution-shift robustness Accuracy drops on new test sets drawn the same way; robustness to such shifts is rare. | ImageNetV2 | 2019–2021 | 2 |
| Broad multitask knowledge benchmarks Test models across dozens to hundreds of academic and reasoning tasks at once. | MMLU | 2020–2022 | 7 |
| True few-shot evaluation Evaluate few-shot ability without a hidden validation set used for tuning. | True few-shot learning | 2021 | 0 |