Paper Lineage
Esc

Benchmarks, evaluation & robustness

How progress is measured, and how models fail under distribution shift.

Concepts in Benchmarks, evaluation & robustness
ConceptIntroduced byYearsPapers
Extractive reading comprehension
Answer a question by selecting a span from a passage.
SQuAD2016–202116
Retriever–reader open-domain QA
Retrieve documents, then read them to extract an answer.
DrQA2017–20181
ImageNet accuracy predicts transfer
Better ImageNet models are, almost linearly, better feature extractors elsewhere.
Do better ImageNet models transfer bette20180
Texture vs shape bias
ImageNet CNNs recognise textures more than shapes, unlike humans.
ImageNet-trained CNNs are biased towards2018–20211
Corruption robustness
Measure accuracy under noise, blur, weather and other common corruptions.
ImageNet-C20191
Natural distribution-shift robustness
Accuracy drops on new test sets drawn the same way; robustness to such shifts is rare.
ImageNetV22019–20212
Broad multitask knowledge benchmarks
Test models across dozens to hundreds of academic and reasoning tasks at once.
MMLU2020–20227
True few-shot evaluation
Evaluate few-shot ability without a hidden validation set used for tuning.
True few-shot learning20210