Object detection & grounding
Finding objects and linking words or phrases to image regions.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Integrated ConvNet localisation & detection One shared ConvNet does classification, localisation and detection with dense sliding windows. | OverFeat | 2013–2015 | 2 |
| Phrase grounding Link each phrase in a caption to the image region it mentions. | Flickr30k Entities | 2015–2022 | 7 |
| Region proposal network Predict object proposals from shared conv features, making two-stage detection near real-time. | Faster R-CNN | 2015–2021 | 14 |
| Box attention for visual relationships Model subject–object pairs inside a standard detector to find relationships. | Box Attention | 2018–2019 | 2 |
| Decoupled box proposal & featurisation Separate where-boxes-are from what-features-to-extract so more label data can be used. | Decoupled box proposals captioning | 2019 | 0 |
| Attention-unified detection head Unify scale, spatial and task awareness in a detector head with three attentions. | Dynamic Head | 2021 | 1 |