Skip to content

Learning Deep Features for Discriminative Localization

Bolei Zhou, Aditya Khosla, Àgata Lapedriza, Aude Oliva, Antonio B. Torralba

2016 · 11,058 citationsOpen access

Abstract

In this work, we revisit the global average pooling layer proposed in [13], and shed light on how it explicitly enables the convolutional neural network (CNN) to have remarkable localization ability despite being trained on imagelevel labels. While this technique was previously proposed as a means for regularizing training, we find that it actually builds a generic localizable deep representation that exposes the implicit attention of CNNs on an image. Despite the apparent simplicity of global average pooling, we are able to achieve 37.1% top-5 error for object localization on ILSVRC 2014 without training on any bounding box annotation. We demonstrate in a variety of experiments that our network is able to localize the discriminative image regions despite just being trained for solving classification task1.

Cite this paper

Zhou, B., Khosla, A., Lapedriza, À., Oliva, A., & Torralba, A. B. (2016). Learning deep features for discriminative localization. 2921–2929. https://doi.org/10.1109/cvpr.2016.319

Read it with every claim anchored

Add this paper to a project, ask questions of it, and get answers that point to the exact passage.

Start free
  1. Deep Residual Learning for Image Recognition2016
  2. ImageNet classification with deep convolutional neural networks2017
  3. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks2016
  4. You Only Look Once: Unified, Real-Time Object Detection2016
  5. Going deeper with convolutions2015

Metadata from OpenAlex (CC0). Citations are generated from the published record.