Skip to content

Pyramid Scene Parsing Network

Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, Jiaya Jia

2017 · 15,919 citations

Abstract

Scene parsing is challenging for unrestricted open vocabulary and diverse scenes. In this paper, we exploit the capability of global context information by different-region-based context aggregation through our pyramid pooling module together with the proposed pyramid scene parsing network (PSPNet). Our global prior representation is effective to produce good quality results on the scene parsing task, while PSPNet provides a superior framework for pixel-level prediction. The proposed approach achieves state-of-the-art performance on various datasets. It came first in ImageNet scene parsing challenge 2016, PASCAL VOC 2012 benchmark and Cityscapes benchmark. A single PSPNet yields the new record of mIoU accuracy 85.4% on PASCAL VOC 2012 and accuracy 80.2% on Cityscapes.

Cite this paper

Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. 6230–6239. https://doi.org/10.1109/cvpr.2017.660

Read it with every claim anchored

Add this paper to a project, ask questions of it, and get answers that point to the exact passage.

Start free
  1. Deep Residual Learning for Image Recognition2016
  2. ImageNet classification with deep convolutional neural networks2017
  3. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks2016
  4. You Only Look Once: Unified, Real-Time Object Detection2016
  5. Going deeper with convolutions2015

Metadata from OpenAlex (CC0). Citations are generated from the published record.