Scinovex
articleTop 1% cited

Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

Kaiming HeXiangyu ZhangShaoqing RenJian Sun

Abstract

Existing deep convolutional neural networks (CNNs) require a fixed-size (e.g., 224 × 224) input image. This requirement is "artificial" and may reduce the recognition accuracy for the images or sub-images of an arbitrary size/scale. In this work, we equip the networks with another pooling strategy, "spatial pyramid pooling", to eliminate the above requirement. The new network structure, called SPP-net, can generate a fixed-length representation regardless of image size/scale. Pyramid pooling is also robust to object deformations. With these advantages, SPP-net should in general improve all CNN-based image classification methods. On the ImageNet 2012 dataset, we demonstrate that SPP-net boosts the accuracy of a variety of CNN architectures despite their different designs. On the Pascal VOC 2007 and Caltech101 datasets, SPP-net achieves state-of-the-art classification results using a single full-image representation and no fine-tuning. The power of SPP-net is also significant in object detection. Using SPP-net, we compute the feature maps from the entire image only once, and then pool features in arbitrary regions (sub-images) to generate fixed-length representations for training the detectors. This method avoids repeatedly computing the convolutional features. In processing test images, our method is 24-102 × faster than the R-CNN method, while achieving better or comparable accuracy on Pascal VOC 2007. In ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2014, our methods rank #2 in object detection and #3 in image classification among all 38 teams. This manuscript also introduces the improvement made for this competition.

Advanced Neural Network ApplicationsAdvanced Image and Video Retrieval TechniquesDomain Adaptation and Few-Shot LearningPoolingPascal (unit)Artificial intelligenceComputer scienceConvolutional neural networkPattern recognition (psychology)Pyramid (geometry)Contextual image classificationObject detectionDeep learning

MeSH terms

AlgorithmsAnimalsHumansImage Processing, Computer-AssistedPattern Recognition, AutomatedNeural Networks, Computer
Citations
11,231
FWCI
213.81
field-weighted impact
References
70
Percentile
100%
vs. same field & year
Citations per year
Cited by
PCANet: A Simple Deep Learning Baseline for Image Classification?
IEEE Transactions on Image Processing · 2015 · 1,337 citations
CornerNet: Detecting Objects as Paired Keypoints
International Journal of Computer Vision · 2019 · 1,115 citations
CE-Net: Context Encoder Network for 2D Medical Image Segmentation
IEEE Transactions on Medical Imaging · 2019 · 2,134 citations
Recent advances in convolutional neural networks
Pattern Recognition · 2017 · 6,130 citations
References
The Pascal Visual Object Classes (VOC) Challenge
International Journal of Computer Vision · 2009 · 19,127 citations
The Pascal Visual Object Classes Challenge: A Retrospective
International Journal of Computer Vision · 2014 · 7,183 citations
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
Backpropagation Applied to Handwritten Zip Code Recognition
Neural Computation · 1989 · 11,706 citations
Distinctive Image Features from Scale-Invariant Keypoints
International Journal of Computer Vision · 2004 · 54,768 citations
ImageNet classification with deep convolutional neural networks
Communications of the ACM · 2017 · 75,550 citations
Object Detection with Discriminatively Trained Part-Based Models
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2009 · 9,994 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.