Scinovex
articleTop 1% cited

Squeeze-and-Excitation Networks

Jie HuLi ShenSamuel AlbanieGang SunEnhua Wu

Abstract

The central building block of convolutional neural networks (CNNs) is the convolution operator, which enables networks to construct informative features by fusing both spatial and channel-wise information within local receptive fields at each layer. A broad range of prior research has investigated the spatial component of this relationship, seeking to strengthen the representational power of a CNN by enhancing the quality of spatial encodings throughout its feature hierarchy. In this work, we focus instead on the channel relationship and propose a novel architectural unit, which we term the "Squeeze-and-Excitation" (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels. We show that these blocks can be stacked together to form SENet architectures that generalise extremely effectively across different datasets. We further demonstrate that SE blocks bring significant improvements in performance for existing state-of-the-art CNNs at slight additional computational cost. Squeeze-and-Excitation Networks formed the foundation of our ILSVRC 2017 classification submission which won first place and reduced the top-5 error to 2.251 percent, surpassing the winning entry of 2016 by a relative improvement of ∼ 25 percent. Models and code are available at https://github.com/hujie-frank/SENet.

Funding

  • Canadian Institute for Advanced Research
  • University of Manchester
  • National Natural Science Foundation of China
  • Chinese Academy of Sciences
  • Universidade de Macau
  • Engineering and Physical Sciences Research Council
  • National Key Research and Development Program of China
Citations
12,333
FWCI
1241.40
field-weighted impact
References
156
Percentile
100%
vs. same field & year
Citations per year
Cited by
Deep Learning for Generic Object Detection: A Survey
International Journal of Computer Vision · 2019 · 2,702 citations
Deep Residual Shrinkage Networks for Fault Diagnosis
IEEE Transactions on Industrial Informatics · 2019 · 1,313 citations
Deep face recognition: A survey
Neurocomputing · 2020 · 934 citations
A Survey on Deep Learning for Multimodal Data Fusion
Neural Computation · 2020 · 720 citations
A Small-Sized Object Detection Oriented Multi-Scale Feature Fusion Approach With Application to Defect Detection
IEEE Transactions on Instrumentation and Measurement · 2022 · 482 citations
A review on the attention mechanism of deep learning
Neurocomputing · 2021 · 3,096 citations
References
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · 52,930 citations
Image Classification with the Fisher Vector: Theory and Practice
International Journal of Computer Vision · 2013 · 1,510 citations
Long Short-Term Memory
Neural Computation · 1997 · 95,078 citations
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
A model of saliency-based visual attention for rapid scene analysis
IEEE Transactions on Pattern Analysis and Machine Intelligence · 1998 · 11,235 citations
Computational modelling of visual attention
Nature reviews. Neuroscience · 2001 · 4,703 citations
ImageNet classification with deep convolutional neural networks
Communications of the ACM · 2017 · 75,550 citations
Places: A 10 Million Image Database for Scene Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · 3,945 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.