Scinovex
article Open AccessTop 1% cited

Res2Net: A New Multi-Scale Backbone Architecture

Shanghua GaoMing‐Ming ChengKai ZhaoXinyu ZhangMing–Hsuan YangPhilip H. S. Torr

Abstract

Representing features at multiple scales is of great importance for numerous vision tasks. Recent advances in backbone convolutional neural networks (CNNs) continually demonstrate stronger multi-scale representation ability, leading to consistent performance gains on a wide range of applications. However, most existing methods represent the multi-scale features in a layer-wise manner. In this paper, we propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer. The proposed Res2Net block can be plugged into the state-of-the-art backbone CNN models, e.g., ResNet, ResNeXt, and DLA. We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.g., CIFAR-100 and ImageNet. Further ablation studies and experimental results on representative computer vision tasks, i.e., object detection, class activation mapping, and salient object detection, further verify the superiority of the Res2Net over the state-of-the-art baseline methods. The source code and trained models are available on https://mmcheng.net/res2net/.

Visual Attention and Saliency DetectionAdvanced Neural Network ApplicationsAdvanced Image and Video Retrieval TechniquesComputer scienceBlock (permutation group theory)Artificial intelligenceResidualBackbone networkConvolutional neural networkObject detectionPattern recognition (psychology)Scale (ratio)Representation (politics)

Funding

  • National Natural Science Foundation of China
  • Natural Science Foundation of Tianjin City
  • Engineering and Physical Sciences Research Council
Citations
3,252
FWCI
116.62
field-weighted impact
References
115
Percentile
100%
vs. same field & year
Citations per year
Cited by
Inf-Net: Automatic COVID-19 Lung Infection Segmentation From CT Images
IEEE Transactions on Medical Imaging · 2020 · 1,199 citations
References
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · 52,930 citations
Salient Object Detection: A Benchmark
IEEE Transactions on Image Processing · 2015 · 1,319 citations
The Pascal Visual Object Classes (VOC) Challenge
International Journal of Computer Vision · 2009 · 19,127 citations
The Pascal Visual Object Classes Challenge: A Retrospective
International Journal of Computer Vision · 2014 · 7,183 citations
Shape matching and object recognition using shape contexts
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2002 · 6,295 citations
Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2015 · 11,231 citations
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
Distinctive Image Features from Scale-Invariant Keypoints
International Journal of Computer Vision · 2004 · 54,768 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.