Deformable Convolutional Networks

Guodong Zhang; Han Hu; Haozhi Qi; Jifeng Dai; Yichen Wei; Yi Li; Yuwen Xiong

arxiv: 1703.06211 · v3 · pith:SBDJ6QY4new · submitted 2017-03-17 · 💻 cs.CV

Deformable Convolutional Networks

Jifeng Dai , Haozhi Qi , Yuwen Xiong , Yi Li , Guodong Zhang , Han Hu , Yichen Wei This is my paper

classification 💻 cs.CV

keywords deformablemodulescnnsconvolutionalnetworksadditionalgeometricoffsets

0 comments

read the original abstract

Convolutional neural networks (CNNs) are inherently limited to model geometric transformations due to the fixed geometric structures in its building modules. In this work, we introduce two new modules to enhance the transformation modeling capacity of CNNs, namely, deformable convolution and deformable RoI pooling. Both are based on the idea of augmenting the spatial sampling locations in the modules with additional offsets and learning the offsets from target tasks, without additional supervision. The new modules can readily replace their plain counterparts in existing CNNs and can be easily trained end-to-end by standard back-propagation, giving rise to deformable convolutional networks. Extensive experiments validate the effectiveness of our approach on sophisticated vision tasks of object detection and semantic segmentation. The code would be released.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Mapped Convolutions
cs.CV 2019-06 unverdicted novelty 7.0

Mapped convolutions generalize standard convolutions by decoupling sampling and weighting, enabling direct convolution on spherical and mesh data with a 17% improvement in spherical depth estimation.
UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis
cs.CV 2026-05 unverdicted novelty 6.0

UniFixer is a universal reference-guided framework that fixes spatial, temporal, and backbone-related degradations in diffusion-based view synthesis via coarse-to-fine modules and achieves zero-shot SOTA results on no...
Reprojection R-CNN: A Fast and Accurate Object Detector for 360{\deg} Images
cs.CV 2019-07 unverdicted novelty 6.0

Reprojection R-CNN is a two-stage detector for 360° images combining a distortion-aware spherical RPN on ERP with a reprojection network on perspective projections, reporting higher mAP than prior methods on two new s...
Rethinking Atrous Convolution for Semantic Image Segmentation
cs.CV 2017-06 unverdicted novelty 6.0

DeepLabv3 improves semantic segmentation by capturing multi-scale context with cascaded or parallel atrous convolutions and adding global context to ASPP, achieving better results on PASCAL VOC 2012 without DenseCRF p...