REVIEW 3 cited by
Pixel-wise Attentional Gating for Parsimonious Pixel Labeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
To achieve parsimonious inference in per-pixel labeling tasks with a limited computational budget, we propose a \emph{Pixel-wise Attentional Gating} unit (\emph{PAG}) that learns to selectively process a subset of spatial locations at each layer of a deep convolutional network. PAG is a generic, architecture-independent, problem-agnostic mechanism that can be readily "plugged in" to an existing model with fine-tuning. We utilize PAG in two ways: 1) learning spatially varying pooling fields that improve model performance without the extra computation cost associated with multi-scale pooling, and 2) learning a dynamic computation policy for each pixel to decrease total computation while maintaining accuracy. We extensively evaluate PAG on a variety of per-pixel labeling tasks, including semantic segmentation, boundary detection, monocular depth and surface normal estimation. We demonstrate that PAG allows competitive or state-of-the-art performance on these tasks. Our experiments show that PAG learns dynamic spatial allocation of computation over the input image which provides better performance trade-offs compared to related approaches (e.g., truncating deep models or dynamically skipping whole layers). Generally, we observe PAG can reduce computation by $10\%$ without noticeable loss in accuracy and performance degrades gracefully when imposing stronger computational constraints.
Forward citations
Cited by 3 Pith papers
-
Improved Techniques for Training Adaptive Deep Networks
Training multi-exit adaptive networks with gradient rescaling, inline logit sharing, and self-distillation improves their accuracy at fixed compute budgets.
-
Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space
A deep neural network learns depth and hierarchical feature maps from monocular video, and camera pose is computed by aligning those features directly, preserving metric scale.
-
Simultaneous Semantic Segmentation and Outlier Detection in Presence of Domain Shift
A two-head segmentation model trained with pasted ImageNet negatives performs dense outlier detection alongside semantic segmentation in one forward pass and sets a new WildDash state of the art.
Discussion (0). Continue with ORCID to comment.