REVIEW 2 cited by
Is the deconvolution layer the same as a convolutional layer?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this note, we want to focus on aspects related to two questions most people asked us at CVPR about the network we presented. Firstly, What is the relationship between our proposed layer and the deconvolution layer? And secondly, why are convolutions in low-resolution (LR) space a better choice? These are key questions we tried to answer in the paper, but we were not able to go into as much depth and clarity as we would have liked in the space allowance. To better answer these questions in this note, we first discuss the relationships between the deconvolution layer in the forms of the transposed convolution layer, the sub-pixel convolutional layer and our efficient sub-pixel convolutional layer. We will refer to our efficient sub-pixel convolutional layer as a convolutional layer in LR space to distinguish it from the common sub-pixel convolutional layer. We will then show that for a fixed computational budget and complexity, a network with convolutions exclusively in LR space has more representation power at the same speed than a network that first upsamples the input in high resolution space.
Forward citations
Cited by 2 Pith papers
-
Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation
Interactive segmentation reveals that feature upsampler choice strongly affects dense prediction quality in frozen DINOv2, with LoftUp giving the best results.
-
LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
A coordinate-based cross-attention transformer, trained with mask-refined and self-distilled pseudo-groundtruth, upsamples VFM features to full resolution and improves downstream task performance.
Discussion (0). Continue with ORCID to comment.