Pith. sign in

REVIEW 10 cited by

How Much Position Information Do Convolutional Neural Networks Encode?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.08248 v1 pith:2QKAH5L6 submitted 2020-01-22 cs.CV cs.LG

classification cs.CVcs.LG
keywords informationnetworkscnnsneuralpositionabsoluteconvolutionaldeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In contrast to fully connected networks, Convolutional Neural Networks (CNNs) achieve efficiency by learning weights associated with local filters with a finite spatial extent. An implication of this is that a filter may know what it is looking at, but not where it is positioned in the image. Information concerning absolute position is inherently useful, and it is reasonable to assume that deep CNNs may implicitly learn to encode this information if there is a means to do so. In this paper, we test this hypothesis revealing the surprising degree of absolute position information that is encoded in commonly used neural networks. A comprehensive set of experiments show the validity of this hypothesis and shed light on how and where this information is represented while offering clues to where positional information is derived from in deep CNNs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RoFormer: Enhanced Transformer with Rotary Position Embedding

    cs.CL 2021-04 accept novelty 8.0 of 10

    RoFormer introduces rotary position embeddings that encode absolute positions via rotation matrices and relative dependencies in attention, outperforming prior position methods on long text classification tasks.

  2. ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ElasticDiT introduces an elastic DiT architecture with adjustable spatial compression and block depth plus Shift Sparse Block Attention and a distilled VAE to enable a single model to cover multiple fidelity-latency p...

  3. Elastic Attention Cores for Scalable Vision Transformers

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    VECA learns effective visual representations using core-periphery attention where patches interact exclusively via a resolution-invariant set of learned core embeddings, achieving linear O(N) complexity while maintain...

  4. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

    cs.CV 2024-10 unverdicted novelty 6.0 of 10

    Sana-0.6B produces high-resolution images with strong text alignment at 20x smaller size and 100x higher throughput than Flux-12B by combining 32x image compression, linear DiT blocks, and a decoder-only LLM text encoder.

  5. A Physics-Flavored Transformer Network for Parametrizing Contraction Dynamics of Engineered Skeletal Muscle Tissues

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A physics-flavored CNN-transformer automatically fits stretched-exponential parameters to engineered skeletal muscle contraction curves, using synthetic pre-training and unsupervised real-data alignment.

  6. A physics-informed neural network for improving surface reconstruction of intracranial saccular aneurysms via variational membrane equilibrium

    math.NA 2026-07 reject novelty 5.0 of 10

    The paper introduces a PINN/B-spline reconstruction pipeline that filters imaging artifacts from aneurysm surfaces by enforcing a Laplace membrane-equilibrium loss, and claims cleaner rupture-risk maps concentrated at...

  7. Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling

    cs.CV 2026-06 conditional novelty 5.0 of 10

    A survey diagnosing panoramic scene understanding as a field that converged on compatibility-preserving geometric adaptation rather than sphere-native modeling, while its evaluation protocols systematically fail to me...

  8. The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology

    q-bio.QM 2025-08 conditional novelty 5.0 of 10

    EAGLE-Net adds spatial encoding and a top-K neighborhood loss to attention-based MIL, achieving modest accuracy gains and more coherent attention maps on pan-cancer histology benchmarks.

  9. ARMA Block: A CNN-Based Autoregressive and Moving Average Module for Long-Term Time Series Forecasting

    cs.LG 2025-09 conditional novelty 4.0 of 10

    A two-branch convolutional ARMA block with residual refinement matches or beats older linear and transformer baselines on several long-term forecasting benchmarks with trend shifts, and appears to encode position information.

  10. Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling

    cs.CV 2026-06 unverdicted novelty 3.0 of 10

    Survey organizing panoramic scene analysis literature by architectural design and training paradigm, identifying the absence of methods achieving both strict spherical equivariance and full reuse of perspective-pretra...

Pith tools