REVIEW 22 cited by
The Intrinsic Dimension of Images and Its Impact on Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
It is widely believed that natural image data exhibits low-dimensional structure despite the high dimensionality of conventional pixel representations. This idea underlies a common intuition for the remarkable success of deep learning in computer vision. In this work, we apply dimension estimation tools to popular datasets and investigate the role of low-dimensional structure in deep learning. We find that common natural image datasets indeed have very low intrinsic dimension relative to the high number of pixels in the images. Additionally, we find that low dimensional datasets are easier for neural networks to learn, and models solving these tasks generalize better from training to test data. Along the way, we develop a technique for validating our dimension estimation tools on synthetic data generated by GANs allowing us to actively manipulate the intrinsic dimension by controlling the image generation process. Code for our experiments may be found here https://github.com/ppope/dimensions.
Forward citations
Cited by 22 Pith papers
-
Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
Transformers perform kernel-based prediction for Hölder regression on manifolds and achieve intrinsic-dimension-dependent minimax rates with sufficient training tasks.
-
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
The subspace intervention framework reveals that pre-training objectives shape how ViTs encode geometric information in compressible low-rank subspaces, with peak precision at intermediate layers.
-
Learning to Strategically Acquire Resources in Competition
A game-theoretic model for multi-agent resource acquisition establishes BNE existence and computability under common priors, convergence conditions for learning dynamics, and simulations on financial data.
-
UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures
UR-JEPA applies uniform rectifiability regularization via a smoothed Carleson square function to JEPA training, producing embeddings with 4-5 order PCA spectral drop at dimension 20-25 and lower seed variance than Gau...
-
CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL
CreFlow combines LTL compositional rewards with credit-aware NFT and corrective reflow losses in online RL to improve embodied video diffusion models, raising downstream task success by 23.8 percentage points on eight...
-
Implicit Neural Optimal Transport via Fixed-Point Optimization
A single-network implicit neural optimal transport method that solves the c-transform via proximal fixed-point iteration for stable, non-adversarial training.
-
Implicit Neural Optimal Transport via Fixed-Point Optimization
A single-network fixed-point formulation for neural optimal transport eliminates adversarial min-max optimization and implicit differentiation while enforcing dual feasibility exactly.
-
Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights
Transformers achieve approximation and generalization error bounds for noisy manifold regression that scale with the intrinsic dimension of the task-level manifold.
-
CASC: Causal Adversarial Subspace Clustering for Multivariate Spatiotemporal Data
CASC uses a U-Net-style adversarial autoencoder with attention and causal-regularized self-expression to cluster multivariate spatiotemporal series into evolving regimes, validated only on internal cluster metrics.
-
Diffusion Models Adapt to Low-Dimensional Structure Under Flexible Coefficient Choices
For a broad class of coefficients, diffusion models achieve Õ(k/ε) iteration complexity for ε-accurate TV sampling under low-dimensional structure, independent of ambient dimension.
-
Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models
SAMG uses spatially adaptive guidance scales derived from a geometric analysis of classifier-free guidance to resolve the detail-artifact dilemma in diffusion-based image and video generation.
-
Emergent Manifold Separability during Reasoning in Large Language Models
Reasoning in LLMs produces a transient geometric pulse in which concept manifolds untangle into linearly separable subspaces immediately before computation and compress afterward.
-
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
MoE Top-k routing equals the k-th elementary symmetric tropical polynomial, making sparsity combinatorial depth that scales capacity by binom(N,k) and gives MoE combinatorial resilience on manifolds.
-
Complexity of Quantum Trajectories
Intrinsic dimension of quantum trajectories serves as an unsupervised probe sensitive to chaos, integrability, and ergodicity breaking in dissipative quantum systems.
-
Density Estimation via Binless Multidimensional Integration
BMTI estimates log-density via integration of neighbor differences on data manifolds using maximum-likelihood weighting, without binning or explicit coordinates.
-
A Multiclass Quantum Aligned Centroid Kernel
A sample-to-centroid fidelity kernel enables linear-scaling multiclass quantum classification; in simulation it beats pure quantum baselines, and untrained 124-qubit hardware results match an RBF kernel.
-
Provable diffusion-based posterior sampling for linear inverse problems via DDIM
A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.
-
ShuffleFlow: Scalable Posterior Inference for Bayesian Inverse Imaging
ShuffleFlow is a variational inference framework that partitions images with pixel-unshuffling and models the joint posterior over sub-images using a shared conditional normalizing flow conditioned on neural field fea...
-
On the Limits of Latent Reuse in Diffusion Models
Reusing source latent spaces in diffusion models under distribution shift produces target score error set by principal-angle misalignment and diffusion-time-amplified ambient noise.
-
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
Training flow matching along sphere geodesics with a curvature-aware loss weight lets standard DiT-B converge on DINOv2 features (FID 3.37 with guidance), contradicting the need for width scaling.
-
A Systematic Analysis of Out-of-Distribution Detection Under Representation and Training Paradigm Shifts
Benchmark across architectures and shift regimes finds OOD detector rankings shift with representation collapse; proposes NC-based shortlist predictor and PCA filter without extra OOD data.
-
Geometric Analysis of Neural Regression Collapse via Intrinsic Dimension
Neural regression collapse occurs when last-layer feature intrinsic dimension falls below target intrinsic dimension, creating over-compressed and under-compressed regimes that govern generalization based on data quan...
Discussion (0). Sign in to comment.