REVIEW 3 cited by
Score Approximation, Estimation and Distribution Recovery of Diffusion Models on Low-Dimensional Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models achieve state-of-the-art performance in various generation tasks. However, their theoretical foundations fall far behind. This paper studies score approximation, estimation, and distribution recovery of diffusion models, when data are supported on an unknown low-dimensional linear subspace. Our result provides sample complexity bounds for distribution estimation using diffusion models. We show that with a properly chosen neural network architecture, the score function can be both accurately approximated and efficiently estimated. Furthermore, the generated distribution based on the estimated score function captures the data geometric structures and converges to a close vicinity of the data distribution. The convergence rate depends on the subspace dimension, indicating that diffusion models can circumvent the curse of data ambient dimensionality.
Forward citations
Cited by 3 Pith papers
-
Denoising growth complexity: Data geometry and certified schedules for diffusion sampling
A new measure, the denoising growth complexity, provides local KL error bounds for Euler diffusion samplers and yields certified, geometry-adaptive schedules.
-
Fast Score-Based Sampling via Log-Concave Reductions
Score-based sampling reduces to a short sequence of strongly log-concave sampling problems, giving √d polylog(1/ε) complexity bounds and logarithmic dependence on the condition number for log-concave targets.
-
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
Adding a reconstruction target that redraws the object region makes a vision-language-action model focus its attention on the right object and manipulate more precisely.
Discussion (0). Sign in to comment.