Pith. sign in

REVIEW 2 cited by

How much is a noisy image worth? Data Scaling Laws for Ambient Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02780 v1 pith:G7RSUN3U submitted 2024-11-05 cs.LG cs.CV

classification cs.LGcs.CV
keywords datacleanmodelsnoisysampledatasetstrainedambient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The quality of generative models depends on the quality of the data they are trained on. Creating large-scale, high-quality datasets is often expensive and sometimes impossible, e.g. in certain scientific applications where there is no access to clean data due to physical or instrumentation constraints. Ambient Diffusion and related frameworks train diffusion models with solely corrupted data (which are usually cheaper to acquire) but ambient models significantly underperform models trained on clean data. We study this phenomenon at scale by training more than $80$ models on data with different corruption levels across three datasets ranging from $30,000$ to $\approx 1.3$M samples. We show that it is impossible, at these sample sizes, to match the performance of models trained on clean data when only training on noisy data. Yet, a combination of a small set of clean data (e.g.~$10\%$ of the total dataset) and a large set of highly noisy data suffices to reach the performance of models trained solely on similar-size datasets of clean data, and in particular to achieve near state-of-the-art performance. We provide theoretical evidence for our findings by developing novel sample complexity bounds for learning from Gaussian Mixtures with heterogeneous variances. Our theoretical model suggests that, for large enough datasets, the effective marginal utility of a noisy sample is exponentially worse than that of a clean sample. Providing a small set of clean samples can significantly reduce the sample size requirements for noisy data, as we also observe in our experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fusion of multi-source precipitation records via coordinate-based generative model

    physics.ao-ph 2025-06 conditional novelty 6.0 of 10

    A coordinate-based diffusion model fuses multi-source precipitation records and corrects biases in unseen operational forecasts.

  2. Noise-Robust Conditional Flow Matching: Generating Clean Samples from Noisy Datasets

    cs.CV 2026-07 reject novelty 5.0 of 10

    NR-CFM corrects noisy-bridge flow matching with Tweedie scores, but its endpoint readout is a conditional mean, so the output distribution is shrunk, not the true clean distribution.

Pith tools