Pith. sign in

REVIEW 2 cited by

A Scaling Law for Synthetic-to-Real Transfer: How Much Is Your Pre-training Effective?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.11018 v3 pith:OJ5Y6JVS submitted 2021-08-25 cs.LG cs.CV

classification cs.LGcs.CV
keywords datalearningperformancescalingsynthetictransferimagesimprove
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthetic-to-real transfer learning is a framework in which a synthetically generated dataset is used to pre-train a model to improve its performance on real vision tasks. The most significant advantage of using synthetic images is that the ground-truth labels are automatically available, enabling unlimited expansion of the data size without human cost. However, synthetic data may have a huge domain gap, in which case increasing the data size does not improve the performance. How can we know that? In this study, we derive a simple scaling law that predicts the performance from the amount of pre-training data. By estimating the parameters of the law, we can judge whether we should increase the data or change the setting of image synthesis. Further, we analyze the theory of transfer learning by considering learning dynamics and confirm that the derived generalization bound is consistent with our empirical findings. We empirically validated our scaling law on various experimental settings of benchmark tasks, model sizes, and complexities of synthetic images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development

    cs.LG 2026-07 conditional novelty 5.0 of 10

    In a stylized model, a proactive flywheel that fixes whole groups of related scenarios needs Θ(K log K) update rounds versus Θ(M log M) for reactive patching.

  2. GUST: Quantifying Free-Form Geometric Uncertainty of Metamaterials Using Small Data

    cs.LG 2025-05 conditional novelty 5.0 of 10

    GUST combines synthetic-data pretraining with transfer learning on a conditional diffusion model to quantify free-form geometric uncertainty in manufactured metamaterials from small real-world datasets.

Pith tools