Pith. sign in

REVIEW 1 cited by

Reducing the Amount of Real World Data for Object Detector Training with Synthetic Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00632 v1 pith:NLRW7H3Z submitted 2022-01-31 cs.CV

classification cs.CV
keywords datarealworlddetectiontrainingdatasetmixedperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A number of studies have investigated the training of neural networks with synthetic data for applications in the real world. The aim of this study is to quantify how much real world data can be saved when using a mixed dataset of synthetic and real world data. By modeling the relationship between the number of training examples and detection performance by a simple power law, we find that the need for real world data can be reduced by up to 70% without sacrificing detection performance. The training of object detection networks is especially enhanced by enriching the mixed dataset with classes underrepresented in the real world dataset. The results indicate that mixed datasets with real world data ratios between 5% and 20% reduce the need for real world data the most without reducing the detection performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 2 citations worldwide. Full citation record

  1. Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fine-tuning on real data after synthetic pretraining beats mixing both data types in most settings, but simple mixing wins for CNNs on large-gap sketch data.

Pith tools