Pith. sign in

REVIEW 1 cited by

GLM Inference with AI-Generated Synthetic Data Using Misspecified Linear Regression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.21968 v2 pith:BBVSHXWY submitted 2025-03-27 stat.ME

classification stat.ME
keywords datasyntheticinferencelinearestimatorglmsmisspecifiedoriginal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Data privacy concerns have led to the growing interest in synthetic data, which strives to preserve the statistical properties of the original dataset while ensuring privacy by excluding real records. Recent advances in deep neural networks and generative artificial intelligence have facilitated the generation of synthetic data. However, although prediction with synthetic data has been the focus of recent research, statistical inference with synthetic data remains underdeveloped. In particular, in many settings, including generalized linear models (GLMs), the estimator obtained using synthetic data converges much more slowly than in standard settings. To address these limitations, we propose a method that leverages summary statistics from the original data. Using a misspecified linear regression estimator, we then develop inference that greatly improves the convergence rate and restores the standard root-$n$ behavior for GLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. General Synthetic-Powered Inference

    stat.ME 2025-09 conditional novelty 6.0 of 10

    GESPI combines real and synthetic data by aggregating three runs of a base inference method and guarantees an error rate of at most alpha+epsilon without any assumptions on the synthetic distribution.

Pith tools