Pith. sign in

REVIEW 2 cited by

Missing value imputation with adversarial random forests -- MissARF

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.15681 v1 pith:H2AE2BUG submitted 2025-07-21 stat.ML cs.AIcs.LG

Missing value imputation with adversarial random forests -- MissARF

classification stat.ML cs.AIcs.LG
keywords imputationmissarfmissingadversarialmultiplerandomvaluefast
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Handling missing values is a common challenge in biostatistical analyses, typically addressed by imputation methods. We propose a novel, fast, and easy-to-use imputation method called missing value imputation with adversarial random forests (MissARF), based on generative machine learning, that provides both single and multiple imputation. MissARF employs adversarial random forest (ARF) for density estimation and data synthesis. To impute a missing value of an observation, we condition on the non-missing values and sample from the estimated conditional distribution generated by ARF. Our experiments demonstrate that MissARF performs comparably to state-of-the-art single and multiple imputation methods in terms of imputation quality and fast runtime with no additional costs for multiple imputation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Testing independence in the presence of missing data: high-dimensional case

    stat.ME 2026-04 unverdicted novelty 6.0

    Two new modifications to a Kendall tau-based test are proposed and analyzed for independence testing in high-dimensional data with missing observations, backed by theory and simulations.

  2. Can synthetic data reproduce real-world findings in epidemiology? A replication study using adversarial random forests

    q-bio.QM 2025-08 conditional novelty 5.0

    Adversarial random forests generate synthetic epidemiological data that aligns with original findings across six replication studies from German and Canadian cohorts.