Pith. sign in

REVIEW 5 cited by

CFMI: Flow Matching for Missing Data Imputation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.09258 v1 pith:DXWMR2H2 submitted 2025-06-10 cs.LG stat.ML

CFMI: Flow Matching for Missing Data Imputation

classification cs.LG stat.ML
keywords dataimputationmethodcfmimatchingtraditionalconditionalflow
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce conditional flow matching for imputation (CFMI), a new general-purpose method to impute missing data. The method combines continuous normalising flows, flow-matching, and shared conditional modelling to deal with intractabilities of traditional multiple imputation. Our comparison with nine classical and state-of-the-art imputation methods on 24 small to moderate-dimensional tabular data sets shows that CFMI matches or outperforms both traditional and modern techniques across a wide range of metrics. Applying the method to zero-shot imputation of time-series data, we find that it matches the accuracy of a related diffusion-based method while outperforming it in terms of computational efficiency. Overall, CFMI performs at least as well as traditional methods on lower-dimensional data while remaining scalable to high-dimensional settings, matching or exceeding the performance of other deep learning-based approaches, making it a go-to imputation method for a wide range of data types and dimensionalities.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PAMF: Prior-Aware Multimodal Fusion for Incomplete Time Series Data

    cs.LG 2026-06 unverdicted novelty 6.0

    PAMF initializes flow matching with missingness-type priors and shares encoder weights between imputation and classification to improve multimodal time-series prediction under incomplete observations.

  2. Overcoming Selection Bias in Statistical Studies With Amortized Bayesian Inference

    stat.ML 2026-04 unverdicted novelty 6.0

    Embedding selection mechanisms into generative simulators enables amortized Bayesian inference to produce debiased, well-calibrated posteriors without tractable likelihoods.

  3. Neural Conditional Simulation for Complex Spatial Processes

    stat.ME 2025-08 conditional novelty 6.0

    Neural conditional simulation trains a masked diffusion model on unconditional spatial field samples to draw from predictive distributions, demonstrated on Gaussian and Brown–Resnick processes.

  4. Flow Matching with Missing Data

    cs.LG 2026-07 conditional novelty 5.0

    Resampling missing coordinates and averaging the flow-matching loss reproduces the complete-data objective exactly under MCAR with oracle completions; one completion per example is optimal for a fixed budget.

  5. Incomplete Data, Complete Dynamics: A Diffusion Approach

    cs.LG 2025-09 unverdicted novelty 5.0

    A conditional diffusion model trained on partitioned incomplete samples for physical dynamics achieves asymptotic convergence to the true generative process under mild conditions and outperforms baselines in imputation.