REVIEW 13 references
FUSE adapts a multimodal diffusion transformer for flow-matching-based posterior estimation and adds inference-time likelihood-guided particle resampling, achieving state-of-the-art posterior fidelity on SBI benchmarks and real-world exoplanet orbital estimation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-07 21:22 UTC pith:ZRLYPFMI
FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
FUSE outperforms state-of-the-art neural SBI baselines (NPE, FMPE, Simformer) on the SBIBM benchmark across seven metrics at 10^5 simulations, and on the real-world β Pictoris b exoplanet orbital estimation task, FUSE with FK-steering recovers complex parameter degeneracies that baselines fail to capture, achieving a Mode L2 Distance of 3.85 versus 75.74 (NPE) and 132.41 (FMPE), while completing inference in ~3 minutes versus ~8.5 hours for PTMCMC.
Load-bearing premise
The FK-steering mechanism depends on the denoised proxy θ̂_t = θ_t − t·v_ϕ(θ_t, t, x) being a faithful approximation of E[θ_0 | θ_t, x] (Eq. 11, §4.3). This proxy is exact only when v_ϕ equals the optimal conditional velocity (Appendix I.2, Eq. 37). When the learned velocity field is imperfect—which is the regime where FK-steering is most needed—the proxy can produce misleading likelihood scores, potentially steering particles toward wrong regions. The paper acknowledges this is a 'tractable likelihood-guided correction rather than an exact posterior sampler' but does not empirically characterize how proxy error degrades FK performance, nor does it evaluate FK-steering across all 10 benchmark tasks (only SLCP and the exoplanet task).
Editorial analysis
A structured set of objections, weighed in public.
Axiom & Free-Parameter Ledger
free parameters (6)
- λ (FK scaling factor) =
not explicitly stated
- α (noise scale) =
0.3
- B (beam width / particle count) =
8
- K (tokens per parameter) =
2
- D (hidden dimension) =
128 (benchmark), 256 (exoplanet)
- Resampling schedule =
every 5 steps, step 20-200
axioms (4)
- domain assumption The learned velocity field v_ϕ approximates the optimal conditional velocity well enough that the denoised proxy θ̂_t = θ_t − t·v_ϕ(θ_t, t, x) is a meaningful estimate of E[θ_0 | θ_t, x]
- domain assumption The simulator likelihood p(x|θ) and prior p(θ) are available and tractable to evaluate at inference time
- standard math The rectified flow interpolation θ_t = (1−t)θ_0 + tϵ with constant velocity target ϵ−θ_0 is an appropriate transport for posterior estimation
- domain assumption Multinomial resampling with stochastic rejuvenation sufficiently mitigates particle depletion for finite B
Cite this review
Pith. "Pith review of FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation." pith.science (2026). https://pith.science/paper/ZRLYPFMI
@misc{pith2026260705252,
author = {Pith},
title = {Pith review of: FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZRLYPFMI}},
note = {Machine review of arXiv:2607.05252}
}
read the original abstract
Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods struggle with effective multimodal modeling. They often rely on brute-force fusion strategies that ignore the structural disparities between parameters and observations, thus limiting estimation fidelity. In this work, we introduce FUSE (Feynman-Kac steered mUlti-modal flow matching for efficient Simulation-based posterior Estimation). Unlike prior work, FUSE employs a dual-track architecture that preserves the distinct features of multimodal inputs while facilitating dynamic interaction. Additionally, we propose an FK-steered sampling strategy that leverages intermediate observation likelihoods to guide the generative trajectories, effectively improving the sample quality during inference. Our approach outperforms state-of-the-art baselines on standard SBI benchmarks, producing posteriors that closely match ground-truth MCMC. Furthermore, in a real-world exoplanet orbital estimation task, FUSE successfully resolves complex parameter degeneracies that challenge existing methods, highlighting its potential to accelerate complex scientific discoveries in astrophysics and beyond.
Figures
Reference graph
Works this paper leans on
-
[1]
URL https://openaccess.thecvf.co m/content/CVPR2023/html/Bao_All_Are_ Worth_Words_A_ViT_Backbone_for_Diffu sion_Models_CVPR_2023_paper.html. Blunt, S., Wang, J. J., Angelo, I., Ngo, H., Cody, D., De Rosa, R. J., Graham, J. R., Hirsch, L., Nagpal, V ., Nielsen, E. L., Pearce, L., Rice, M., and Tejada, R. or- bitize!: A comprehensive orbit-fitting software ...
-
[2]
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
doi: 10.1109/CVPR52729.2023.01764. URL https://openaccess.thecvf.com/conten t/CVPR2023/html/Brooks_InstructPix2P ix_Learning_To_Follow_Image_Editing_ Instructions_CVPR_2023_paper.html. Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y ., Wang, Z., Kwok, J., Luo, P., Lu, H., and Li, Z. PixArt-α: Fast training of diffusion transformer for photorealistic tex...
-
[3]
URL https://ojs.aaai.org/index.php/AAA I/article/view/30018
doi: 10.1609/aaai.v38i18.30018. URL https://ojs.aaai.org/index.php/AAA I/article/view/30018. Cuturi, M. Sinkhorn distances: Lightspeed computation of optimal transport. InAdvances in Neural Information Processing Systems, volume 26, 2013. URL https: //papers.nips.cc/paper_files/paper/2 013/hash/af21d0c97db2e27e13572cbf59e b343d-Abstract.html. Dax, M., Gre...
-
[4]
URL https: //doi.org/10.1007/978-1-4684-9393-1
doi: 10.1007/978-1-4684-9393-1. URL https: //doi.org/10.1007/978-1-4684-9393-1. Dhariwal, P. and Nichol, A. Diffusion Models Beat GANs on Image Synthesis. InAdvances in Neural Information Processing Systems, volume 34, pp. 8780–8794. Curran Associates, Inc., 2021. URL https://proceeding s.neurips.cc/paper/2021/hash/49ad23d 1ec9fa4bd8d77d02681df5cfa-Abstra...
-
[5]
Classifier-Free Diffusion Guidance
doi: 10.1086/670067. URL https://doi.or g/10.1086/670067. Geffner, T., Papamakarios, G., and Mnih, A. Composi- tional score modeling for simulation-based inference. InProceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Ma- chine Learning Research. PMLR, 2023. URL https: //proceedings.mlr.press/v202/geffner 23a...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.1086/670067 2023
-
[6]
Rectified Flow: A Marginal Preserving Approach to Optimal Transport
URL https://openreview.net/forum ?id=PqvMRDCJT9t. Liu, J., Liu, G., Liang, J., Li, Y ., Liu, J., Wang, X., Wan, P., Zhang, D., and Ouyang, W. Flow-GRPO: Training flow matching models via online RL. InAdvances in Neural Information Processing Systems, volume 38, 2025. URL https://proceedings.neurips.cc/paper _files/paper/2025/hash/3a10c46572628 d58cb44fb70...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.1073/pnas.0306899100 2025
-
[7]
URL https://openaccess.thecvf.co m/content/ICCV2023/html/Peebles_Scal able_Diffusion_Models_with_Transform ers_ICCV_2023_paper.html. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In...
-
[8]
Ruth, M.Exoplanet Orbital Characterization Using Simulation-Based Inference
URL https://openaccess.thecvf.co m/content/CVPR2023/html/Ruiz_DreamBo oth_Fine_Tuning_Text-to-Image_Diffusi on_Models_for_Subject-Driven_Generati on_CVPR_2023_paper.html. Ruth, M.Exoplanet Orbital Characterization Using Simulation-Based Inference. PhD thesis, University of Li`ege, Li`ege, Belgium, June 2024. URL https://ma theo.uliege.be/handle/2268.2/203...
work page 2024
-
[9]
Song, J., Meng, C., and Ermon, S
URL https://proceedings.mlr.pres s/v267/singhal25b.html. Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. InInternational Conference on Learning Representations, 2021a. URL https://openrevi ew.net/forum?id=St1giarCHLP. Song, Y ., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Er- mon, S., and Poole, B. Score-based generative modeling ...
-
[10]
Joint Multimodal Learning with Deep Generative Models
arXiv:1611.01891 [stat]. Tavar´e, S., Balding, D. J., Griffiths, R. C., and Donnelly, P. Inferring Coalescence Times From DNA Sequence Data.Genetics, 145(2):505–518, February 1997. ISSN 1943-2631. doi: 10.1093/genetics/145.2.505. URL 13 FUSE: FK-Steered Multi-Modal Flow Matching https://academic.oup.com/genetics/ar ticle/145/2/505/6018089. Tejero-Cantero,...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.1093/genetics/145.2.505 1997
-
[11]
URL https: //doi.org/10.21105/joss.02505
doi: 10.21105/joss.02505. URL https: //doi.org/10.21105/joss.02505. V ousden, W. D., Farr, W. M., and Mandel, I. Dynamic temperature selection for parallel tempering in Markov chain Monte Carlo simulations.Monthly Notices of the Royal Astronomical Society, 455(2):1919–1937, January
-
[12]
doi: 10.1093/mnras/stv2422. Wang, Q., Kulkarni, S. R., and Verd ´u, S. Divergence es- timation for multidimensional densities via k-nearest- neighbor distances.IEEE Transactions on Information Theory, 55(5):2392–2405, 2009. doi: 10.1109/TIT.2009 .2016060. Wildberger, J., Dax, M., Buchholz, S., Green, S. R., Macke, J. H., and Sch ¨olkopf, B. Flow matching ...
-
[13]
To ensure numerical stability across different tasks, we normalize C by the maximum squared distance found within the reference samples. The distance is then obtained by solving an entropy-regularized optimal transport problem: Sε = X i,j PijCij,(33) where P is the optimal coupling matrix and ε is the regularization parameter. In our implementation, we so...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.