Pith. sign in

REVIEW 3 major objections 5 minor 90 references

CausticFlow maps irregular binary-microlensing light curves to flexible posterior samples in under a second and, after brief local polish, recovers most real events.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 10:57 UTC pith:DKFNUIB7

load-bearing objection Solid methods paper that cleanly joins Neural CDEs and expressive flows for binary microlensing NPE; the 7/10 real recovery is useful but still a thin pillar for survey-scale claims. the 3 major comments →

arxiv 2607.04955 v1 pith:DKFNUIB7 submitted 2026-07-06 astro-ph.IM astro-ph.EPastro-ph.GAastro-ph.SR

CausticFlow: An Efficient Machine Learning Framework Combining Neural Differential Equations and Normalizing Flows for Binary Microlensing Parameter Inference

classification astro-ph.IM astro-ph.EPastro-ph.GAastro-ph.SR
keywords gravitational microlensingbinary lens microlensingneural networkstime series analysisnormalizing flowsneural differential equationssimulation-based inference
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Binary microlensing events carry strong but hard-to-search degeneracies, and traditional global modeling is too slow for the coming survey flood. CausticFlow trains a neural controlled differential equation on irregular light curves so that a normalizing flow can output a full surrogate posterior over the main binary parameters in a fraction of a second. On simulated KMTNet-like data the maximum-a-posteriori point already reaches roughly 17 percent precision in mass ratio and 3 percent in projected separation; drawing a few samples and polishing them with ordinary local optimizers improves those figures to under 5 percent and under 1 percent and recovers the true model chi-squared for about 80 percent of events. Applied to ten real binary events that include higher-order effects, uneven cadence, and real noise, the same workflow recovers parameters, light-curve shape, and lens geometry for seven events in roughly ten CPU minutes each. The paper therefore presents the learned posterior as a practical proposal engine that can replace exhaustive grid searches for systematic analysis of large microlensing samples.

Core claim

A single trained network that pairs a neural controlled differential equation with a normalizing flow produces usable posterior samples for binary microlensing parameters from irregular, gappy light curves. Used as a proposal distribution for local optimization, it recovers model chi-squared for about 80 percent of simulated events at high precision and, after simple refinement, recovers parameters, morphology, and geometry for 7 of 10 real events despite clear mismatches between training simulations and real data.

What carries the argument

CausticFlow: a Neural CDE that compresses an irregular light curve (via log-signatures) into a fixed conditioning vector, followed by a Masked Autoregressive Flow that transforms a simple base density into a multimodal surrogate posterior over (t_E, u_0, ρ, q, s, α). The CDE absorbs irregular sampling and gaps; the flow preserves close/wide and other degeneracies so that a handful of draws seed local optimizers.

Load-bearing premise

The network is trained only on white-noise simulations without parallax or orbital motion and with mass ratio above one-thousandth, yet those proposals are assumed to remain useful for real light curves that have correlated noise, higher-order effects, and parameters outside that range.

What would settle it

Run the identical trained model and the same polishing budget on a larger, independently chosen set of real binary events that deliberately include strong orbital motion or parallax and events near or below the training mass-ratio floor; if recovery falls well below 7/10 and polished chi-squared systematically misses published solutions by hundreds, the claim that the learned posterior is a reliable real-world proposal fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • MAP precisions of roughly 17 percent in mass ratio and 3 percent in separation are already usable for statistical studies of high-mass-ratio binaries.
  • Ten local polishes seeded by the learned posterior recover model chi-squared for about 80 percent of simulated events at under 5 percent and under 1 percent precision in q and s.
  • The same amortized engine recovers 7 of 10 real events after about 10 CPU minutes of refinement each, making systematic modeling of large archives feasible.
  • Physically relevant multimodality (close/wide, trajectory-angle flip) is retained in the surrogate posterior and can be recovered by polishing rather than by exhaustive grids.
  • The workflow is positioned as a first-pass proposal stage for high-volume surveys such as Roman, CSST, and ET.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Injecting simulated signals into real baseline photometry would likely raise recovery on weak-anomaly and out-of-prior events that currently fail.
  • Lowering the training mass-ratio floor would turn the same architecture into a proposal engine for planetary microlensing, the main driver of space surveys.
  • Feeding photometric uncertainties and mild higher-order effects into training would shrink the residual simulation–reality gap without changing the core design.
  • Once amortized, the posterior could rank archival events by predicted caustic topology before any human modeling, enabling automated triage of large catalogs.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents CausticFlow, a neural posterior estimation framework that embeds irregular microlensing light curves with Neural RDEs (log-signature control paths) and models the conditional posterior over binary-lens parameters (t_E, u_0, ρ, q, s, α) with a conditional MAF of rational-quadratic splines. Trained on 2 imes10^6 simulated KMTNet-like light curves (q≥10^{-3}, white Gaussian noise, no higher-order effects, χ^{2}_PSPL/dof>2 cut), the network produces MAP estimates with MAD ~0.07–0.1 dex in log q and log s; drawing N=10 samples and polishing with local optimizers recovers model χ^{2} for ~80% of a held-out simulated set and tightens precisions to <5% in q and <1% in s. On 10 real KMTNet events from Han et al. (2026) that include higher-order effects, out-of-prior ho/t_E, and real noise, the same workflow recovers parameters, morphology, and geometry for 7/10 events in ~10 CPU minutes after polishing. The authors position the method as a fast proposal engine for systematic binary-lens modeling in large surveys.

Significance. If the claimed recovery rates hold under broader real-event tests, CausticFlow would supply a practical, amortized proposal distribution that substantially reduces the cost of exhaustive binary-lens modeling—currently a bottleneck for stellar-binary and remnant population studies and for the expected event rates of Roman, CSST, and ET. Strengths include an architecture well-matched to irregular sampling and multimodal posteriors, transparent reporting of both successes and failures on real data (Table 1, Figs. 7–8), explicit comparison to MCMC on a multimodal example (Fig. 4), and a clear separation between the amortized network and conventional local polishing. The work is a concrete advance over earlier NPE microlensing efforts that either assumed regular sampling or used less flexible posterior models.

major comments (3)
  1. Section 5 and Table 1: the central claim that CausticFlow is a “fast and robust proposal engine” for large-scale surveys rests on a 7/10 recovery rate for a hand-selected sample of 10 events with log q > −3. The three failures (KMT-2023-BLG-1246, -1056, -2427) coincide exactly with the acknowledged training–reality mismatches (weak anomaly + real noise; t_E/ ho outside prior; strong orbital motion). With N=10 and literature-assisted preprocessing, it is not yet clear whether ~70% is representative or an optimistic upper bound. A larger, less curated real-event test set (or an explicit statement of the intended domain of applicability) is needed before the survey-scale claim can be regarded as demonstrated.
  2. Sections 3.2–3.3 and 5: training uses white Gaussian noise, no annual parallax or lens orbital motion, and a hard χ^{2}_PSPL/dof>2 cut that shifts the sample toward strong caustic features. Real events routinely contain correlated noise, higher-order effects that alter light-curve morphology, and weaker anomalies. The paper correctly notes these gaps, but the only quantitative evidence that the learned posterior remains a useful proposal under such shifts is the 10-event test. Without either (i) injection of simulated signals into real baseline light curves or (ii) a controlled ablation that quantifies degradation as higher-order effects or noise realism increase, the generalization argument remains under-supported relative to the abstract’s claim.
  3. Section 4 / Fig. 6: recovery is defined by Δχ^{2}<100 (sim) or <50 (real) relative to the input or literature model, and success is reported after polishing N=10–20 posterior draws. These thresholds and the polishing budget are free parameters of the evaluation. The paper should show how the recovery fraction and parameter MAD change with N and with a stricter Δχ^{2} cut, and should clarify whether the reported ~80% / 7/10 figures remain stable under modest changes of these choices; otherwise the performance numbers are difficult to compare with traditional grid-search pipelines.
minor comments (5)
  1. Figure 3 vs. Figure 5: the improvement from MAP to polished solutions is clear, but the text would benefit from a short quantitative summary (e.g., MAD ratios for each parameter) rather than leaving the reader to extract numbers from the panels.
  2. Section 3.1: the template-matching procedure used to recover t_0 is described only briefly. A short statement of its failure rate on the held-out set would strengthen confidence that the six-parameter network plus post-hoc t_0 search is robust.
  3. Section 2.1: the log-signature depth k=4 and window W=10 are stated without a sensitivity check. Even a one-sentence note that nearby (k,W) pairs give comparable validation NLL would help.
  4. Table 1: the column “Outside training range” is useful; adding a brief note on whether the polished solutions for the recovered events still lie inside the prior volume would clarify how often the network is extrapolating.
  5. References: the software stack is thoroughly cited; a short data/code availability statement (even if weights are released later) would aid reproducibility.

Circularity Check

1 steps flagged

No significant circularity: amortized NPE trained on simulated (theta, D) pairs is evaluated on held-out simulations and independent real events; mild literature-assisted preprocessing does not force the recovery rates.

specific steps
  1. self citation load bearing [Section 5, real-event preprocessing paragraph]
    "We start from the aligned light curves from C. Han (private comm.); this light curve alignment is in general independent of the microlensing model, although for this test we have made use of the best-fit models of C. Han et al. (2026) for simplicity."

    The only mild circularity is that real-event light-curve alignment for the 10-event test set uses the same literature best-fit models that later serve as the reference solutions. Alignment is stated to be generally model-independent, and the network + polishing stage can (and does) still fail, so the dependence does not force the reported 7/10 recovery rate. It is a convenience choice, not a definitional reduction of the central claim.

full rationale

CausticFlow is a standard neural posterior estimation pipeline. Parameters are drawn from explicit priors (Eq. 10), light curves are simulated with microlux under a KMTNet-like cadence and white-noise model, and the network is trained by minimizing the expected negative log-likelihood of the surrogate posterior (Eq. 3). MAP and polished-solution metrics are reported on a held-out set of 10^5 simulated light curves never seen during training; recovery fractions (Delta chi^2 < 100 for ~80% with N=10 starts) are therefore genuine out-of-sample performance, not tautologies. The real-event test uses 10 events from Han et al. (2026) whose literature solutions are transformed into the paper's coordinate convention and refined with a static binary model; these serve as external reference solutions. Preprocessing of the real light curves does make use of the Han et al. best-fit models for alignment 'for simplicity,' which introduces a mild dependence on the same literature, but the subsequent network prediction and local polishing are still free to fail (and do fail for 3/10 events). No equation equates a claimed precision or recovery rate to a fitted input by construction, no uniqueness theorem is imported from the authors' prior work to forbid alternatives, and the architecture citations (Neural CDEs, MAF/RQS) are external. The result is therefore self-contained against its stated benchmarks; the residual score of 1 reflects only the non-load-bearing literature-assisted alignment step.

Axiom & Free-Parameter Ledger

7 free parameters · 4 axioms · 1 invented entities

The central performance claims rest on standard microlensing physics, standard SBI/NPE training, and several modeling choices that define the training distribution and success metrics. No new physical entities are postulated; free parameters are mostly network/training hyperparameters and selection/recovery thresholds that shape reported success rates.

free parameters (7)
  • log-signature truncation depth k and window W
    Set to k=4, W=10 for 1000-point series; controls the Neural RDE control path and is chosen for speed/stability rather than derived.
  • Neural CDE / ResNet widths (d_z=512, d_cond=256, 3 identity blocks)
    Architecture sizes chosen by authors; affect capacity of the embedding and thus posterior quality.
  • MAF depth and RQS knots (10 layers, 32 knots)
    Expressivity hyperparameters for the conditional density estimator; not fixed by theory.
  • Photometric noise model (σ_sys=0.006, m0=21.755, background scaling)
    Empirical KMTS-like noise curve used to generate all training data; mismatches real correlated noise.
  • Single-lens rejection cut χ²_PSPL/dof > 2
    Hand-chosen filter that reshapes the training distribution toward strong binary signatures (Fig. 2).
  • Parameter priors (especially q≥10^{-3}, ρ, t_E ranges)
    Define the support of the learned posterior; several real failures sit outside these ranges.
  • Recovery thresholds Δχ² < 100 (sim) / < 50 (real) and N=10 or 20 starts
    Success fractions (~80%, 7/10) depend on these operational definitions of recovery.
axioms (4)
  • domain assumption Standard static binary-lens magnification model with center-of-magnification coordinates and flux normalization (Eqs. 4–8) adequately describes the light curves of interest when higher-order effects are weak.
    Used throughout simulation and polishing; real events with strong parallax/orbital motion are acknowledged mismatches (Section 5).
  • domain assumption Neural CDEs with log-signature controls yield sampling-invariant summaries of irregular microlensing time series (Kidger et al.; Morrill et al.).
    Core inductive bias of the embedding (Section 2.1).
  • standard math Minimizing expected NLL of a conditional normalizing flow yields a useful amortized approximation to the Bayesian posterior (NPE / SBI).
    Training objective Eq. (3); standard in the SBI literature cited.
  • ad hoc to paper Local polishing from a small number of flow samples can recover global modes when the surrogate posterior covers the relevant branches.
    Operational claim underlying the 80% / 7-of-10 recovery numbers (Sections 4–5).
invented entities (1)
  • CausticFlow architecture (Neural RDE embedding + conditional MAF/RQS) no independent evidence
    purpose: Amortized proposal distribution for binary microlensing parameters from irregular light curves.
    Named framework combining existing components; not a new physical object. Independent evidence is empirical performance on held-out sim and real events, not an external physical prediction.

pith-pipeline@v1.1.0-grok45 · 23568 in / 3422 out tokens · 30919 ms · 2026-07-11T10:57:05.810511+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of CausticFlow: An Efficient Machine Learning Framework Combining Neural Differential Equations and Normalizing Flows for Binary Microlensing Parameter Inference." pith.science (2026). https://pith.science/paper/DKFNUIB7

@misc{pith2026260704955,
  author       = {Pith},
  title        = {Pith review of: CausticFlow: An Efficient Machine Learning Framework Combining Neural Differential Equations and Normalizing Flows for Binary Microlensing Parameter Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DKFNUIB7}},
  note         = {Machine review of arXiv:2607.04955}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce CausticFlow, a machine learning framework that combines neural controlled differential equations with normalizing flows to infer binary microlensing parameters. This architecture naturally handles irregularly sampled time series and data gaps while flexibly capturing strongly correlated and multimodal posterior distributions. Trained on simulated KMTNet-like light curves, CausticFlow generates posterior samples in a fraction of a second, with maximum-a-posteriori estimates achieving typical precisions of $\sim17\%$ for the mass ratio $q$ and $\sim3\%$ for the projected separation $s$. When used as a proposal distribution for downstream local optimization, the framework improves these precisions to $<5\%$ and $<1\%$, respectively, and recovers model $\chi^2$ for $\sim80\%$ of simulated events. We test the generalizability of the framework on 10 real binary lensing events characterized by higher-order effects, varied cadences, and real-world noise. Despite these mismatches between simulation and reality, CausticFlow successfully recovers the model parameters, light-curve morphology, and lensing geometry for 7 of the 10 events after simple local refinement, achieving precision levels comparable to those found for simulated data in 10 CPU minutes per event. These results demonstrate that CausticFlow acts as a fast and robust proposal engine, bridging the gap between the rapid influx of data and the need for systematic modeling in large-scale microlensing surveys such as Roman, CSST, and ET.

Figures

Figures reproduced from arXiv: 2607.04955 by Haibin Ren, Wei Zhu.

Figure 1
Figure 1. Figure 1: Overview of the CausticFlow architecture. Given an irregularly sampled light curve D = {(ti, xi)}, the framework first applies a log-signature transformation to construct a compact control path X. An initial network maps the initial value of the control path to the initial hidden state of the Neural CDE. The Neural CDE then evolves this hidden state along the full control path and produces the final summar… view at source ↗
Figure 2
Figure 2. Figure 2: Distribution of simulated binary lens events re￾tained after the single-lens rejection cut, χ 2 PSPL/dof > 2, in the (log s, log q) plane. The selection shifts the sample toward parameter regions with stronger binary signatures, while still preserving broad coverage of the close, resonant, and wide caustic topologies. The two black curves mark the boundaries between resonant and close/wide caustic geome￾tr… view at source ↗
Figure 16
Figure 16. Figure 16: Although the comparison is not strictly fair [PITH_FULL_IMAGE:figures/full_fig_p006_16.png] view at source ↗
Figure 3
Figure 3. Figure 3: True versus MAP-predicted parameters on the held-out simulated test set for u0, tE, α, log ρ, log q, and log s. The red dashed lines indicate the one-to-one relation, and the RMSE and MAD are listed in each panel. The off-diagonal structures in the α and log s panels trace the learned physical degeneracies, while the broader dispersion in log ρ reflects weaker finite-source constraints. dicted parameter sp… view at source ↗
Figure 4
Figure 4. Figure 4: Posterior comparison for a representative multimodal event exhibiting the close/wide degeneracy from the simulated dataset. The left corner plot compares the posterior structure inferred from MCMC (red) with the surrogate posterior from CausticFlow (blue) in the six-dimensional parameter subspace (u0, tE, α, log ρ, log q, log s). The red point and lines mark the true input parameters, while the yellow and … view at source ↗
Figure 5
Figure 5. Figure 5: True versus polished parameters on the held-out simulated test set. For each event, N = 10 initial points are drawn from the learned posterior and refined by local optimization, and the best-polished solution is shown. The layout is the same as in [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Cumulative distributions of the best-fit χ 2 val￾ues obtained after local minimization from N initial points drawn from the learned posterior. Different curves corre￾spond to different numbers of starting points, up to N = 10. The vertical dashed line marks χ 2 = 1100. failed cases show larger deviations in at least some of the model parameters. For the seven recovered events, the Fisher-matrix uncertainti… view at source ↗
Figure 7
Figure 7. Figure 7: Comparison between the reference parameters and the corresponding polished parameters obtained from the learned posterior for the real-event test sample. The reference parameters are obtained by transforming the literature solutions into our parameter convention and refining them with a static binary lens model, so they are not the same as given in the discovery paper (C. Han et al. 2026). The red dashed l… view at source ↗
Figure 8
Figure 8. Figure 8: Light-curve and lens-geometry comparison for the real-event test sample. Each row shows the aligned light curve and the corresponding source trajectory relative to the caustic structure. ACKNOWLEDGMENTS The authors thank Cheongho Han for providing the real￾event data. We also thank Qingru Hu and Hongjing Yang for assistance with early tests and relevant discus￾sion. This work is supported by the National N… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

90 extracted references · 16 canonical work pages · 1 internal anchor

  1. [1]

    2000, title Detection of rotation in a binary microlens: PLANET photometry of MACHO 97-BLG-41, The Astrophysical Journal, 534, 894

    Albrow, M., Beaulieu, J.-P., Caldwell, J., et al. 2000, title Detection of rotation in a binary microlens: PLANET photometry of MACHO 97-BLG-41, The Astrophysical Journal, 534, 894

  2. [2]

    An , J. H. 2005, title Gravitational lens under perturbations: symmetry of perturbing potentials with invariant caustics , , 356, 1409, 10.1111/j.1365-2966.2004.08581.x

  3. [3]

    H., & Han , C

    An , J. H., & Han , C. 2002, title Effect of a Wide Binary Companion to the Lens on the Astrometric Behavior of Gravitational Microlensing Events , , 573, 351, 10.1086/340557

  4. [4]

    P., Anderson , J., & Gaudi , B

    Bennett , D. P., Anderson , J., & Gaudi , B. S. 2007, title Characterization of Gravitational Microlensing Planetary Host Stars , , 660, 781, 10.1086/513013

  5. [5]

    Bozza, V. 2010, title Microlensing with an Advanced Contour Integration Algorithm: Green 's Theorem to Third Order, Error Control, Optimal Sampling and Limb Darkening: Advanced Contour Integration in Microlensing, , 408, 2188, 10.1111/j.1365-2966.2010.17265.x

  6. [6]

    2024, title RTModel: A platform for real-time modeling and massive analyses of microlensing events , , 688, A83, 10.1051/0004-6361/202450450

    Bozza , V. 2024, title RTModel: A platform for real-time modeling and massive analyses of microlensing events , , 688, A83, 10.1051/0004-6361/202450450

  7. [7]

    2025, title VBMicroLensing: Three algorithms for multiple lensing with contour integration , , 694, A219, 10.1051/0004-6361/202452648

    Bozza , V., Saggese , V., Covone , G., Rota , P., & Zhang , J. 2025, title VBMicroLensing: Three algorithms for multiple lensing with contour integration , , 694, A219, 10.1051/0004-6361/202452648

  8. [8]

    2018, JAX : composable transformations of P ython+ N um P y programs, 0.3.13 http://github.com/google/jax

    Bradbury, J., Frostig, R., Hawkins, P., et al. 2018, JAX : composable transformations of P ython+ N um P y programs, 0.3.13 http://github.com/google/jax

  9. [9]

    A., Coleman, T

    Branch, M. A., Coleman, T. F., & Li, Y. 1999, title A subspace, interior, and conjugate gradient method for large-scale bound-constrained minimization problems, SIAM Journal on Scientific Computing, 21, 1

  10. [10]

    2008, title Microlensing constraints on the Galactic bulge initial mass function , , 480, 723, 10.1051/0004-6361:20078439

    Calchi Novati , S., de Luca , F., Jetzer , P., Mancini , L., & Scarpetta , G. 2008, title Microlensing constraints on the Galactic bulge initial mass function , , 480, 723, 10.1051/0004-6361:20078439

  11. [11]

    2005, title Properties of Central Caustics in Planetary Microlensing , , 630, 535, 10.1086/432048

    Chung , S.-J., Han , C., Park , B.-G., et al. 2005, title Properties of Central Caustics in Planetary Microlensing , , 630, 535, 10.1086/432048

  12. [12]

    2020, title The frontier of simulation-based inference, Proceedings of the National Academy of Sciences, 117, 30055

    Cranmer, K., Brehmer, J., & Louppe, G. 2020, title The frontier of simulation-based inference, Proceedings of the National Academy of Sciences, 117, 30055

  13. [13]

    2026, title Introduction to the Chinese Space Station Survey Telescope (CSST) , Science China Physics, Mechanics, and Astronomy, 69, 239501, 10.1007/s11433-025-2809-0

    CSST Collaboration , Gong , Y., Miao , H., et al. 2026, title Introduction to the Chinese Space Station Survey Telescope (CSST) , Science China Physics, Mechanics, and Astronomy, 69, 239501, 10.1007/s11433-025-2809-0

  14. [14]

    2020, The D eep M ind JAX E cosystem, http://github.com/google-deepmind

    DeepMind, Babuschkin, I., Baumli, K., et al. 2020, The D eep M ind JAX E cosystem, http://github.com/google-deepmind

  15. [15]

    2025, title Simulation-based inference: A practical guide, arXiv preprint arXiv:2508.12939

    Deistler, M., Boelts, J., Steinbach, P., et al. 2025, title Simulation-based inference: A practical guide, arXiv preprint arXiv:2508.12939

  16. [16]

    1996, title Do Microlensing Events Repeat? , , 457, 93, 10.1086/176713

    Di Stefano , R., & Mao , S. 1996, title Do Microlensing Events Repeat? , , 457, 93, 10.1086/176713

  17. [17]

    1998, title Galactic microlensing with rotating binaries , , 329, 361, 10.48550/arXiv.astro-ph/9702039

    Dominik , M. 1998, title Galactic microlensing with rotating binaries , , 329, 361, 10.48550/arXiv.astro-ph/9702039

  18. [18]

    1999, title The binary gravitational lens and its extreme cases , , 349, 108, 10.48550/arXiv.astro-ph/9903014

    Dominik , M. 1999, title The binary gravitational lens and its extreme cases , , 349, 108, 10.48550/arXiv.astro-ph/9903014

  19. [19]

    A., et al

    Droettboom, M., Hunter, J., Caswell, T. A., et al. 2016, matplotlib: matplotlib v1.5.1, v1.5.1 Zenodo, 10.5281/zenodo.44579

  20. [20]

    2019, title Neural spline flows, Advances in neural information processing systems, 32

    Durkan, C., Bekasov, A., Murray, I., & Papamakarios, G. 2019, title Neural spline flows, Advances in neural information processing systems, 32

  21. [21]

    2016, title corner.py: Scatterplot matrices in Python, The Journal of Open Source Software, 1, 24, 10.21105/joss.00024

    Foreman-Mackey, D. 2016, title corner.py: Scatterplot matrices in Python, The Journal of Open Source Software, 1, 24, 10.21105/joss.00024

  22. [22]

    W., Lang , D., & Goodman , J

    Foreman-Mackey , D., Hogg , D. W., Lang , D., & Goodman , J. 2013, title emcee: The MCMC Hammer , , 125, 306, 10.1086/670067

  23. [23]

    2012, title Implementing the Nelder-Mead Simplex Algorithm with Adaptive Parameters, Computational Optimization and Applications, 51, 259, 10.1007/s10589-010-9329-3

    Gao, F., & Han, L. 2012, title Implementing the Nelder-Mead Simplex Algorithm with Adaptive Parameters, Computational Optimization and Applications, 51, 259, 10.1007/s10589-010-9329-3

  24. [24]

    Gaudi , B. S. 2012, title Microlensing Surveys for Exoplanets , , 50, 411, 10.1146/annurev-astro-081811-125518

  25. [25]

    S., & Gould , A

    Gaudi , B. S., & Gould , A. 1997, title Planet Parameters in Microlensing Events , , 486, 85, 10.1086/304491

  26. [26]

    2022, title ET White Paper: To Find the First Earth 2.0 , arXiv e-prints, arXiv:2206.06693, 10.48550/arXiv.2206.06693

    Ge , J., Zhang , H., Zang , W., et al. 2022, title ET White Paper: To Find the First Earth 2.0 , arXiv e-prints, arXiv:2206.06693, 10.48550/arXiv.2206.06693

  27. [27]

    1992, title Extending the MACHO Search to approximately 10 6 M sub sun , , 392, 442, 10.1086/171443

    Gould , A. 1992, title Extending the MACHO Search to approximately 10 6 M sub sun , , 392, 442, 10.1086/171443

  28. [28]

    2000, title A Natural Formalism for Microlensing , , 542, 785, 10.1086/317037

    Gould , A. 2000, title A Natural Formalism for Microlensing , , 542, 785, 10.1086/317037

  29. [29]

    1992, title Discovering Planetary Systems through Gravitational Microlenses , , 396, 104, 10.1086/171700

    Gould , A., & Loeb , A. 1992, title Discovering Planetary Systems through Gravitational Microlenses , , 396, 104, 10.1086/171700

  30. [30]

    2021, title Masses for free-floating planets and dwarf planets , Research in Astronomy and Astrophysics, 21, 133, 10.1088/1674-4527/21/6/133

    Gould , A., Zang , W.-C., Mao , S., & Dong , S.-B. 2021, title Masses for free-floating planets and dwarf planets , Research in Astronomy and Astrophysics, 21, 133, 10.1088/1674-4527/21/6/133

  31. [31]

    2019, title Automatic posterior transformation for likelihood-free inference, in International conference on machine learning, PMLR, 2404--2414

    Greenberg, D., Nonnenmacher, M., & Macke, J. 2019, title Automatic posterior transformation for likelihood-free inference, in International conference on machine learning, PMLR, 2404--2414

  32. [32]

    1998, title The Use of High-Magnification Microlensing Events in Discovering Extrasolar Planets , , 500, 37, 10.1086/305729

    Griest , K., & Safizadeh , N. 1998, title The Use of High-Magnification Microlensing Events in Discovering Extrasolar Planets , , 500, 37, 10.1086/305729

  33. [33]

    Hampel, F. R. 1974, title The influence curve and its role in robust estimation, Journal of the american statistical association, 69, 383

  34. [34]

    2006, title Properties of Planetary Caustics in Gravitational Microlensing , , 638, 1080, 10.1086/498937

    Han , C. 2006, title Properties of Planetary Caustics in Gravitational Microlensing , , 638, 1080, 10.1086/498937

  35. [35]

    Candidate Microlensing Brown Dwarfs in Binary Lens Systems from the 2023--2025 Observing Seasons

    Han , C., Udalski , A., Bond , I. A., et al. 2026, title Candidate Microlensing Brown Dwarfs in Binary Lens Systems from the 2023--2025 Observing Seasons , arXiv e-prints, arXiv:2604.07932, 10.48550/arXiv.2604.07932

  36. [36]

    R., Millman , K

    Harris , C. R., Millman , K. J., van der Walt , S. J., et al. 2020, title Array programming with NumPy , , 585, 357, 10.1038/s41586-020-2649-2

  37. [37]

    2016 a , title Deep residual learning for image recognition, in Proceedings of the IEEE conference on computer vision and pattern recognition, 770--778

    He, K., Zhang, X., Ren, S., & Sun, J. 2016 a , title Deep residual learning for image recognition, in Proceedings of the IEEE conference on computer vision and pattern recognition, 770--778

  38. [38]

    2016 b , title Identity mappings in deep residual networks, in European conference on computer vision, Springer, 630--645

    He, K., Zhang, X., Ren, S., & Sun, J. 2016 b , title Identity mappings in deep residual networks, in European conference on computer vision, Springer, 630--645

  39. [39]

    Hunter, J. D. 2007, title Matplotlib: A 2D Graphics Environment, Computing in Science & Engineering, 9, 90, 10.1109/MCSE.2007.55

  40. [40]

    K., Zang , W., Han , C., et al

    Jung , Y. K., Zang , W., Han , C., et al. 2022, title Systematic KMTNet Planetary Anomaly Search. VI. Complete Sample of 2018 Sub-prime-field Planets , , 164, 262, 10.3847/1538-3881/ac9c5c

  41. [41]

    2022, title On Neural Differential Equations , arXiv e-prints, arXiv:2202.02435, 10.48550/arXiv.2202.02435

    Kidger , P. 2022, title On Neural Differential Equations , arXiv e-prints, arXiv:2202.02435, 10.48550/arXiv.2202.02435

  42. [42]

    2021, title E quinox: neural networks in JAX via callable P y T rees and filtered transformations, Differentiable Programming workshop at Neural Information Processing Systems 2021

    Kidger, P., & Garcia, C. 2021, title E quinox: neural networks in JAX via callable P y T rees and filtered transformations, Differentiable Programming workshop at Neural Information Processing Systems 2021

  43. [43]

    2021, title S ignatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU , in International Conference on Learning Representations

    Kidger, P., & Lyons, T. 2021, title S ignatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU , in International Conference on Learning Representations

  44. [44]

    2020, title Neural controlled differential equations for irregular time series, Advances in neural information processing systems, 33, 6696

    Kidger, P., Morrill, J., Foster, J., & Lyons, T. 2020, title Neural controlled differential equations for irregular time series, Advances in neural information processing systems, 33, 6696

  45. [45]

    Kim , S.-L., Lee , C.-U., Park , B.-G., et al. 2016, title KMTNET: A Network of 1.6 m Wide-Field Optical Telescopes Installed at Three Southern Observatories , Journal of Korean Astronomical Society, 49, 37, 10.5303/JKAS.2016.49.1.37

  46. [46]

    P., Salimans, T., Jozefowicz, R., et al

    Kingma, D. P., Salimans, T., Jozefowicz, R., et al. 2016, title Improved variational inference with inverse autoregressive flow, Advances in neural information processing systems, 29

  47. [47]

    2016, title Jupyter Notebooks -- a publishing format for reproducible computational workflows, in Positioning and Power in Academic Publishing: Players, Agents and Agendas, ed

    Kluyver, T., Ragan-Kelley, B., P \'e rez, F., et al. 2016, title Jupyter Notebooks -- a publishing format for reproducible computational workflows, in Positioning and Power in Academic Publishing: Players, Agents and Agendas, ed. F. Loizides & B. Scmidt (IOS Press), 87--90. https://eprints.soton.ac.uk/403913/

  48. [48]

    2015, title The complete catalogue of light curves in equal-mass binary microlensing , , 450, 1565, 10.1093/mnras/stv733

    Liebig , C., D'Ago , G., Bozza , V., & Dominik , M. 2015, title The complete catalogue of light curves in equal-mass binary microlensing , , 450, 1565, 10.1093/mnras/stv733

  49. [49]

    2019, title Decoupled Weight Decay Regularization, in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 (OpenReview.net)

    Loshchilov, I., & Hutter, F. 2019, title Decoupled Weight Decay Regularization, in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 (OpenReview.net). https://openreview.net/forum?id=Bkg6RiCqY7

  50. [50]

    J., Bassetto, G., et al

    Lueckmann, J.-M., Goncalves, P. J., Bassetto, G., et al. 2017, title Flexible statistical inference for mechanistic models of neural dynamics, Advances in neural information processing systems, 30

  51. [51]

    H., Gunn , J

    Lupton , R. H., Gunn , J. E., & Szalay , A. S. 1999, title A Modified Magnitude System that Produces Well-Behaved Magnitudes, Colors, and Errors Even for Low Signal-to-Noise Ratio Measurements , , 118, 1406, 10.1086/301004

  52. [52]

    1991, title Gravitational Microlensing by Double Stars and Planetary Systems , , 374, L37, 10.1086/186066

    Mao, S., & Paczynski, B. 1991, title Gravitational Microlensing by Double Stars and Planetary Systems , , 374, L37, 10.1086/186066

  53. [53]

    2020, title A Generalised Signature Method for Multivariate Time Series Feature Extraction , arXiv e-prints, arXiv:2006.00873, 10.48550/arXiv.2006.00873

    Morrill , J., Fermanian , A., Kidger , P., & Lyons , T. 2020, title A Generalised Signature Method for Multivariate Time Series Feature Extraction , arXiv e-prints, arXiv:2006.00873, 10.48550/arXiv.2006.00873

  54. [54]

    2021, title Neural Controlled Differential Equations for Online Prediction Tasks , arXiv e-prints, arXiv:2106.11028, 10.48550/arXiv.2106.11028

    Morrill , J., Kidger , P., Yang , L., & Lyons , T. 2021, title Neural Controlled Differential Equations for Online Prediction Tasks , arXiv e-prints, arXiv:2106.11028, 10.48550/arXiv.2106.11028

  55. [55]

    2021, title Neural Rough Differential Equations for Long Time Series, in Proceedings of Machine Learning Research, Vol

    Morrill, J., Salvi, C., Kidger, P., & Foster, J. 2021, title Neural Rough Differential Equations for Long Time Series, in Proceedings of Machine Learning Research, Vol. 139, Proceedings of the 38th International Conference on Machine Learning, ed. M. Meila & T. Zhang (PMLR), 7829--7838. https://proceedings.mlr.press/v139/morrill21b.html

  56. [56]

    2017, title No large population of unbound or wide-orbit Jupiter-mass planets , , 548, 183, 10.1038/nature23276

    Mr \'o z , P., Udalski , A., Skowron , J., et al. 2017, title No large population of unbound or wide-orbit Jupiter-mass planets , , 548, 183, 10.1038/nature23276

  57. [57]

    2019, title Microlensing Optical Depth and Event Rate toward the Galactic Bulge from 8 yr of OGLE-IV Observations , , 244, 29, 10.3847/1538-4365/ab426b

    Mr \'o z , P., Udalski , A., Skowron , J., et al. 2019, title Microlensing Optical Depth and Event Rate toward the Galactic Bulge from 8 yr of OGLE-IV Observations , , 244, 29, 10.3847/1538-4365/ab426b

  58. [58]

    A., & Mead, R

    Nelder, J. A., & Mead, R. 1965, title A Simplex Method for Function Minimization , The Computer Journal, 7, 308, 10.1093/comjnl/7.4.308

  59. [59]

    Oliveira , R. A. P., Poleski , R., Mr \'o z , P., et al. 2025, title Automated Detection and Modeling of Binary Microlensing Events in OGLE-IV data. I. Events with Well-Separated Bumps , , 75, 27, 10.32023/0001-5237/75.1.2

  60. [60]

    1986, title Gravitational Microlensing by the Galactic Halo , , 304, 1, 10.1086/164140

    Paczynski , B. 1986, title Gravitational Microlensing by the Galactic Halo , , 304, 1, 10.1086/164140

  61. [61]

    2016, title Fast -free inference of simulation models with bayesian conditional density estimation, Advances in neural information processing systems, 29

    Papamakarios, G., & Murray, I. 2016, title Fast -free inference of simulation models with bayesian conditional density estimation, Advances in neural information processing systems, 29

  62. [62]

    J., Mohamed, S., & Lakshminarayanan, B

    Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., & Lakshminarayanan, B. 2021, title Normalizing flows for probabilistic modeling and inference, Journal of Machine Learning Research, 22, 1

  63. [63]

    2017, title Masked autoregressive flow for density estimation, Advances in neural information processing systems, 30

    Papamakarios, G., Pavlakou, T., & Murray, I. 2017, title Masked autoregressive flow for density estimation, Advances in neural information processing systems, 30

  64. [64]

    2013, title On the difficulty of training recurrent neural networks, in International conference on machine learning, Pmlr, 1310--1318

    Pascanu, R., Mikolov, T., & Bengio, Y. 2013, title On the difficulty of training recurrent neural networks, in International conference on machine learning, Pmlr, 1310--1318

  65. [65]

    T., Gaudi , B

    Penny , M. T., Gaudi , B. S., Kerins , E., et al. 2019, title Predictions of the WFIRST Microlensing Survey. I. Bound Planet Detection Rates , , 241, 3, 10.3847/1538-4365/aafb69

  66. [66]

    2025, title A Differentiable Binary Microlensing Model Using Adaptive Contour Integration Method , , 169, 170, 10.3847/1538-3881/adb1b2

    Ren , H., & Zhu , W. 2025, title A Differentiable Binary Microlensing Model Using Adaptive Contour Integration Method , , 169, 170, 10.3847/1538-3881/adb1b2

  67. [67]

    2015, title Variational inference with normalizing flows, in International conference on machine learning, PMLR, 1530--1538

    Rezende, D., & Mohamed, S. 2015, title Variational inference with normalizing flows, in International conference on machine learning, PMLR, 1530--1538

  68. [68]

    2008, title MOA-cam3: a wide-field mosaic CCD camera for a gravitational microlensing survey in New Zealand , Experimental Astronomy, 22, 51, 10.1007/s10686-007-9082-5

    Sako , T., Sekiguchi , T., Sasaki , M., et al. 2008, title MOA-cam3: a wide-field mosaic CCD camera for a gravitational microlensing survey in New Zealand , Experimental Astronomy, 22, 51, 10.1007/s10686-007-9082-5

  69. [69]

    Savitzky, A., & Golay, M. J. 1964, title Smoothing and differentiation of data by simplified least squares procedures., Analytical chemistry, 36, 1627

  70. [70]

    2016, title The frequency of snowline-region planets from four years of OGLE-MOA-Wise second-generation microlensing , , 457, 4089, 10.1093/mnras/stw191

    Shvartzvald , Y., Maoz , D., Udalski , A., et al. 2016, title The frequency of snowline-region planets from four years of OGLE-MOA-Wise second-generation microlensing , , 457, 4089, 10.1093/mnras/stw191

  71. [71]

    2011, title Binary microlensing event OGLE-2009-BLG-020 gives verifiable mass, distance, and orbit predictions, The Astrophysical Journal, 738, 87

    Skowron, J., Udalski, A., Gould, A., et al. 2011, title Binary microlensing event OGLE-2009-BLG-020 gives verifiable mass, distance, and orbit predictions, The Astrophysical Journal, 738, 87

  72. [72]

    2025, title Transformer Embeddings for Fast Microlensing Inference , arXiv e-prints, arXiv:2512.11687, 10.48550/arXiv.2512.11687

    Smyth , N., Perreault-Levasseur , L., & Hezaveh , Y. 2025, title Transformer Embeddings for Fast Microlensing Inference , arXiv e-prints, arXiv:2512.11687, 10.48550/arXiv.2512.11687

  73. [73]

    2015, title Wide-field infrarred survey telescope-astrophysics focused telescope assets WFIRST-AFTA 2015 report, ArXiv e-prints, arXiv

    Spergel, D., Gehrels, N., Baltay, C., et al. 2015, title Wide-field infrarred survey telescope-astrophysics focused telescope assets WFIRST-AFTA 2015 report, ArXiv e-prints, arXiv

  74. [74]

    P., Sumi , T., et al

    Suzuki , D., Bennett , D. P., Sumi , T., et al. 2016, title The Exoplanet Mass-ratio Function from the MOA-II Survey: Discovery of a Break and Likely Peak at a Neptune Mass , , 833, 145, 10.3847/1538-4357/833/2/145

  75. [75]

    K., Anderson , J., Beichman , C

    Terry , S. K., Anderson , J., Beichman , C. A., et al. 2026, title An HST Wide Field Survey of the Galactic Bulge: Overview, Strategy, and First Results , arXiv e-prints, arXiv:2605.06778. 2605.06778

  76. [76]

    K., & Szyma \'n ski , G

    Udalski , A., Szyma \'n ski , M. K., & Szyma \'n ski , G. 2015, title OGLE-IV: Fourth Phase of the Optical Gravitational Lensing Experiment , , 65, 1. 1504.05966

  77. [77]

    C., & Varoquaux, G

    van der Walt, S., Colbert, S. C., & Varoquaux, G. 2011, title The NumPy Array: A Structure for Efficient Numerical Computation, Computing in Science & Engineering, 13, 22, 10.1109/MCSE.2011.37

  78. [78]

    E., et al

    Virtanen , P., Gommers , R., Oliphant , T. E., et al. 2020, title SciPy 1.0: fundamental algorithms for scientific computing in Python , Nature Methods, 17, 261, 10.1038/s41592-019-0686-2

  79. [79]

    2025, FlowJAX: Distributions and Normalizing Flows in Jax, 17.2.1 10.5281/zenodo.10402073

    Ward, D. 2025, FlowJAX: Distributions and Normalizing Flows in Jax, 17.2.1 10.5281/zenodo.10402073

  80. [80]

    2017, title The Initial Mass Function of the Inner Galaxy Measured from OGLE-III Microlensing Timescales , , 843, L5, 10.3847/2041-8213/aa794e

    Wegg , C., Gerhard , O., & Portail , M. 2017, title The Initial Mass Function of the Inner Galaxy Measured from OGLE-III Microlensing Timescales , , 843, L5, 10.3847/2041-8213/aa794e

Showing first 80 references.