Pith. sign in

REVIEW 2 major objections 5 minor 29 references

A 2% local blue residual does not contaminate a Gaia WD–MS binary catalog; bulk failure starts only above 10–20%.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 06:14 UTC pith:KMNMNJYI

load-bearing objection Solid, reproducible audit: at the realistic ~2% local BP residual the contamination null holds; the useful product is the turn-on curve and the unit conversion, not a purity number for the full catalog. the 2 major comments →

arxiv 2607.08856 v1 pith:KMNMNJYI submitted 2026-07-09 astro-ph.IM astro-ph.SR

A Calibration Audit of a Gaia XP White-Dwarf Main-Sequence Binary Catalog: How Much BP-Band Residual it Takes to Manufacture Contamination

classification astro-ph.IM astro-ph.SR
keywords Gaia XP spectraWD-MS binariesBP-band residualcalibration auditinjection testsimulation-based calibrationGALEX UV excessamortized posterior
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

A catalog of roughly 30,000 white-dwarf main-sequence binary candidates was built from Gaia XP spectra by a classifier that returns only a probability. A natural worry is that a leftover blue-flux residual of about 2% in the BP band, which points the same way as a hot white dwarf, could manufacture false candidates. This audit shows that it does not. The key is units: 2% of a star’s total flux dumped into the narrow blue band is a median 55% local excess on a red main-sequence star, about 27 times a 2% local residual. When the residual is injected at its realistic local amplitude into real single-star spectra, the spurious rate barely rises (0.08 on a 0.05 baseline) and an amortized posterior keeps most of its coverage. The selection only fails in bulk once the local excess reaches tens of percent—the raw, uncorrected regime already removed by the catalog’s magnitude cut and the XP correction. The open question is the faint half of the catalog, where that residual has not been measured. Independently, the cleanest reliability check is geometric: sources fit off the white-dwarf cooling sequence are ultraviolet-deficient, a signal that survives a distance control.

Core claim

At the 1–2% local BP-band residual that survives XP correction, the systematic is harmless to the WD–MS selection: the injection spurious rate is 0.08 ± 0.01 against a 0.05 baseline, and an amortized posterior keeps 0.84 of its 90% coverage against a clean 0.88. Bulk manufacturing of candidates, and coverage collapse, begin only above a 10–20% local excess and reach a 0.96 spurious rate near 50%—the uncorrected bias already handled by the correction and the B < 18 cut.

What carries the argument

A local-fractional blue-excess injection into real Gaia XP single-MS spectra, scored by a 95th-percentile renormalized Δχ² threshold that stands in for the catalog classifier, together with the amplitude conversion that 2% of total flux deposited in the blue is a median 55% local excess (a factor of ~27). That unit distinction and the resulting turn-on curve carry the null result.

Load-bearing premise

The paper never runs the released Gaussian-process classifier on the injected spectra; all spurious rates come from a Δχ² threshold stand-in whose fidelity to the actual catalog boundary is assumed.

What would settle it

Measure the post-correction local BP residual specifically for the BP > 17.5 half of the catalog; if that residual reaches the 10–20% local range, the injection turn-on curve predicts the selection begins to manufacture candidates in bulk.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript audits whether the residual ~1–2% local blue-flux excess remaining in corrected Gaia DR3 XP BP spectra can manufacture false WD–MS binary candidates in the Li et al. catalog of ~30,000 sources. Using real single-MS XP spectra, a leave-one-out template library, and a multiplicative local-excess injection, the authors show that a 2% local residual yields only a modest rise in spurious rate (0.08 vs 0.05 baseline) under a 95th-percentile renormalized Δχ² threshold, while bulk contamination appears only above 10–20% local excess (reaching ~0.96 near 50%). An amortized neural posterior exhibits the same amplitude dependence (90% coverage 0.84 at 2% vs clean 0.88). External cross-matches (SDSS/LAMOST spectroscopy, GALEX FUV) identify off-cooling-sequence sources as a clean UV-validated contamination indicator (FUV detection 0.19 vs 0.50) and mark the prior-driven Δχ²<0 majority as externally unverified. The central claim is therefore a null at the realistic residual, with a conditional caution for the unmeasured faint (BP>17.5) half of the catalog.

Significance. If the result holds, the paper supplies a concrete, transferable amplitude threshold that catalog builders need when deciding whether a given XP correction is sufficient for WD–MS selection. The careful unit conversion (2% of total flux deposited in the blue equals a median 55% local excess on red MS stars) and the multi-seed turn-on curve are immediately useful. The spectrally specific failure mode of the amortized posterior (companion-fraction coverage collapse under blue residual while fit quality elsewhere remains clean) is a documented caution for neural-posterior successors. Strengths include full reproducibility from public code and fixed seeds, leave-one-out control of library incompleteness, a discrete-rank SBC null (0.882), and an external UV test with a distance-matched control. The work is a solid calibration audit rather than a new catalog, but the quantitative “how much residual it takes” result is of lasting practical value.

major comments (2)
  1. §3 and §7: the production pipeline scores a 95th-percentile renormalized Δχ² threshold (Eq. 5) on a 50+50 leave-one-out template library rather than Li et al.’s released Gaussian-process classifier (prob_binary>0.8), which is never executed on the injected spectra. The authors correctly flag this as a stand-in, yet the central claim is framed as applying to “the WD–MS selection.” A short quantitative bridge—e.g., correlation of Δχ² with published prob_binary on the real catalog, or a limited re-run of the GP on a subset of injected spectra—would make the proxy claim load-bearing rather than plausible. Without it the null remains well-supported for the spectral feature the GP was trained on, but not strictly for the catalog boundary itself.
  2. §3 (final paragraphs) and Appendix A.2: the model-comparison gate and the Δχ² inversion are evaluated at WD flux shares 0.05–0.23 and at the catalog median (~0.01). The gate loses power and the statistics invert precisely at the catalog-typical share. Because the main injection rates of Figure 1 are not themselves resolved by WD share, it remains unclear how much of the 0.08 spurious rate at 2% local excess is driven by the faint-share regime that dominates the real catalog. A share-binned version of the top panel of Figure 1 (or an explicit statement that the rates are share-averaged) would close this gap.
minor comments (5)
  1. Figure 1 caption and §3: the green line marking Huang’s 2% residual is clear, but the grey band for the “raw uncorrected” ~50% regime could be labeled with the corresponding Riello/Huang references for readers who skip the text.
  2. Appendix A.1, Eqs. (1)–(3): the half-cosine injection template and the severity-to-local conversion are carefully defined; a one-sentence reminder in the main text of §3 that the plotted amplitude is the multiplicative local excess a (not the total-flux severity s) would reduce the chance of unit confusion.
  3. Table 1: the Gentile Fusillo match is correctly interpreted as potentially indicating lone white dwarfs; a parenthetical note that the 220 matches are therefore an upper bound on contamination rather than a purity floor would help casual readers.
  4. §5 and Figure 3: the basis caveat (leading MS component carries only ~5.5% of its loading below 500 nm) is stated, but the figure itself does not annotate which parameters are blue-weighted; a brief legend note would make the spectral-specificity claim self-contained.
  5. Data availability: the Zenodo and GitHub links are given; confirming that the exact configuration files and seeds used for the three-seed rates and four-seed SBC runs are tagged would further strengthen the reproducibility claim already made in the text.

Circularity Check

1 steps flagged

No load-bearing circularity; the contamination null is an independent injection test against a clean baseline plus orthogonal external labels, with only a non-essential methodological self-citation to the author's prior X-ray audit.

specific steps
  1. self citation load bearing [Section 1, paragraph 5]
    "This is the optical counterpart of a test we ran on X-ray spectra [1], where a 3% detector gain shift slips past every per-spectrum trust check and the evidence check alike..."

    The sentence cites the author's own prior work solely as methodological precedent. The citation is not load-bearing: none of the injection amplitudes, Δχ² thresholds, spurious rates, gate AUCs, or SBC coverages in the present paper are taken from or forced by [1]; the analogy can be deleted without altering any numerical claim.

full rationale

The paper's central claim (2% local BP residual yields spurious rate 0.08 on a 0.05 baseline; bulk failure only above 10-20% local excess) is obtained by forward-injecting a half-cosine blue taper into real leave-one-out Gaia XP single-MS spectra, refitting under a 50+50 template library, and scoring a 95th-percentile renormalized Δχ² threshold that is fixed on clean data alone (Appendix A.1–A.2, Eqs. 3–5, Figure 1). That construction is not self-definitional: the threshold is set once on uncontaminated singles, the injection amplitude is an external calibration residual taken from Huang et al., and the measured rate is an empirical outcome, not a fitted parameter renamed as a prediction. Reliability statements rest on cross-matches to SDSS/LAMOST spectroscopy and GALEX FUV (orthogonal to the optical score and to the MS–MS flag). The amortized-posterior SBC is trained exclusively on clean real-template PCA draws; the systematic appears only at inference. The sole self-reference is the parenthetical analogy to the author's earlier X-ray gain-shift audit [1]; it supplies no uniqueness theorem, no ansatz, and no numerical input used in any equation or rate. The acknowledged stand-in character of the Δχ² threshold for Li et al.'s GP classifier is a limitation of scope, not a circular reduction. Consequently the derivation chain is self-contained against its stated inputs and external benchmarks.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 1 invented entities

The central null rests on standard spectral-fitting practice plus a few modeling choices: a constructed blue-taper injection shape, a Δχ² threshold as classifier proxy, a capped real-template library under leave-one-out, SNR=30 noise, and Huang’s published residual amplitude as the realistic scale. No new physical entity is postulated. Free parameters are operational (library size, SNR, amplitude grid, BIC k counts), not fitted to force the null. Domain assumptions about XP sampling, NNLS non-negativity, and GALEX as a hot-WD floor are standard in the subfield.

free parameters (6)
  • local excess amplitude grid (a)
    Hand-chosen sweep (2%, 5%, 10%, 20%, 30%, ~50%) that defines the turn-on curve; not fitted to data but chosen to bracket Huang residual and Riello raw bias.
  • SNR = 30 (default noise scale)
    Spectrum-level Gaussian noise scale used for all injection and SBC runs; fixed by hand, not estimated per object from real XP uncertainties.
  • fit library cap 50 MS + 50 WD
    Random subsample size for NNLS binary search; larger libraries hold 120 per class but fitting is capped for cost.
  • binary decision threshold θ = Q0.95(d_s)
    95th percentile of renormalized Δχ² on clean singles, fixing a 5% FPR by construction; stand-in for Li’s prob_binary>0.8 cut.
  • WD flux share prior box [0.05, 0.30] for SBC
    Training/prior range for amortized posterior; catalog median share is ~0.01, so coverage numbers are on brighter binaries than typical catalog members (share-resolved gate is separate).
  • PCA component counts (k_MS=2, k_WD=3)
    Chosen to capture >99% / ~96% library variance; basis choice affects apparent MS immunity to blue residual.
axioms (6)
  • domain assumption Huang et al. corrected-XP residual is better than ~2% local in 336–400 nm for the validated magnitude range, and this is the realistic post-correction amplitude to test.
    Sets the null amplitude in §3; invoked as the green line in Figure 1 and throughout the conclusion.
  • ad hoc to paper A 95th-percentile renormalized single-vs-binary Δχ² threshold is a sufficient proxy for the behavior of Li et al.’s Gaussian-process classifier under BP injection.
    Stated in §3 and §7; the released classifier is never executed on injected spectra.
  • domain assumption Leave-one-out exclusion of the generating template prevents trivial perfect fits and does not itself manufacture the observed spurious rates.
    §3 and Appendix A.2; zero-amplitude injection recovers the 0.05 baseline.
  • domain assumption GALEX FUV detection is a lower bound on hot white-dwarf presence, independent of the optical classifier score.
    §4 reliability audit and Figure 2; cool WDs are FUV-faint so the indicator is one-sided.
  • domain assumption Noise is independent Gaussian per pixel with constant σ = mean flux / SNR across the 61-pixel grid.
    Appendix A.1 noise model used for all χ² and SBC draws.
  • domain assumption Non-negative least-squares amplitudes on mean-normalized templates adequately represent single and binary XP fits for the audit.
    Appendix A.2; standard spectral unmixing assumption.
invented entities (1)
  • Half-cosine BP injection template R_INJ (and mismatched linear gate basis R_GATE) no independent evidence
    purpose: To model the wavelength-localized blue residual as a controlled additive/multiplicative excess without using an oracle matched shape in the primary gate.
    Constructed in Appendix A.1–A.2 to track Huang’s short-wavelength residual; not a new physical component, but a paper-specific systematic model whose shape choice could affect quantitative rates at large amplitude.

pith-pipeline@v1.1.0-grok45 · 20659 in / 4236 out tokens · 47517 ms · 2026-07-13T06:14:05.406263+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of A Calibration Audit of a Gaia XP White-Dwarf Main-Sequence Binary Catalog: How Much BP-Band Residual it Takes to Manufacture Contamination." pith.science (2026). https://pith.science/paper/KMNMNJYI

@misc{pith2026260708856,
  author       = {Pith},
  title        = {Pith review of: A Calibration Audit of a Gaia XP White-Dwarf Main-Sequence Binary Catalog: How Much BP-Band Residual it Takes to Manufacture Contamination},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KMNMNJYI}},
  note         = {Machine review of arXiv:2607.08856}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

A Gaussian-process classifier on Gaia DR3 XP spectra produced $\sim$30{,}000 white-dwarf main-sequence (WD--MS) binary candidates, each with a probability but no likelihood or goodness-of-fit. The corrected XP BP band keeps a local blue-flux residual of about 2\%, where a hot white dwarf also adds flux. We asked whether that residual contaminates the selection at its realistic amplitude. It does not. The answer turns on one unit: 2\% of total flux, deposited in the narrow blue band, is a median 55\% local excess on a red MS star, 27 times the same number read locally. Injected as a 2\% local excess it gives a spurious rate of 0.08 on a 0.05 baseline through the $\Delta\chi^2$ threshold standing in for the classifier, and an amortized posterior keeps 0.84 of its 90\% coverage against a clean 0.88. The selection fails only above a 10--20\% local excess and reaches 0.96 near 50\%, the raw uncorrected bias that correction and the $B<18$ cut remove. A model-comparison gate certifies binarity only above a WD flux share near 0.05; at the catalog's median share the statistics invert and the spurious carry the larger $\Delta\chi^2$ improvement. The clean reliability signal is off-cooling-sequence UV deficiency (GALEX FUV detection 0.19 against 0.50), robust to a distance control. At the residual the correction leaves, the contamination hypothesis is a null; the open regime is the faint half, where that residual is unmeasured. The audit gives the amplitude it would take, and the shape of the failure past it.

Figures

Figures reproduced from arXiv: 2607.08856 by Karan Akbari.

Figure 1
Figure 1. Figure 1: Where the BP-band excess starts to matter. Top: the spurious WD–MS rate from injecting a blue excess into real single￾MS Gaia XP spectra and refitting under leave-one-out, as a function of the local fractional excess in the blue band (three seeds, production pipeline). At Huang’s validated 2% local residual (green line) the rate is 0.08, on the 0.05 baseline (dashed). It crosses 0.5 near a 20% local excess… view at source ↗
Figure 2
Figure 2. Figure 2: GALEX FUV-detection fraction, a floor on the fraction of each subset that hosts a hot white dwarf, with Wilson 95% intervals and annotated subset sizes. FUV detection rises monotonically across prob_binary terciles, from 0.28 in the lowest third to 0.59 in the highest, so the score does rank true white-dwarf presence, while even the top tercile is only 59% UV-confirmed. The cleanest contamination signal is… view at source ↗
Figure 3
Figure 3. Figure 3: Where a calibration residual lands in the posterior, at a 10% local excess (four seeds, mean 90% coverage per parameter, null 0.882 dashed). A blue BP-shaped residual leaves the leading main-sequence component near the null and drives the companion fraction (log10_frac) to about 0.08, while a red-band excess instead hammers the main-sequence component, and matched extra noise degrades every parameter about… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

29 extracted references · 6 canonical work pages · 1 internal anchor

  1. [1]

    What an Amortized X-ray Posterior Cannot See: Gain Shifts, Silent Miscalibration, and the Limits of the Evidence Check

    Karan Akbari. What an amortized x-ray posterior cannot see: Gain shifts, silent miscalibration, and the limits of the evidence check.arXiv e-prints, 2026. doi: 10.48550/arXiv.2606.17098

  2. [2]

    Tests for model misspecification in simulation- based inference: From local distortions to global model checks.Physical Review D, 111(8):083013, 2025

    Noemi Anau Montel, James Alvey, and Christoph Weniger. Tests for model misspecification in simulation- based inference: From local distortions to global model checks.Physical Review D, 111(8):083013, 2025. doi: 10.1103/PhysRevD.111.083013

  3. [3]

    Calibrating gravitational-wave search al- gorithms with conformal prediction.Physical Review D, 109(12):123027, 2024

    Gregory Ashton, Nicolo Colombo, Ian Harry, and Surabhi Sachdev. Calibrating gravitational-wave search al- gorithms with conformal prediction.Physical Review D, 109(12):123027, 2024. doi: 10.1103/PhysRevD.109. 123027

  4. [4]

    Revised catalog of GALEX ultraviolet sources

    Luciana Bianchi, Bernie Shiao, and David Thilker. Revised catalog of GALEX ultraviolet sources. I. the all-sky survey: GUVcat_AIS.The Astrophysical Journal Supplement Series, 230(2):24, 2017. doi: 10.3847/1538-4365/ aa7053

  5. [5]

    UltraNest – a robust, general purpose Bayesian inference engine.Journal of Open Source Software, 6(60):3001, 2021

    Johannes Buchner. UltraNest – a robust, general purpose Bayesian inference engine.Journal of Open Source Software, 6(60):3001, 2021. doi: 10.21105/joss.03001. 12

  6. [6]

    Patrick Cannon, Daniel Ward, and Sebastian M. Schmon. Investigating the impact of model misspecification in neural simulation-based inference.arXiv e-prints, 2022. doi: 10.48550/arXiv.2209.01845

  7. [7]

    De Angeli, M

    F. De Angeli, M. Weiler, P. Montegriffo, D. W. Evans, M. Riello, R. Andrae, J. M. Carrasco, G. Busso, P. W. Burgess, C. Cacciari, et al. Gaia data release 3. processing and validation of BP/RP low-resolution spectral data. Astronomy & Astrophysics, 674:A2, 2023. doi: 10.1051/0004-6361/202243680

  8. [9]

    Vallenari, A

    Gaia Collaboration, A. Vallenari, A. G. A. Brown, T. Prusti, J. H. J. de Bruijne, F. Arenou, C. Babusiaux, et al. Gaia data release 3. summary of the content and survey properties.Astronomy & Astrophysics, 674:A1, 2023. doi: 10.1051/0004-6361/202243940

  9. [10]

    A random forest spectral classification of the Gaia 500-pc white dwarf population.Astronomy & Astrophysics, 699: A3, 2025

    Enrique Miguel García-Zamora, Santiago Torres, Alberto Rebassa-Mansergas, and Aina Ferrer-Burjachs. A random forest spectral classification of the Gaia 500-pc white dwarf population.Astronomy & Astrophysics, 699: A3, 2025. doi: 10.1051/0004-6361/202554414

  10. [11]

    N. P. Gentile Fusillo, P.-E. Tremblay, E. Cukanovaite, A. V orontseva, R. Lallement, M. Hollands, B. T. Gänsicke, K. B. Burdge, J. McCleery, and S. Jordan. A catalogue of white dwarfs in Gaia EDR3.Monthly Notices of the Royal Astronomical Society, 508(3):3877, 2021. doi: 10.1093/mnras/stab2672

  11. [12]

    A trust crisis in simulation-based inference? Your posterior approximations can be unfaithful.Transactions on Machine Learning Research, 2022

    Joeri Hermans, Arnaud Delaunoy, François Rozet, Antoine Wehenkel, V olodimir Begy, and Gilles Louppe. A trust crisis in simulation-based inference? Your posterior approximations can be unfaithful.Transactions on Machine Learning Research, 2022. doi: 10.48550/arXiv.2110.06581

  12. [13]

    A comprehensive correction of the Gaia DR3 XP spectra.The Astrophysical Journal Supplement Series, 271(1):13, 2024

    Bowen Huang, Haibo Yuan, Maosheng Xiang, Yang Huang, Kai Xiao, Shuai Xu, Ruoyi Zhang, Lin Yang, Zexi Niu, and Hongrui Gu. A comprehensive correction of the Gaia DR3 XP spectra.The Astrophysical Journal Supplement Series, 271(1):13, 2024. doi: 10.3847/1538-4365/ad18b1

  13. [14]

    Ying Jin and Emmanuel J. Candès. Selection by prediction with conformal p-values.Journal of Machine Learning Research, 24(244):1–41, 2023

  14. [15]

    Green, and Xiangyu Zhang

    Jiadong Li, Hans-Walter Rix, Yuan-Sen Ting, Johanna Müller-Horn, Kareem El-Badry, Chao Liu, Rhys See- burger, Gregory M. Green, and Xiangyu Zhang. Millions of main-sequence binary stars from Gaia BP/RP spectra.Astronomy & Astrophysics, 704:A126, 2025. doi: 10.1051/0004-6361/202556362

  15. [16]

    Green, David W

    Jiadong Li, Yuan-Sen Ting, Hans-Walter Rix, Gregory M. Green, David W. Hogg, Juan-Juan Ren, Johanna Müller-Horn, and Rhys Seeburger. Identification of 30,000 white dwarf–main-sequence binary candidates from Gaia DR3 BP/RP (XP) low-resolution spectra.The Astrophysical Journal Supplement Series, 279(2):47, 2025. doi: 10.3847/1538-4365/addf3a

  16. [17]

    Christopher Martin, James Fanson, David Schiminovich, Patrick Morrissey, Peter G

    D. Christopher Martin, James Fanson, David Schiminovich, Patrick Morrissey, Peter G. Friedman, Tom A. Bar- low, Tim Conrow, Robert Grange, Patrick N. Jelinsky, et al. The Galaxy Evolution Explorer: A space ultraviolet survey mission.The Astrophysical Journal, 619(1):L1, 2005. doi: 10.1086/426387

  17. [18]

    Montegriffo, F

    P. Montegriffo, F. De Angeli, R. Andrae, M. Riello, E. Pancino, N. Sanna, M. Bellazzini, D. W. Evans, J. M. Carrasco, R. Sordo, et al. Gaia data release 3. external calibration of BP/RP low-resolution spectroscopic data. Astronomy & Astrophysics, 674:A3, 2023. doi: 10.1051/0004-6361/202243880

  18. [19]

    Prasanta K. Nayak. Revealing unresolved white dwarf-main sequence binaries using Gaia DR3 and GALEX. I. A volume-limited study of 100 pc.Astronomy & Astrophysics, 709:A114, 2026. doi: 10.1051/0004-6361/ 202452939

  19. [20]

    Finding white dwarfs’ hidden companions using an unsupervised machine learning technique.The Astrophysical Journal, 988(1):51, 2025

    Xabier Pérez-Couto, Minia Manteiga, and Eva Villaver. Finding white dwarfs’ hidden companions using an unsupervised machine learning technique.The Astrophysical Journal, 988(1):51, 2025. doi: 10.3847/1538-4357/ addfd7. 13

  20. [21]

    Rebassa-Mansergas, J

    A. Rebassa-Mansergas, J. J. Ren, S. G. Parsons, B. T. Gänsicke, M. R. Schreiber, E. García-Berro, X.-W. Liu, and D. Koester. The SDSS spectroscopic catalogue of white dwarf–main-sequence binaries: New identifications from DR 9–12.Monthly Notices of the Royal Astronomical Society, 458(4):3808, 2016. doi: 10.1093/mnras/stw554

  21. [22]

    Brown, Steven G

    Alberto Rebassa-Mansergas, Enrique Solano, Alex J. Brown, Steven G. Parsons, Raquel Murillo-Ojeda, Roberto Raddi, Maria Camisassa, Santiago Torres, and Jan van Roestel. A magnitude-limited catalogue of unresolved white dwarf-main sequence binaries from Gaia DR3.Astronomy & Astrophysics, 699:A153, 2025. doi: 10.1051/ 0004-6361/202554700

  22. [23]

    J.-J. Ren, A. Rebassa-Mansergas, S. G. Parsons, X.-W. Liu, A.-L. Luo, X. Kong, and H.-T. Zhang. White dwarf– main-sequence binaries from LAMOST: the DR5 catalogue.Monthly Notices of the Royal Astronomical Society, 477(4):4641, 2018. doi: 10.1093/mnras/sty805

  23. [24]

    Riello, F

    M. Riello, F. De Angeli, D. W. Evans, et al. Gaia early data release 3: Photometric content and validation. Astronomy & Astrophysics, 649:A3, 2021. doi: 10.1051/0004-6361/202039587

  24. [25]

    Triage of the Gaia DR3 astrometric orbits

    Sahar Shahaf, Dolev Bashi, Tsevi Mazeh, Simchon Faigler, Frédéric Arenou, Kareem El-Badry, and Hans-Walter Rix. Triage of the Gaia DR3 astrometric orbits. I. A sample of binaries with probable compact companions. Monthly Notices of the Royal Astronomical Society, 518(2):2991, 2023. doi: 10.1093/mnras/stac3290

  25. [26]

    Triage of the Gaia DR3 astrometric orbits

    Sahar Shahaf, Na’ama Hallakoun, Tsevi Mazeh, Sagi Ben-Ami, Prajwal Rekhi, Kareem El-Badry, and Silvia Toonen. Triage of the Gaia DR3 astrometric orbits. II. A census of white dwarfs.Monthly Notices of the Royal Astronomical Society, 529(4):3729, 2024. doi: 10.1093/mnras/stae773

  26. [27]

    Validating Bayesian infer- ence algorithms with simulation-based calibration.arXiv e-prints, 2018

    Sean Talts, Michael Betancourt, Daniel Simpson, Aki Vehtari, and Andrew Gelman. Validating Bayesian infer- ence algorithms with simulation-based calibration.arXiv e-prints, 2018. doi: 10.48550/arXiv.1804.06788

  27. [28]

    Gonçalves, David S

    Alvaro Tejero-Cantero, Jan Boelts, Michael Deistler, Jan-Matthis Lueckmann, Conor Durkan, Pedro J. Gonçalves, David S. Greenberg, and Jakob H. Macke. sbi: A toolkit for simulation-based inference.Journal of Open Source Software, 5(52):2505, 2020. doi: 10.21105/joss.02505

  28. [29]

    Neural posterior estimation for white dwarf spectroscopic characterization.arXiv e-prints, 2025

    Olivier Vincent, Patrick Dufour, and Pierre Bergeron. Neural posterior estimation for white dwarf spectroscopic characterization.arXiv e-prints, 2025. doi: 10.48550/arXiv.2510.16261

  29. [30]

    Daniel Ward, Patrick Cannon, Mark Beaumont, Matteo Fasiolo, and Sebastian M. Schmon. Robust neural pos- terior estimation and statistical model criticism.Advances in Neural Information Processing Systems, 35, 2022. doi: 10.48550/arXiv.2210.06564. 14