Pith. sign in

REVIEW 4 major objections 5 minor 74 references

The paper argues that several stellar streams can jointly pin down the Milky Way's gravitational potential: train once on simulations, then combine per-stream posteriors by adding their score fields, so future datasets need no new inference

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 09:32 UTC pith:HTGDKJFG

load-bearing objection A real methods contribution with honest diagnostics, but the paper's own MMD test undercuts the Gaia parameter values as measurements. the 4 major comments →

arxiv 2607.20725 v1 pith:HTGDKJFG submitted 2026-07-22 astro-ph.GA

Hierarchical Bayesian inference with compositional score modeling for stellar streams

classification astro-ph.GA
keywords stellar streamsMilky Way gravitational potentialdark matter halo shapesimulation-based inferencescore-based diffusion modelshierarchical Bayesian inferencecompositional score modelingGaia
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The authors aim to establish that the Milky Way's gravitational potential can be inferred from a population of stellar streams without paying the inference cost afresh for each new dataset. They set up a hierarchical Bayesian model that separates stream-specific progenitor parameters from the potential parameters shared by all streams, train score-based diffusion networks on a library of a million simulated streams, and combine the per-stream posteriors at inference time by adding score fields. On held-out simulations, the per-stream and composed posteriors are well calibrated and accurate, and the composition tightens the global constraints relative to any single stream. Applied to three Gaia streams (Pal 5, NGC 3201, M 68) plus the circular velocity curve, the combined posterior favours a mildly oblate dark halo (axis ratio q = 0.77), a disk mass of 4.3 x 10^10 solar masses, and a local dark matter density consistent with recent stream-based analyses. The paper also flags that its forward model fails a quantitative misspecification test against the real data, so the Gaia numbers carry that caveat.

Core claim

Central claim: a hierarchical model separating shared potential parameters from per-stream progenitor parameters factorizes the joint posterior into per-stream terms, so one amortized network learns each stream's score field and the multi-stream posterior is assembled at inference time by summing scores. On held-out simulations the per-stream and composed posteriors are well calibrated and accurate, and composition tightens the global constraints. On Gaia data the composed posterior favours a mildly oblate halo (q_NFW = 0.77, a_NFW = 9.3 kpc, disk mass 4.3 x 10^10 M_sun); the virial mass is low because the streams probe only R < 20 kpc. The paper's own misspecification test rejects the forwa

What carries the argument

The load-bearing object is the compositional score: the per-stream posterior score fields learned by one global diffusion network are added at inference time with a prior-correction term and a time-dependent damping factor, so density products become score sums. This turns the hierarchical factorization of the joint posterior into post-training arithmetic, so multi-stream inference costs a single reverse-diffusion integration rather than a new sampling run, and the training budget stays linear in the number of streams. Supporting machinery: a particle-spray stream simulator with a Gaia-like observational layer (magnitude-dependent noise, sky-window selection, missing line-of-sight velocities

Load-bearing premise

The load-bearing premise is that the particle-spray simulator over a simple axisymmetric parametric potential, together with the modelled Gaia selection and error functions, faithfully reproduces the real Milky Way; the paper's own MMD test (Section 3.5) rejects that premise, so the Gaia-data posterior would be biased even though the networks pass simulation-based calibration and recovery tests.

What would settle it

The paper supplies the falsifying direction itself: its posterior-predictive rotation curve puts only 62% of the measured points in the central 68% band and underpredicts the independent outer data beyond 30 kpc, and its MMD test places the real streams outside the generative model's support. A decisive check is to add a tracer anchoring the potential beyond 30 kpc: if the composed posterior still demands the low-mass compact halo implied by q_NFW = 0.77 and a_NFW = 9.3 kpc, and independent data rule that halo out, the Gaia-data claim fails even though the composition machinery may be sound.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Amortization becomes real: updated member lists, improved astrometry from Gaia DR4, or new line-of-sight velocities tighten the composed posterior at the cost of one reverse-diffusion sampling step instead of a fresh MCMC exploration.
  • Composition resolves inter-stream disagreement: Pal 5 alone favours a prolate halo (q ~ 1.33), M 68 an oblate one (q ~ 0.71), and their combination settles on a mildly oblate halo with q ~ 0.77.
  • The Gaia analysis yields a concrete mass model, with disk mass 4.3 x 10^10 M_sun, local dark matter density 0.0115 M_sun pc^-3, and mass within 20 kpc of 1.9 x 10^11 M_sun, in agreement with a recent 29-stream analysis of the same catalogue.
  • The inferred virial mass (0.59 x 10^12 M_sun, roughly half of the comparison value) is interpreted as an extrapolation artefact, since the three streams constrain the potential only inside roughly 20 kpc.
  • The paper reports its own caveats on the real-data application: an MMD-based test rejects the forward model, the predictive rotation curve underpredicts data beyond 30 kpc, M 68's declination is systematically underestimated, the composed posterior shows a small calibration deviation for the inner halo slope, and hyperparameters were chosen to favour a conservative posterior after unstable alterna

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the per-stream posteriors are addable, one could compose arbitrary subsets or reweight streams post hoc, so a consistency diagnostic would naturally flag streams whose individual posteriors barely overlap the composed one, as Pal 5 and M 68 already do on halo flattening.
  • The misspecification signature (rotation curve low beyond 30 kpc, M 68 track offset) points to the simplified axisymmetric potential as the binding constraint; a direct test is to re-run the same pipeline with a radially varying halo or multi-component disk and check whether the MMD null is no longer rejected and the outer curve enters the 68% band.
  • The real gate at population scale is simulation and training cost: with hundreds of known streams, the method's payoff is largest for repeated re-analysis of a fixed stream library as astrometry and error models improve, not for one-off fits, since newly discovered streams require regenerating the simulation library.
  • An independent stream or tracer reaching beyond 30 kpc would be the sharpest external check: if its posterior-predictive track is incompatible with the composed potential, or if including it does not pull the virial mass toward literature values, the low outer mass is a forward-model artefact rather than a measurement.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a hierarchical, amortized simulation-based inference framework for Milky Way potential reconstruction from stellar streams. Global potential parameters are shared across three streams (Pal 5, NGC 3201, M 68); per-stream progenitor parameters are modeled separately. Score-based diffusion networks are trained on 10^6 Agama particle-spray simulations with realistic Gaia-like noise and selection, and streams are combined at inference time via compositional score modeling (Eq. 14). The authors validate the method on held-out simulations using simulation-based calibration and recovery tests, then apply it to Gaia DR3 data plus the Zhou et al. (2023) circular-velocity curve, reporting q_NFW=0.77, a_NFW=9.3 kpc, disk mass 4.3e10 Msun, and rho_NFW,odot=0.01153 Msun/pc^3. The paper also presents posterior predictive checks and a maximum mean discrepancy (MMD) model-misspecification test.

Significance. If the methodological claim holds, this is a valuable contribution to amortized SBI for Galactic dynamics: per-stream posteriors are learned once and then combined at inference time, avoiding new MCMC runs when streams are added. The public code, held-out SBC/recovery diagnostics, and explicit comparison to classical likelihood-based analyses (Ibata et al. 2024; Palau & Miralda-Escudé 2023) are strengths. However, the applied Gaia-data claim is not supported by the paper's own diagnostics. Section 3.5 rejects the null hypothesis that the generative model matches the real data, stating that 'the real observations lie outside the support of our generative model'; Section 3.4 shows an offset in the M68 posterior-predictive track and a low-biased outer rotation curve (62% of points in the central 68% band). SBC and recovery under the same simulator cannot certify a posterior for data the simulator cannot reproduce. Thus the methodological contribution is credible, but the reported Milky Way parameter measurements should not be presented as valid inference unless the model misspecification is resolved.

major comments (4)
  1. [§3.5 & §3.3] The MMD test in §3.5 explicitly rejects the null: 'the real observations lie outside the support of our generative model.' In spite of this, §3.3 and the abstract report q_NFW=0.77, a_NFW=9.3 kpc, etc. as 'the model favours...'. SBC (§3.1) and recovery (§3.2) are computed under the same simulator and therefore cannot validate a posterior for data that the simulator cannot reproduce. The real-data parameter estimates in Table 4 are conditional on a model that is rejected by the paper's own diagnostic; this is a load-bearing unsupported claim. The authors should either modify the forward model until the MMD test is not rejected and the PPCs pass, or drop/reframe the Gaia parameter measurements as a conditional demonstration rather than a measurement.
  2. [§4 (hyperparameter selection)] The discussion states that during Optuna tuning 'the posterior prediction on Gaia data with similar performance would return incompatible posteriors... We therefore selected the hyperparameters that returned a more conservative posterior.' This is model selection using the real data that are later used to report the Science results. It invalidates the frequentist calibration of the final posterior and can introduce bias. Selection must be based solely on held-out simulation validation, or the sensitivity across the Pareto front should be reported so the reader can assess the spread.
  3. [§3.4 & Fig. 12] The posterior predictive rotation-curve check is quantitative: only 62% of the Zhou et al. data fall in the central 68% band and the prediction is systematically low beyond 30 kpc against the independent Huang et al. curve. This is direct evidence that the inferred potential underestimates the outer mass, consistent with the low M200 in Table 4. This is not a minor visual detail: it undercuts the virial-mass and circular-velocity claims and should be treated as a failing predictive check, not 'mild agreement'.
  4. [§2.7.3, Eq. (14)] The compositional score is exact only at t=0 and t=1; the damping constant d1=1/6 is chosen from validation behavior. The paper acknowledges this approximation, and the SBC/recovery tests do provide empirical support. However, the calibration deviation for γ_NFW at r≈0.9 (Fig. 7) and the approximate nature of Eq. (14) should be stated in the abstract/conclusions; the phrase 'well calibrated' should be qualified. This is not a blocker for the method, but it is a load-bearing caveat for the methodological claim.
minor comments (5)
  1. [§2.4.1] The phrase 'interpolate over the value of Lindegren et al. (2021, Table 4)' is awkward; it should be 'interpolate the values in Lindegren et al. (2021, Table 4)'.
  2. [References] Viterbo & Buck 2024a and 2024b list the same arXiv identifier (2411.17269); one is likely a typo.
  3. [Fig. 12] The x-axis has a gap between 30 kpc and 40 kpc; the tick spacing is inconsistent and should be made uniform.
  4. [Table 4] The local-parameter rows for the Compositional column are empty; it would be clearer to state '—' with a footnote explaining that local parameters are per-stream and not combined.
  5. [§2.7.3] The damping factor d(t)=d0·exp(...) is introduced but d0 is not explicitly defined; state d0=1.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

The inference is only as good as the simulator that produced the training set. The paper's own MMD test (§3.5) shows the simulator does not populate the summary space of the real data, and the final hyperparameters were selected partly on the basis of the Gaia posterior behavior (§4). The method contributes machinery, but the physical measurement borrows heavily from assumptions about potential shape, disruption physics, and Gaia selection effects.

free parameters (5)
  • Compositional score terminal damping d1 = 1/6
    Set empirically in §2.7.3 to balance accuracy vs calibration of the composed posterior; controls width/shape of the final inference and is not derived from first principles.
  • Final hyperparameter run selected from Optuna Pareto front = not reported
    §2.8 tunes on a validation set; §4 admits that among runs with similar validation performance, the one giving a conservative Gaia posterior was chosen. This is post-hoc selection informed by the target data.
  • Prior hyperparameters for the 7 global parameters = see Table 1
    Hand-chosen ranges, e.g. γ_NFW ~ U(-2,2), q_NFW ~ U(0.5,1.5), Σ_D ~ U(1e7,1.5e9). Some runs collapse at prior edges, and z_D is effectively unconstrained.
  • Fixed progenitor masses, scale radii, and backward-integration times = taken from Palau & Miralda-Escudé (2023)
    They enter the simulator as fixed auxiliary inputs (§2.3); if inaccurate, every generated training stream is biased.
  • Bulge parameters = fixed to Cautun et al. (2020) best fit
    All bulge parameters are fixed because the streams do not constrain the inner few kpc; this choice still shapes the inferred disk/halo balance.
axioms (6)
  • domain assumption The axisymmetric single-disk generalized-NFW plus fixed-bulge potential (Sec. 2.1) can represent the real Milky Way mass distribution.
    All posterior statements are conditional on this parametric family; §4 identifies the potential parametrization as the most likely cause of the MMD rejection.
  • domain assumption The particle-spray prescription of Chen et al. (2025) produces stream debris equivalent to self-consistent tidal disruption for inference purposes.
    Used as the forward simulator in §2.3; §4 lists 'the forward model itself' as a possible source of misspecification.
  • domain assumption The attention-masked SetTransformer training with N_pool=300, magnitude-dependent noise, and mean-filling of missing V_R (Sec. 2.4) faithfully models the Gaia catalogues and their selection effects.
    The empirical noise/selection model is built from Ibata et al. (2024) and Lindegren et al. (2021); mismatches here feed the MMD rejection.
  • ad hoc to paper Equation (14), the annealed sum of per-stream scores with d1=1/6, provides a valid approximation of the joint compositional posterior, exact only at t=0 and t=1.
    The authors state the intermediate-time score fields are approximate; the damping factor is chosen empirically and a calibration deviation remains for γ_NFW.
  • standard math Standard score-matching/reverse-SDE theory (Vincent 2011; Song et al. 2021; Karras et al. 2022) is taken as background.
    Used to justify the DSM loss and posterior sampling; not re-derived in this paper.
  • standard math Streams are conditionally independent given (η_MW, θ_j), making Eq. (9) factorize.
    Graphical model of §2.6; required for Eq. (13) composition; plausible but unverifiable on real data.

pith-pipeline@v1.3.0-alltime-deepseek · 25689 in / 14653 out tokens · 133290 ms · 2026-08-01T09:32:10.189762+00:00 · methodology

0 comments
read the original abstract

Context: Stellar streams trace the gravitational potential of the Milky Way over a wide range of Galactocentric radii. Since different streams sample different regions of the Galaxy, combining several of them can constrain the global mass distribution more tightly than modeling any single stream in isolation. Most of the existing multi-stream analyses rely on likelihood-based methods that require a new inference run whenever additional streams or kinematic measurements become available. Aims: We aim to infer the Milky Way potential from multiple stellar streams combined with an additional constraint through the Galactic circular velocity curve within a single hierarchical framework. Methods: We model the problem hierarchically, separating parameters that are common to all streams from parameters that are specific to each progenitor. We train score-based neural posterior estimators on a library of simulated streams and combine information from different streams through compositional score modeling as a post-training step. Results: Tests on independent simulations show that the inferred posteriors are well calibrated and accurate. Combining several streams reduces the uncertainties on the global potential parameters relative to single-stream analyses. Applied to Gaia data, the model favours a mildly oblate dark matter halo with axis ratio $q_{NFW} = 0.77$, scale radius $a_{NFW} = 9.3$ kpc, a disk mass of $4.3 \times 10^{10} M_\odot$, and a local dark matter density $\rho_{NFW,\odot} = 0.01153 M_\odot pc^{-3}$ consistent with recent stream-based studies. Conclusions: This hierarchical framework provides a practical way to combine information from multiple stellar streams without repeating the full inference procedure for each new dataset. The method is adaptable to new datasets, like future Gaia DR4, or new spectroscopic surveys, with minimal computational cost.

Figures

Figures reproduced from arXiv: 2607.20725 by Giuseppe Viterbo, Jonas Arruda, Tobias Buck.

Figure 1
Figure 1. Figure 1: Magnitude distribution in the Gaia G-band of the ob [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: Prior predictive checks for the three streams M 68 (left panel) NGC 3201(centre panel) and Pal 5(right panel): synthetic [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Overview of the hierarchical inference pipeline. We draw [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Training pipeline. Upper panel: simulated stream stars in [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Illustration of compositional score modeling, where we [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Simulation-based calibration of the global potential pos [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Recovery of the global potential parameters on held-out test simulations. We plot the inferred posterior median (point) and [PITH_FULL_IMAGE:figures/full_fig_p010_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Recovery of the four local progenitor parameters [PITH_FULL_IMAGE:figures/full_fig_p010_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Posterior samples over the seven global parameters [PITH_FULL_IMAGE:figures/full_fig_p011_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Posterior predictive check of the three streams. Observed member stars (scatter) are overlaid with mock streams generated [PITH_FULL_IMAGE:figures/full_fig_p012_11.png] view at source ↗
Figure 14
Figure 14. Figure 14: Model-misspecification check in the GFN summary [PITH_FULL_IMAGE:figures/full_fig_p013_14.png] view at source ↗
Figure 13
Figure 13. Figure 13: Posterior on the virial radius R200 and virial mass M200 derived from the compositional global posterior for the MW. The vertical dashed line indicates the 16th-84th percentiles, while the solid line indicates the median. and [PITH_FULL_IMAGE:figures/full_fig_p013_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

74 extracted references · 16 linked inside Pith

  1. [1]

    2019, in The 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2623–2631

    Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. 2019, in The 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2623–2631

  2. [2]

    Arruda, J., Bracher, N., Köthe, U., Hasenauer, J., & Radev, S. T. 2025, arXiv e-prints, arXiv:2512.20685

  3. [3]

    2026a, arXiv preprint arXiv:2604.18319

    Arruda, J., Chervet, S., Staudt, P., et al. 2026a, arXiv preprint arXiv:2604.18319

  4. [4]

    2026b, in The Fourteenth International Conference on Learning Representations Astropy Collaboration, Price-Whelan, A

    Arruda, J., Pandey, V ., Sherry, C., et al. 2026b, in The Fourteenth International Conference on Learning Representations Astropy Collaboration, Price-Whelan, A. M., Lim, P. L., et al. 2022, ApJ, 935, 167

  5. [5]

    2019, MNRAS, 482, 5138

    Baumgardt, H., Hilker, M., Sollima, A., & Bellini, A. 2019, MNRAS, 482, 5138

  6. [6]

    Bonaca, A., Geha, M., Küpper, A. H. W., et al. 2014, ApJ, 795, 94

  7. [7]

    & Hogg, D

    Bonaca, A. & Hogg, D. W. 2018, ApJ, 867, 101

  8. [8]

    K., & Kallivayalil, N

    Bovy, J., Bahmanyar, A., Fritz, T. K., & Kallivayalil, N. 2016, ApJ, 833, 31

  9. [9]

    Bowden, A., Belokurov, V ., & Evans, N. W. 2015, MNRAS, 449, 1391

  10. [10]

    M., Mandel, K

    Boyd, B. M., Mandel, K. S., Grayling, M., et al. 2026, arXiv e-prints, arXiv:2603.11165

  11. [11]

    H., & Buder, S

    Buck, T., Günes, B., Viterbo, G., Oliver, W. H., & Buder, S. 2025, A&A, 702, A184

  12. [12]

    M., Cautun, M., Deason, A

    Callingham, T. M., Cautun, M., Deason, A. J., et al. 2019, MNRAS, 484, 5453

  13. [13]

    J., et al

    Cautun, M., Benítez-Llambay, A., Deason, A. J., et al. 2020, MNRAS, 494, 4291

  14. [14]

    Y ., & Ash, N

    Chen, Y ., Valluri, M., Gnedin, O. Y ., & Ash, N. 2025, ApJS, 276, 32

  15. [15]

    2020, Proceedings of the National Academy of Science, 117, 30055

    Cranmer, K., Brehmer, J., & Louppe, G. 2020, Proceedings of the National Academy of Science, 117, 30055

  16. [16]

    R., Gair, J., et al

    Dax, M., Green, S. R., Gair, J., et al. 2025, Nature, 639, 49

  17. [17]

    2025, arXiv preprint arXiv:2508.12939

    Deistler, M., Boelts, J., Steinbach, P., et al. 2025, arXiv preprint arXiv:2508.12939

  18. [18]

    M., Sanders, J

    Dillamore, A. M., Sanders, J. L., Belokurov, V ., & Zhang, H. 2025, MNRAS, 541, 214

  19. [19]

    W., Rix, H.-W., & Ness, M

    Eilers, A.-C., Hogg, D. W., Rix, H.-W., & Ness, M. K. 2019, ApJ, 871, 120

  20. [20]

    Erkal, D., Belokurov, V ., Laporte, C. F. P., et al. 2019, MNRAS, 487, 2685

  21. [21]

    D., Wildberger, J., Dax, M., et al

    Gebhard, T. D., Wildberger, J., Dax, M., et al. 2025, Astronomy & Astrophysics, 693, A42

  22. [22]

    2022, arXiv e-prints, arXiv:2209.14249

    Geffner, T., Papamakarios, G., & Mnih, A. 2022, arXiv e-prints, arXiv:2209.14249

  23. [23]

    B., Stern, H

    Gelman, A., Carlin, J. B., Stern, H. S., et al. 2013, Bayesian Data Analysis (3rd Edition) (Chapman and Hall/CRC)

  24. [24]

    Gibbons, S. L. J., Belokurov, V ., & Evans, N. W. 2014, MNRAS, 445, 3788

  25. [25]

    Gloeckler, M., Toyota, S., Fukumizu, K., & Macke, J. H. 2025, in The Thirteenth International Conference on Learning Representations

  26. [26]

    M., Ting, Y .-S., & Kamdar, H

    Green, G. M., Ting, Y .-S., & Kamdar, H. 2023, ApJ, 942, 26

  27. [27]

    2025, arXiv e-prints, arXiv:2507.05060

    Gunes, B., Buder, S., & Buck, T. 2025, arXiv e-prints, arXiv:2507.05060

  28. [28]

    Harris, W. E. 1996, AJ, 112, 1487

  29. [29]

    Harris, W. E. 2010, arXiv e-prints, arXiv:1012.3224

  30. [30]

    2020, Denoising Diffusion Probabilistic Models, arXiv:2006.11239 [cs]

    Ho, J., Jain, A., & Abbeel, P. 2020, Denoising Diffusion Probabilistic Models, arXiv:2006.11239 [cs]

  31. [31]

    2016, MNRAS, 463, 2623

    Huang, Y ., Liu, X.-W., Yuan, H.-B., et al. 2016, MNRAS, 463, 2623

  32. [32]

    2024, ApJ, 967, 89

    Ibata, R., Malhan, K., Tenachi, W., et al. 2024, ApJ, 967, 89

  33. [33]

    Jing, Y . P. & Suto, Y . 2002, ApJ, 574, 538

  34. [34]

    2021, arXiv preprint arXiv:2105.14080

    Jolicoeur-Martineau, A., Li, K., Piché-Taillefer, R., Kachman, T., & Mitliagkas, I. 2021, arXiv preprint arXiv:2105.14080

  35. [35]

    2022, arXiv e-prints, arXiv:2206.00364 Knöll, N., Buck, T., Branca, L., & Viterbo, G

    Karras, T., Aittala, M., Aila, T., & Laine, S. 2022, arXiv e-prints, arXiv:2206.00364 Knöll, N., Buck, T., Branca, L., & Viterbo, G. 2026, arXiv e-prints, arXiv:2606.10762

  36. [36]

    E., Rix, H.-W., & Hogg, D

    Koposov, S. E., Rix, H.-W., & Hogg, D. W. 2010, ApJ, 712, 260

  37. [37]

    & Gilmore, G

    Kuijken, K. & Gilmore, G. 1991, ApJ, 367, L9 Küpper, A. H. W., Balbinot, E., Bonaca, A., et al. 2015, ApJ, 803, 80 Kühmichel, L., Huang, J. M., Pratz, V ., et al. 2026, arXiv preprint arXiv:2602.07098

  38. [38]

    Law, D. R. & Majewski, S. R. 2010, ApJ, 714, 229

  39. [39]

    2019, Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks, arXiv:1810.00825 [cs]

    Lee, J., Lee, Y ., Kim, J., et al. 2019, Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks, arXiv:1810.00825 [cs]

  40. [40]

    2024, arXiv preprint arXiv:2409.04332

    Li, C., Vehtari, A., Bürkner, P.-C., et al. 2024, arXiv preprint arXiv:2409.04332

  41. [41]

    A., Hernández, J., et al

    Lindegren, L., Klioner, S. A., Hernández, J., et al. 2021, A&A, 649, A2

  42. [42]

    L., & Rodrigues, P

    Linhart, J., Cardoso, G., Gramfort, A., Corff, S. L., & Rodrigues, P. L. C. 2026, Transactions on Machine Learning Research

  43. [43]

    & Hutter, F

    Loshchilov, I. & Hutter, F. 2017, arXiv e-prints, arXiv:1711.05101

  44. [44]

    C., Starkenburg, T., Ho, M., et al

    Lovell, C. C., Starkenburg, T., Ho, M., et al. 2025, MNRAS, 544, 3949

  45. [45]

    & Ibata, R

    Malhan, K. & Ibata, R. A. 2019, MNRAS, 486, 2995

  46. [46]

    2023, MNRAS, 520, 5225

    Mateu, C. 2023, MNRAS, 520, 5225

  47. [47]

    McClure-Griffiths, N. M. & Dickey, J. M. 2016, The Astrophysical Journal, 831, 124

  48. [48]

    McMillan, P. J. 2017, MNRAS, 465, 76

  49. [49]

    2022, ApJ, 940, 22

    Nibauer, J., Belokurov, V ., Cranmer, M., Goodman, J., & Ho, S. 2022, ApJ, 940, 22

  50. [50]

    & Bonaca, A

    Nibauer, J. & Bonaca, A. 2025, ApJ, 985, L22

  51. [51]

    H., Elahi, P

    Oliver, W. H., Elahi, P. J., Lewis, G. F., & Buck, T. 2024, MNRAS, 530, 2637

  52. [52]

    Palau, C. G. & Miralda-Escudé, J. 2023, MNRAS, 524, 2124

  53. [53]

    Plummer, H. C. 1911, MNRAS, 71, 460

  54. [54]

    M., Hogg, D

    Price-Whelan, A. M., Hogg, D. W., Johnston, K. V ., & Hendel, D. 2014, ApJ, 794, 4

  55. [55]

    M., Mateu, C., Iorio, G., et al

    Price-Whelan, A. M., Mateu, C., Iorio, G., et al. 2019, AJ, 158, 223

  56. [56]

    M., Sanderson, R

    Reino, S., Rossi, E. M., Sanderson, R. E., et al. 2021, MNRAS, 502, 4170

  57. [57]

    2026, ApJ, 1000, L16 Säilynoja, T., Bürkner, P.-C., & Vehtari, A

    Ruzza, A., Lodato, G., Rosotti, G., et al. 2026, ApJ, 1000, L16 Säilynoja, T., Bürkner, P.-C., & Vehtari, A. 2022, Statistics and Computing, 32, 32

  58. [58]

    Salimans, T. & Ho, J. 2022, arXiv e-prints, arXiv:2202.00512

  59. [59]

    S., Kawata, D., Makinen, T

    Sante, A., Font, A. S., Kawata, D., Makinen, T. L., & Grand, R. J. J. 2026, MN- RAS

  60. [60]

    A., Piras, D., Jeffrey, N., et al

    Saoulis, A. A., Piras, D., Jeffrey, N., et al. 2025, MNRAS[arXiv:2505.21215]

  61. [61]

    Schmitt, M., Bürkner, P.-C., Köthe, U., & Radev, S. T. 2021, arXiv e-prints, arXiv:2112.08866

  62. [62]

    2024, in Proceedings of the 41st International Conference on Machine Learning, ICML’24 (JMLR.org)

    Sharrock, L., Simons, J., Liu, S., & Beaumont, M. 2024, in Proceedings of the 41st International Conference on Machine Learning, ICML’24 (JMLR.org)

  63. [63]

    & Baumgardt, H

    Sollima, A. & Baumgardt, H. 2017, MNRAS, 471, 3668

  64. [64]

    P., et al

    Song, Y ., Sohl-Dickstein, J., Kingma, D. P., et al. 2021, in International Confer- ence on Learning Representations (ICLR)

  65. [65]

    2018, Val- idating Bayesian Inference Algorithms with Simulation-Based Calibration, arXiv:1804.06788 [stat]

    Talts, S., Betancourt, M., Simpson, D., Vehtari, A., & Gelman, A. 2018, Val- idating Bayesian Inference Algorithms with Simulation-Based Calibration, arXiv:1804.06788 [stat]

  66. [66]

    V ., Dutton, A

    Tollet, E., Macciò, A. V ., Dutton, A. A., et al. 2016, MNRAS, 456, 3542

  67. [67]

    2021, MNRAS, 501, 2279

    Vasiliev, E., Belokurov, V ., & Erkal, D. 2021, MNRAS, 501, 2279

  68. [68]

    2011, Neural Computation, 23, 1661

    Vincent, P. 2011, Neural Computation, 23, 1661

  69. [70]

    & Buck, T

    Viterbo, G. & Buck, T. 2024b, arXiv e-prints, arXiv:2411.17269

  70. [71]

    & Buck, T

    Viterbo, G. & Buck, T. 2025, arXiv e-prints, arXiv:2511.22468

  71. [72]

    & Buck, T

    Viterbo, G. & Buck, T. 2026, A&A, 707, A363

  72. [73]

    2019, MNRAS, 485, 3296

    Wegg, C., Gerhard, O., & Bieth, M. 2019, MNRAS, 485, 3296

  73. [74]

    2023, ApJ, 946, 73

    Zhou, Y ., Li, X., Huang, Y ., & Zhang, H. 2023, ApJ, 946, 73

  74. [75]

    2026, A&A, 706, A193 Article number, page 15 A&A proofs:manuscript no

    Zhu, L., Cai, R., Kang, X., et al. 2026, A&A, 706, A193 Article number, page 15 A&A proofs:manuscript no. main Appendix A: Prior and posterior predictive checks We show in this appendix, for each stream, the prior and poste- rior predictive checks in the full Gaia observational space, com- plementing the summary shown in Fig. 11. For each stream we show t...