{"id":"932f3fea-498b-4d4a-a67a-d49d70a0cf7f","arxiv_id":"2505.05590","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"OTI recovers the known vertical acceleration of a FIRE-2 simulated Milky Way-like disk within 3 sigma in 15 of 16 solar-analog volumes, though its error bars likely understate model systematics.","lead":"Researchers applied the Orbital Torus Imaging (OTI) technique to a realistic supercomputer model of the Milky Way to test whether it can recover the galaxy's vertical gravitational pull. OTI matched the simulation's known acceleration in most of 16 tested solar-neighborhood regions, providing a benchmark for applying the method to real survey data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recovery fractions use precision-only uncertainties that the authors admit exclude model error, so the 94%/75% claim as stated is unsupported.","rationale":"The paper provides an external benchmark for OTI against FIRE-2's stored accelerations, and the method's qualitative success (15/16 volumes) is encouraging. However, the quantitative headline—the 94%/75% recovery fractions—is computed against error bars that the authors themselves describe as precision-only (Section 5.1). The OTI model is knowingly misspecified for this system: it assumes axisymmetry, steady state, and R-z separability, and drops the radial term of the CBE; the authors note the separability assumption may fail beyond |z| >= 1.5 kpc while data are fitted out to zmax ~ 3.5 kpc (Section 4.1). Model misspecification is therefore not negligible, and it is not represented in the MCMC+bootstrap uncertainties. The statement that 12/16 within 1σ matches the naive statistical expectation is circular because the same incomplete σ defines both the observed counts and the expectation. A fair benchmark should include a systematic error term, e.g., estimated from the observed scatter of residuals across volumes. I agree with the reader's CONDITIONAL verdict; the central claim is plausible but the stated precision-based success metric is not a valid accuracy metric. The R-z separability issue, which the reader identified as the weakest assumption, is related but secondary: it is a likely source of the model error that the uncertainty budget fails to capture.","tokens_in":27105,"tokens_out":8907,"duration_ms":100077,"concrete_test":"Add a systematic error term to the total z-profile uncertainty. For each volume, compute the residuals (FIRE - OTI) over the plotted z range, and estimate a systematic variance from the scatter across volumes (e.g., the 1.5 x MAD of normalized residuals shown in Figure 7, or the MAD of the median residual profile). Recompute the 1σ/3σ recovery fractions with total error = sqrt(sigma_MCMC+boot^2 + sigma_sys^2). If V14 and V13 remain the only volumes failing at 3σ, the central claim is robust with a caveat; if additional volumes (e.g., V2, V3, V16) also fail, the 94%/75% statement is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline recovery fractions (15/16 within 3σ, 12/16 within 1σ) are computed using uncertainties that combine MCMC sampling and bootstrapping only. The authors state in Section 5.1 that this error \"reflects the precision of our measurement but not model inaccuracies\" and that the symmetry assumptions are violated in m12i. Because the OTI model is knowingly misspecified (neglect of the radial CBE term, R-z separability, axisymmetry; see the footnote in Section 3 and Section 4.1), the quoted σ excludes the dominant error source for exactly the volumes (V13, V14) with large scale heights and older populations. The recovery statement is therefore not a measure of accuracy; it is a measure of whether precision-only error bars happen to cover the true profile. The conclusion in Section 7 that \"12/16 within 1σ is almost what we observe\" (expecting 11/16) is circular, since the same incomplete σ defines both the observed counts and the statistical expectation. The paper's Figure 7 shows a systematic residual trend (OTI overpredicts above the midplane, underpredicts below, with median residual ~15%), confirming that bias is present. A fair benchmark must either quote total uncertainty (statistical + systematic) or compare using an error metric that includes model misspecification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper benchmarks the Orbital Torus Imaging (OTI) method on the FIRE-2 cosmological zoom-in simulation m12i. The authors select 16 solar-analog volumes at R = 8 kpc, fit OTI models to the binned mean [Fe/H] distribution in (z, vz) vertical phase space, and compare the inferred vertical acceleration profiles and total surface mass densities at z = 1.1 kpc against the accelerations and densities stored in the simulation. They report that the OTI-inferred az profiles match the known FIRE-2 profiles within 3σ for 15/16 volumes and within 1σ for 12/16 volumes, and that Σ☉(z = 1.1 kpc) agrees within 3σ for 15/16 volumes. The paper also compares OTI precision with Jeans modeling, identifies volume properties (scale height, median stellar age, surface density) that correlate with fit quality, and discusses implications of OTI for distinguishing dark matter models.","tokens_in":27382,"tokens_out":3694,"duration_ms":48274,"significance":"If the headline recovery fractions were supported by a complete error budget, this would be a valuable validation of OTI in a realistic, live, out-of-equilibrium simulated disk. The comparison is genuinely external: the FIRE-2 stored accelerations played no role in the OTI fit, so the benchmark is not circular. The use of 16 spatially separated volumes within one cosmological simulation is a meaningful step beyond the earlier single-volume N-body test, and the authors are commendably transparent about their assumptions and residual trends. However, the central quantitative claim currently rests on an error budget that the authors themselves state excludes model inaccuracies, and the paper's own Figure 7 shows a systematic residual pattern. The significance of the result is therefore conditional on re-framing or supplementing the uncertainty assessment.","major_comments":[{"comment":"The central claim that OTI recovers the true vertical acceleration 'within 3σ/1σ for 94%/75% of volumes' is not supported by the quoted uncertainties. The errors shown in Figure 6 combine MCMC sampling and bootstrap resampling only. The text in §5.1 explicitly states that this combined uncertainty 'reflects the precision of our measurement but not model inaccuracies' and that OTI's symmetry assumptions are violated in m12i. Since the true FIRE-2 profiles are known, the appropriate accuracy statement would use a total error budget that includes model misspecification, or report residuals directly. Figure 7 shows a systematic trend: OTI overpredicts above the midplane and underpredicts below it, with median percent differences of 15% and 14%. With a precision-only σ, the 15/16 and 12/16 counts measure whether the precision error bars happen to cover the true profile, not whether the method is accurately recovering it. The abstract and conclusions should be revised to either quote residual-based accuracy metrics or to explicitly qualify the σ as precision-only, ideally with an added systematic component.","section":"§5.1, Figure 6, and abstract"},{"comment":"The argument that '12/16 within 1σ is almost what we observe' (expected 11/16 under pure statistical errors) is circular. The expected count of 11/16 is derived from the same incomplete σ that defines the observed count, so the near-agreement between expectation and observation carries no information about whether systematic errors dominate. A meaningful test would compare the observed residuals to a σ that includes model error, or would compare the magnitude of the systematic residual to the statistical precision independently. As written, this bullet over-interprets the agreement between an observation and an expectation that were constructed from the same input.","section":"§7, first bullet"},{"comment":"The paper's own model assumptions are acknowledged to break down in precisely the regimes where the failures occur: R–z separability and neglect of the radial term of the collisionless Boltzmann equation 'may not be valid for regions far (|z| ≥ 1.5 kpc) from the midplane' (footnote in §3), and §4.1 restricts the analysis accordingly. Volumes V13 and V14, which have the largest scale heights and oldest, kinematically heated populations, are the ones with the worst recovery, consistent with this misspecification. Because the model error from separability is not quantified, the reader cannot tell how much of the recovery success in the other volumes is due to the method working versus the volumes being close enough to the thin-disk regime that the misspecification is small. The authors should add a sensitivity test or comparable systematic-error estimate that varies the model assumptions (e.g., including the radial term or fitting a non-separable model) and report how the recovery fractions change.","section":"§3, §4.1, and Figure 6, volumes V13/V14"}],"minor_comments":[{"comment":"There is a typo: 'We employ the the 1D vertical collisionless Boltzmann equation' should read 'We employ the 1D vertical collisionless Boltzmann equation'.","section":"§3, first paragraph"},{"comment":"The statement that OTI provides a more precise az estimate 'by ~85% across the disk' is not defined precisely: it is unclear whether this is the median ratio of Jeans to OTI error bars, averaged over z, or computed in some other way. A brief definition of the metric would improve reproducibility.","section":"§5.1, paragraph after Figure 6"},{"comment":"The 26 free parameters are listed, but the bounds and priors used in the MCMC sampling are not given. For a paper whose headline is about uncertainty quantification, specifying the priors and the chain convergence checks (e.g., R-hat, effective sample size) would be important for reproducibility.","section":"§4.2 and Appendix D"},{"comment":"The Pearson correlation coefficients are reported to two decimals, but no significance or confidence intervals are given. Given only 16 volumes, reporting p-values or bootstrap confidence intervals on r would help the reader judge whether the claimed weak/moderate correlations are meaningful.","section":"§6, Figure 9"},{"comment":"The phrase 'recently quiescent' is used for m12i, but the quantitative definition is deferred to a citation. Stating the last major merger time and the current snapshot time explicitly would make the 'dynamically-evolving' context easier to interpret.","section":"§1 and §2"}],"recommendation":"major_revision","confidential_remarks":"The core experiment is sound and the external comparison is a real strength, but the headline recovery fractions need to be reconciled with the acknowledged model misspecification and the systematic residual trend. I think this is fixable within the paper's scope: re-cast the abstract and conclusions around residual-based accuracy or add a systematic error component. If the authors choose to retain the precision-only σ, they should clearly state that the 15/16 and 12/16 numbers are precision-coverage fractions, not accuracy metrics. I would also encourage them to include the MCMC priors and convergence diagnostics in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid benchmark paper with an honest central caveat that should keep it from being oversold. The OTI-inferred vertical accelerations are compared against FIRE-2 m12i's stored accelerations — genuinely external, since those accelerations play no role in the fit — across 16 solar-analog volumes. That is new and useful, and it gives the community a selection-function-free route to vertical mass profiles with quantified success rates.\n\nThe real issue is the numeracy of the headline. The 94%/75% claim is computed against uncertainties that combine MCMC and bootstrap only. The authors themselves say in Section 5.1 that this error \"reflects the precision of our measurement but not model inaccuracies.\" So the fraction of volumes whose true profile falls inside the error bars is a statement about precision-only coverage, not full accuracy. And the Section 7 claim that 12/16 is \"almost what we observe\" compared to an expected 11/16 is circular, because the same incomplete sigma defines both the observed counts and the statistical expectation. The systematic residual trend in Figure 7 — median ~15% offset, OTI overpredicting above the midplane and underpredicting below — reinforces that model bias is present. This does not sink the paper; the authors are transparent, and the benchmark is still informative. But the abstract and conclusions should be rephrased to say \"precision-only error bars,\" and a systematic error budget or a model-agnostic agreement metric should be added.\n\nWhat is genuinely good: the multi-volume test in a full-hydro cosmological simulation is new; prior tests used analytic potentials and a single N-body merger. The sensitivity analysis (scale height, median age, total density) is thoughtful and matches the physical expectation from the R-z separability assumption. Public code and public simulation data make the work reproducible.\n\nThe weakest structural assumption — neglecting the radial collisionless Boltzmann term and assuming R-z separability — is explicitly flagged in a footnote in Sections 3 and 4.1. The failures in V13 and V14 are exactly the large-scale-height, older-population volumes where this assumption breaks down. That is a coherent story, but it means the method's quoted precision is not its accuracy in those regimes. Minor point: all volumes sit at R = 8 kpc, so the test does not probe radial variation of the potential, and the Jeans comparison is about precision, not accuracy — the authors do acknowledge that.\n\nWho should read it: anyone doing local vertical mass modeling, Gaia/spectroscopic chemodynamics, or OTI applications. It deserves a serious referee, but I would send it back for a revision that separates precision from accuracy in the headline claims.","headline":"Useful FIRE-2 benchmark for OTI, but the 94%/75% recovery fractions lean on precision-only error bars the authors themselves flag as incomplete.","tokens_in":28001,"tokens_out":2528,"would_cite":true,"duration_ms":28743,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Run on a realistic, out-of-equilibrium simulated galaxy, Orbital Torus Imaging recovers the true vertical acceleration within $3\\sigma$ for 15 of 16 solar-analog volumes and within $1\\sigma$ for 12 of 16.","keywords":["Orbital Torus Imaging","Galactic potential","vertical acceleration","FIRE-2 simulation","disequilibrium","vertical metallicity gradient","surface mass density","Milky Way disk"],"falsifier":"Re-run the analysis on the four volumes that missed the $1\\sigma$ threshold (V2, V13, V14, V16) using only stars with $|z| < 1.5$ kpc. The paper's own explanation predicts these volumes should then agree within $3\\sigma$; if they still disagree, the proposed failure mechanism, breakdown of vertical-radial separability, is wrong.","tokens_in":26866,"feed_emoji":"🌌","tokens_out":11828,"duration_ms":103129,"temperature":0.7,"pith_summary":"Orbital Torus Imaging (OTI) infers a galaxy's gravitational potential from the way a stellar label, here the mean iron abundance $\\langle\\mathrm{[Fe/H]}\\rangle$, is arranged in the vertical phase space of height $z$ and vertical velocity $v_z$, without fitting a parametric potential or integrating orbits. The paper asks whether this steady-state machinery still works in a galaxy that is obviously not in equilibrium, the situation the Milky Way is actually in, and answers by running OTI on the FIRE-2 cosmological simulation m12i, whose true vertical accelerations are stored. It finds that OTI recovers the simulated acceleration profiles within $3\\sigma$ for 15 of 16 and within $1\\sigma$ for 12 of 16 solar-analog volumes, and that its estimates are roughly 85% more precise than Jeans modeling on the same data. Accuracy degrades in volumes with thick, old, kinematically heated stellar populations, pointing to the model's load-bearing assumption that vertical and radial motions decouple.","feed_headline":"Orbit-map method nails galaxy gravity in 15 of 16 simulated zones","feed_subtitle":"The technique recovers vertical gravity within 3σ in 94 percent of test volumes, despite a churning disk.","key_machinery":"The engine of the method is the vertical metallicity gradient: stars near the midplane carry higher $\\langle\\mathrm{[Fe/H]}\\rangle$ than stars at large $|z|$, so contours of constant mean abundance in the $(z, v_z)$ plane outline the shapes of stellar orbits. OTI parameterizes those contours as ellipses distorted by $m = 2$ and $m = 4$ Fourier terms whose coefficients are splines of a proxy vertical action $r_z$, fits the 26-parameter model to binned abundance data by gradient descent through automatic differentiation, and then reads the vertical acceleration $a_z$ off the best-fit orbital shapes through the one-dimensional vertical collisionless Boltzmann equation with the radial term dropped. The proxy-action splines carry the argument: they let the model represent the transition from elliptical orbits near the midplane to pinched, diamond-like orbits at large heights without ever committing to a parameterized mass model, which is why the technique is insensitive to the selection function.","core_discovery":"The paper's central claim is that OTI, despite assuming axisymmetry and steady state, recovers the true vertical acceleration profile of a realistic Milky-Way-mass galaxy that is out of equilibrium. Across 16 solar-analog volumes at $R = 8$ kpc, the inferred $a_z$ matches the simulation's stored acceleration within $3\\sigma$ in 15 volumes and within $1\\sigma$ in 12, and the total surface mass density at $z = 1.1$ kpc agrees within $3\\sigma$ in the same 15 volumes. The method's failures concentrate in the volumes with the largest scale heights (above roughly 1.5 kpc) and oldest stars, exactly where the assumed separability of vertical and radial motion breaks down; notably, vertical asymmetry of the density distribution does not hurt accuracy. A previously published OTI estimate of the Milky Way's surface mass density at $z = 1.1$ kpc falls squarely within the range of true simulated values, which the authors take as evidence that the method can be trusted on real survey data.","pith_inferences":["The four volumes that miss the $1\\sigma$ bar (V2, V13, V14, V16) are precisely those with large scale heights and old median ages, so a cheap pre-analysis cut on those two observables could flag survey volumes where OTI results are likely biased before any potential fitting is done.","The residual pattern the paper reports, overpredicting $a_z$ above the midplane and underpredicting it below by about 15%, is a coherent fingerprint of the dropped radial term; if the same antisymmetric signature appears in real Milky Way data, it would identify the separability approximation, not measurement noise, as the limiting error.","Because the paper finds that vertical asymmetry barely affects accuracy, OTI may be resilient to the warps and breathing modes the Milky Way is currently experiencing; this resilience is testable directly by running the same protocol on mock Gaia-like catalogs drawn from other FIRE-2 snapshots.","The natural stress test is to repeat the 16-volume protocol on FIRE-2 galaxies that have undergone recent mergers; recovery fractions falling well below 94% would map out how much disequilibrium OTI can tolerate and where it should not be trusted."],"forward_implications":["OTI can be applied to Milky Way-like disks that are mildly out of equilibrium, with a quantified error budget: 94% of volumes within $3\\sigma$ and 75% within $1\\sigma$ for the vertical acceleration, and the same 94% for the surface mass density at $z = 1.1$ kpc.","The Milky Way's OTI-inferred surface mass density from earlier work lands inside the range of true FIRE-2 values, so existing OTI-based Galactic constraints gain an independent plausibility check.","With uncertainties from MCMC sampling and bootstrapping, OTI's vertical acceleration estimates are about 85% more precise than Jeans modeling, making it the sharper tool for mapping the disk's vertical mass profile.","OTI is most trustworthy in thin, dense, young volumes; regions with scale heights above about 1.5 kpc or old stellar populations should be flagged before interpreting their inferred potentials.","The error bars derived from the simulation establish a touchstone for interpreting results from current and forthcoming surveys such as SDSS-V, Gaia, WEAVE, and 4MOST."],"supporting_citations":[{"why":"Introduces Orbital Torus Imaging, the idea that element-abundance gradients in phase space trace orbital structure and constrain the potential.","marker":"Price-Whelan et al. 2021"},{"why":"Supplies the flexible orbit-shape model, the torusimaging code, and the equations for the vertical acceleration and proxy action that this paper fits to FIRE-2 data.","marker":"Price-Whelan et al. 2025"},{"why":"Provides the Milky Way OTI estimate of the total surface mass density at z = 1.1 kpc that the paper compares against its simulated values, plus the surface-density equation used.","marker":"Horta et al. 2024"},{"why":"Defines the FIRE-2 feedback physics that produced the simulated galaxy m12i used as the ground-truth testbed.","marker":"Hopkins et al. 2018"},{"why":"Establishes the Latte suite and m12i initial conditions and the cosmological parameters of the simulation.","marker":"Wetzel et al. 2016"},{"why":"Shows that the vertical acceleration gradient distinguishes SIDM from CDM and motivates the R = 8 kpc volume choice and the dark-matter discussion.","marker":"Arora et al. 2024b"},{"why":"Supplies the multi-galaxy FIRE-2 vertical metallicity profiles used to check that m12i's gradient is typical.","marker":"Graf et al. 2025"},{"why":"Defines the total surface mass density at z = 1.1 kpc benchmark that the paper uses as its traditional surface-density comparison.","marker":"Kuijken & Gilmore 1991"},{"why":"Provides the M1 merger simulation that the earlier single-volume OTI test ran on, the comparison case this work extends to cosmology with feedback.","marker":"Hunt et al. 2021"}],"fun_headline_variants":["Orbit-map method recovers galaxy gravity in 94% of test zones","Orbital torus imaging holds up against chaotic galaxy simulation","Galaxy gravity from stellar orbits: passes 15 of 16 on FIRE","OTI method recovers gravity in churning galaxy disk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes a star's up-down motion is independent of its in-plane motion, which holds only for stars on nearly circular, dynamically cold orbits and fails at heights above roughly 1.5 kpc.","fun_headline_variants_meta":{"raw":{"variants":["Orbit-map method recovers galaxy gravity in 94% of test zones","Orbital torus imaging holds up against chaotic galaxy simulation","Galaxy gravity from stellar orbits: passes 15 of 16 on FIRE","OTI method recovers gravity in churning galaxy disk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2969,"prompt_tokens":1053,"completion_tokens":1916,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":1839}},"tokens_in":669,"tokens_out":1916,"duration_ms":15929,"temperature":1.0,"reasoning_tokens":1839,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:02:49.466730+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis on the four volumes that missed the $1\\sigma$ threshold (V2, V13, V14, V16) using only stars with $|z| < 1.5$ kpc. The paper's own explanation predicts these volumes should then agree within $3\\sigma$; if they still disagree, the proposed failure mechanism, breakdown of vertical-radial separability, is wrong.","supporting_citations":[],"review_version":1}