{"id":"42d4331e-fddb-4818-8b59-c2fe543ab7ff","arxiv_id":"2506.13877","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Applying their Zingularity neural network to real 2017 EHT data, the authors infer M87* to be a retrograde, high-spin, magnetically arrested system and Sgr A* to be a high-spin, low-inclination system, with future array upgrades reducing parameter errors.","lead":"The authors used a Bayesian neural network trained on thousands of simulated telescope images to estimate black hole spin and accretion parameters for M87* and Sgr A* from 2017 Event Horizon Telescope data. The inferred values favor a retrograde, magnetically arrested flow around M87* and a high-spin, low-inclination Sgr A*, and they predict a future telescope in Africa will sharpen constraints on non-Kerr gravity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The M87* posterior peaks at the training grid boundary (Rhigh=160) and at two neighboring spin grid values (-0.5 and -0.94); the paper concedes these may be grid artifacts, so the headline M87* spin/Rhigh claim is not yet established.","rationale":"The reader's weakest assumption identifies the finite, discretely sampled GRMHD library as the load-bearing premise. My concern sharpens this to a specific, observable failure mode: for M87*, the inferred parameters sit at the very edge of the library (Rhigh=160) and at two neighboring grid points in spin. The paper itself acknowledges both effects in Section 5, so the concern is not speculative. If the M87* posterior were a true Bayesian posterior over the continuous parameter space, it would not be pinned to training-grid labels, especially with a network designed to interpolate. The proposed test directly checks whether the inference moves when the grid is extended and refined. If it moves, the abstract's M87* claim is a grid artifact; if it does not, the concern is resolved. Because this is a concrete, addressable limitation rather than a fundamental flaw, the reader's CONDITIONAL verdict remains appropriate, and no change to that verdict is needed. I do not find a comparably load-bearing issue for the Sgr A* inference, which lies interior to the training grid, nor for the AMT validation-error comparison, which is based on held-out synthetic data and is less sensitive to domain shift.","tokens_in":15320,"tokens_out":3929,"duration_ms":42475,"concrete_test":"Run new GRMHD simulations at a* = -0.7, -0.9 and Rhigh = 320 (plus an interior point, e.g., a* = -0.7, Rhigh = 160), ray-trace them with the same Symba/MeqSilhouette pipeline used for the training library, add them to the training set, retrain the M87* BANN, and re-run inference on the April 11 2017 data. If the posterior shifts away from the old boundary/grid peaks (e.g., to a* = -0.7 or to Rhigh > 160), the published M87* parameters are grid artifacts; if it remains at the old labels, the boundary concern is resolved and the current claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 (Fig. 4) reports an M87* posterior with Rhigh = 159.91 ± 0.34, and Section 5 states that 'the training data had a maximum Rhigh value of 160, so it might be that a model with an even higher Rhigh would describe the data better.' A posterior whose peak sits at the edge of the training range cannot support the claim that M87* is 'best described by' Rhigh=160; at most it supports a lower limit. Similarly, the spin posterior is bimodal at a* = -0.5 and -0.94, which Section 5 explicitly says 'likely peaks at those particular values because they are two neighboring values in our GRMHD training data grid space.' The network is an interpolator within a discrete library; when the data drive it to the library boundary or to a gap between grid points, the point estimate and the uncertainty are artifacts of the sampling, not inferences about the source. The abstract's 'spin between 0.5 and 0.94' is the span of two training labels, not a measured range. This does not invalidate the Sgr A* inference, whose parameters fall interior to the grid, but it removes the M87* half of the paper's central claim unless the grid is refined.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies Bayesian neural networks (the Zingularity framework) trained on a large library of ray-traced GRMHD synthetic EHT observations to the 2017 EHT data for M87* and Sgr A*. It uses 1000 bootstrapped realizations of the data with simulated gain errors, polarization leakage, thermal noise, and Faraday rotation, and reports posteriors for spin, electron heating ratio Rhigh, magnetic state, inclination, and position angle. The main results are that M87* is best described by a retrograde MAD accretion flow with Rhigh near the training maximum of 160 and spin between -0.5 and -0.94, while Sgr A* prefers high spin (~0.8-0.9), low inclination, position angle 106-137 degrees, and an inconclusive MAD/SANE state. The paper also predicts that the proposed Africa Millimeter Telescope (AMT) extension will reduce parameter inference errors by a factor of three for non-Kerr models.","tokens_in":15623,"tokens_out":5051,"duration_ms":54262,"significance":"If the M87* inference were clean, this would be a significant methodological demonstration: full-Stokes Bayesian neural network inference on EHT visibilities, with explicit bootstrapping of known systematics and a substantially larger synthetic library than earlier work. The paper has real strengths: reproducible data reduction scripts, public availability of the training data and code, careful validation diagnostics, and unusually candid acknowledgment of the limitations of the discrete GRMHD grid. The Sgr A* results are plausible and largely interior to the training grid, and the AMT forecasting exercise is useful. However, the M87* headline parameters sit at or between discrete training grid values, and the paper itself concedes these may be artifacts; as a result the M87* half of the central claim is not yet established at the level claimed in the abstract.","major_comments":[{"comment":"The M87* Rhigh posterior peaks at Rhigh = 159.91 with the training maximum at 160, and the text concedes that a model with an even higher Rhigh might describe the data better. A posterior peaking at the edge of the training range cannot support the abstract's statement that M87* is 'best described by' Rhigh=160; at most it supports a lower limit under the current library. This should be fixed by extending the Rhigh grid or by explicitly reporting the result as censored at the training boundary.","section":"Section 4, Fig. 4, and Section 5"},{"comment":"The M87* spin posterior is bimodal at a* = -0.5 and -0.94, which Section 5 identifies as two neighboring values in the GRMHD training grid. The abstract's 'spin between 0.5 and 0.94' is therefore the span of two discrete labels, not a measured credible interval. The paper needs a refined spin grid to test whether the bimodality persists, or the claim must be downgraded to 'the data are consistent with retrograde spins in this range, with the posterior location currently driven by grid discreteness.'","section":"Section 4, Fig. 4, and Section 5"},{"comment":"The claim that the inference is 'without being impacted by the unknown foreground Faraday screens and data calibration biases' is supported by internal validation on the synthetic library and by the RM de-rotation test, but that test only varies the constant RM component. The paper states that applying a time-variable RM changes the results because the network uses internal time-variability of Q and U phases as discriminating features. Since the training library is the only source of that variability, the posterior widths for Sgr A* inclination and spin may be under-calibrated if real Faraday variability differs from the simulated one. A coverage test on held-out simulations with perturbed Faraday and calibration parameters would directly address this concern.","section":"Section 3, Fig. 3, and Section 4"}],"minor_comments":[{"comment":"The header and abstract run 'Zingularityresults' together; there should be a space between 'Zingularity' and 'results'.","section":"Title"},{"comment":"The sentence 'For Kerr Sgr A∗, we two equally viable models' is missing a verb and should read 'we have two equally viable models'.","section":"Section 3"},{"comment":"The sentence 'The low Rhigh∼ 14 value corresponds to a SANE MAD = 0.36+0.25−0.20' is unclear; it should specify that this is the inferred value associated with a particular magnetic-state classification, not a definition of Rhigh.","section":"Section 4"},{"comment":"The phrase 'The small 10◦ difference in inclination angle' does not specify which two inferences are being compared; please clarify whether it refers to the two fiducial Sgr A* networks or the fiducial versus dilaton network.","section":"Section 5"},{"comment":"The discrete parameter grid (spin values, Rhigh values, ilos and PA step sizes) is not fully stated in this paper; it should be listed explicitly so that boundary and grid-discreteness effects can be assessed without consulting the companion papers.","section":"Sections 4 and 5"}],"recommendation":"major_revision","confidential_remarks":"The M87* grid-boundary issue is the main risk to the paper's headline claim. If the authors cannot extend the GRMHD grid, they should soften the abstract and conclusions to present M87* as a lower limit on Rhigh and a grid-induced bimodality in spin. The Sgr A* results and the AMT forecasting exercise are publishable and useful. The paper relies heavily on two companion papers; that is acceptable, but this manuscript should be self-contained enough for the grid claims to be evaluated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this is the first application of the full-Stokes Zingularity BANN to real 2017 EHT data. The Sgr A* posterior is the part that holds up, because its parameters fall inside the discrete GRMHD training grid. The M87* half does not. The posterior peaks at Rhigh=160, the edge of the training range, and at spin values -0.94 and -0.5, which are neighboring grid points. The authors say so themselves in Section 5. The abstract's 'spin between 0.5 and 0.94' is therefore the span of two training labels, not a measured range.\n\nCredit where due: the paper is unusually transparent. The bootstrapping of gain errors, polarization leakage, and thermal noise is careful, and the validation on withheld GRMHD data is standard practice with reported errors. The code and data availability are exemplary, with containerized Rpicard and Symba pipelines. The AMT forecast is a clean simulation experiment; the factor-of-three reduction in validation error for the dilaton models is a useful projection, though it is a forecast, not an observation.\n\nThe soft spots are real and central. The load-bearing weakness is the discrete GRMHD library. The paper concedes that the M87* spin peak may be a grid artifact and that Rhigh might be higher than 160, so the M87* result is at best a lower limit. The abstract's Faraday-robustness claim is stronger than what Section 5 supports: constant RM de-rotation leaves posteriors unchanged, but time-variable RM changes them. That is robustness to constant foreground screens, not to Faraday rotation in general. The Sgr A* MAD/SANE state is inconclusive, which the paper admits, and the inclination values are tied to 20-degree training sampling.\n\nThe stress-test note holds up on reading. There is no equation-level circularity; the network is a learned map from visibilities to labels. The issue is sampling resolution, not fitting.\n\nWho gets value: anyone working on GRMHD parameter inference from EHT data, and anyone using deep learning on sparsely sampled simulation libraries. It deserves a serious referee, but the referee should require the M87* claims to be reframed as grid-boundary artifacts, and the abstract's Faraday claim to be matched to the actual test.\n\nRecommendation: send it to peer review after that revision.","headline":"First full-Stokes BANN inference on 2017 EHT data; Sgr A* is credible, M87* spin/Rhigh claims sit on training-grid boundaries and need reframing.","tokens_in":16178,"tokens_out":3086,"would_cite":true,"duration_ms":29443,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network trained on simulated EHT images finds M87* in a retrograde magnetically arrested state, Sgr A* at high spin, and predicts the Africa Millimeter Telescope cuts non-Kerr errors threefold.","keywords":["black hole parameter inference","Event Horizon Telescope","GRMHD simulations","Bayesian neural networks","M87*","Sgr A*","very long baseline interferometry","general relativity tests"],"falsifier":"Apply the identical trained network to an independent EHT epoch for Sgr A* and to a later M87* campaign; if the inferred parameters shift outside the stated posteriors, the synthetic library is not representative of real data. A spectroscopic or polarimetric detection of a nonthermal electron distribution in M87*'s inner flow would similarly falsify the thermal temperature-ratio premise behind the jet-synchrotron interpretation.","tokens_in":15112,"feed_emoji":"🕳️","tokens_out":16969,"duration_ms":150624,"temperature":0.7,"pith_summary":"The paper reports the first direct application of a Bayesian neural network to infer black-hole and accretion parameters from 2017 Event Horizon Telescope visibility measurements. The network, trained on a large library of ray-traced GRMHD simulations with simulated calibration errors, returns narrow posteriors that place M87* in a retrograde magnetically arrested accretion state with spin magnitude 0.5 to 0.94 and strong jet synchrotron emission, and Sgr A* at high spin (about 0.8 to 0.9), low inclination, and position angle 106 to 137 degrees, with no clear magnetic-state classification. The same machinery predicts that adding the Africa Millimeter Telescope will reduce non-Kerr parameter-inference errors by a factor of three, sharpening future tests of general relativity. If these inferences hold, they provide a fast, reproducible route from raw EHT visibilities to physical parameters and a quantitative guide for array design.","feed_headline":"Deep learning pegs M87* with retrograde spin, Sgr A* high spin","feed_subtitle":"Bayesian network on 2017 EHT data finds magnetically arrested M87*; predicts Africa telescope cuts errors 3x.","key_machinery":"Zingularity: a Bayesian artificial neural network that combines a residual network with variational fully connected layers, trained on synthetic EHT observations generated from ray-traced GRMHD simulations. The network maps full-Stokes visibility amplitudes and phases to posterior distributions over spin, electron-proton temperature ratio, MAD/SANE magnetic state, inclination, and position angle. Inference is repeated over 1000 bootstrap realizations of the observational data with simulated polarization leakage, gain and gain-curve errors, and thermal noise, so the reported posteriors include data systematics.","core_discovery":"The paper's central claim is that, under its GRMHD model library, the 2017 EHT data are best explained by a retrograde magnetically arrested disk around M87* with spin magnitude between 0.5 and 0.94 and an electron-proton temperature ratio at the training-grid maximum, implying jet-dominated synchrotron emission; and by a high-spin, low-inclination, prograde flow around Sgr A* with position angle near 106 to 137 degrees, in a state beyond the standard MAD/SANE dichotomy. The authors present this as the first application of a Bayesian neural network trained on a synthetic GRMHD library directly to EHT visibilities, using full Stokes information and bootstrapping of known systematics. They also claim the same network predicts a threefold reduction in non-Kerr parameter-inference errors when the Africa Millimeter Telescope is added to the array. They interpret the M87* spin posterior as peaking at neighboring grid values and therefore argue the true spin is likely intermediate between the two, and they emphasize that the posteriors are insensitive to constant Faraday rotation and calibration gain biases.","pith_inferences":["Beyond the paper: the method treats the synthetic GRMHD library as the prior, so the reported posteriors are conditional on that library; retraining on a library that includes nonthermal electron distributions, tilted disks, or different magnetic field polarities could shift the inferred parameters.","Beyond the paper: the demonstrated insensitivity to constant Faraday rotation could be exported to other polarized VLBI sources, letting the network separate foreground screen rotation from intrinsic source structure without explicit rotation-measure fitting.","Beyond the paper: the failure to train on Kerr-Newman models points to a spin-charge degeneracy in current baseline coverage; a network that recovers both spin and charge would provide a direct, quantitative test of whether future arrays can break the degeneracy.","Beyond the paper: a cheap validation would be to apply the same trained networks to the 2018 or 2021 EHT epochs; stable posteriors across epochs would strengthen the claim that the 2017 result is not an artifact of that particular observing run."],"forward_implications":["If M87* is truly in a retrograde magnetically arrested state, its powerful jet and counter-rotation fit a merger history, and the inferred parameters satisfy jet-power constraints measured on larger scales.","If Sgr A* has high spin and a spin axis nearly aligned with our line of sight, the polarization-loop direction from GRAVITY favors the inclination solution that makes the accretion flow rotate clockwise on the sky.","Because the network is insensitive to the constant Faraday rotation measure, the posteriors do not require assumptions about where the rotation measure originates, removing a known systematic in previous GRMHD scoring.","Adding the Africa Millimeter Telescope should reduce non-Kerr parameter-inference errors by roughly a factor of three, primarily through improved northeast-southwest resolution and short-baseline calibration tracks.","Producing new GRMHD simulations at the inferred interpolated parameters would allow direct model-data comparisons of accretion rate, jet power, and broadband spectral energy distributions."],"supporting_citations":[{"why":"Supplies the library of synthetic EHT observations and the improved CASA calibration data that define the training set.","marker":"Janssen et al. (2025a)"},{"why":"Supplies the Zingularity Bayesian neural network architecture, training, and bootstrap-based uncertainty estimation used for inference.","marker":"Janssen et al. (2025b)"},{"why":"Provides the 2017 M87* and Sgr A* EHT visibility data that the network is applied to.","marker":"Event Horizon Telescope Collaboration et al. 2019b, 2022b"},{"why":"Defines the MAD/SANE magnetic-state classes and the previous GRMHD scoring results that frame the inference.","marker":"Event Horizon Telescope Collaboration et al. 2019c, 2022d"},{"why":"Introduces the electron-proton temperature-coupling parameter whose posterior is a central output.","marker":"Moscibrodzka et al. 2016"},{"why":"Provided the Symba pipeline that generates synthetic VLBI training data matching the calibrated observations.","marker":"Roelofs et al. 2020"},{"why":"Proposed the Africa Millimeter Telescope whose addition is predicted to reduce non-Kerr errors threefold.","marker":"Backes et al. 2016"},{"why":"Analyzed the resolution gains from the Africa Millimeter Telescope that the paper invokes to explain the error reduction.","marker":"La Bella et al. 2023"},{"why":"Supplied the dilaton GRMHD models used for the beyond-Kerr parameter-inference tests.","marker":"Mizuno et al. 2018; Röder et al. 2023"}],"fun_headline_variants":["Deep learning: M87* retrograde spin, Sgr A* high spin","AI on EHT: M87* retrograde, Sgr A* fast","Neural net pegs M87* retrograde, Sgr A* high spin","Bayesian AI: M87* retrograde, Sgr A* high spin","Zingularity: M87* retrograde, Sgr A* spins fast"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands on the assumption that the finite set of simulated black-hole images used for training adequately represents the real 2017 EHT data, including its calibration errors, polarization leakage, and Faraday rotation; if the real data contain physics absent from the simulations, the inferred posteriors would not follow.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning: M87* retrograde spin, Sgr A* high spin","AI on EHT: M87* retrograde, Sgr A* fast","Neural net pegs M87* retrograde, Sgr A* high spin","Bayesian AI: M87* retrograde, Sgr A* high spin","Zingularity: M87* retrograde, Sgr A* spins fast"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00067,"raw_usage":{"total_tokens":3166,"prompt_tokens":1173,"completion_tokens":1993,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":789,"completion_tokens_details":{"reasoning_tokens":1885}},"tokens_in":789,"tokens_out":1993,"duration_ms":16135,"temperature":1.0,"reasoning_tokens":1885,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:27:16.432765+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the identical trained network to an independent EHT epoch for Sgr A* and to a later M87* campaign; if the inferred parameters shift outside the stated posteriors, the synthetic library is not representative of real data. A spectroscopic or polarimetric detection of a nonthermal electron distribution in M87*'s inner flow would similarly falsify the thermal temperature-ratio premise behind the jet-synchrotron interpretation.","supporting_citations":[{"cited_title":"M., et al","cited_arxiv_id":null,"evidence_quote":"Supplied the dilaton GRMHD models used for the beyond-Kerr parameter-inference tests."}],"review_version":1}