{"id":"c7d34d7e-4f6e-4019-903e-99d99fb33804","arxiv_id":"2506.13875","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Bayesian neural networks trained on synthetic EHT observations of Sgr A* and M87* recover spin and magnetic state well in cross-code tests, but give overconfident wrong estimates for temperature ratio and inclination when the training grid is sparse.","lead":"This astronomy methods paper presents Zingularity, an open-source Bayesian neural network framework for estimating black hole spin and accretion state from Event Horizon Telescope observations, trained on large libraries of simulated observations. It matters because it tests whether deep learning plus full polarization data can replace slower grid-based model scoring, while also revealing when the trained network gives overconfident wrong answers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-grid bhac tests in §6.3 yield confidently wrong posteriors (Rhigh 1→34.5, ilos 70→106), contradicting the text's 'small errors' and undermining the 'trustworthy uncertainties' claim.","rationale":"The reader's weakest assumption centers on sparse grid coverage of the GRMHD parameter space, with the Sgr A* MAD bhac test as the failure example. My reading confirms that concern but identifies a more specific and more damaging problem: even for test models whose parameters lie inside the training grid, the posteriors are confidently wrong under a change of simulation code. The Sgr A* SANE model (Rhigh=1, ilos=70) is a within-grid case, yet the inferred Rhigh and ilos are far from truth and their credible intervals exclude the truth. Similarly, the M87* SANE model (Rhigh=40) yields Rhigh≈8.8. Thus the issue is not only missing parameter combinations but the network's sensitivity to code-level details of the forward model. Since real observational data are generated by whatever the true accretion physics is, and not by kharma with the Rhigh prescription, this mismatch is a first-order threat to the 'trustworthy uncertainties' and 'reliable results from observational data' claims. The paper's own text in §6.3 ('small Rhigh and ilos errors') is contradicted by its Figure 5, an internal inconsistency that should be corrected regardless. The proposed coverage test would settle whether the posterior widths are calibrated under this kind of code variation; if coverage is far below nominal, the conclusion that the network generalizes well to observational data is not supported. I nevertheless keep the reader's CONDITIONAL verdict: the framework is open-source, reproducible, and the authors are transparent about many failure modes, so the issues are addressable by rescoping claims and adding calibration diagnostics rather than by rejecting the work outright.","tokens_in":32391,"tokens_out":5465,"duration_ms":55939,"concrete_test":"Run an empirical coverage test on the existing bhac test suite: for each test model and each inferred parameter, compute the percentile rank of the ground-truth value under the network's posterior, using the reported 1000 bootstraps × 1000 posterior draws. If the in-grid truths (Sgr A* Rhigh=1 and ilos=70; M87* Rhigh=40) fall outside the nominal 68% or 95% credible intervals, the uncertainties are miscalibrated under code variation. Extend to a larger set of bhac models with parameters on the kharma grid (all available combinations of a*, Rhigh, ilos) and compare empirical coverage to nominal coverage; also reconcile the §6.3 'small errors' statement with the Figure 5 posteriors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6.3 reports the only out-of-distribution validation: bhac/raptor test data fed through kharma-trained networks. For the Sgr A* SANE model with parameters inside the training grid (a*=0.94, Rhigh=1, ilos=70), the posterior mean is Rhigh=34.5 (68% interval roughly 26–43) and ilos=105.95 (68% interval roughly 102–109); both ground truths lie far outside the credible intervals. The M87* SANE model (Rhigh=40) is inferred as Rhigh≈8.8. The text states 'We ascribe the small Rhigh and ilos errors to the aforementioned differences in ray-tracing,' which directly contradicts Figure 5. These are not small errors, and the posterior widths do not expand to reflect the simulator mismatch. The central claim—that the Bayesian nature gives 'trustworthy uncertainties' and that the networks 'can generalize well so that reliable results can be obtained from observational data'—requires the posterior to be calibrated under forward-modeling uncertainty that will inevitably be present for real EHT data. This test shows that a modest code-level change (kharma→bhac) produces confident, wrong posteriors for parameters inside the training grid; the failure is therefore not solely due to the sparse parameter grid, as the paper asserts, but to the network latching onto code-specific features. The paper's own diagnostics do not detect this failure from the posterior alone.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents Zingularity, an open-source TensorFlow-based framework for Bayesian neural network inference on very long baseline interferometry data, applied here to synthetic Event Horizon Telescope observations of Sgr A* and M87*. The training library is built from a large set of kharma GRMHD-GRRT model images processed through the Symba signal-path simulator, and the network maps full-Stokes visibilities to posteriors over MAD/SANE classification, spin, Rhigh, inclination, and position angle. The paper reports training diagnostics, hyperparameter surveys, and validation tests on both held-out kharma data and independent bhac-raptor test data, with the stated conclusion that the Bayesian networks yield trustworthy uncertainties and generalize well enough for reliable inference on observational data. The companion paper is said to apply the trained networks to real EHT observations.","tokens_in":32682,"tokens_out":2424,"duration_ms":30120,"significance":"If the central claims hold, this would be a substantial practical contribution: a publicly available, containerized, reproducible deep-learning pipeline that uses full-Stokes visibility information and a very large synthetic training set to infer GRMHD parameters orders of magnitude faster than conventional likelihood-based methods. The paper's strengths include the unusually large and detailed training library, the use of full polarization products, the careful treatment of signal-path corruption effects, the explicit hyperparameter stability surveys, and the availability of a Docker container with configuration files for reproduction. The inclusion of an out-of-distribution bhac-raptor test is commendable and provides the only genuinely independent evidence in the paper. However, the interpretation of that test is where the paper's central claim breaks down: the sole independent out-of-distribution test produces confidently wrong posteriors for parameters that lie inside the training grid, and the paper's own text incorrectly describes these as small errors.","major_comments":[{"comment":"The only out-of-distribution validation test with in-grid parameters is the Sgr A* SANE bhac test with ground truth a*=0.94, Rhigh=1, ilos=70. The inferred posterior is a*=0.94+0.03/-0.05, Rhigh=34.49+8.49/-2.26, and ilos=105.95+3.43/-4.14. The true Rhigh and ilos lie far outside the 68% credible intervals. The text states that \"We ascribe the small Rhigh and ilos errors to the aforementioned differences in ray-tracing,\" but these are not small errors in any meaningful sense, and the posterior widths do not expand to reflect the simulator mismatch. Since this is the only test that does not share the training forward model, it directly contradicts the abstract and Section 7 claims that the Bayesian nature of the networks gives trustworthy uncertainties and that the networks generalize well enough for reliable results on observational data. The paper needs either a calibration/tempering procedure for forward-model uncertainty or a substantially weakened statement of what the posteriors mean.","section":"§6.3, Fig. 5"},{"comment":"The low validation errors shown in Figure 4 are computed on held-out samples drawn from the same kharma GRMHD-GRRT library used for training. This is an interpolation check within one forward-modeling family, not evidence of external reliability. The paper repeatedly uses these low validation errors as evidence of generalization, but the only independent test, the bhac-raptor data in Section 6.3, shows that the network latches onto code-specific features: in-grid parameters are misestimated with high confidence. The manuscript should explicitly state that Figure 4 validates interpolation within the training library, and it should not be cited as evidence that the network generalizes across forward-modeling assumptions.","section":"§4, §6.1"},{"comment":"The paper explains the Sgr A* SANE model misidentification as being due to \"the limited grid of model parameters in the training data.\" This explanation is insufficient and partly contradicted by the paper's own test: the Sgr A* SANE model has a*=0.94, Rhigh=1, ilos=70, all inside the training grid, yet the posterior is concentrated at Rhigh=34.5 and ilos=105.95. If the failure were purely a grid-density problem, an in-grid test should not fail this badly. The visibility comparison in Figure 6 shows that the network, by design, finds the kharma model with the most similar visibilities, which is a different code's model; this demonstrates sensitivity to code-specific nuisance features rather than a mere sparsity effect. The manuscript should either provide a kharma-based counter-test at the true parameter values or acknowledge that the posterior cannot be expected to signal this form of failure.","section":"§6.3"},{"comment":"The bootstrapping procedure in Section 5.2 resamples only the known corruption effects (D-terms, gains, gain curves, thermal noise). It does not marginalize over the GRMHD-GRRT forward-modeling uncertainty, which the bhac-raptor test shows is a dominant source of systematic error. Section 8 lists alternative electron prescriptions, nonideal MHD, and other extensions as future work, but the claims in the abstract and Section 7 that \"uncertainties in the data are accurately taken into account\" and that the posteriors are \"trustworthy\" require the model-form uncertainty to be either incorporated or explicitly excluded from the scope. The scope restriction is acceptable, but it must be stated prominently in the abstract and conclusions rather than only in the outlook.","section":"§5.2, §7"}],"minor_comments":[{"comment":"The phrase \"We carried out out supervised learning\" contains a duplicated word and should be corrected.","section":"Abstract"},{"comment":"The sentence \"Additionally, showed that a sufficiently large training dataset is needed...\" is missing a subject and should read \"Additionally, we showed that...\".","section":"§7"},{"comment":"The word \"unlabled\" should be \"unlabeled.\"","section":"§5.1.1"},{"comment":"The caption describes the panels as corner plots, but the displayed figures appear to be marginal posterior distributions rather than full corner plots; the caption should match the actual figure content.","section":"Fig. 5 caption"},{"comment":"The metric discussion states that parameter-dependent performance will be evident in the posteriors because regions of poor network performance will produce wide posteriors. This is not guaranteed for misspecified forward models, as the bhac test in Section 6.3 demonstrates; a caveat should be added here.","section":"§5.1.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering contribution with excellent reproducibility practices, and the negative bhac-raptor result is scientifically valuable. My main concern is that the abstract and conclusions overstate the reliability of the posteriors relative to what the paper's own Figure 5 shows. If the authors substantially revise the claims, add a calibration discussion, and either temper the posteriors for forward-model uncertainty or clearly scope the method to interpolation within the training library, the paper would be publishable. The current mismatch between the text and the paper's own independent validation is too large to accept as a minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: Zingularity is a real engineering contribution. The containerized code, deterministic seeding, and the large full-Stokes synthetic library are genuinely useful for the EHT/ngEHT community. The bhac/raptor out-of-distribution test is the right thing to do, and it exposes exactly where the paper's claims go too far.\n\nWhat's new: the framework itself, and the scale of training data with full-Stokes visibilities. The ablation-style tests (Stokes I only vs full polarization, thermal noise only vs full corruption) are instructive and well presented.\n\nSoft spots: the central claim that the Bayesian nature gives trustworthy uncertainties and that the networks can generalize to observational data is not supported by Figure 5. For the Sgr A* SANE bhac test with parameters inside the training grid, the network returns Rhigh=34.5 (truth 1) and ilos=105.95 (truth 70), with narrow 68% intervals that exclude the truth. The text ascribes these to 'small errors' from ray-tracing differences. These are not small, and the posterior widths do not expand to reflect the simulator mismatch. The stress-test note is right: this is not purely a sparse-grid problem; the network latches onto code-specific features, and the paper's own explanation (limited grid) is contradicted by the fact that the parameters are in-grid. The M87* SANE test (Rhigh=40 inferred as ~8.8) reinforces the point.\n\nAlso, hyperparameters were selected using validation data, so the reported validation errors are optimistic. The paper would be stronger with seed-to-seed variability reported and a calibration check for the BNN posteriors.\n\nThat said, the paper is honest about model dependence and failure modes (the multimodal posterior for the out-of-grid M87* model is a good example), and the framework itself is solid enough to warrant serious refereeing. The fix is a re-scoping of the claims and a correction of Section 6.3's characterization. The companion paper applies the network to real data; this paper should not claim reliable observational inference without that analysis or without explicit hedging.\n\nWho should read it: anyone building ML-based inference pipelines for EHT/ngEHT. It deserves peer review; I would push for major revision rather than desk reject.","headline":"Useful, reproducible BANN framework for EHT parameter inference, but the paper's 'trustworthy uncertainties' claim is contradicted by its own out-of-distribution test.","tokens_in":33289,"tokens_out":1555,"would_cite":true,"duration_ms":17026,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian neural network trained on synthetic full-Stokes EHT visibilities can infer the spin, magnetic state, and temperature ratio of Sgr A* and M87*, with uncertainties that survive validation tests.","keywords":["Bayesian neural networks","Event Horizon Telescope","very long baseline interferometry","GRMHD simulations","black hole spin","Sgr A*","M87*","variational inference"],"falsifier":"One decisive test is to add the misidentified Sgr A* model (SANE, a*=0.94, Rhigh=1, ilos=70) and its parameter-space neighbors to the training library and retrain; if the network still returns a narrow posterior near Rhigh=34 and ilos=106, the claim of trustworthy uncertainties for out-of-grid data fails. A complementary test is to apply the trained networks to the real 2017 EHT observations and check whether the posterior modes contradict independently established multiwavelength constraints on Sgr A* and M87*.","tokens_in":32113,"feed_emoji":"🕳️","tokens_out":11891,"duration_ms":107083,"temperature":0.7,"pith_summary":"This paper aims to establish that a Bayesian deep neural network, trained on synthetic Event Horizon Telescope observations built from general-relativistic magnetohydrodynamic (GRMHD) simulations, can recover physical parameters of Sgr A* and M87* directly from the interferometric visibilities. The target parameters are black hole spin, the magnetic state of the accretion flow, the ion-to-electron temperature ratio, and the viewing geometry. If this holds, parameter estimation that currently requires expensive scoring of many simulation snapshots becomes near-instant and comes with posterior uncertainties. The authors present the open-source Zingularity framework, validate it on held-out synthetic data, cross-code test datasets, and bootstrapped observational noise, and conclude that reliable inference on real EHT data is achievable.","feed_headline":"Bayesian network maps EHT visibilities to black hole parameters","feed_subtitle":"Full-Stokes EHT visibilities become posteriors on spin, magnetic state, and temperature ratio.","key_machinery":"The central object is the Bayesian artificial neural network (BANN), whose weights are trainable probability distributions rather than point values. A convolutional ResNet stack compresses the full time-baseline visibility array into salient features, and dense variational layers then output a stochastic posterior over the physical parameters, with a softmax head for the magnetically arrested disk (MAD) versus standard and normal evolution (SANE) classification and linear heads for spin, temperature ratio, and viewing angle. Variational inference approximates the intractable weight posterior through an evidence-lower-bound objective, so repeated forward passes sample the predictive distribution. The other half of the machinery is the training set: hundreds of thousands of synthetic observations produced by ray-traced GRMHD models with thermal noise, gain errors, polarization leakage, atmospheric and scattering effects, with the same corruption effects bootstrapped onto the observational data at inference time.","core_discovery":"The paper's central claim is that a Bayesian artificial neural network built from convolutional ResNet blocks feeding variational layers learns a mapping from 10-second-sampled, full-polarization EHT visibilities to the GRMHD model parameters, with the network's predictive spread serving as a trustworthy posterior. The authors show that polarization information is essential, that realistic forward modeling of the signal path is necessary to avoid overconfident but wrong inference, and that the Bayesian posterior can expose failure modes: an out-of-grid M87* spin produces a multimodal posterior, while an out-of-grid Sgr A* model is confidently misidentified as its nearest training-data neighbor. Their conclusion is that, with enough training samples and honest forward modeling, the trained networks generalize well enough to give reliable parameter constraints on observational data.","pith_inferences":["Applied to real data, the posterior width should not be read as the full systematic uncertainty; a coverage audit of the training grid, such as distance to the nearest training models, should accompany any astrophysical conclusion.","A natural extension is to add explicit out-of-distribution detection to the pipeline, so that data resembling the misidentified Sgr A* model are flagged before a confident posterior is quoted.","With the denser baseline coverage of next-generation EHT arrays, the same architecture should tighten the spin and temperature posteriors, but only if the simulation library is expanded at the same time; otherwise the network will keep interpolating across gaps.","The closeness of the misidentified model to its training-data neighbor suggests the network is interpolating the GRMHD library, which implies that a finer parameter grid could turn today's confident misidentifications into broadened posteriors."],"forward_implications":["Inference becomes nearly instantaneous: once trained, producing 100 posterior samples from 100 bootstrapped datasets takes about 20 seconds, compared with the heavy computational cost of scoring many GRMHD snapshots.","Polarization is required: networks trained on Stokes I alone barely train, so the full polarization content of the visibilities is what makes spin and magnetic-state inference possible.","Simplified forward modeling is dangerous: when only thermal noise is added to the synthetic data, validation errors drop and the network would overfit the real corruption effects present in observational data.","The posterior is informative about out-of-distribution data: an out-of-grid M87* spin gives a multimodal posterior, while an out-of-grid Sgr A* model is confidently assigned to a wrong region of parameter space.","The training-grid density becomes a measurable systematic: the Sgr A* misidentification is attributed directly to the limited grid of model parameters in the GRMHD library."],"supporting_citations":[{"why":"Companion paper that defines the GRMHD-GRRT synthetic data library and the modeled EHT signal path used as the training set.","marker":"Janssen et al. (2025a)"},{"why":"Supplies the GRMHD simulation code used to generate the training library's ground-truth models.","marker":"Wong et al. (2022)"},{"why":"Supplies the alternative GRMHD code whose runs form the cross-code test datasets.","marker":"Porth et al. (2017)"},{"why":"Supplies the ray-tracing code used to create the cross-code test visibilities.","marker":"Bronzwaer et al. (2018)"},{"why":"Supplies the signal-path simulator that converts model images into synthetic EHT observations with corruption effects.","marker":"Roelofs et al. (2020)"},{"why":"Supplies the probabilistic programming library used for variational inference in the Bayesian layers.","marker":"Dillon et al. (2017)"},{"why":"Provides the Sgr A* model scoring and multiwavelength constraints against which the trained network's inferences are interpreted.","marker":"Event Horizon Telescope Collaboration et al. (2022e)"},{"why":"Provides the M87* GRMHD scoring baseline that motivates training over the full parameter space despite some regions being disfavored.","marker":"Event Horizon Telescope Collaboration et al. (2019d)"}],"fun_headline_variants":["Bayesian deep learning reads black hole spin from EHT data","Zingularity framework: Bayesian nets infer black hole parameters","EHT visibilities become black hole posteriors via Bayesian NN","Neural network yields trustworthy uncertainties for black hole spin","From EHT data to black hole spin and magnetic state with Bayesian nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the sparse grid of GRMHD training models, all assuming ideal magnetohydrodynamics and the same electron-temperature prescription, spans the true parameter space of Sgr A* and M87* closely enough that the network's learned mapping and its posterior remain valid for actual EHT data.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian deep learning reads black hole spin from EHT data","Zingularity framework: Bayesian nets infer black hole parameters","EHT visibilities become black hole posteriors via Bayesian NN","Neural network yields trustworthy uncertainties for black hole spin","From EHT data to black hole spin and magnetic state with Bayesian nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000469,"raw_usage":{"total_tokens":2375,"prompt_tokens":1023,"completion_tokens":1352,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":1280}},"tokens_in":639,"tokens_out":1352,"duration_ms":11112,"temperature":1.0,"reasoning_tokens":1280,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:28:42.227907+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One decisive test is to add the misidentified Sgr A* model (SANE, a*=0.94, Rhigh=1, ilos=70) and its parameter-space neighbors to the training library and retrain; if the network still returns a narrow posterior near Rhigh=34 and ilos=106, the claim of trustworthy uncertainties for out-of-grid data fails. A complementary test is to apply the trained networks to the real 2017 EHT observations and check whether the posterior modes contradict independently established multiwavelength constraints on Sgr A* and M87*.","supporting_citations":[{"cited_title":"2020, , 636, A5","cited_arxiv_id":null,"evidence_quote":"Supplies the signal-path simulator that converts model images into synthetic EHT observations with corruption effects."}],"review_version":1}