{"id":"27c7426c-67ac-43c3-8dae-013527d8f9bd","arxiv_id":"2506.13873","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors upgraded the EHT calibration pipeline and generated 962,000 synthetic Sgr A* and M87* datasets that emulate real telescope measurements for training deep learning inference models.","lead":"This paper improves how Event Horizon Telescope data are calibrated and releases a library of 962,000 simulated observations of the two supermassive black holes the telescope studies. These simulations mimic the telescope's real errors and noise, giving researchers a large testbed for training machine learning models to extract black hole properties.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The library's realism is undercut by the explicit §4.9 assumption that ALMA calibration errors are negligible; a network trained on synthetic data with uncorrupted model fluxes will be miscalibrated on real data carrying ALMA gain errors.","rationale":"The reader's weakest_assumption identifies the right risk: the synthetic library is only as realistic as its corruption model. I focus on the most concrete and explicitly stated instance, Section 4.9, which removes ALMA calibration errors from the network-calibration step of synthetic data generation. This is not an external disagreement with EHT consensus; it is a limitation acknowledged in the manuscript itself, and it directly affects the central claim that the library supports calibrated inference on real data. A quantitative check is straightforward because the Synba pipeline and the companion network exist; the test I propose would settle whether the omission is benign or load-bearing. Other concerns, such as the lack of uncertainties on the Figure 1 detection-count comparison or the 'upon reasonable request' data availability, are real but secondary to the paper's primary deliverable of a training library. The paper is otherwise a careful, well-documented methods paper with a large parameter space and a first-principles forward model; the stated limitation does not invalidate it but does justify the reader's CONDITIONAL verdict. I therefore recommend no change to that verdict.","tokens_in":23452,"tokens_out":5646,"duration_ms":61175,"concrete_test":"Regenerate a validation subset (e.g., 10,000 datasets from standard Sgr A* and M87* models) with the same Symba/Rpicard pipeline but with ALMA gain errors drawn from the §2.2 model (1% relative scatter, 10% static offsets, plus optionally time-varying D-terms), keeping all other corruption parameters as in Table 2. Then take the trained network from Janssen et al. (2025a), or retrain a small network on the existing library, and evaluate predictive calibration on this ALMA-corrupted validation set versus a validation set without ALMA errors: report expected calibration error and coverage of 68%/95% credible intervals on spin, Rhigh, phi_mag, and inclination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the paper's central claim to hold, the synthetic library must be a faithful training distribution for Bayesian neural-network inference on real EHT data. Section 4.9 states: 'For the synthetic data, we assumed that ALMA calibration errors are negligible for the network calibration, that is, we used the uncorrupted model fluxes.' This is a direct limitation: real-data network calibration uses ALMA and SMA total-flux measurements to set the absolute gain scale, and §2.2 assigns those gains errors at the ~1% relative and ~10% static-offset level. Because ALMA is the most sensitive EHT station, its gain errors propagate into every calibrated visibility amplitude. The synthetic data, lacking this corruption, produce a training distribution whose amplitude scatter is narrower and centered on the true model flux; when the trained network is applied to real data with a realization of the ALMA gain error, point estimates will be biased and posterior intervals will undercover. The effect is comparable in size to physical model differences: Fig. A.1 shows that telescope gain errors cause amplitude losses of the same order as model features, and §6 identifies gain-induced polarization signals. Section 4.9 also ignores higher-order noise contributions (spillover, astronomical source), and §4.6 restricts D-terms to be constant over tracks and band. These omissions mean the library is not yet demonstrated to reproduce the real-data corruption statistics on which the downstream inference claim depends.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper (Series I) presents both an updated EHT calibration pathway and a large synthetic data library for Sgr A* and M87*. The calibration update combines parallel-hand correlation products and full bandwidth in fringe fitting and applies frequency-resolved system temperatures before fringe fitting, yielding more detections at intermediate S/N. The library contains 962,000 synthetic datasets built from Kerr, Kerr-Newman, and dilaton GRMHD-GRRT models, with forward modeling through Symba/MeqSilhouette/Rpicard that includes atmospheric turbulence, pointing errors, thermal noise, polarization leakage, scattering, and gain errors. The paper argues that the synthetic data match 2017 EHT baseline coverage and noise properties and can support Bayesian neural-network parameter inference in companion papers.","tokens_in":23718,"tokens_out":7613,"duration_ms":75881,"significance":"If its realism assumptions hold, the library is a valuable community resource: its scale and parameter coverage are unprecedented for EHT model comparison; the workflow is containerized and run on grid infrastructure; and the paper identifies concrete, falsifiable feature predictions (e.g., the shift of the total-intensity visibility minimum for Kerr-Newman charge in Figure 5). The library's utility, however, hinges on the corruption model matching real EHT systematics, and two of the paper's positive claims—'realistic synthetic data' and 'considerably better quality'—are stronger than the evidence presented.","major_comments":[{"comment":"The assumption that ALMA calibration errors are negligible for the network calibration is a load-bearing simplification for the library's realism. Section 2.2 assigns gain uncertainties of typically 1% relative plus static polarization-independent offsets at the ~10% level, and Section 4.9 states that the network-calibration technique uses ALMA and SMA total-flux measurements to set the absolute gain scale; as ALMA is the most sensitive station, its gain errors propagate into every calibrated visibility amplitude. The synthetic data, which 'used the uncorrupted model fluxes' for calibration, therefore have a narrower and incorrectly centered amplitude scatter relative to real data. A Bayesian neural network trained on this library can be expected to produce biased point estimates and undercovering posterior intervals when applied to real data unless the companion papers explicitly model this mismatch. The limitation is acknowledged, but the abstract's 'realistic synthetic data' claim and the library's fitness as a training distribution are not yet demonstrated under this assumption.","section":"Section 4.9"},{"comment":"The claim that the newly reduced EHT datasets have 'considerably better quality' (Abstract) rests on Figure 1, which compares cumulative detection counts between reductions without any uncertainties or statistical test. The differences at signal-to-noise around 5 could be within Poisson counting noise or systematic choices in the detection threshold; no error bars, bootstrap, or independent validation metric (e.g., scatter in closure quantities, gain stability, or astrometric consistency) is presented. The authors should either add uncertainty estimates and a significance statement for the detection-count difference or temper the abstract and Section 7 to say 'more detections at some S/N' rather than 'considerably better quality.'","section":"Section 2.3 / Figure 1"},{"comment":"Beyond the ALMA gain issue, the corruption model assumes D-terms constant over entire tracks and frequency bands (Section 4.6) and ignores higher-order noise contributions such as spillover and the astronomical source contribution (Section 4.9). These simplifications are stated, but their quantitative impact on the synthetic data is not assessed. Appendix A validates synthetic data against ground-truth model visibilities, not against the statistical properties of real EHT data; a comparison of, e.g., the distribution of residual gains, closure-phase scatter, or visibility-amplitude scatter between synthetic and real data would be needed to support the claim that the library encompasses the noise properties of EHT observations.","section":"Section 4.6 / 4.9 / Appendix A"}],"minor_comments":[{"comment":"The text says 'Each of the 14 models' but the preceding list contains 15 spin-charge combinations (2+3+3+3+3+1). If the intended number is 14, one entry is mislisted; if 15, the subsequent image count (16,632) should be 17,820 for 198 frames and six Rhigh values.","section":"Section 3.8.1"},{"comment":"The phrase 'amount of arimass toward the horizon' should read 'amount of airmass toward the horizon.'","section":"Section 4.9"},{"comment":"The sentence 'We have used the ... Symba Docker container to generate the synthetic date presented in this work' contains a typo: 'date' should be 'data.'","section":"Section 5"},{"comment":"The y-axis label 'Baseline detections - f(ξ)' with f(ξ)=280 log(ξ)−305 is difficult to interpret; the caption should explain why counts are plotted minus this arbitrary function, or the raw cumulative counts with uncertainties should be shown.","section":"Figure 1"},{"comment":"For a resource paper, providing permanent archival DOIs for the synthetic data (rather than 'access upon reasonable request') would better match the reproducibility emphasis of the workflow description.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"This is a foundational resource paper for the series, and my concerns are about calibration between its claims and its evidence rather than about its overall direction. The ALMA gain-error omission is openly disclosed, and the companion papers may handle it, but the abstract and conclusions currently overstate the realism and quality improvements. I would urge the authors to add uncertainty to Figure 1, quantify the impact of the Section 4.9 assumption, correct the Section 3.8.1 count discrepancy, and deposit the data in a public archive before final acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this is a well-executed data-release and calibration paper, and the central resource—the 962,000-dataset synthetic library—is real and carefully constructed. The calibration upgrades (combined-band, combined-polarization fringe fitting, pre-fringe Tsys) are incremental but sensible, and the forward-modeling chain (Symba, MeqSilhouette, Rpicard) is described in enough detail to reproduce. The paper earns credit for shipping containers, DOI'd intermediate products, and a full parameter table. The case studies produce at least one genuinely useful result: the closure-phase veto of a MAD a*=0.5 Sgr A* model that previously passed the stationary EHT constraints.\n\nSoft spots are real but mostly addressable. The headline claim of 'considerably better quality' rests on Figure 1, a cumulative detection-count diagram without uncertainties. That is a thin plank for the abstract's assertion. More importantly, Section 4.9 explicitly assumes ALMA calibration errors are negligible for network calibration and uses uncorrupted model fluxes. The stress-test note is correct: since ALMA is the most sensitive station, a network trained on synthetic data without ALMA gain errors will be partly miscalibrated on real data, and the effect is comparable to model differences. The paper discloses this and points to the companion inference papers as the real test, which is the right way to handle a known limitation, but it means the library's realism is not yet demonstrated for the downstream claim. The same section also lists higher-order noise contributions (spillover, source) as ignored. I don't think this is fatal; it is a condition that can be fixed by regenerating with ALMA gain corruptions or by marginalizing over them in training. The D-term constancy (Sect. 4.6) is a minor caveat.\n\nThe qualitative validation in Appendix A is adequate for a data-release paper, though a statistical comparison to real data statistics would have strengthened the realism argument. Data availability 'upon reasonable request' is a mild knock for a library meant to be foundational.\n\nMy verdict matches the reader's: conditional. This paper deserves serious refereeing. The math and modeling look solid, the circularity burden is low, and the writing is clear. Who reads it: anyone working on EHT parameter inference, simulation-to-observation pipelines, or ML training sets for VLBI. I would bring it to reading group and would cite it if I were in that space. Recommend a proper peer review with attention to the ALMA gain assumption and to the quantification of the Figure 1 claim.","headline":"A solid, carefully documented methods and data-release paper; the headline claims are mostly fair, with the main caveat that the synthetic library omits ALMA gain errors, which could bias downstream inference but is disclosed and addressable.","tokens_in":24338,"tokens_out":1927,"would_cite":true,"duration_ms":19714,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By emulating the full EHT signal path, 962,000 synthetic datasets make black hole parameter inference a tractable machine-learning problem.","keywords":["synthetic data library","Event Horizon Telescope","very long baseline interferometry","GRMHD simulations","general relativistic ray tracing","black hole parameter inference","data calibration","Bayesian neural network"],"falsifier":"Measure ALMA's absolute gain errors on the 2017 EHT tracks by comparing ALMA's measured flux of Sgr A* against independent total-flux monitoring at the same epoch; if the errors exceed the few-percent level assumed here, a network trained on this library, which used uncorrupted model fluxes for network calibration, will be miscalibrated in amplitude on real data.","tokens_in":23245,"feed_emoji":"🕳️","tokens_out":6521,"duration_ms":64323,"temperature":0.7,"pith_summary":"This paper builds the training foundation for using deep learning to turn Event Horizon Telescope (EHT) observations into measurements of black hole properties. It claims to have produced a library of 962,000 synthetic datasets for Sgr A* and M87*, each generated by taking simulated black hole images and pushing them through a detailed emulation of the EHT signal path, including atmospheric turbulence, telescope gain errors, polarization leakage, and the calibration process itself. It also claims that upgrades to the calibration pipeline yield real EHT data of considerably better quality than previous reductions, with more fringe detections across the array. If these claims hold, a neural network trained on the library can recover parameters such as black hole spin, magnetic flux state, electron temperature ratio, and inclination directly from real EHT visibilities, with intrinsic model variability rather than data quality becoming the main obstacle.","feed_headline":"962,000 simulated black hole datasets ready for AI","feed_subtitle":"Mock EHT observations trained from the full signal path let networks recover spin and accretion state.","key_machinery":"The load-bearing mechanism is the end-to-end forward-modeling chain. GRMHD simulations of the accretion flow are ray-traced in all Stokes parameters, then the forward-modeling pipeline, calibrated to per-antenna parameters, injects interstellar scattering, antenna pointing errors, tropospheric phase turbulence with a Kolmogorov power law, thermal noise, D-term polarization leakage, and static gain errors; finally the same calibration pipeline used for real data (which gains sensitivity by combining frequency bands and polarization channels) is applied to the corrupted visibilities. The result is that synthetic and real data share the same corruption and calibration statistics, so the synthetic visibilities can be compared directly with observed ones.","core_discovery":"The central claim is that a forward-modeled synthetic data library can be made realistic enough to serve as training data for machine-learning-based inference from EHT observations. From a broad parameter space of GRMHD-GRRT models, including Kerr, Kerr-Newman, and dilaton spacetimes, the authors generated 962,000 synthetic visibility datasets that match the baseline coverage and noise properties of the 2017 EHT observations of Sgr A* and M87*, as well as future arrays. The key validation is at the level of data products: closure phases, which are robust to calibration errors, preserve ground-truth model differences, while polarization amplitudes are dominated by simulated corruption effects such as gain errors and D-terms. The paper further claims that the updated calibration, which combines all polarization channels over the full bandwidth before fringe fitting, improves fringe sensitivity by about 10 percent and recovers detections that previous reductions missed, so the real data products are also cleaner.","pith_inferences":["The same library could be repurposed as a benchmark for VLBI image reconstruction, since every synthetic dataset has a known ground-truth movie that a reconstruction can be compared against.","The forward-modeling recipe transfers to other millimeter-VLBI targets; the same pipeline could generate libraries for a future global array without re-deriving the corruption model.","A cheap test of the calibration improvements: inject known gain errors into real data and verify that the new pipeline's closure phases are unchanged while amplitude-based products shift as expected.","The paper's case studies imply a quantitative prediction: over multi-year monitoring, SANE and MAD accretion states should separate in closure-phase variability statistics, even where single triangles lack discriminative power."],"forward_implications":["A Bayesian neural network trained on the library should recover GRMHD-GRRT parameters (spin, magnetic flux state, electron temperature ratio, inclination) from real EHT observations, as demonstrated in the follow-up papers.","Corruption-insensitive products are identified: closure phases and total-intensity visibility minima are reliable discriminators, while polarization amplitudes should be downweighted in inference.","Intrinsic model variability, not thermal noise, sets the ultimate limit on single-epoch parameter inference, making long-term monitoring of M87* and Sgr A* a scientific requirement.","Planned array extensions such as the Africa Millimeter Telescope or the next-generation EHT will tighten parameter constraints, and the library already contains datasets with those configurations.","The upgraded calibration pipeline, with its higher fringe-detection counts at a given signal-to-noise ratio, becomes the new reference reduction for EHT observations."],"supporting_citations":[{"why":"Supplies the Symba forward-modeling pipeline that turns GRRT images into synthetic visibilities.","marker":"Roelofs et al. 2020"},{"why":"Provides MeqSilhouette v2 with Cholesky-factorized tropospheric phase turbulence and corruption effects used in the forward modeling.","marker":"Natarajan et al. 2022"},{"why":"Describes the Rpicard calibration pipeline applied both to real EHT data and to the synthetic data.","marker":"Janssen et al. 2019b"},{"why":"Provides the ipole ray-tracing code used to produce full-Stokes GRRT images from the GRMHD simulations.","marker":"Mościbrodzka & Gammie 2018"},{"why":"Defines the standard kHARMA/patoka model library that supplies the fiducial Kerr GRMHD simulations.","marker":"Wong et al. 2022"},{"why":"Establishes the 2017 M87* calibration methodology that this paper updates and emulates.","marker":"Event Horizon Telescope Collaboration et al. 2019c"},{"why":"Introduces the network-calibration technique used to constrain co-located telescope gains from total flux measurements.","marker":"Blackburn et al. 2019"},{"why":"Provides the interstellar scattering screen parameters used to corrupt Sgr A* synthetic images.","marker":"Johnson et al. 2018"}],"fun_headline_variants":["Black hole AI trained on 962K synthetic datasets","962K synthetic black hole datasets to train AI","Forward-modeled EHT data: 962,000 datasets for AI","Improved calibration and 962K synthetic EHT datasets for AI","Simulated EHT library: 962K realistic mock observations for AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire enterprise rests on the assumption that the simulation of telescope corruption, especially the treatment of ALMA's calibration errors as negligible, matches how the real EHT actually corrupts the signal.","fun_headline_variants_meta":{"raw":{"variants":["Black hole AI trained on 962K synthetic datasets","962K synthetic black hole datasets to train AI","Forward-modeled EHT data: 962,000 datasets for AI","Improved calibration and 962K synthetic EHT datasets for AI","Simulated EHT library: 962K realistic mock observations for AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001928,"raw_usage":{"total_tokens":7588,"prompt_tokens":1028,"completion_tokens":6560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":6475}},"tokens_in":644,"tokens_out":6560,"duration_ms":47320,"temperature":1.0,"reasoning_tokens":6475,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:26:40.867007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure ALMA's absolute gain errors on the 2017 EHT tracks by comparing ALMA's measured flux of Sgr A* against independent total-flux monitoring at the same epoch; if the errors exceed the few-percent level assumed here, a network trained on this library, which used uncorrupted model fluxes for network calibration, will be miscalibrated in amplitude on real data.","supporting_citations":[{"cited_title":"2020, A&A, 636, A5","cited_arxiv_id":null,"evidence_quote":"Supplies the Symba forward-modeling pipeline that turns GRRT images into synthetic visibilities."},{"cited_title":"2022, MNRAS, 512, 490","cited_arxiv_id":null,"evidence_quote":"Provides MeqSilhouette v2 with Cholesky-factorized tropospheric phase turbulence and corruption effects used in the forward modeling."}],"review_version":1}