{"id":"784678c5-d82d-4580-8f56-32fff765b7fd","arxiv_id":"2606.28603","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Machine learning on simulated images identifies that flux eruption events cause more diffuse, polarized, lower-flux millimeter emission with decreased Q-U loop rotation rate, achieving ~80% accuracy with random forests on summary statistics.","lead":"Machine learning models trained on simulated black hole images find that flux eruption events produce more diffuse emission, higher linear polarization, lower total flux, and slower Q-U loop rotation in millimeter observations. A smart generalist might read this to see how future Event Horizon Telescope data could flag magnetic reconnection events near black holes.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"CNN trained only on uncorrupted images may assign FEE labels using features erased by EHT sampling/noise, so the 80% RF accuracy on summary statistics may not indicate real-world detectability.","rationale":"The reader's weakest assumption (simulation fidelity and generalization) is the same load-bearing point; the uncorrupted CNN training supplies a concrete, testable mechanism by which that assumption could fail. This justifies moving from UNVERDICTED to CONDITIONAL pending the corruption test.","tokens_in":1884,"tokens_out":356,"duration_ms":27301,"concrete_test":"Take the same simulated image library, corrupt a held-out subset with realistic EHT effects (array response, thermal noise, interstellar scattering via eht-imaging or equivalent), re-apply the CNN labeling step, retrain the random forest on the same summary statistics, and measure the change in class-weighted accuracy; a drop below ~65% would indicate the signatures are not robust to actual observations.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The pipeline trains a CNN on ideal simulated images to learn FEE representations, then applies it to label a larger set before training the random forest and logistic regression on observable summary statistics. Because the initial CNN never sees sparse uv-coverage, thermal noise, or scattering, its learned features (and therefore the derived labels) could rely on simulation details that real EHT data lack. The abstract already states that the reported signatures are weak relative to ordinary variability; if the labels themselves are contaminated by non-observable structure, both the identified trends (diffuse emission, polarization fraction, Q-U rotation rate) and the 80% class-weighted accuracy lose their claimed connection to observable EHT signatures.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper trains a CNN on ideal (uncorrupted) simulated millimeter images of black hole accretion flows to label flux eruption events (FEEs), then fits interpretable models (random forest, logistic regression) to observable summary statistics derived from those labels. It reports that FEEs are associated with more diffuse emission, higher linear polarization, lower total flux, and slower Q-U loop rotation rates, though these trends are weak relative to ordinary variability; the random forest achieves ~80% class-weighted accuracy on the summary statistics, implying the CNN captures additional FEE structure.","tokens_in":2063,"tokens_out":564,"duration_ms":23822,"significance":"If the pipeline generalizes, the work supplies concrete, testable guidance on which image properties (size, polarization fraction) could flag candidate FEEs in EHT data and demonstrates that traditional summary statistics do not fully capture the CNN-learned representation. The empirical, simulation-driven approach and the explicit statement that signatures are weak compared with normal variability are strengths.","major_comments":[{"comment":"Abstract and Methods (pipeline description): The CNN is trained exclusively on uncorrupted simulated images. Because the subsequent labels are used both to identify the reported trends (diffuse emission, polarization, Q-U rate) and to train the random forest that reaches ~80% accuracy, any FEE-discriminating features that are erased by realistic EHT uv-coverage, thermal noise, or scattering would render both the signatures and the accuracy claim non-observable. A concrete test (e.g., re-training or evaluating the CNN on forward-modeled EHT images) is required to establish that the reported observational signatures survive the instrument response.","section":"Abstract / Methods"},{"comment":"Results (80% accuracy claim): No information is supplied on training/validation splits, hyperparameter selection, class-imbalance handling, or error bars/statistical significance tests for the random-forest performance. Without these details the headline accuracy figure cannot be evaluated and the claim that the CNN learns structure “not fully mapped onto these traditional summary statistics” remains unsupported.","section":"Results"}],"minor_comments":[{"comment":"Clarify the exact definition of the summary statistics fed to the random forest and logistic regression (e.g., how image size, polarization fraction, and Q-U rotation rate are computed from the images).","section":"Methods"},{"comment":"The abstract states the signatures are “weak for most FEEs”; quantify this statement with effect sizes or overlap metrics relative to the non-FEE variability distribution.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive and detailed report. The comments highlight important limitations in the current presentation of our methods and results. We respond point-by-point below, indicating where revisions will be made to strengthen the manuscript.","responses":[{"response":"We agree that training exclusively on ideal images means the reported signatures and accuracy are not yet demonstrated to be observable. The manuscript intentionally isolates intrinsic simulation features before instrumental effects, as stated in the abstract and methods. In revision we will add explicit language in the abstract, methods, and discussion clarifying that the signatures are derived from uncorrupted images and constitute potential rather than guaranteed observables. We will also include a qualitative assessment of how uv-coverage, noise, and scattering are expected to affect diffuse emission and polarization fraction. A full forward-modeling test lies outside the present computational scope but will be noted as future work. This revision makes the scope of the claims transparent without overstating current results.","revision_made":"partial","referee_comment":"[Abstract / Methods] Abstract and Methods (pipeline description): The CNN is trained exclusively on uncorrupted simulated images. Because the subsequent labels are used both to identify the reported trends (diffuse emission, polarization, Q-U rate) and to train the random forest that reaches ~80% accuracy, any FEE-discriminating features that are erased by realistic EHT uv-coverage, thermal noise, or scattering would render both the signatures and the accuracy claim non-observable. A concrete test (e.g., re-training or evaluating the CNN on forward-modeled EHT images) is required to establish that the reported observational signatures survive the instrument response."},{"response":"The omission of these implementation details was an oversight. In the revised manuscript we will insert a new subsection (likely in Results or an expanded Methods) that specifies: (i) the train/validation split and any temporal blocking used to avoid leakage, (ii) hyperparameter search procedure (grid or random search with cross-validation), (iii) class-imbalance treatment via class weights or resampling, and (iv) uncertainty quantification via bootstrap or k-fold estimates together with a statistical comparison (e.g., McNemar test or permutation test) against a null model. These additions will allow readers to evaluate the ~80% figure and the claim that the CNN captures structure beyond the summary statistics.","revision_made":"yes","referee_comment":"[Results] Results (80% accuracy claim): No information is supplied on training/validation splits, hyperparameter selection, class-imbalance handling, or error bars/statistical significance tests for the random-forest performance. Without these details the headline accuracy figure cannot be evaluated and the claim that the CNN learns structure “not fully mapped onto these traditional summary statistics” remains unsupported."}],"tokens_in":1531,"tokens_out":588,"duration_ms":28236,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The two things to know are that the work reports a drop in Q-U loop rotation rate during flux eruption events, which runs counter to some joint flare-loop pictures, and that a random forest on observables reaches roughly 80% class-weighted accuracy after the CNN supplies labels.\n\nWhat is new is the two-stage pipeline: CNN on ideal images to learn FEE structure, followed by random forest and logistic regression on flux, polarization fraction, image size, and loop rotation to pull out interpretable signatures. The abstract is straightforward that most of those signatures are weak compared with ordinary variability.\n\nThe paper does a reasonable job of making the CNN output more usable by mapping it back to standard observables and by flagging that high-dynamic-range images will still be required for confirmation.\n\nThe main soft spot is exactly the one in the stress-test note. The CNN never sees sparse uv-coverage, noise, or scattering, so its learned features and the labels it assigns could depend on simulation details that disappear in actual EHT data. If that happens, both the reported trends and the 80% accuracy lose their claimed connection to observable signatures. The abstract gives no training splits, hyperparameter details, or significance tests, which leaves the accuracy number hard to evaluate.\n\nThis is for the EHT polarization and accretion-flow modeling crowd. A reader already running MAD simulations or looking at time-variable linear polarization might pick up the Q-U rotation result to test.\n\nIt deserves a serious referee because the targeted application is timely and the method is a clear step beyond pure simulation work, even though the generalization question will need direct attention in revision.\n\nRecommendation: send it out, but ask reviewers to focus on whether the CNN labels remain stable once realistic observational effects are added.","headline":"The paper finds weaker Q-U loop rotation during FEEs and gets 80% RF accuracy on summary stats, but training the CNN only on clean simulations makes the real-world link to EHT data shaky.","tokens_in":2578,"tokens_out":426,"would_cite":false,"duration_ms":28403,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Machine learning identifies diffuse emission, higher polarization, and lower flux as signatures of flux eruption events in black hole accretion flows, though these are weak compared to normal variability.","keywords":["flux eruption events","black hole accretion flows","Event Horizon Telescope","machine learning","linear polarization","millimeter observations","Q-U loops","magnetic reconnection"],"falsifier":"Comparing the predicted changes in image diffuseness, polarization fraction, and Q-U rotation during observed flares with EHT data against the absence or presence of such signatures in the simulations.","tokens_in":2784,"feed_emoji":"","tokens_out":626,"duration_ms":41278,"temperature":0.7,"pith_summary":"The paper uses machine learning to identify observable signatures of flux eruption events in simulated millimeter images of supermassive black hole accretion flows. A convolutional neural network first learns to detect FEEs in clean simulations, after which interpretable models extract features like image diffuseness and polarization. These signatures could allow the Event Horizon Telescope to spot transient magnetic reconnection events near the event horizon. However, the changes are often subtle and may require high-resolution imaging to distinguish from ordinary variability.","feed_headline":"ML spots subtle marks of black hole flux eruptions","feed_subtitle":"Diffuse emission and higher polarization appear during these events in simulations, but often weaker than normal variability.","key_machinery":"A two-stage machine learning approach: a convolutional neural network trained on uncorrupted simulated images to learn FEE representations, followed by random forest and logistic regression models on summary statistics for interpretable signatures.","core_discovery":"During a flux eruption event, simulated images tend toward more diffuse emission, higher linear polarization, and lower total fluxes, with the Q-U loop rotation rate decreasing, contrary to a picture in which FEEs cause both loops and flares. A random forest trained on observable summary statistics achieves about 80% class-weighted accuracy, indicating that the CNN learns FEE structure not fully captured by these traditional statistics. The results imply that image size and polarization fraction can flag candidate FEEs, but high-resolution, high-dynamic range images remain important for confirmation.","pith_inferences":["If applicable to real data, EHT observations could search for magnetic reconnection in accretion flows using polarization and image properties.","The gap between CNN performance and summary statistics suggests deep learning may reveal new aspects of accretion dynamics.","Statistical stacking of multiple observations may be necessary to detect these weak signals amid variability.","This method could be applied to identify other transient phenomena in future higher-sensitivity black hole images."],"forward_implications":["Image size and polarization fraction can flag candidate FEEs.","High-resolution and high-dynamic range images are needed to confirm FEEs.","FEEs decrease the Q-U loop rotation rate and do not jointly cause both loops and flares.","Machine learning captures FEE features beyond traditional summary statistics.","These signatures are weak for most FEEs relative to usual time variability."],"fun_headline_variants":["ML spots FEEs in black hole accretion flows","Diffuse emission rises during flux eruption events","Higher polarization and lower flux signal FEEs","Q-U rotation slows in black hole FEE simulations"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The simulated accretion flows with strong magnetic fields accurately represent the physical conditions around real supermassive black holes that the Event Horizon Telescope can observe, and the machine learning models trained on these simulations generalize to actual observational data without being dominated by simulation-specific artifacts.","fun_headline_variants_meta":{"raw":{"variants":["ML spots FEEs in black hole accretion flows","Diffuse emission rises during flux eruption events","Higher polarization and lower flux signal FEEs","Q-U rotation slows in black hole FEE simulations"]},"model":"grok-4.3","cost_usd":0.005879,"raw_usage":{"total_tokens":2844,"prompt_tokens":770,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":58787000,"prompt_tokens_details":{"text_tokens":770,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2017,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":770,"tokens_out":57,"duration_ms":22515,"temperature":1.0,"reasoning_tokens":2017,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T00:41:12.906491+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Comparing the predicted changes in image diffuseness, polarization fraction, and Q-U rotation during observed flares with EHT data against the absence or presence of such signatures in the simulations.","supporting_citations":[],"review_version":1}