{"id":"2ca92c70-ca7d-4dea-9a60-58c46ec2ef28","arxiv_id":"2507.03613","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The 52 IceCube-associated blazar candidates are statistically compatible with the general blazar population, with a mild excess of radiatively efficient, radio-bright (HERG-like) objects.","lead":"Astronomers measured the optical spectra of 52 blazars that may have emitted IceCube neutrinos to see whether they differ from ordinary blazars. They find the candidates are mostly ordinary, but slightly more likely to have hot accretion disks and powerful radio jets, which gives a clue about how neutrinos are made.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed HERG/radio-bright preference lacks a parent-catalog control: P1.4 GHz is Peto-logrank compatible with 5BZCat (p=0.233), so the 63%/61% fractions may simply reflect the parent catalog.","rationale":"The reader's weakest assumption focused on contamination of the 52-source sample by spurious neutrino associations. My concern is related but sharper and partly independent: even if every association were genuine, the paper's 'preference' conclusion would still lack a valid baseline, because the only reported comparison with the parent 5BZCat catalog shows compatibility (Peto p=0.233 for P1.4 GHz), and the significant differences are against S12, a sample with known redshift/radio selection bias. Thus the central 'HERG-like tendency' is vulnerable on two fronts: sample purity and the absence of an appropriate control. The paper is otherwise careful and transparent—it presents new optical spectra, acknowledges censoring issues, and uses survival-analysis checks—and the broad 'overall compatible' statement is well supported. However, the more specific claim of a 'preference' for HERG-like, radio-bright objects does not currently clear the control-sample bar. Since the reader's verdict is already CONDITIONAL, my analysis does not change the verdict, but it identifies the missing check that should be required: a matched parent-catalog comparison. I am not raising a formal inconsistency or questioning the data reduction; the concern is about the interpretative step from absolute threshold fractions to a population preference.","tokens_in":44746,"tokens_out":5543,"duration_ms":68160,"concrete_test":"Compute the fraction of the full 5BZCat, and of a control sample matched in redshift, Fermi-LAT detection status, and radio flux, lying above P1.4 GHz = 1e26 W/Hz and LBLR/LEdd = 5e-4 using the same line-luminosity recipe as Section 5.2; if these fractions are within Poisson error of the candidate sample's 63% and 61%, the 'preference' claim reduces to 'compatible with 5BZCat'. A sharper version: re-run the Peto logrank test of P1.4 GHz for the candidate sample against a redshift-matched 5BZCat subsample; if p > 0.05, the claimed HERG-like radio preference is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference—that candidate neutrino-emitter blazars 'prefer' radiatively efficient accretion and powerful radio jets—is based on absolute fractions relative to fixed thresholds: ~61% above LBLR/LEdd = 5e-4 and ~63% above P1.4 GHz = 1e26 W/Hz. No control sample drawn from the 5BZCat parent population is used to define the expected fraction. The only direct parent-catalog comparison reported in Table A.3 (P1.4 GHz: candidate sample vs. BZCat) gives a Peto logrank p-value of 0.233, i.e., no evidence that the candidate sample differs from the very catalog from which it was selected. The nominally significant comparisons (P1.4 GHz and redshift vs. S12) are against a sample whose selection is explicitly skewed toward lower redshifts and lower radio powers (Section 3.2), so they cannot establish a physical 'preference' over the parent population. Moreover, the paper itself states that a non-negligible fraction of the 52 associations is expected to be spurious (Sections 1 and 3.1); under that expectation, the candidate sample should statistically resemble 5BZCat, which is exactly what the Peto logrank test finds for the radio power. The LBLR/LEdd preference is likewise not benchmarked against 5BZCat, so the claimed HERG-like tendency is not yet separated from parent-catalog selection or chance-association contamination.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes the physical properties of 52 blazars previously proposed as candidate IceCube neutrino emitters (Buson et al. 2022a, 2023) using optical spectroscopy (much of it newly acquired), radio luminosities, and gamma-ray luminosities. The authors derive BLR luminosities, black hole masses, Eddington ratios, and jet-power proxies, then compare the sample with reference blazar samples (Sbarrato et al. 2012; Paliya et al. 2017, 2021; 5BZCat; 4LAC-DR3) using Anderson-Darling and Peto logrank tests, including simulations of censored-data effects. They conclude that the candidate sample is overall compatible with reference blazar populations but shows a mild tendency toward HERG-like, radiatively efficient accretion and high radio power, and they report three comparisons at ≳3σ (black hole mass vs. P17; redshift and radio power vs. S12).","tokens_in":45061,"tokens_out":6359,"duration_ms":78379,"significance":"If the central claim were robust, this would be a useful first physical characterization of candidate neutrino-emitting blazars and would support theoretical models favoring neutrino production in radiation-field-rich environments. The paper's strengths are the careful optical spectroscopic analysis, the public presentation of new spectra, and the explicit simulation-based warning about the fragility of censored-data tests. However, the 'preference for HERG-like objects' claim is not benchmarked against the parent 5BZCat catalog, and the one direct parent-catalog comparison (P1.4 GHz vs. BZCat, Table A.3) shows no difference. The paper's own acknowledgement that a non-negligible fraction of the 52 associations may be spurious (Sections 1 and 3.1) makes this missing control especially important. The MBH result also relies on a p-value correction in Appendix D that appears internally inconsistent. The central inference therefore needs additional work before it can be considered established.","major_comments":[{"comment":"The central inference that the candidates 'prefer' HERG-like objects and powerful radio jets is based on absolute fractions above fixed thresholds (~63% above P1.4 GHz = 1e26 W/Hz and ~61% above LBLR/LEdd = 5e-4), with no control sample drawn from the parent 5BZCat catalog defining the expected fractions. The only direct parent-catalog comparison reported, P1.4 GHz vs. BZCat, gives a Peto logrank p-value of 0.233 (Table A.3), i.e., no evidence that the candidate sample differs from the very catalog from which it was selected. Given the paper's own statement that a non-negligible number of the 52 associations are expected to be spurious (Sections 1 and 3.1), the observed fractions should be compared against the 5BZCat expectation before any 'preference' claim is made.","section":"Section 6.2, Table A.3"},{"comment":"The significant differences found for redshift and P1.4 GHz relative to the S12 sample cannot support a physical HERG-like preference, because the S12 sample is explicitly selected toward lower redshifts and lower radio powers (Section 3.2). A statistically significant difference against such a biased comparison sample only shows that the candidate sample is not like S12; it does not show that the sample is unusual relative to the blazar parent population. The authors should either construct a radio/redshift-matched control from 5BZCat or substantially soften the corresponding conclusion in Section 7.","section":"Section 6.2, Table A.3"},{"comment":"The upper-limit correction used to retain the MBH vs. P17 result as ≳3σ is internally inconsistent. The text states that simulations with a censoring fraction closest to the observed case (38% ULs) 'conservatively result in a p-value of > 10^-2', and then divides the observed p = 1.70e-6 by 0.01 to obtain 1.70e-4. If censoring alone can produce null p-values above 10^-2, the conservative censoring-corrected p-value is at least of order 10^-2, not 1.7e-4. Under that reading the MBH comparison would not survive the Benjamini-Hochberg critical value at rank 3 (~2.8e-4), and the MBH ≳3σ claim should be removed or re-derived with a properly calibrated correction.","section":"Section 6.2, Appendix D"},{"comment":"The Benjamini-Hochberg implementation uses an FDR level of Q = 0.003 for m = 32 tests. This is a nonstandard choice that effectively enforces a 3σ threshold for all discoveries, and the paper does not justify why this particular Q was selected. Because the 'post-trial' significance claims in Section 6.2 depend on this choice, the authors should either justify Q = 0.003 on a priori grounds or use a conventional FDR level and report the corresponding conclusions.","section":"Section 6.1"}],"minor_comments":[{"comment":"There is a typo in the introduction: 'hereafer Paper II' should read 'hereafter Paper II'.","section":"Section 1"},{"comment":"The quoted fractions (~63% for P1.4 GHz and ~61% for LBLR/LEdd) are presented without explicitly stating whether upper limits are treated as detections or excluded; the later statement that ~87% of objects excluding upper limits occupy the HERG region shows that the definition matters and should be stated in the text.","section":"Section 6.2"},{"comment":"The simulation results are only summarized as 'p > 10^-2' in the text; reporting the actual p-value floors for the specific sample sizes and censoring patterns (e.g., for the MBH comparison with 32 measurements and 20 upper limits) would make the robustness assessment more quantitative and reproducible.","section":"Appendix D, Figs. D.1-D.2"},{"comment":"The table is dense and several redshift entries rely on footnotes that are not all self-explanatory; a column with the redshift reference or an explicit 'this work' flag for newly measured values would improve readability.","section":"Table A.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is built on the authors' own previously published candidate list, and the paper itself acknowledges that a non-negligible fraction of the 52 associations may be spurious. In this situation, a parent-catalog control analysis is not a refinement but a necessary condition for the central physical claim. The current comparison against S12, which is selected in a way that biases toward lower redshift and radio power, cannot substitute for it. If the authors cannot provide a 5BZCat-based control, the conclusions should be explicitly limited to 'the sample is broadly compatible with the general blazar population,' with the HERG/radio-bright preference removed or heavily qualified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful observational paper, not a discovery paper. The genuinely new bits are the spectra and the population characterization: twelve new observations across VLT/Gemini/GTC, including a new z=0.57 for 5BZB J0035+1515 that moves it from blue-FSRQ candidate to LERG, plus the first systematic comparison of the 52 Buson candidates with S12/P17/P21. The spectral-line work is careful, the simulations of logrank/Peto behavior are a real service, and the authors are admirably explicit that the 52 include spurious associations and that their tests are sensitive to censoring.\n\nThe soft spot is the central \"mild preference for HERG-like objects.\" As stated, ~61% above LBLR/LEdd and ~63% above P1.4GHz are absolute fractions, not fractions relative to what 5BZCat would give. The one direct parent-catalog comparison in Table A.3—radio power vs. BZCat—has Peto p=0.233, i.e., no detectable difference from the catalog the sample was drawn from. The significant S12 comparison is against a sample the authors themselves say is skewed to low z and low radio power. So if there is a HERG-like excess over the parent population, this paper has not established it. Given the expected contamination, the null result against BZCat is pretty much what you would expect. The LBLR/LEdd \"preference\" is not benchmarked against BZCat at all. The authors mostly use hedged language, but the conclusion section still says they \"are more HERG-like in terms of jet power\"—that overstates what the data show.\n\nThe other issues are minor: uniform gamma-ray upper limits at a nominal 10-yr sensitivity, no propagation of factor 2-4 systematics into the survival tests, and the p/0.01 correction. The correction is ad hoc but conservative and documented; I would rather see a simulation-calibrated version, but it is not a red flag. The MBH-vs-P17 discrepancy may be real, but it is not the paper's main point.\n\nWho is this for? People working on blazar-neutrino associations and on censored-data methods in AGN samples. The new spectra and the careful comparison table justify publication. A referee should push for a proper 5BZCat control (or at least an explicit fraction comparison) and softer wording. I would send it to review; the observational core is solid enough to deserve referee time.","headline":"Useful new spectra and a careful censored-data study; the HERG-preference claim needs a parent-catalog control before it can carry weight.","tokens_in":45648,"tokens_out":2757,"would_cite":true,"duration_ms":34050,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The 52 candidate neutrino-emitter blazars show a mild but consistent preference for radiatively efficient accretion and powerful radio jets, tying neutrino production to strong external radiation fields.","keywords":["neutrino astrophysics","blazars","IceCube","HERG/LERG classification","accretion modes","optical spectroscopy","radio jets","multimessenger astronomy"],"falsifier":"Take a control sample of 52 blazars drawn randomly from 5BZCat with the same redshift and Fermi-detection distribution as the candidates, measure the same LBLR/LEdd and P1.4GHz fractions, and repeat the Peto logrank comparisons; if the control shows the same roughly 61-63% HERG-like and radio-loud fractions, the paper's population-level conclusion would not stand. Alternatively, when individual IceCube associations are confirmed through real-time alerts, the confirmed subset's HERG fraction should remain elevated; if it drops to the parent population level, the claimed preference is an artifact of the candidate selection.","tokens_in":44547,"feed_emoji":"🧊","tokens_out":5149,"duration_ms":54015,"temperature":0.7,"pith_summary":"Blazars are prime suspects for the sources of IceCube's astrophysical neutrinos, but which ones actually emit neutrinos, and why, remains unsettled. This paper takes the 52 blazars that a previous positional cross-correlation matched to IceCube hotspots and gives them a first multiwavelength physical characterization, using optical spectra to measure their accretion regime and radio and gamma-ray data to gauge jet power. The central result is that, while the sample remains statistically compatible with reference blazar populations, it shows a mild preference for sources with intense radiation fields, radiatively efficient accretion, and radio powers above the HERG-LERG divide: roughly 61% sit in the high-excitation (HERG) region of LBLR/LEdd, and 63% have radio power above about $10^{26}$ W/Hz. If genuine, this would tie neutrino production to environments rich in external photon fields, matching proton-photon interaction scenarios, and it would argue against treating historical BL Lac/FSRQ classifications as physically meaningful in neutrino stacking studies. The paper also demonstrates that survival-analysis tests commonly used with censored data can produce spurious significances as the fraction of upper limits grows, a caution that applies to much of the literature.","feed_headline":"Blazar neutrino candidates skew toward radio-loud, efficient accretors","feed_subtitle":"A first optical-radio-gamma census of 52 IceCube-associated blazars finds a mild HERG-like preference, with limits on spurious matches.","key_machinery":"The load-bearing object is the physically driven HERG/LERG taxonomy applied to blazars, in which the Eddington-scaled broad-line-region luminosity LBLR/LEdd (threshold near 5e-4 for radiatively efficient accretion), the Eddington-scaled gamma-ray luminosity Lgamma/LEdd (threshold near 0.1), and the 1.4 GHz radio power P1.4GHz (threshold near $10^{26}$ W/Hz) replace the traditional equivalent-width split between BL Lacs and FSRQs. The machinery that carries the argument is optical spectroscopy: the authors measure or place upper limits on broad-line region lines (H-$\\alpha$, H-$\\beta$, Mg II, C IV), convert line luminosities to LBLR through the relative line ratios of a composite spectrum, estimate virial black hole masses and Eddington luminosities, and derive disk luminosities assuming a 10% BLR covering factor. On the statistical side, the paper employs the Peto logrank test with Kaplan-Meier estimators to handle censored data, after simulations showing that the test's false-positive rate grows with the fraction of upper limits and with sample size. That machinery lets the authors compare their 52 candidates with three literature samples and isolate which apparent differences are reliable.","core_discovery":"The paper claims that the 52 candidate neutrino-emitter ('PeVatron') blazars selected from IceCube hotspots are, as a population, mildly shifted toward high-excitation radio galaxy (HERG) characteristics: radiatively efficient accretion and powerful radio jets. After re-estimating black hole masses, broad-line-region luminosities, Eddington ratios, and radio powers from optical spectra and catalog data, the authors find that about 61% of the sample sits above the LBLR/LEdd ~ 5e-4 accretion-efficiency threshold and about 63% above P1.4GHz ~ $10^{26}$ W/Hz, with median values of LBLR/LEdd = $10^{-3}$ and Lgamma/LEdd = 9.73e-2. Only three comparisons survive correction for multiple trials at the roughly 3-$\\sigma$ level (black hole mass versus one reference sample, and redshift and radio power versus another), so the authors do not claim a decisive separation from the general blazar population. Instead, they argue for a mild tendency: candidate neutrino emitters preferentially show efficient accretion, strong radiation fields, and high radio power, consistent with hadronic models in which protons interact with external photon fields. The paper also reclassifies four historically 'masquerading BL Lac' objects: three are HERG-like, while a newly analyzed GTC spectrum shows that 5BZB J0035+1515 is a LERG, undermining stacking studies that place such objects in the BL Lac subsample.","pith_inferences":["A testable extension: stack IceCube events on HERG-classified versus LERG-classified blazars using this physical taxonomy rather than BLL/FSRQ labels; the model predicts excess emission preferentially on the HERG-like subset.","If the HERG preference is real, 'changing-look' blazars should appear disproportionately among future neutrino candidates, since their apparent line variability masks a stable efficient-accretion engine.","The censoring-sensitivity result suggests that re-analysis of older flux-limit comparisons in AGN surveys, using mock samples with matched upper-limit fractions, would be a worthwhile methodological companion to any future catalog-level neutrino association study."],"forward_implications":["If the 52 candidates are genuinely neutrino-linked, their mild HERG-like shift implies that sources with radiatively efficient accretion and strong external radiation fields are favored neutrino factories, consistent with proton-photon interactions.","The reclassification of TXS 0506+056, PKS 1424+240, and 5BZB J0630-2406 as HERG-like means previous stacking limits on BL Lac neutrino flux contributions should be revisited, since these objects were counted in the wrong subsample.","The demonstration that Peto logrank p-values shrink as censoring fractions grow implies that published claims of differences between blazar subpopulations based on such tests need case-by-case simulation before being accepted.","The 24 Fermi-detected candidates span gamma-ray luminosities from about 10^42 to 10^48 erg/s, so gamma-ray brightness alone does not single out neutrino emitters; accretion and radio properties carry additional discriminating information."],"supporting_citations":[{"why":"Supplies the southern-hemisphere positional cross-correlation of IceCube hotspots with 5BZCat that defines part of the 52-candidate sample.","marker":"Buson et al. (2022a)"},{"why":"Supplies the northern-hemisphere cross-correlation analysis that completes the 52-candidate sample.","marker":"Buson et al. (2023)"},{"why":"Provides one of the three reference blazar samples and the methods for estimating LBLR, Lgamma, and Eddington ratios.","marker":"Sbarrato et al. (2012)"},{"why":"Provides a second reference sample of gamma-loud and gamma-quiet blazars with optical-spectroscopy-based physical properties.","marker":"Paliya et al. (2017)"},{"why":"Provides the large Fermi-detected reference sample used for the statistical comparisons.","marker":"Paliya et al. (2021)"},{"why":"Supplies the LBLR/LEdd and Lgamma/LEdd thresholds that define the HERG-like accretion regime.","marker":"Ghisellini et al. (2011b)"},{"why":"Supplies the radio-power threshold P1.4GHz ~ 10^26 W/Hz separating HERG- and LERG-dominated populations.","marker":"Best & Heckman (2012)"},{"why":"Provides the survival-analysis framework (Kaplan-Meier with logrank/Peto logrank tests) whose censoring sensitivity the paper investigates.","marker":"Feigelson & Nelson (1985)"},{"why":"Provides the theoretical expectation that intense external radiation fields enhance proton-photon neutrino production in blazar jets.","marker":"Dermer et al. (2014)"},{"why":"Provides the previous stacking limits on blazar contributions to the diffuse neutrino flux that the paper argues may be affected by misclassification.","marker":"Aartsen et al. (2017a)"}],"fun_headline_variants":["IceCube blazar candidates mildly favor efficient accretion","52 blazars linked to IceCube show mild accretion lean","Blazar neutrino emitters skew to radio-loud, efficient engines","Optical census of 52 blazars shows mild HERG preference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 52 candidate blazars, which the paper itself says include a non-negligible number of chance associations, still represent neutrino-emitting blazars as a population; if most of them are spurious matches, the observed HERG-like tendency merely reflects the parent 5BZCat catalog.","fun_headline_variants_meta":{"raw":{"variants":["IceCube blazar candidates mildly favor efficient accretion","52 blazars linked to IceCube show mild accretion lean","Blazar neutrino emitters skew to radio-loud, efficient engines","Optical census of 52 blazars shows mild HERG preference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3335,"prompt_tokens":1155,"completion_tokens":2180,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":771,"completion_tokens_details":{"reasoning_tokens":2109}},"tokens_in":771,"tokens_out":2180,"duration_ms":20796,"temperature":1.0,"reasoning_tokens":2109,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:06:08.244758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a control sample of 52 blazars drawn randomly from 5BZCat with the same redshift and Fermi-detection distribution as the candidates, measure the same LBLR/LEdd and P1.4GHz fractions, and repeat the Peto logrank comparisons; if the control shows the same roughly 61-63% HERG-like and radio-loud fractions, the paper's population-level conclusion would not stand. Alternatively, when individual IceCube associations are confirmed through real-time alerts, the confirmed subset's HERG fraction should remain elevated; if it drops to the parent population level, the claimed preference is an artifact of the candidate selection.","supporting_citations":[{"cited_title":"2012, , 421, 1764","cited_arxiv_id":null,"evidence_quote":"Provides one of the three reference blazar samples and the methods for estimating LBLR, Lgamma, and Eddington ratios."},{"cited_title":"S., Marcotulli , L., Ajello , M., et al","cited_arxiv_id":null,"evidence_quote":"Provides a second reference sample of gamma-loud and gamma-quiet blazars with optical-spectroscopy-based physical properties."},{"cited_title":"S., Dom \\' nguez , A., Ajello , M., Olmo-Garc \\' a , A., & Hartmann , D","cited_arxiv_id":null,"evidence_quote":"Provides the large Fermi-detected reference sample used for the statistical comparisons."}],"review_version":1}