{"id":"40c23e20-e420-4a8e-9cc6-ecc80120a2d7","arxiv_id":"2512.06584","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"RM²–galaxy cross-correlations can isolate the magnetic field contribution from specific redshifts and may be detectable with SKA-era RM catalogs.","lead":"This paper proposes measuring the cross-correlation between the squared Faraday rotation measure of background radio sources and foreground galaxies to map cosmic magnetic fields. It uses Illustris-TNG simulations and forecasts that future SKA data could detect the signal tomographically across redshift.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Forecast amplitude is set by TNG300-3, the lowest-resolution run, and Appendix C shows the RM2-halo signal decreases with resolution; the claimed current-data and near-term detection significances are therefore not converged.","rationale":"The paper's central contribution is a new estimator, and it honestly documents its main limitation in Appendix C. But the strongest claim — that the statistic is detectable now and at high significance with SKA — is a quantitative statement controlled by the amplitude of wRM2,g, and that amplitude comes from the lowest-resolution simulation in the suite. The resolution test in Figure 11 shows systematic suppression at higher resolution, so the reported SNRs are not anchored to a converged prediction. This is exactly the load-bearing weakness identified by the reader. I considered whether the unproven replacement PB→P_tildeB in Eq. 3.2 is more fundamental; however, the simulation signal itself, not the analytic approximation, sets the forecast amplitude on the scales that dominate the SNR, and the analytic extrapolation enters mainly at large separations. The concrete test would settle the issue by replacing the low-res signal with the high-res signal and recomputing the SNRs. If the high-res signal remains large enough, the concern is non-fatal; if not, the detectability claim needs to be softened to a conditional one. Thus the reader's CONDITIONAL verdict is appropriate and no adjustment is needed.","tokens_in":28601,"tokens_out":11498,"duration_ms":124637,"concrete_test":"Recompute the coeval RM2–halo cross-correlation at z≈0.1 and z≈0.5 using the public TNG300-1 and TNG300-2 snapshots with the same 0.41 Mpc/h gridding and stacking pipeline used for TNG300-3. Then rerun the §4.2 forecast for the NRM=1e3, 1e5, and 1e7 cases with the TNG300-1 signal in place of TNG300-3, keeping σtot=6.6 rad/m^2 and Ngal=1e6. If the z=0.1 SNR for the current-sample case falls below ~3 (or the 1e5 case below ~5), the paper's near-term detectability claim is overestimated; also refit K from the high-resolution runs and check whether the resulting forecasts change by more than ~30%.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative claim that ⟨RM2×g⟩ is detectable with current surveys and at high significance with SKA (§4, Fig. 9) uses the TNG300-3 signal as its amplitude. TNG300-3 has 64× fewer resolution elements than TNG300-1, and Appendix C (Fig. 11) shows that wRM2,g at z=0 decreases by roughly a factor of 2–4 from TNG300-3 to TNG300-1, attributed to better-resolved magnetic-field reversals canceling along the line of sight. The paper's own z=0.1 forecast SNRs — 36.5 for 1e5 RMs, 3.65 for 1e3 RMs, and 3.9 for the NVSS-like sample (§4.3) — scale linearly with signal amplitude. A factor ~3 lower high-resolution amplitude would move the current-sample SNR below ~1.3 and reduce the 1e5-RM forecast to ~12, weakening the 'within striking distance' claim. The analytic model (Eq. 3.3) used to extrapolate the forecasts to large scales inherits the same bias: the calibration factor K=3 is fitted to TNG300-3 and Appendix C notes it is held fixed, so the resolution bias propagates directly into the forecast. The estimator concept is not invalidated, but the central detectability significance is not robust until the amplitude is shown to converge.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new estimator for extragalactic magnetic fields: the cross-correlation between the squared Faraday rotation measure, RM^2, toward background sources and the projected galaxy overdensity, ⟨RM^2 × g⟩. The authors derive its relation to a bispectrum involving two electron-density–weighted line-of-sight magnetic field factors and one galaxy overdensity (Eq. 2.6), work in the flat-sky/Limber approximations, and develop an approximate analytic model in which the scale dependence follows the electron–galaxy cross-power and the amplitude is set by an integral over the electron-density–weighted magnetic field power spectrum (Eq. 3.3). They calibrate a multiplicative factor K=3 against Illustris-TNG TNG300-3 simulations, check the smoothing-scale dependence, quantify the noise bias of ⟨|RM| × g⟩, and make forecasts for current and SKA-era RM catalogs. The central claimed result is that ⟨RM^2 × g⟩ avoids the sign cancellation and noise bias of |RM|-based estimators and can tomographically probe the redshift evolution of cosmic magnetic fields.","tokens_in":28975,"tokens_out":3278,"duration_ms":36585,"significance":"If the simulation-calibrated amplitude is correct, the proposed estimator is a genuinely new and potentially powerful probe: it is immune to zero-mean RM noise bias in the average signal, it has a cleaner connection to the underlying bispectrum than |RM| correlators, and it offers a tomographic redshift decomposition by binning foreground galaxies. The derivation in Appendix A is a useful contribution, and the comparison with TNG simulations is a reasonable first modeling step. However, the quantitative forecasting rests on an analytic model with an explicit heuristic replacement of P_B by P_tildeB and a fitted factor K, and the calibration is performed on the lowest-resolution TNG300 run. Since Appendix C shows the signal decreases substantially with resolution, the predicted detection significances are not yet robust. The conceptual method is valuable, but the numerical forecasts need to be either re-derived at converged resolution or presented with a much larger systematic uncertainty.","major_comments":[{"comment":"The forecast amplitudes are not converged with simulation resolution. The forecasts use TNG300-3, which has 64× fewer resolution elements than TNG300-1. Appendix C, Fig. 11, shows that w_RM2,g at z=0 decreases by roughly a factor of 2–4 as resolution increases, attributed to better-resolved magnetic field reversals cancelling along the line of sight. Since all SNR estimates in §4.2 scale linearly with signal amplitude, the headline numbers (SNR 36.5 for 10^5 RMs, 3.9 for the NVSS-like sample at z=0.1) would drop to roughly 9–18 and 1–2 if the TNG300-1 amplitude is closer to the true signal. The statement in Appendix C that the main conclusions are robust is therefore contradicted by the figure it accompanies. The authors should either rerun the forecasts with TNG300-1-calibrated amplitudes, or explicitly marginalize over the resolution uncertainty and soften the detectability claims.","section":"§4.2, Fig. 9, and Appendix C"},{"comment":"The analytic model used to interpret the simulations and to extrapolate forecasts to scales beyond the simulation box is built on an unproven replacement of P_B by P_tildeB and a calibration factor K=3 fitted to the same TNG300-3 simulations. The paper itself describes this as a “heuristic step which we do not attempt to rigorously justify.” Since Eq. (3.3) is then used in §4 to extend the correlation function to r⊥ > 20 Mpc/h and to predict SNRs, the systematic error in K and in the P_B→P_tildeB replacement propagates directly into all detectability claims. A comparison against TNG300-1 or another MHD simulation is needed to establish that K is not resolution-dependent; alternatively, the forecasts should be presented as conditional on the TNG300-3 calibration, with the calibration uncertainty included in the error budget.","section":"§3.2, Eq. (3.2)–(3.3), and Appendix A.3"},{"comment":"The forecast covariance assumes that both the RM field and the galaxy field are noise-dominated and that the error bars in different radial bins are uncorrelated. The paper acknowledges the neglect of sample variance, but the resulting error formula, Eq. (4.2), is used to claim high-significance detections. Given that the signal itself is only calibrated to a single simulation at a single resolution, the forecast SNR values should be interpreted as upper limits modulo these simplifications. At minimum, the authors should state explicitly how much of the claimed SNR comes from the assumed noise-dominated covariance and how much could be affected by sample variance on the scales where the signal is actually measured.","section":"§4.2 and Appendix B"}],"minor_comments":[{"comment":"Typo: “notoation” should be “notation” in the sentence introducing Eq. (A.3).","section":"Appendix A.1"},{"comment":"The text says “Our general aim is to calculate…” and later “in our forecasts (§4.3)”; the forecasts of the RM^2 statistic are in §4.2, while §4.3 concerns |RM|. Please correct the cross-reference.","section":"Appendix B"},{"comment":"The caption states that “the main conclusions are robust to simulation resolution,” but the figure itself shows a factor–2–4 decrease in amplitude. Either soften the claim or provide a quantitative convergence test, e.g., on the total SNR under TNG300-1.","section":"Fig. 11 caption"},{"comment":"The discussion of smoothing is helpful, but the comparison between the 3D mesh smoothing and the angular smoothing of real RM catalogs is only qualitative. Since the signal depends so strongly on the smoothing scale, a more explicit treatment of the mapping between 3D voxel size and effective angular beam would improve the paper's readiness for data application.","section":"§3.4"}],"recommendation":"major_revision","confidential_remarks":"The core idea—RM^2–galaxy cross-correlation as a noise-bias-free, tomographic probe—is worth publishing if the quantitative forecasts are made credible. The main blocker is that the calibration and forecast use TNG300-3, and the paper's own resolution test shows the signal is not converged. This is not a fatal flaw in the estimator, but it is a load-bearing issue for the central detectability claims. I would require the authors to either redo the forecasts with the highest-resolution run or explicitly fold the resolution systematic into the SNR, and to temper the abstract and conclusion accordingly. The heuristic P_B→P_tildeB replacement is also an admitted gap; a comparison of K across resolutions would at least establish its empirical stability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zhang & Lidz propose something genuinely new: cross-correlating RM² with galaxy overdensities to get a tomographic look at cosmic magnetic fields. The estimator is a real departure from the |RM| approach, it avoids the noise-bias problem cleanly, and the flat-sky/Limber derivation of the bispectrum relation in §2.2/Appendix A is compact and honest. The analogy to the kSZ projected-fields estimator is apt and they don't oversell it.\n\nCredit where it's due: they also do the right thing by reporting the resolution test. Appendix C shows the RM²-halo amplitude drops by roughly a factor of a few from TNG300-3 to TNG300-1. That is load-bearing for the forecasts, since the SNRs in §4 scale linearly with that amplitude. The current-sample SNR of ~3.9 and the SKA-era forecasts are therefore not converged until the amplitude is shown to stabilize with resolution. The text says the conclusions are robust to resolution, but the figure points the other way for the specific detectability claims.\n\nThe analytic model has a second soft spot: K=3 is fitted to the simulations, and the P_B → P_tildeB replacement is admitted to be a heuristic. Because K is calibrated on the low-res run, the forecast inherits the same resolution bias. That said, these are caveats rather than fatal flaws. The statistic is well-defined, the noise-bias immunity is exact, and the qualitative redshift evolution (strong rise toward low z) is probably robust. What is soft is the quantitative amplitude and the derived significance numbers.\n\nWho is this for? People planning RM cross-correlation analyses, cosmic magnetism modelers, and anyone comparing forecasting methods. It deserves a serious referee, but with a clear request: validate the amplitude with higher-resolution runs or multiple simulations, propagate resolution uncertainty into the forecasts, and tune down the 'within striking distance' phrasing until that is done.\n\nMy recommendation: send to review, conditional. Not a desk reject, but the revision needs to confront the resolution dependence directly.","headline":"The RM²-galaxy statistic is genuinely new and the derivation is clean, but the detectability forecasts lean on a low-resolution simulation run and a fitted fudge factor, so the headline SNRs should be treated as provisional.","tokens_in":29446,"tokens_out":1882,"would_cite":true,"duration_ms":21750,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes that cross-correlating the squared Faraday rotation measure, RM², with the foreground galaxy density isolates cosmic magnetic fields near the galaxies' redshifts, and forecasts that the signal is measurable with current","keywords":["Faraday rotation measure","cosmic magnetic fields","RM-squared galaxy cross-correlation","tomography","bispectrum","magnetic field evolution","projected fields estimator","kinetic Sunyaev-Zel'dovich effect"],"falsifier":"Measure ⟨RM²×g⟩ at z≈0.1 using a few thousand background RMs and a wide-area galaxy catalog; the paper forecasts signal-to-noise around 4 for such a sample, so a null detection at that level would indicate the simulated amplitude or the statistic's assumptions need revision.","tokens_in":28447,"feed_emoji":"🧲","tokens_out":6278,"duration_ms":58516,"temperature":0.7,"pith_summary":"The paper introduces a new statistic, the cross-correlation between the squared Faraday rotation measure (RM²) of background radio sources and the projected density of foreground galaxies, and argues that it cleanly extracts the magnetic field contribution arising at the redshifts of those galaxies. This matters because direct RM–galaxy correlations cancel out—magnetic fields can point toward or away from us—while the absolute-value variant is biased by measurement noise. The RM²-based estimator is noise-free and can be written as a bispectrum of two electron-density–weighted magnetic fields and one galaxy overdensity, so it can be modeled and tomographically binned in galaxy redshift. Using cosmological magnetohydrodynamic simulations, the paper finds the signal grows by roughly three orders of magnitude from redshift 3 to 0 and forecasts high-significance detections with upcoming radio and galaxy surveys.","feed_headline":"RM²–galaxy correlation traces cosmic magnetic field growth","feed_subtitle":"Squaring Faraday rotation removes noise bias and lets galaxy redshifts date magnetic field amplification.","key_machinery":"The central object is the squared rotation measure field A = RM² and its two-point cross-correlation with the projected galaxy density, ⟨RM²×g⟩, built in direct analogy to the 'projected fields' estimator used for the kinetic Sunyaev-Zel'dovich effect. The paper shows ⟨RM²×g⟩ equals a line-of-sight projection of a bispectrum of two copies of the electron-density–weighted line-of-sight magnetic field and one galaxy overdensity, and introduces a heuristic approximation (involving a simulation-calibrated factor K) in which the correlation's shape is the electron–galaxy cross-correlation and its amplitude is the projected power of the electron-density–weighted magnetic field. That approximation","core_discovery":"The central claim is that ⟨RM²×g⟩ measures the line-of-sight integral of the electron-density–weighted magnetic field produced near foreground galaxy redshifts, free of contamination from Milky Way, source-intrinsic, and other line-of-sight RM contributions. The statistic is related to a bispectrum involving two copies of the electron-density–weighted magnetic field and one galaxy overdensity; in the paper's approximate model the shape of the correlation is set by the electron–galaxy two-point function while the amplitude is set by the projected power of the electron-weighted line-of-sight magnetic field. In the simulations the effective field strength is dominated by the inner regions of ha","pith_inferences":["If the statistic works as claimed, the redshift-binned amplitude provides a measurement of cosmic magnetic energy density evolution that does not require assuming a particular magnetogenesis scenario, since the electron-density weighting can be calibrated separately.","The strong dependence of the signal on the RM smoothing scale suggests that denser future RM catalogs could probe the small-scale cutoff of the magnetic field power spectrum; conversely, a detection at a given smoothing scale carries information about field coherence length.","Comparing RM² and |RM| estimators on the same data would isolate the non-Gaussian tail of the electron-density–weighted field, since RM² weights rare high-RM pixels quadratically while |RM| weights them linearly.","A natural extension is to cross-correlate RM² with galaxy lensing or with higher-order galaxy statistics to separate the halo-mass dependence of magnetic field amplification from the pure redshift evolution."],"forward_implications":["The RM²–galaxy estimator is immune to the noise bias that suppresses |RM|–galaxy correlations, because RM² noise terms are uncorrelated with the galaxy field.","Splitting the foreground galaxies into redshift bins turns the statistic into a tomographic probe of the electron-density–weighted magnetic field strength across cosmic time.","With ~10⁷ RM measurements (as expected from future radio surveys) and ~10⁶ galaxies per redshift bin, the signal is forecast to be detectable at high significance over roughly 0.5–50 Mpc/h scales and redshifts up to z≈1.","Even with current-size samples (~10³ RMs), the paper forecasts a marginal detection at low redshift, making pilot measurements worthwhile now.","The shape of the correlation largely follows the electron–galaxy clustering, so it can be calibrated by other probes such as fast radio burst dispersion measure and kSZ cross-correlations, leaving the amplitude as a cleaner measure of magnetic field strength."],"fun_headline_variants":["RM²–galaxy correlation yields tomographic cosmic B-field map","Noise-free RM²–galaxy probe dates magnetic field growth","Squared RM cross-correlation with galaxies traces cosmic magnetism","Tomographic magnetic field map via RM²–galaxy cross-correlation","RM²–galaxy cross-correlation avoids bias to trace cosmic B-fields"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The predicted signal amplitude and detection significance depend on the magnetic field statistics of the simulations being representative; the paper itself shows that higher-resolution runs—which resolve more magnetic field reversals—produce a weaker projected signal, so the forecasts could be optimistic if the real field is as tangled as those runs suggest.","fun_headline_variants_meta":{"raw":{"variants":["RM²–galaxy correlation yields tomographic cosmic B-field map","Noise-free RM²–galaxy probe dates magnetic field growth","Squared RM cross-correlation with galaxies traces cosmic magnetism","Tomographic magnetic field map via RM²–galaxy cross-correlation","RM²–galaxy cross-correlation avoids bias to trace cosmic B-fields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001421,"raw_usage":{"total_tokens":5637,"prompt_tokens":871,"completion_tokens":4766,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":4673}},"tokens_in":615,"tokens_out":4766,"duration_ms":29951,"temperature":1.0,"reasoning_tokens":4673,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:06:59.521095+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure ⟨RM²×g⟩ at z≈0.1 using a few thousand background RMs and a wide-area galaxy catalog; the paper forecasts signal-to-noise around 4 for such a sample, so a null detection at that level would indicate the simulated amplitude or the statistic's assumptions need revision.","supporting_citations":[],"review_version":1}