REVIEW 3 major objections 5 minor 6 references
Verifying the Australian MWA EoR pipeline II: fundamental limits of the AusEoRPipe and the impact of instrumental effects
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper argues that recovering the 21-cm signal with sky-based calibration requires subtracting unresolved point-source flux down to about 1 mJy apparent ($10^{-3}$ Jy, 98.5% of the apparent flux in the adopted sky model), and that the…
desk verdict Quantitative systematics ranking for MWA EoR, useful and mostly sound, but internal numbers conflict and the claim only covers direction-independent effects. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the apparent-flux threshold on the discrete sky model: sources are weighted by the MWA primary beam at 182 MHz, and three sub-models cut at $10^{-1}$, $10^{-2}$, and $10^{-3}$ Jy are subtracted from the test bed. The power-spectrum machinery is CHIPS with a wedge cut that excludes modes with $k_\parallel < 0.08\,h\,\mathrm{Mpc}^{-1}$ and $k_\perp > 0.06\,h\,\mathrm{Mpc}^{-1}$ plus the horizon line. The mechanism that carries the calibration result is per-channel, direction-independent calibration, whose small spectral ripples transfer smooth foreground power into the EoR window; the same catalogue used for calibration is also used for subtraction. The diffuse model is an EDA2 m-mode map upgraded and smoothed to MWA resolution, which contains discrete sources and therefore double-counts flux when combined with the discrete model, a point the paper accepts as long as total power matches real data.
What would settle it
Run the same WODEN/AusEoRPipe test with the sky model extended below the $10^{-3}$ Jy apparent-flux cut (or with direction-dependent peeling of the brightest sources) and measure the residual power in the EoR window; the paper's claim predicts the signal is recovered once 98.5% of apparent flux is subtracted, so any simulation in which the 21-cm signal still fails to appear above that residual would falsify the threshold result.
Extended reading notes
Core claim
In a single-observation test bed with the discrete and 21-cm sky models but no instrumental effects, the paper shows that an apparent-flux cut at $10^{-3}$ Jy (35% of the sources, 98.5% of the apparent flux) is necessary and sufficient for the 21-cm signal to be recovered after direct subtraction in visibility space. Adding the diffuse sky model makes low-$k$ modes unrecoverable without diffuse emission removal. When calibration is included, percent-level frequency-dependent amplitude fluctuations in the per-channel gain solutions couple low-$k$ foreground power into higher $k$-modes, masking the signal; applying the same calibration solutions to the 21-cm model alone does not bias that signal. Across the 15-observation test bed with noise, the incomplete sky model is the single greatest cause of leakage, while cable reflections, flagged coarse-band channels, and tile gain errors each add comparable power and less than the sky model. Averaging calibration solutions over the 30-minute pointing reduces window leakage in both simulations and real data.
Load-bearing premise
The simulated data must reproduce the real MWA signal chain closely enough that the ranking of systematics transfers to reality; the paper itself reports that simulations leak less into the window than matched real data, so unmodelled effects such as ionospheric refraction, direction-dependent errors, or missing dipoles could change the priority order and the required flux threshold.
Editorial extensions
If this is right
- Sky-based calibration and power-spectrum recovery of the 21-cm signal requires removing unresolved point sources down to about $10^{-3}$ Jy apparent flux (more than 90% of apparent flux; 98.5% in the adopted model), not just the bright calibrators.
- With diffuse emission in the simulation, some $k$-modes are never accessible from 30 minutes of zenith data, so diffuse foreground removal is a necessary part of the pipeline, not an optional refinement.
- The ordering of systematics is concrete: an incomplete sky model causes more EoR-window leakage than flagged coarse-band channels, cable reflections, or tile gain errors at the tested levels, so sky-model development should take priority.
- Averaging calibration solutions over the 15 snapshots of a pointing reduces window leakage in simulations and real data, pointing to a cheap sensitivity gain.
Reading between the lines
- Inference: if the ranking transfers to real data, extending the southern-sky catalogue below 1 mJy should buy more EoR sensitivity per unit effort than any other single pipeline change; early SKA-Low arrays may find their detection limited by the same sky-model gap until their catalogues catch up.
- Inference: because the diffuse map already contains the discrete sources, the residual diffuse power at high $k$ is an upper limit in this test; a cleaner diffuse/discrete separation would turn 'some k-modes cannot be accessed' into a quantitative statement.
- Inference: the paper tests only direction-independent effects; ionospheric refraction and per-tile primary-beam variation (missing dipoles) are explicitly left out and could plausibly inject leakage at or above the sky-model level, changing the priority order.
- Inference: the calibration-averaging result suggests a direct upgrade: enforce spectral smoothness of gain solutions (LOFAR-style regularization) to attack the leakage mechanism at its source rather than after the fact.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper uses the WODEN simulator to generate realistic MWA EoR0 zenith-pointing observations containing a 21-cm model, a discrete foreground model, and a diffuse foreground model, then runs them through the AusEoRPipe analysis pipeline. It tests four direction-independent instrumental effects (edge/centre channel flagging, tile gain errors, cable reflections, and thermal noise) both in isolation and in combination, compares the resulting power spectra to matched real MWA data, and ranks the contributions of these effects to leakage into the EoR window. The headline claims are that an apparent-flux cut near 1 mJy (claimed in the abstract as 10 mJy) is needed to recover the 21-cm signal, that diffuse emission prevents access to some k-modes, that an incomplete sky model is the single greatest cause of leakage, and that the other tested effects are comparable to each other and subdominant.
Significance. If the results hold, the paper provides a concrete, simulation-based prioritisation for MWA EoR analysis: invest in deeper/higher-resolution source catalogues, implement gain spectral smoothness, and develop diffuse-emission treatment. The strengths are the controlled simulation experiment, the use of external sky catalogues for the discrete model and EDA2-derived diffuse map, the matched comparison to real observations, and the open-source WODEN simulator. The paper is also honest about its own limitations, explicitly stating in Section 5.5 that the simulations show less leakage than real data and in Section 2.2 that only direction-independent effects are considered. The central ranking among the simulated effects is a solid contribution, but the quantitative threshold in the abstract and conclusion is inconsistent with the body, and the 'single greatest cause' claim is only established within the simulated direction-independent effect set, which the paper itself acknowledges.
major comments (3)
- [Abstract, Section 4.1, Table 1, Section 8] The abstract and conclusion state that more than 90% of unresolved point-source flux down to 10 mJy apparent must be subtracted, but Section 4.1 and Table 1 show that the necessary cut is at 10^-3 Jy (1 mJy), which removes 98.5% of apparent flux. This is an order-of-magnitude discrepancy in the headline quantitative result. The abstract and conclusion (including the bullet 'sources down to less than 10 mJy need to be removed') must be corrected to match the body's 1 mJy / 98.5% threshold, or the body must be re-examined if 10 mJy was intended.
- [Section 2.2, Section 5.5, Section 7] The claim that 'the single greatest cause of leakage is an incomplete sky model' is established only among the set of direction-independent effects simulated. Section 2.2 explicitly restricts the study to 'instrumental effects that are independent of direction upon the sky', and Section 5.5 states 'there is clearly less leakage into the window from the simulations, indicating further unmodelled systematics'. Direction-dependent effects such as the ionosphere and per-tile primary-beam errors are deferred (Section 7). Because the ionosphere is known to introduce spectral structure and direction-dependent phase errors that leak foreground power, the ranking may not transfer to real data. Please qualify the headline claim as 'among the tested direction-independent effects' and attach the same qualifier to the flux-cut threshold derived from this ranking.
- [Section 3.3, Section 4.1, Figure 4] The diffuse sky model contains all discrete sources, so simulations combining the diffuse and discrete models double-count flux. The paper acknowledges this ('will result in some double counting of flux') but does not quantify the resulting spurious power at high k-modes, even though Figure 4 notes residual power there 'may be false power'. Since the full sky model is used for the ranking experiments in Section 5, the double-counting could in principle change the relative sizes of the leakage terms that underpin the 'incomplete sky model dominates' conclusion. Please quantify the double-counted component, for example by comparing a simulation with the diffuse map alone against one with the discrete model alone in the same k-range, and show that the ranking is robust to this effect.
minor comments (5)
- [Section 5.6] The first-person phrasing in Section 5.6 ('In these tests on real data, I have assumed all dipoles are alive') should be converted to formal third-person wording, and the caveat that this may introduce a different calibration systematic should be integrated into the main analysis rather than left as a parenthetical aside.
- [Figure 6 caption] The caption says 'All real and simulated data have been calibrated using 10,000 sources, with the real data calibrated using 8000 sources' but Section 5.5 states that simulations use 'the apparent 10,000 brightest sources'; please clarify whether the real data use 8000 or 10,000 sources and reconcile the caption with the text.
- [Section 4.1, Table 1] The notation '10-1, 10-2, 10-3 Jy' in the text should be typeset with superscripts, e.g., 10^-1 Jy, and Table 1 caption should use the same formatting.
- [Section 6] The reference 'Jordan et al, 2024, submitted' is incomplete; either provide a full citation or remove the reference marker.
- [Throughout] The manuscript inconsistently uses 'sky model' and 'skymodel' (e.g., Section 2.1 vs. Section 5.6); please standardise the spelling.
Circularity Check
No significant circularity: the paper's conclusions are outcomes of controlled simulations using external sky catalogues and independently measured instrumental effects; self-citations identify tools and inputs rather than carrying the argument.
full rationale
The derivation chain is a forward simulation and processing experiment, not an analytic derivation whose conclusion is built into its premises. The authors use WODEN to generate visibilities from external sky catalogues (GLEAM, LoBES, an EDA2-derived diffuse map) and an injected 21-cm model, then apply direction-independent calibration and subtraction, inject specific instrumental effects in isolation, and measure the leakage into the EoR window. The headline flux-cut result is obtained by explicitly generating three sub-sky-model cuts (10^-1, 10^-2, 10^-3 Jy), subtracting each from the simulated test bed, and observing which cut allows recovery of the injected 21-cm signal; the threshold is an outcome of that comparison, not a fitted parameter renamed as a prediction, and the authors explicitly caution that it is specific to their discrete sky model. Similarly, the claim that the incomplete sky model is the single greatest cause of leakage is a ranking established within the paper's suite of independently parameterised effects (channel flagging, tile gain errors, cable reflections, thermal noise), not a restatement of the simulation inputs. Self-citations appear throughout, but they are load-bearing only as references to tools and prior models: WODEN (Line 2022), the Paper I 21-cm model and AusEoRPipe description (Line et al. 2024), and the shapelet/Gaussian component machinery in other first-author works. These are external, code-reproduced, or empirically anchored inputs rather than self-imported uniqueness theorems or ansatze that predetermine the present conclusions. The cable-reflection amplitude range is attributed to a Barry private communication, but that is an empirical input to the simulation, not a circular derivation of the leakage ranking. The paper itself identifies unmodelled direction-dependent systematics (ionosphere, missing dipoles, directional RFI) and notes that simulations leak less than real data, but these are acknowledged limitations on external validity, not circularity. There is no equation in which the claimed prediction is equal to an input by construction, and no fitted parameter is renamed as a predicted quantity. Score 0.
Assumptions & free parameters
free parameters (6)
- Cable reflection amplitude range =
U(0.02, 0.1)
- Cable reflection phase =
U(0, 180 deg)
- Tile gain amplitude range =
U(0.7, 1.3)
- Gain phase slope range =
max phase offset U(-60, 60 deg)
- Dipole leakage alignment errors =
Psi ~ U(0, 0.02 deg), chi ~ U(0, 0.05 deg)
- Power spectrum wedge cut =
k_parallel > 0.08 h/Mpc, k_perp > 0.06 h/Mpc plus horizon line
assumptions (6)
- domain assumption Visibilities from different sky components are additive, allowing separate simulation and summation of 21-cm, discrete, and diffuse components.
- domain assumption Direction-independent instrumental effects can be added to simulated visibilities post-hoc without changing sky response.
- domain assumption The EDA2-derived diffuse map, although it includes discrete sources and double-counts flux, is a sufficient proxy for diffuse power.
- domain assumption The injected 21-cm sky model is a realistic realization of the EoR signal, so recovering it validates the pipeline.
- domain assumption The calibration catalogue is incomplete exactly at the tested flux levels; fainter unmodeled sources are negligible in simulation.
- domain assumption The thermal noise model with Tsky=228 K, Trec=50 K, Aeff=20.35 m^2 describes MWA noise in the test observations.
Cite this review
Pith. "Pith review of Verifying the Australian MWA EoR pipeline II: fundamental limits of the AusEoRPipe and the impact of instrumental effects." pith.science (2026). https://pith.science/paper/XRTZRYZU
@misc{pith2026250107004,
author = {Pith},
title = {Pith review of: Verifying the Australian MWA EoR pipeline II: fundamental limits of the AusEoRPipe and the impact of instrumental effects},
year = {2026},
howpublished = {\url{https://pith.science/paper/XRTZRYZU}},
note = {Machine review of arXiv:2501.07004}
}
read the original abstract
Detection of the weak cosmological signal from high-redshift hydrogen demands careful data analysis and an understanding of the full instrument signal chain. Here we use the WODEN simulation pipeline to produce realistic data from the Murchison Widefield Array Epoch of Reionisation experiment, and test the effects of different instrumental systematics through the AusEoRPipe analysis pipeline. The simulations include a realistic full sky model, direction-independent calibration, and both random and systematic instrumental effects. Results are compared to matched real observations. We find that, (i) with a sky-based calibration and power spectrum approach we have need to subtract more than 90% of all unresolved point source flux (10mJy apparent) to recover 21-cm signal in the absence of instrumental effects; (ii) when including diffuse emission in simulations, some k-modes cannot be accessed, leading to a need for some diffuse emission removal; (iii) the single greatest cause of leakage is an incomplete sky model; (iv) other sources of errors, such as cable reflections, flagged channels and gain errors, impart comparable systematic power to one another, and less power than the incomplete skymodel.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
https://doi.org/10.22323/1.215.0001
April. https://doi.org/10.22323/1.215.0001. arXiv: 1505.07568 [astro-ph.CO]. Kriele, Michael A., Randall B. Wayth, Mark J. Bentum, Budi Juswardy, and Cathryn M. Trott. 2022. Imaging the southern sky at 159 MHz using spherical harmonics with the engineering development array 2. PASA 39 (April): e017. https://doi.org/10.1017/pasa.2022.2. arXiv: 2201.04281 [...
arXiv 2022
-
[7]
http : / / arxiv . org / abs / 1206 . 6945 % 20http : / / www . journals . cambridge.org/abstract_S1323358012000070. Trott, Cathryn M., C. H. Jordan, S. Midgley, N. Barry, B. Greig, B. Pindor, J. H. Cook, et al. 2020. Deep multiredshift limits on Epoch of Reion- ization 21 cm power spectra from four seasons of Murchison Wide- field Array observations. MNR...
arXiv 2020
-
[34]
http://arxiv.org/abs/ 1601.02073%20http://dx.doi.org/10.3847/0004-637X/818/2/139
https://doi.org/10.3847/0004-637X/818/2/139. http://arxiv.org/abs/ 1601.02073%20http://dx.doi.org/10.3847/0004-637X/818/2/139. Wayth, R B, E Lenc, M E Bell, J R Callingham, K S Dwarakanath, T. M. O. Franzen, B.-Q. For, et al. 2015. GLEAM: The GaLactic and Extragalac- tic All-Sky MW A Survey.PASA 32 (June): e025. ISSN: 1323-3580. https: //doi.org/10.1017/p...
-
[124]
https://doi.org/10.3847/1538-4357/acaf 50. Hurley-Walker, N., J. R. Callingham, P. J. Hancock, T. M. O. Franzen, L. Hindson, A. D. Kapińska, J. Morgan, et al. 2017. GaLactic and Extra- galactic All-sky Murchison Widefield Array (GLEAM) survey – I. A low-frequency extragalactic catalogue. MNRAS 464 (1): 1146–1167. ISSN: 0035-8711. https : / / doi . org / 1...
-
[141]
https://doi.org/10.3847/1538- 4357/ab55e4
ISSN: 1538-4357. https://doi.org/10.3847/1538- 4357/ab55e4. arXiv: 1911.10216. https://iopscience.iop.org/article/10.3847/1538- 4357/ab55e4. Line, J. L. B., D. A. Mitchell, B. Pindor, J. L. Riding, B. McKinley, R. L. Webster, C. M. Trott, N. Hurley-Walker, and A. R. Offringa. 2020. Modelling and peeling extended sources with shapelets: A Fornax A case stu...
arXiv 1911
-
[3987]
ISSN: 0035-8711. https : / / doi . org / 10 . 1093 / mnras / stx1797. http : / / academic . oup . com / mnras / article / 471 / 4 / 3974 / 3979467 / Characterization-of -the-ionosphere-above-the. 12 J. L. B. Line et al. Joye, W. A., and E. Mandel. 2003. New Features of SAOImage DS9. In Astronomical data analysis software and systems xii, edited by H. E. P...
work page 2003
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.