Pith. sign in

REVIEW 3 major objections 5 minor 6 references

Verifying the Australian MWA EoR pipeline II: fundamental limits of the AusEoRPipe and the impact of instrumental effects

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that recovering the 21-cm signal with sky-based calibration requires subtracting unresolved point-source flux down to about 1 mJy apparent ($10^{-3}$ Jy, 98.5% of the apparent flux in the adopted sky model), and that the…

desk verdict Quantitative systematics ranking for MWA EoR, useful and mostly sound, but internal numbers conflict and the claim only covers direction-independent effects. read the letter →

arxiv 2501.07004 v1 pith:XRTZRYZU submitted 2025-01-13 astro-ph.CO astro-ph.IM

classification astro-ph.COastro-ph.IM
keywords epochofreionisation21-cmpowerspectrumMWAEoRpipelineskymodelcompletenessradiointerferometriccalibrationinstrumentalsystematicsWODENsimulations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish where the Australian MWA EoR pipeline (AusEoRPipe) loses the ability to detect the cosmological 21-cm signal, by simulating realistic MWA data with WODEN and running known instrumental effects through the pipeline in isolation. It finds that sky-model incompleteness dominates: without subtracting unresolved point sources down to an apparent flux of $10^{-3}$ Jy (98.5% of the apparent flux in the adopted sky model), calibration and subtraction errors inject enough power to mask the 21-cm signal even when no instrumental effects are present. When diffuse emission is added, some $k$-modes are never accessible, so diffuse foreground removal is also required. The payoff is a concrete priority list for the experiment: fix the sky model first, then add gain smoothness, cable-reflection fitting, and diffuse treatment. That is what a detection attempt should spend its effort on.

What carries the argument

The central object is the apparent-flux threshold on the discrete sky model: sources are weighted by the MWA primary beam at 182 MHz, and three sub-models cut at $10^{-1}$, $10^{-2}$, and $10^{-3}$ Jy are subtracted from the test bed. The power-spectrum machinery is CHIPS with a wedge cut that excludes modes with $k_\parallel < 0.08\,h\,\mathrm{Mpc}^{-1}$ and $k_\perp > 0.06\,h\,\mathrm{Mpc}^{-1}$ plus the horizon line. The mechanism that carries the calibration result is per-channel, direction-independent calibration, whose small spectral ripples transfer smooth foreground power into the EoR window; the same catalogue used for calibration is also used for subtraction. The diffuse model is an EDA2 m-mode map upgraded and smoothed to MWA resolution, which contains discrete sources and therefore double-counts flux when combined with the discrete model, a point the paper accepts as long as total power matches real data.

What would settle it

Run the same WODEN/AusEoRPipe test with the sky model extended below the $10^{-3}$ Jy apparent-flux cut (or with direction-dependent peeling of the brightest sources) and measure the residual power in the EoR window; the paper's claim predicts the signal is recovered once 98.5% of apparent flux is subtracted, so any simulation in which the 21-cm signal still fails to appear above that residual would falsify the threshold result.

Watch

Extended reading notes

Core claim

In a single-observation test bed with the discrete and 21-cm sky models but no instrumental effects, the paper shows that an apparent-flux cut at $10^{-3}$ Jy (35% of the sources, 98.5% of the apparent flux) is necessary and sufficient for the 21-cm signal to be recovered after direct subtraction in visibility space. Adding the diffuse sky model makes low-$k$ modes unrecoverable without diffuse emission removal. When calibration is included, percent-level frequency-dependent amplitude fluctuations in the per-channel gain solutions couple low-$k$ foreground power into higher $k$-modes, masking the signal; applying the same calibration solutions to the 21-cm model alone does not bias that signal. Across the 15-observation test bed with noise, the incomplete sky model is the single greatest cause of leakage, while cable reflections, flagged coarse-band channels, and tile gain errors each add comparable power and less than the sky model. Averaging calibration solutions over the 30-minute pointing reduces window leakage in both simulations and real data.

Load-bearing premise

The simulated data must reproduce the real MWA signal chain closely enough that the ranking of systematics transfers to reality; the paper itself reports that simulations leak less into the window than matched real data, so unmodelled effects such as ionospheric refraction, direction-dependent errors, or missing dipoles could change the priority order and the required flux threshold.

Editorial extensions

If this is right

  • Sky-based calibration and power-spectrum recovery of the 21-cm signal requires removing unresolved point sources down to about $10^{-3}$ Jy apparent flux (more than 90% of apparent flux; 98.5% in the adopted model), not just the bright calibrators.
  • With diffuse emission in the simulation, some $k$-modes are never accessible from 30 minutes of zenith data, so diffuse foreground removal is a necessary part of the pipeline, not an optional refinement.
  • The ordering of systematics is concrete: an incomplete sky model causes more EoR-window leakage than flagged coarse-band channels, cable reflections, or tile gain errors at the tested levels, so sky-model development should take priority.
  • Averaging calibration solutions over the 15 snapshots of a pointing reduces window leakage in simulations and real data, pointing to a cheap sensitivity gain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: if the ranking transfers to real data, extending the southern-sky catalogue below 1 mJy should buy more EoR sensitivity per unit effort than any other single pipeline change; early SKA-Low arrays may find their detection limited by the same sky-model gap until their catalogues catch up.
  • Inference: because the diffuse map already contains the discrete sources, the residual diffuse power at high $k$ is an upper limit in this test; a cleaner diffuse/discrete separation would turn 'some k-modes cannot be accessed' into a quantitative statement.
  • Inference: the paper tests only direction-independent effects; ionospheric refraction and per-tile primary-beam variation (missing dipoles) are explicitly left out and could plausibly inject leakage at or above the sky-model level, changing the priority order.
  • Inference: the calibration-averaging result suggests a direct upgrade: enforce spectral smoothness of gain solutions (LOFAR-style regularization) to attack the leakage mechanism at its source rather than after the fact.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper uses the WODEN simulator to generate realistic MWA EoR0 zenith-pointing observations containing a 21-cm model, a discrete foreground model, and a diffuse foreground model, then runs them through the AusEoRPipe analysis pipeline. It tests four direction-independent instrumental effects (edge/centre channel flagging, tile gain errors, cable reflections, and thermal noise) both in isolation and in combination, compares the resulting power spectra to matched real MWA data, and ranks the contributions of these effects to leakage into the EoR window. The headline claims are that an apparent-flux cut near 1 mJy (claimed in the abstract as 10 mJy) is needed to recover the 21-cm signal, that diffuse emission prevents access to some k-modes, that an incomplete sky model is the single greatest cause of leakage, and that the other tested effects are comparable to each other and subdominant.

Significance. If the results hold, the paper provides a concrete, simulation-based prioritisation for MWA EoR analysis: invest in deeper/higher-resolution source catalogues, implement gain spectral smoothness, and develop diffuse-emission treatment. The strengths are the controlled simulation experiment, the use of external sky catalogues for the discrete model and EDA2-derived diffuse map, the matched comparison to real observations, and the open-source WODEN simulator. The paper is also honest about its own limitations, explicitly stating in Section 5.5 that the simulations show less leakage than real data and in Section 2.2 that only direction-independent effects are considered. The central ranking among the simulated effects is a solid contribution, but the quantitative threshold in the abstract and conclusion is inconsistent with the body, and the 'single greatest cause' claim is only established within the simulated direction-independent effect set, which the paper itself acknowledges.

major comments (3)
  1. [Abstract, Section 4.1, Table 1, Section 8] The abstract and conclusion state that more than 90% of unresolved point-source flux down to 10 mJy apparent must be subtracted, but Section 4.1 and Table 1 show that the necessary cut is at 10^-3 Jy (1 mJy), which removes 98.5% of apparent flux. This is an order-of-magnitude discrepancy in the headline quantitative result. The abstract and conclusion (including the bullet 'sources down to less than 10 mJy need to be removed') must be corrected to match the body's 1 mJy / 98.5% threshold, or the body must be re-examined if 10 mJy was intended.
  2. [Section 2.2, Section 5.5, Section 7] The claim that 'the single greatest cause of leakage is an incomplete sky model' is established only among the set of direction-independent effects simulated. Section 2.2 explicitly restricts the study to 'instrumental effects that are independent of direction upon the sky', and Section 5.5 states 'there is clearly less leakage into the window from the simulations, indicating further unmodelled systematics'. Direction-dependent effects such as the ionosphere and per-tile primary-beam errors are deferred (Section 7). Because the ionosphere is known to introduce spectral structure and direction-dependent phase errors that leak foreground power, the ranking may not transfer to real data. Please qualify the headline claim as 'among the tested direction-independent effects' and attach the same qualifier to the flux-cut threshold derived from this ranking.
  3. [Section 3.3, Section 4.1, Figure 4] The diffuse sky model contains all discrete sources, so simulations combining the diffuse and discrete models double-count flux. The paper acknowledges this ('will result in some double counting of flux') but does not quantify the resulting spurious power at high k-modes, even though Figure 4 notes residual power there 'may be false power'. Since the full sky model is used for the ranking experiments in Section 5, the double-counting could in principle change the relative sizes of the leakage terms that underpin the 'incomplete sky model dominates' conclusion. Please quantify the double-counted component, for example by comparing a simulation with the diffuse map alone against one with the discrete model alone in the same k-range, and show that the ranking is robust to this effect.
minor comments (5)
  1. [Section 5.6] The first-person phrasing in Section 5.6 ('In these tests on real data, I have assumed all dipoles are alive') should be converted to formal third-person wording, and the caveat that this may introduce a different calibration systematic should be integrated into the main analysis rather than left as a parenthetical aside.
  2. [Figure 6 caption] The caption says 'All real and simulated data have been calibrated using 10,000 sources, with the real data calibrated using 8000 sources' but Section 5.5 states that simulations use 'the apparent 10,000 brightest sources'; please clarify whether the real data use 8000 or 10,000 sources and reconcile the caption with the text.
  3. [Section 4.1, Table 1] The notation '10-1, 10-2, 10-3 Jy' in the text should be typeset with superscripts, e.g., 10^-1 Jy, and Table 1 caption should use the same formatting.
  4. [Section 6] The reference 'Jordan et al, 2024, submitted' is incomplete; either provide a full citation or remove the reference marker.
  5. [Throughout] The manuscript inconsistently uses 'sky model' and 'skymodel' (e.g., Section 2.1 vs. Section 5.6); please standardise the spelling.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's conclusions are outcomes of controlled simulations using external sky catalogues and independently measured instrumental effects; self-citations identify tools and inputs rather than carrying the argument.

full rationale

The derivation chain is a forward simulation and processing experiment, not an analytic derivation whose conclusion is built into its premises. The authors use WODEN to generate visibilities from external sky catalogues (GLEAM, LoBES, an EDA2-derived diffuse map) and an injected 21-cm model, then apply direction-independent calibration and subtraction, inject specific instrumental effects in isolation, and measure the leakage into the EoR window. The headline flux-cut result is obtained by explicitly generating three sub-sky-model cuts (10^-1, 10^-2, 10^-3 Jy), subtracting each from the simulated test bed, and observing which cut allows recovery of the injected 21-cm signal; the threshold is an outcome of that comparison, not a fitted parameter renamed as a prediction, and the authors explicitly caution that it is specific to their discrete sky model. Similarly, the claim that the incomplete sky model is the single greatest cause of leakage is a ranking established within the paper's suite of independently parameterised effects (channel flagging, tile gain errors, cable reflections, thermal noise), not a restatement of the simulation inputs. Self-citations appear throughout, but they are load-bearing only as references to tools and prior models: WODEN (Line 2022), the Paper I 21-cm model and AusEoRPipe description (Line et al. 2024), and the shapelet/Gaussian component machinery in other first-author works. These are external, code-reproduced, or empirically anchored inputs rather than self-imported uniqueness theorems or ansatze that predetermine the present conclusions. The cable-reflection amplitude range is attributed to a Barry private communication, but that is an empirical input to the simulation, not a circular derivation of the leakage ranking. The paper itself identifies unmodelled direction-dependent systematics (ionosphere, missing dipoles, directional RFI) and notes that simulations leak less than real data, but these are acknowledged limitations on external validity, not circularity. There is no equation in which the claimed prediction is equal to an input by construction, and no fitted parameter is renamed as a predicted quantity. Score 0.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on the fidelity of WODEN simulations and on the chosen amplitudes of injected systematics. The free parameters are the instrument-effect distributions (gain errors, cable reflections, leakage) and the analysis choices (wedge cut). The axioms are standard radio-astronomy domain assumptions: additivity of visibilities, direction-independent effect injection, EDA2 diffuse map as a proxy, and the injected 21-cm model as ground truth. No invented entities are introduced.

free parameters (6)
  • Cable reflection amplitude range = U(0.02, 0.1)
    Drawn for each tile in Section 5.3; the ranking of cable reflections below the incomplete-sky-model effect depends on this range.
  • Cable reflection phase = U(0, 180 deg)
    Random phase offset per tile, Section 5.3.
  • Tile gain amplitude range = U(0.7, 1.3)
    Applied per tile and polarization, Section 5.2; the relative impact of gain errors is set by this spread.
  • Gain phase slope range = max phase offset U(-60, 60 deg)
    Frequency-dependent phase slope per tile, Section 5.2.
  • Dipole leakage alignment errors = Psi ~ U(0, 0.02 deg), chi ~ U(0, 0.05 deg)
    Used to construct Dx and Dy leakage terms, Equations (2) and (3).
  • Power spectrum wedge cut = k_parallel > 0.08 h/Mpc, k_perp > 0.06 h/Mpc plus horizon line
    Modes excluded before 1D binning; the apparent flux threshold needed for recovery depends on this cut (Section 4.1).
assumptions (6)
  • domain assumption Visibilities from different sky components are additive, allowing separate simulation and summation of 21-cm, discrete, and diffuse components.
    Underlies the staged simulation and visibility-space subtraction in Sections 3 and 4.1.
  • domain assumption Direction-independent instrumental effects can be added to simulated visibilities post-hoc without changing sky response.
    Motivates the testing strategy in Section 2.2 and rules out direction-dependent effects such as ionosphere.
  • domain assumption The EDA2-derived diffuse map, although it includes discrete sources and double-counts flux, is a sufficient proxy for diffuse power.
    Stated in Section 3.3; the diffuse-mode conclusions rest on this proxy.
  • domain assumption The injected 21-cm sky model is a realistic realization of the EoR signal, so recovering it validates the pipeline.
    Used in Section 3.1 and throughout; the paper measures recovery of this model, not the true cosmological signal.
  • domain assumption The calibration catalogue is incomplete exactly at the tested flux levels; fainter unmodeled sources are negligible in simulation.
    The Section 4.1 claim that a 10^-3 Jy cut recovers the signal assumes no significant contribution from sources below that cut.
  • domain assumption The thermal noise model with Tsky=228 K, Trec=50 K, Aeff=20.35 m^2 describes MWA noise in the test observations.
    Used for noise injection and calibration-noise behavior in Sections 5.4 and 6.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Verifying the Australian MWA EoR pipeline II: fundamental limits of the AusEoRPipe and the impact of instrumental effects." pith.science (2026). https://pith.science/paper/XRTZRYZU

@misc{pith2026250107004,
  author       = {Pith},
  title        = {Pith review of: Verifying the Australian MWA EoR pipeline II: fundamental limits of the AusEoRPipe and the impact of instrumental effects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XRTZRYZU}},
  note         = {Machine review of arXiv:2501.07004}
}
read the original abstract

Detection of the weak cosmological signal from high-redshift hydrogen demands careful data analysis and an understanding of the full instrument signal chain. Here we use the WODEN simulation pipeline to produce realistic data from the Murchison Widefield Array Epoch of Reionisation experiment, and test the effects of different instrumental systematics through the AusEoRPipe analysis pipeline. The simulations include a realistic full sky model, direction-independent calibration, and both random and systematic instrumental effects. Results are compared to matched real observations. We find that, (i) with a sky-based calibration and power spectrum approach we have need to subtract more than 90% of all unresolved point source flux (10mJy apparent) to recover 21-cm signal in the absence of instrumental effects; (ii) when including diffuse emission in simulations, some k-modes cannot be accessed, leading to a need for some diffuse emission removal; (iii) the single greatest cause of leakage is an incomplete sky model; (iv) other sources of errors, such as cable reflections, flagged channels and gain errors, impart comparable systematic power to one another, and less power than the incomplete skymodel.

Figures

Figures reproduced from arXiv: 2501.07004 by the authors.

Figure 1
Figure 1. All-sky orthographic projections of the sky models, centred at RA,Dec = 0 ◦, –27◦ (-27◦ is zenith for the MWA), where: a) shows a slice of the 21-cm sky model at 167MHz; b) shows the positions of all sources in the discrete sky model; c) shows the diffuse sky model at 200MHz [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Comparison of a real zenith pointed 2 minute snapshot to simulated data. Both real and simulated data were averaged to 8 s, 80 kHz, and the full bandwidth imaged via WSClean. The top row were both imaged with Briggs 0 weighting, where a) shows the real data after calibration, and b) shows the simulated data with no calibration. The bottom row were both imaged with natural weighting, where c) shows the real data afte… view at source ↗
Figure 3
Figure 3. shows that the apparent flux cut at 10–3Jy is necessary to reliably recover the 21-cm signal. Note this is a flux cut on this particular discrete sky model; this is not saying that leaving all sources below 10–3Jy in real data will allow a detection. This result is limited by the com￾pleteness of this discrete sky model, and these simulations do not include confusion noise. It should also be noted that as we cut by … view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: 1D PS depicting data from a single zenith EoR0 observation where both discrete and diffuse foregrounds were included. from the calibrated visibilitiesa . To summarise, we calibrate using an incomplete discrete sky model, and then subtract the entire discrete sky model …
Figure 5
Figure 5. Figure 5: 1D PS depicting data from a single zenith EoR0 observation, show￾ing the effects of calibration on leakage into the window. Note that the dark purple and brown lines showing uncalibrated simulations are colour-coded to lines showing the same simulations in Figures 3 an…
Figure 6
Figure 6. Figure 6: Comparison of an integration over 15 real two-minute zenith observations to simulated data. All real and simulated data have been calibrated using 10,000 sources, with the real data calibrated using 8000 sources. In the ratios, blue means more power in the real data, r…
Figure 7
Figure 7. Figure 7: Comparison of different instrumental effects and their manifestation in the window. These 1D PS were made from 15 zenith observations, the 2D PS of which are shown in [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Comparison of calibrating and subtracting with various numbers of sources, with simulated data on left, real data on the right, when instrumental errors are included. Power spectra are made from integrating 15 EoR0 zenith observations. The simulations are noiseless as …
Figure 9
Figure 9. Figure 9: Calibration amplitudes showing the effects of averaging calibration solutions over 15 observations. Calibrations for a single tile are shown; other tiles show similar behaviours. Left column shows a single observation, right the average over 15 observations. Top row sh…
Figure 10
Figure 10. Figure 10: Effects of averaging calibration solutions over 15 observations for real data (top row) and simulated data (bottom row). The data in the simulation contain both discrete and diffuse sky models, but do not contain noise, in an attempt to better reveal any systematic bi…
Figure 11
Figure 11. Figure 11: Calibration gain amplitudes and phases from a simulated two minute snapshot. These demonstrate the constant gain and flat phase slopes added to the simulation. The underlying simulation was of both the diffuse and discrete sky models, and only contained gain errors wi…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

6 extracted references · 2 canonical work pages

  1. [1]

    https://doi.org/10.22323/1.215.0001

    April. https://doi.org/10.22323/1.215.0001. arXiv: 1505.07568 [astro-ph.CO]. Kriele, Michael A., Randall B. Wayth, Mark J. Bentum, Budi Juswardy, and Cathryn M. Trott. 2022. Imaging the southern sky at 159 MHz using spherical harmonics with the engineering development array 2. PASA 39 (April): e017. https://doi.org/10.1017/pasa.2022.2. arXiv: 2201.04281 [...

  2. [7]

    org / abs / 1206

    http : / / arxiv . org / abs / 1206 . 6945 % 20http : / / www . journals . cambridge.org/abstract_S1323358012000070. Trott, Cathryn M., C. H. Jordan, S. Midgley, N. Barry, B. Greig, B. Pindor, J. H. Cook, et al. 2020. Deep multiredshift limits on Epoch of Reion- ization 21 cm power spectra from four seasons of Murchison Wide- field Array observations. MNR...

  3. [34]

    http://arxiv.org/abs/ 1601.02073%20http://dx.doi.org/10.3847/0004-637X/818/2/139

    https://doi.org/10.3847/0004-637X/818/2/139. http://arxiv.org/abs/ 1601.02073%20http://dx.doi.org/10.3847/0004-637X/818/2/139. Wayth, R B, E Lenc, M E Bell, J R Callingham, K S Dwarakanath, T. M. O. Franzen, B.-Q. For, et al. 2015. GLEAM: The GaLactic and Extragalac- tic All-Sky MW A Survey.PASA 32 (June): e025. ISSN: 1323-3580. https: //doi.org/10.1017/p...

  4. [124]

    Hurley-Walker, N., J

    https://doi.org/10.3847/1538-4357/acaf 50. Hurley-Walker, N., J. R. Callingham, P. J. Hancock, T. M. O. Franzen, L. Hindson, A. D. Kapińska, J. Morgan, et al. 2017. GaLactic and Extra- galactic All-sky Murchison Widefield Array (GLEAM) survey – I. A low-frequency extragalactic catalogue. MNRAS 464 (1): 1146–1167. ISSN: 0035-8711. https : / / doi . org / 1...

  5. [141]

    https://doi.org/10.3847/1538- 4357/ab55e4

    ISSN: 1538-4357. https://doi.org/10.3847/1538- 4357/ab55e4. arXiv: 1911.10216. https://iopscience.iop.org/article/10.3847/1538- 4357/ab55e4. Line, J. L. B., D. A. Mitchell, B. Pindor, J. L. Riding, B. McKinley, R. L. Webster, C. M. Trott, N. Hurley-Walker, and A. R. Offringa. 2020. Modelling and peeling extended sources with shapelets: A Fornax A case stu...

  6. [3987]

    https : / / doi

    ISSN: 0035-8711. https : / / doi . org / 10 . 1093 / mnras / stx1797. http : / / academic . oup . com / mnras / article / 471 / 4 / 3974 / 3979467 / Characterization-of -the-ionosphere-above-the. 12 J. L. B. Line et al. Joye, W. A., and E. Mandel. 2003. New Features of SAOImage DS9. In Astronomical data analysis software and systems xii, edited by H. E. P...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.