REVIEW 1 major objections 3 references
A catalogue of 293 HI sources has been extracted from MIGHTEE survey data cubes in the COSMOS field, with detection completeness quantified through tests of four source-finding algorithms on injected simulated galaxies.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-29 11:05 UTC pith:SDODIO4T
load-bearing objection This paper releases a 293-source HI catalogue for COSMOS from MIGHTEE plus a direct comparison of four source finders on that data, with the main open question being how well the simulation-derived completeness transfers to real sources. the 1 major comments →
MIGHTEE-HI: HI catalogue of 293 sources for the COSMOS field and comparative study of 3-dimensional source finding methods
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
We present a catalogue of 293 HI sources extracted from the MIGHTEE survey data cubes covering the COSMOS field. The catalogue contains 293 sources in the redshift range of 0.004 < z < 0.093. In addition to HI masses and velocity widths, the catalogue includes optical through near-infrared photometry and inferred stellar masses and star-formation rates. The quantity of sources in the HI catalogue acquired through untargeted source finding is greatly influenced by the source finding methods used. This study therefore also provides a well-characterised expected completeness of the detected sample of galaxies based on their properties, inferred through a comparative study of different source fi
What carries the argument
Injection of simulated galaxies into real MeerKAT data cubes, binned narrowly by HI mass, inclination and distance, to benchmark recovery performance of PyBDSF, ProFound, SoFiA and LESHI.
Load-bearing premise
The performance of source finders on injected simulated galaxies divided into narrow bins of mass, inclination and distance accurately predicts their performance on real HI sources in the survey data cubes.
What would settle it
A direct count of how many real sources each of the four finders detects in the COSMOS data cubes, compared against the recovery fractions measured from the same finders on the injected simulated population.
If this is right
- The detected HI sample will have different completeness limits and selection biases depending on which source finder is applied.
- Source-finding strategies for the full MIGHTEE survey and upcoming SKAO surveys can be chosen or combined using the measured completeness curves.
- The catalogue supplies a well-characterised local HI sample whose selection function is now quantified as a function of galaxy properties.
Where Pith is reading between the lines
- Combining outputs from multiple finders may increase overall completeness while allowing users to track which sources are recovered by which algorithm.
- The same injection-and-recovery framework could be used to test source finders on other HI survey fields before final catalogue construction.
- Measurements of the local HI mass function derived from this catalogue will carry finder-dependent systematic uncertainties that can now be estimated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a catalogue of 293 HI sources extracted from MIGHTEE survey data cubes in the COSMOS field (redshift range 0.004 < z < 0.093), including HI masses, velocity widths, optical/NIR photometry, and derived stellar masses and star-formation rates. It also reports a comparative evaluation of four source-finding algorithms (PyBDSF, ProFound, SoFiA, and the new LESHI) by injecting simulated galaxies, binned by mass, inclination, and distance, into real MeerKAT data cubes to measure recovery rates and characterize expected completeness and detection biases for the MIGHTEE and future SKAO surveys.
Significance. The catalogue contributes a useful sample of low-redshift HI detections with multiwavelength ancillary data. The source-finder comparison, if the injected-model results transfer to real sources, supplies practical guidance on detection biases and strategies for large HI surveys. The use of real data cubes for the injection tests is a methodological strength that controls for survey-specific noise properties.
major comments (1)
- [Simulation injection and completeness estimation sections] The central claim that the simulation-based study 'provides a well-characterised expected completeness' and 'informs source finding strategies' for MIGHTEE and SKAO depends on the injected galaxies statistically representing real HI morphologies, kinematics, and environments. The manuscript reports per-bin recovery fractions from the narrow mass/inclination/distance bins but does not include any direct cross-check or validation against the properties of the 293 real detections (e.g., comparing recovered vs. non-recovered real sources or testing for residual cube artefacts). This assumption is load-bearing for the completeness interpretation and is not independently verified in the results or discussion sections.
Simulated Author's Rebuttal
We thank the referee for their constructive review and for recognising the value of the catalogue and the use of real data cubes in the injection tests. We address the major comment below.
read point-by-point responses
-
Referee: [Simulation injection and completeness estimation sections] The central claim that the simulation-based study 'provides a well-characterised expected completeness' and 'informs source finding strategies' for MIGHTEE and SKAO depends on the injected galaxies statistically representing real HI morphologies, kinematics, and environments. The manuscript reports per-bin recovery fractions from the narrow mass/inclination/distance bins but does not include any direct cross-check or validation against the properties of the 293 real detections (e.g., comparing recovered vs. non-recovered real sources or testing for residual cube artefacts). This assumption is load-bearing for the completeness interpretation and is not independently verified in the results or discussion sections.
Authors: We agree that the statistical representativeness of the injected galaxies is central to the interpretation. The simulations were constructed using empirical HI scaling relations (mass-size, mass-velocity width, and inclination distributions) drawn from the literature (primarily ALFALFA and other blind HI surveys) and placed at the observed redshifts and positions within the real MeerKAT cubes; this approach is intended to capture the dominant dependencies while incorporating the actual noise and artefact properties of the data. We acknowledge that an explicit consistency check against the real detections would strengthen the manuscript. In the revised version we will add a short subsection to the discussion that (i) compares the HI mass and velocity-width distributions of the 293 catalogue sources with the completeness-weighted expectations obtained by applying the measured recovery fractions to a model population, and (ii) summarises the visual-inspection criteria already used to reject residual artefacts. We note that a direct identification of 'non-recovered real sources' is not possible without an independent, complete census of all HI in the volume, which lies outside the scope of the present work. revision: partial
Circularity Check
No circularity: completeness follows directly from injection tests
full rationale
The paper reports a catalogue of 293 HI sources and measures recovery fractions by injecting simulated galaxies (binned by mass, inclination, distance) into the real MeerKAT cubes, then running PyBDSF, ProFound, SoFiA and LESHI. No equation or claim reduces a result to a fitted parameter defined from the same data; the completeness curves are direct empirical outputs of the injection procedure. No self-citation chain, uniqueness theorem, or ansatz is invoked to justify the method. The derivation chain is therefore self-contained against the survey cubes and the injected models.
Axiom & Free-Parameter Ledger
read the original abstract
We present a catalogue of HI sources extracted from the MIGHTEE survey data cubes covering the COSMOS field. The catalogue contains 293 sources in the redshift range of 0.004 < z < 0.093. In addition to HI masses and velocity widths, the catalogue includes optical through near-infrared photometry and inferred stellar masses and star-formation rates. The quantity of sources in the HI catalogue acquired through untargeted source finding is greatly influenced by the source finding methods used. This study therefore also provides a well-characterised expected completeness of the detected sample of galaxies based on their properties, informing of any detection biases, inferred through a comparative study of different source finding algorithms. We have tested the performance of widely-used source finders: PyBDSF, ProFound and SoFiA, along with new source finder LESHI, focusing exclusively on HI source detection rather than source characterisation in the first instance. The source finders were tested by injecting a sample of simulated galaxies divided into narrow bins of mass, inclination and distance into a MeerKAT data cube. The results inform the source finding strategies for the MeerKAT International GigaHertz Tiered Extragalactic Exploration (MIGHTEE) survey, as well as upcoming SKAO surveys.
Figures
Reference graph
Works this paper leans on
-
[1]
AdamsN.J.,BowlerR.A.A.,JarvisM.J.,HäußlerB.,LagosC.D.P.,2021, MNRAS, 506, 4933 Adams E. A. K., et al., 2022, A&A, 667, A38 Aihara H., et al., 2019, PASJ, 71, 114 Alam S., et al., 2015, ApJS, 219, 12 Astropy Collaboration et al., 2022, ApJ, 935, 167 Barkai J. A., Verheijen M. A. W., Talavera E., Wilkinson M. H. F., 2023, A&A, 670, A55 Barnes D. G., et al.,...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.22323/1.277.0004 2021
-
[2]
Looking At the Distant Universe with the MeerKAT Array (LADUMA)
pp 496–499 (arXiv:1109.5605), doi:10.1017/S1743921312009702 Hotan A. W., et al., 2021, Publ. Astron. Soc. Australia, 38, e009 Hunter J. D., 2007, Computing in Science and Engineering, 9, 90 Jaffé Y. L., Poggianti B. M., Verheijen M. A. W., Deshev B. Z., van Gorkom J. H., 2013, MNRAS, 431, 2111 Jarvis M., et al., 2016, in MeerKAT Science: On the Pathway to...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.1017/s1743921312009702 2021
-
[3]
low confidence
for the SoFiA, PyBDSF and ProFound source finders, derived analogously to Section 3.2. Figures B1, B2 and B3 contain plots showing the fitted completeness function analogously to Fig. 6 and Fig. 7 for the SoFiA, PyBDSF and ProFound source finders. Qualitatively, the fitted completeness functions for all of the sourcefinders are very similar. Fitted𝑀𝑚𝑖𝑛 pa...
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.