Pith. sign in

REVIEW 1 major objections 3 references

A catalogue of 293 HI sources has been extracted from MIGHTEE survey data cubes in the COSMOS field, with detection completeness quantified through tests of four source-finding algorithms on injected simulated galaxies.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-29 11:05 UTC pith:SDODIO4T

load-bearing objection This paper releases a 293-source HI catalogue for COSMOS from MIGHTEE plus a direct comparison of four source finders on that data, with the main open question being how well the simulation-derived completeness transfers to real sources. the 1 major comments →

arxiv 2605.28731 v1 pith:SDODIO4T submitted 2026-05-27 astro-ph.GA

MIGHTEE-HI: HI catalogue of 293 sources for the COSMOS field and comparative study of 3-dimensional source finding methods

classification astro-ph.GA
keywords HI sourcesMIGHTEE surveyCOSMOS fieldsource findingMeerKATneutral hydrogengalaxy catalogueradio astronomy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper extracts and releases an HI catalogue of 293 sources detected across the COSMOS field in the redshift range 0.004 to 0.093, complete with HI masses, velocity widths, and matched optical-to-near-infrared photometry plus derived stellar masses and star-formation rates. To understand how the choice of algorithm shapes the final sample, the authors inject simulated galaxies into the actual MeerKAT data cubes and measure recovery rates for PyBDSF, ProFound, SoFiA and the new LESHI finder, binned by HI mass, inclination and distance. This comparative exercise shows that the number of sources recovered, and the properties of the galaxies that are missed, depend strongly on which finder is used. The resulting completeness curves therefore supply a practical guide for choosing or combining finders in the ongoing MIGHTEE survey and in future SKAO observations.

Core claim

We present a catalogue of 293 HI sources extracted from the MIGHTEE survey data cubes covering the COSMOS field. The catalogue contains 293 sources in the redshift range of 0.004 < z < 0.093. In addition to HI masses and velocity widths, the catalogue includes optical through near-infrared photometry and inferred stellar masses and star-formation rates. The quantity of sources in the HI catalogue acquired through untargeted source finding is greatly influenced by the source finding methods used. This study therefore also provides a well-characterised expected completeness of the detected sample of galaxies based on their properties, inferred through a comparative study of different source fi

What carries the argument

Injection of simulated galaxies into real MeerKAT data cubes, binned narrowly by HI mass, inclination and distance, to benchmark recovery performance of PyBDSF, ProFound, SoFiA and LESHI.

Load-bearing premise

The performance of source finders on injected simulated galaxies divided into narrow bins of mass, inclination and distance accurately predicts their performance on real HI sources in the survey data cubes.

What would settle it

A direct count of how many real sources each of the four finders detects in the COSMOS data cubes, compared against the recovery fractions measured from the same finders on the injected simulated population.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The detected HI sample will have different completeness limits and selection biases depending on which source finder is applied.
  • Source-finding strategies for the full MIGHTEE survey and upcoming SKAO surveys can be chosen or combined using the measured completeness curves.
  • The catalogue supplies a well-characterised local HI sample whose selection function is now quantified as a function of galaxy properties.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Combining outputs from multiple finders may increase overall completeness while allowing users to track which sources are recovered by which algorithm.
  • The same injection-and-recovery framework could be used to test source finders on other HI survey fields before final catalogue construction.
  • Measurements of the local HI mass function derived from this catalogue will carry finder-dependent systematic uncertainties that can now be estimated.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 0 minor

Summary. The paper presents a catalogue of 293 HI sources extracted from MIGHTEE survey data cubes in the COSMOS field (redshift range 0.004 < z < 0.093), including HI masses, velocity widths, optical/NIR photometry, and derived stellar masses and star-formation rates. It also reports a comparative evaluation of four source-finding algorithms (PyBDSF, ProFound, SoFiA, and the new LESHI) by injecting simulated galaxies, binned by mass, inclination, and distance, into real MeerKAT data cubes to measure recovery rates and characterize expected completeness and detection biases for the MIGHTEE and future SKAO surveys.

Significance. The catalogue contributes a useful sample of low-redshift HI detections with multiwavelength ancillary data. The source-finder comparison, if the injected-model results transfer to real sources, supplies practical guidance on detection biases and strategies for large HI surveys. The use of real data cubes for the injection tests is a methodological strength that controls for survey-specific noise properties.

major comments (1)
  1. [Simulation injection and completeness estimation sections] The central claim that the simulation-based study 'provides a well-characterised expected completeness' and 'informs source finding strategies' for MIGHTEE and SKAO depends on the injected galaxies statistically representing real HI morphologies, kinematics, and environments. The manuscript reports per-bin recovery fractions from the narrow mass/inclination/distance bins but does not include any direct cross-check or validation against the properties of the 293 real detections (e.g., comparing recovered vs. non-recovered real sources or testing for residual cube artefacts). This assumption is load-bearing for the completeness interpretation and is not independently verified in the results or discussion sections.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for their constructive review and for recognising the value of the catalogue and the use of real data cubes in the injection tests. We address the major comment below.

read point-by-point responses
  1. Referee: [Simulation injection and completeness estimation sections] The central claim that the simulation-based study 'provides a well-characterised expected completeness' and 'informs source finding strategies' for MIGHTEE and SKAO depends on the injected galaxies statistically representing real HI morphologies, kinematics, and environments. The manuscript reports per-bin recovery fractions from the narrow mass/inclination/distance bins but does not include any direct cross-check or validation against the properties of the 293 real detections (e.g., comparing recovered vs. non-recovered real sources or testing for residual cube artefacts). This assumption is load-bearing for the completeness interpretation and is not independently verified in the results or discussion sections.

    Authors: We agree that the statistical representativeness of the injected galaxies is central to the interpretation. The simulations were constructed using empirical HI scaling relations (mass-size, mass-velocity width, and inclination distributions) drawn from the literature (primarily ALFALFA and other blind HI surveys) and placed at the observed redshifts and positions within the real MeerKAT cubes; this approach is intended to capture the dominant dependencies while incorporating the actual noise and artefact properties of the data. We acknowledge that an explicit consistency check against the real detections would strengthen the manuscript. In the revised version we will add a short subsection to the discussion that (i) compares the HI mass and velocity-width distributions of the 293 catalogue sources with the completeness-weighted expectations obtained by applying the measured recovery fractions to a model population, and (ii) summarises the visual-inspection criteria already used to reject residual artefacts. We note that a direct identification of 'non-recovered real sources' is not possible without an independent, complete census of all HI in the volume, which lies outside the scope of the present work. revision: partial

Circularity Check

0 steps flagged

No circularity: completeness follows directly from injection tests

full rationale

The paper reports a catalogue of 293 HI sources and measures recovery fractions by injecting simulated galaxies (binned by mass, inclination, distance) into the real MeerKAT cubes, then running PyBDSF, ProFound, SoFiA and LESHI. No equation or claim reduces a result to a fitted parameter defined from the same data; the completeness curves are direct empirical outputs of the injection procedure. No self-citation chain, uniqueness theorem, or ansatz is invoked to justify the method. The derivation chain is therefore self-contained against the survey cubes and the injected models.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

The paper introduces no new theoretical entities, axioms, or free parameters; it applies existing source-finding software to survey data and standard simulation-injection tests.

pith-pipeline@v0.9.1-grok · 5850 in / 1147 out tokens · 52367 ms · 2026-06-29T11:05:13.588756+00:00 · methodology

0 comments
read the original abstract

We present a catalogue of HI sources extracted from the MIGHTEE survey data cubes covering the COSMOS field. The catalogue contains 293 sources in the redshift range of 0.004 < z < 0.093. In addition to HI masses and velocity widths, the catalogue includes optical through near-infrared photometry and inferred stellar masses and star-formation rates. The quantity of sources in the HI catalogue acquired through untargeted source finding is greatly influenced by the source finding methods used. This study therefore also provides a well-characterised expected completeness of the detected sample of galaxies based on their properties, informing of any detection biases, inferred through a comparative study of different source finding algorithms. We have tested the performance of widely-used source finders: PyBDSF, ProFound and SoFiA, along with new source finder LESHI, focusing exclusively on HI source detection rather than source characterisation in the first instance. The source finders were tested by injecting a sample of simulated galaxies divided into narrow bins of mass, inclination and distance into a MeerKAT data cube. The results inform the source finding strategies for the MeerKAT International GigaHertz Tiered Extragalactic Exploration (MIGHTEE) survey, as well as upcoming SKAO surveys.

Figures

Figures reproduced from arXiv: 2605.28731 by Anastasia A. Ponomareva, Andreea A. V\u{a}r\u{a}\c{s}teanu, Ben Maughan, Catherine Hale, Hengxing Pan, Ian Heywood, Maarten Baes, Marcin Glowacki, Matt J. Jarvis, Michalina Maksymowicz-Maciata, Natasha Maddox, Seoyoung Lyla Jung, Sushma Kurapati, Tobias Westmeier, Tom G. Hardy.

Figure 1
Figure 1. Figure 1: Three-dimensional view of a data cube for the frequency range of 1.3685-1.3961 GHz (spanning 1055 channels) used for source injection in this work (left) and an extracted small volume centred on an example of a real source (right), both adapted from SAOImageDS9 application (Joye 2019) visualization. The data cube extract covers the COSMOS field with a mosaic of 15 pointings with a total area spanning ∼ 4 d… view at source ↗
Figure 2
Figure 2. Figure 2: Hi emission channel frames and spectra for a sample of simulated galaxies. The top two panels show galaxies of similar mass and distance, but varying inclination, the middle two panels show galaxies of similar inclination and distance, but varying mass, and the bottom two panels show galaxies of similar inclination and mass, but varying distance. Contours enclose the 3𝜎 level of the simulated Hi emission a… view at source ↗
Figure 3
Figure 3. Figure 3: Flowchart visualising each step of the LESHI script. See the main text for more details. MNRAS 000, 1–22 (2026) [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: An array of two-dimensional histograms colour-coded by the completeness for each bin of mass (x-axis) and cosine of inclination (y-axis) of injected galaxies, achieved by each of the LESHI, SoFiA, PyBDSF and ProFound source finders (accordingly first, second, third and fourth row of histograms) for each simulated luminosity distance of 50, 100, 150 and 350 Mpc (accordingly first, second, third and fourth c… view at source ↗
Figure 5
Figure 5. Figure 5: An array of two-dimensional histograms colour-coded by the completeness for each bin of mass (x-axis) and luminosity distance (y-axis) of injected galaxies, achieved by each of the LESHI, SoFiA, PyBDSF and ProFound source finders (accordingly first, second, third and fourth histogram). MNRAS 000, 1–22 (2026) [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Left panel: inclination averaged completeness achieved by the LESHI source finder vsthe injected Hi mass for different luminosity distances, with the best-fitting error function represented by the dashed lines. Right panels: best-fitting parameters characterising the fitted error function vs the luminosity distance. 6 7 8 9 10 11 log10(MHI/M ) 0.0 0.2 0.4 0.6 0.8 1.0 completeness DL = 50 Mpc DL = 100 Mpc D… view at source ↗
Figure 7
Figure 7. Figure 7: Completeness achieved by the LESHI source finder vs the injected Hi mass for different luminosity distances for face-on sources (represented by coloured diamonds) and edge-on sources (represented by coloured diamonds with black edge) and the best-fitting error functions represented by the dashed lines (for face-on sources) and solid lines (for edge-on sources). We have attempted to find a form of the compl… view at source ↗
Figure 8
Figure 8. Figure 8: showcases a comparison between the source finders: their average completeness (counting only the runs where at least one source finder had at least 5% completeness), average number of outputted false positives (characterising reliability), average frag￾mentation (defined as the average number of found sources output by the source finder that belonged to one injected source) and average runtime of the sourc… view at source ↗
Figure 9
Figure 9. Figure 9: shows the memory usage of each source finder run on the 89.3 GB data cube. It is worth noting that since PyBDSF, Pro￾Found and LESHI work on the channel images separately (with LESHI working on the integrated images), their memory footprints are mostly influenced by the number of parallel threads used, rather than the frequency range of the data cube. Their memory footprint can be therefore adjusted based … view at source ↗
Figure 10
Figure 10. Figure 10: Venn diagram of injected sources that were found by each source finder, representing the overlap between the different source finders (11328 sources were found by at least one source finder out of total of 24000 injected sources). MNRAS 000, 1–22 (2026) [PITH_FULL_IMAGE:figures/full_fig_p009_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Comparison between Hi mass (left panel) and velocity width (right panel) of measured emission (y-axes) and injected emission (x-axes), colour-coded by the observed Hi flux. compare the measured Hi masses to the true ones. As can be seen on the left panel of [PITH_FULL_IMAGE:figures/full_fig_p011_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Top left: Hi moment-0 map with contours over-plotted on top for an example galaxy from our sample, marking column densities of 3.0, 4.3, 5.7, 7.0, 8.4, 9.7 x 1020 cm−2 with the size of the synthesised beam marked by a dashed circle. The row of 9 channel maps at the bottom show the Hi emission for the individual channels, which are marked by points on the spectrum below with their colour matching the borde… view at source ↗
Figure 13
Figure 13. Figure 13: Comparison between stellar mass (left panel) and star-formation rate (right panel), determined through SED fitting in this work (x-axes) and DESI (y-axes) for crossmatched galaxies present in both catalogues. In [PITH_FULL_IMAGE:figures/full_fig_p014_13.png] view at source ↗
Figure 15
Figure 15. Figure 15: Hi mass of the catalogue galaxies plotted against their luminosity distance represented by white points, with most edge-on sources plotted in red and sample distribution histograms plotted on top and on the right of the figure. The theoretical completeness map calculated using Eq. 1 is plotted as the background in colour with the detection threshold (where the expected completeness drops to zero given by … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

3 extracted references · 2 canonical work pages · 2 internal anchors

  1. [1]

    AdamsN.J.,BowlerR.A.A.,JarvisM.J.,HäußlerB.,LagosC.D.P.,2021, MNRAS, 506, 4933 Adams E. A. K., et al., 2022, A&A, 667, A38 Aihara H., et al., 2019, PASJ, 71, 114 Alam S., et al., 2015, ApJS, 219, 12 Astropy Collaboration et al., 2022, ApJ, 935, 167 Barkai J. A., Verheijen M. A. W., Talavera E., Wilkinson M. H. F., 2023, A&A, 670, A55 Barnes D. G., et al.,...

  2. [2]

    Looking At the Distant Universe with the MeerKAT Array (LADUMA)

    pp 496–499 (arXiv:1109.5605), doi:10.1017/S1743921312009702 Hotan A. W., et al., 2021, Publ. Astron. Soc. Australia, 38, e009 Hunter J. D., 2007, Computing in Science and Engineering, 9, 90 Jaffé Y. L., Poggianti B. M., Verheijen M. A. W., Deshev B. Z., van Gorkom J. H., 2013, MNRAS, 431, 2111 Jarvis M., et al., 2016, in MeerKAT Science: On the Pathway to...

  3. [3]

    low confidence

    for the SoFiA, PyBDSF and ProFound source finders, derived analogously to Section 3.2. Figures B1, B2 and B3 contain plots showing the fitted completeness function analogously to Fig. 6 and Fig. 7 for the SoFiA, PyBDSF and ProFound source finders. Qualitatively, the fitted completeness functions for all of the sourcefinders are very similar. Fitted𝑀𝑚𝑖𝑛 pa...