{"id":"f0eddc65-3686-44a6-9379-0e52078744ec","arxiv_id":"2501.09662","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new 2.1 GHz survey of 27 fields yields 2,294 radio sources, with source counts to 150 microjansky and a far-infrared-radio correlation for 394 galaxies.","lead":"This paper presents the SHORES survey, a new 2.1 GHz radio catalog of 2,294 sources detected with the Australia Telescope Compact Array across 27 fields in the Herschel-ATLAS Southern Galactic Field. It provides source counts down to 150 microjansky and a radio-infrared correlation for 394 matched galaxies, useful for studying star formation and active galactic nuclei in faint radio populations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 95% completeness/reliability claims are based on BLOBCAT simulations on a single field and exclude the manual visual-inspection stage, so they may not hold for the final published catalog.","rationale":"The paper has genuine strengths: the primary beam correction is measured with a dedicated calibrator, the detection pipeline is cross-checked with NVSS and RACS, multiple source-extraction tools are compared, and the simulations use realistic UV coverage. I do not dispute the existence of the 2294 detections or the broad shape of the source counts. The load-bearing gap is that the statistical statements '95% complete' and '95% reliable' are attached to the final catalog, while the simulation validating them excludes the subjective visual-inspection stage and is run on one field. This is exactly the gap the reader identified as the weakest assumption, and it matters most for the faintest source-count bin, where the completeness correction is largest. The concern is not that the catalog is wrong, but that its published completeness/reliability error bars are not yet established for the final product. A CONDITIONAL verdict remains appropriate: the catalog can be used, but the completeness-corrected faint counts should not be taken at face value until the visual-inspection loss and field-to-field variation are explicitly simulated.","tokens_in":29381,"tokens_out":9797,"duration_ms":100614,"concrete_test":"Redo the §4.5 injection/recovery experiment on all 27 shallow fields (or at least on fields bracketing the noise range, e.g. s2242-3241 and s2344-3039) and pass the output through the full §5 visual-inspection procedure with an independent, blinded inspector. For each flux bin, compute the retention fraction R(S) = N_final / N_BLOBCAT for injected sources and multiply the published completeness C(S) by R(S). If the completeness at 0.5 mJy falls below 95%, or if the 0.15–0.68 mJy source-count bin shifts by more than its quoted Poisson error, the catalog-level completeness correction and the faint-end counts must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central catalog claims — 95% completeness above 0.5 mJy and 95% reliability at SNR≥4.5 for the final 2294-source catalog — rest on §4.5 simulations that validate BLOBCAT extraction on one representative field, s2242-3241, selected in §4.4 by pixel-histogram similarity. Those simulations do not include the manual visual-inspection procedure described in §5, which pruned 2649 SNR≥4.5 BLOBCAT detections down to 2294 published sources. Since completeness is defined as N_det/N_inj, any genuine source removed by inspection lowers the true completeness; the paper gives no breakdown of the 355 removed objects into artifacts versus real sources, so the size and even the sign of the resulting bias are unquantified. The reliability estimate derived from negative maps likewise ignores the inspection step, so it certifies the extractor rather than the published catalog. Field-to-field noise varies by roughly 35% (σB from 31 to 46 μJy in Table 5), yet no injection tests are run on other fields; the representative-field choice is not validated at the faintest fluxes, where the source-count correction 1/C(S) is largest. Because the faint-end counts at 150 μJy are corrected with this C(S), an overestimated completeness would bias the quoted counts and weaken the claimed agreement with Mancuso et al. (2017).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the SHORES survey, a 2.1 GHz ATCA pencil-beam survey of 27 shallow fields covering ~26 deg² within the H-ATLAS Southern Galactic Field. The authors describe the observations, calibration, imaging, source extraction with three tools (BLOBCAT, AEGEAN, PySE), a custom primary-beam correction measured from dedicated observations, and a final catalog of 2294 sources detected with BLOBCAT at SNR ≥ 4.5. They report 95% completeness above 0.5 mJy and 95% reliability at SNR ≥ 4.5 from simulations on a single representative field, derive 2.1 GHz Euclidean source counts down to 150 μJy after completeness and effective-area corrections, and compare the counts with the Mancuso et al. (2017) model and other surveys. They also cross-match with H-ATLAS to compute the FIR-radio correlation parameter q_FIR for 394 sources and update the status of 20 candidate lensed galaxies.","tokens_in":29696,"tokens_out":4249,"duration_ms":41618,"significance":"If the catalog-level completeness and reliability claims hold, SHORES would provide a valuable new data set at a poorly explored frequency (2.1 GHz) and flux range (sub-mJy to ~1 Jy), with a public catalog, independent cross-checks against NVSS and RACS, and a directly measured ATCA primary beam. The paper's strengths include the multi-tool extraction comparison, the explicit validation against external surveys, the release of the full catalog as supplementary material, and the transparent description of the custom primary-beam and noise-profile corrections. The main significance rests on the faint-end source counts, which are used to test models of the sub-mJy radio population; therefore the completeness correction applied to those counts is the key load-bearing element.","major_comments":[{"comment":"The completeness and reliability simulations validate BLOBCAT extraction on one representative field, but the published catalog is the product of BLOBCAT plus two stages of visual inspection described in §5, which reduced the 2649 SNR ≥ 4.5 detections to the 2294 published sources. The simulations do not include these inspection steps; if any genuine source was removed during inspection, the true completeness of the final catalog is lower than the quoted 95%. The paper does not quantify how many of the 355 removed objects were artifacts versus real sources contaminated by bright neighbors. Because the source counts in §6 are corrected with 1/C(S) derived from these simulations, this issue directly affects the validity of the faint-end counts. Please extend the completeness/reliability analysis to the full catalog-production pipeline, or provide an explicit bound on the number of real sources removed by the visual inspections.","section":"§4.5 and §5"},{"comment":"The completeness simulation is run only on field s2242-3241, chosen because its pixel-histogram distribution is closest to the median of the 27 fields (Figure 10). Table 5 shows that the per-field rms noise varies by about 35% (σ_B from 31 to 46 μJy), and the completeness correction 1/C(S) is largest at the faint fluxes that dominate the 150 μJy counts. The representativeness of s2242-3241 is not validated at the faint end, where source density, confusion, and calibration artifacts could differ between fields. Please run injection tests on at least a few fields spanning the observed σ_B range, or conservatively inflate the completeness uncertainty used in the source-count error budget.","section":"§4.4 and §4.5"},{"comment":"The completeness simulations inject sources drawn from the Mancuso et al. (2017) model, and the same model is used as the reference for the derived 2.1 GHz counts in Figure 20. Since several co-authors of the present paper are also authors of Mancuso et al. (2017), the agreement of the corrected counts with that model is not fully independent. The circularity is not absolute—the completeness correction is not forced to fit the model—but a mismatch between the true source population and the injection model would bias C(S) and hence the corrected counts. A concrete test would be to repeat the completeness simulations with an independent source-count model (e.g., one of the other literature counts shown in Figure 20) and report the change in the corrected counts; this should be added.","section":"§4.5 and §6"},{"comment":"The completeness simulations inject only point sources, while the final catalog contains 376 extended and 62 multi-component sources (§5). Detection efficiency for extended sources at fixed peak SNR is generally lower than for point sources, so the quoted 95% completeness above 0.5 mJy may not apply to the extended subset. Please state explicitly whether the completeness claim applies to the full catalog or only to point sources, and if the latter, provide separate completeness estimates for the extended and multi-component populations.","section":"§4.5 and §5"}],"minor_comments":[{"comment":"Typo: \"coverign\" should be \"covering\" in the description of leakage calibrator observations.","section":"§3.1"},{"comment":"The header of Table 1 reads \"T able 1\"; it should be \"Table 1\".","section":"Table 1"},{"comment":"The title and abstract contain \"A TLAS\" and \"A TCA\", which should be \"ATLAS\" and \"ATCA\".","section":"Abstract and title"},{"comment":"The target name \"HATLASJ005132.8-01848\" appears incomplete; it likely should be \"HATLASJ005132.8-301848\" or similar.","section":"§2"},{"comment":"The number of BLOBCAT SNR ≥ 3 detections is inconsistent: Table 2 lists 13662, while §4.1 and §5 state 13688. Please unify these values.","section":"§4.1 and §5"},{"comment":"The text refers to \"2297 SHORES sources\" in the counts paragraph, but the catalog has 2294 sources; this appears to be a typo.","section":"§6"},{"comment":"The left panel label \"S=467.7 Jy\" should presumably be \"467.7 μJy\" (or 0.468 mJy), consistent with the 95% completeness threshold quoted in the text.","section":"Figure 15"},{"comment":"The column header \"S2.5 2.1GHzdN/dS\" is garbled; it should read \"S^2.5 dN/dS\" with the units correctly formatted.","section":"Table 7"}],"recommendation":"major_revision","confidential_remarks":"The author overlap with Mancuso et al. (2017) is worth watching in revision: the model comparison in Figure 20 is presented as an external check, but the injection model for completeness is the same model. The requested alternative-model test should resolve this. The paper is otherwise well within the scope of a radio-astronomy survey paper, and the catalog release is a positive feature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this paper delivers what it promises — a new 2.1 GHz catalog of 2294 sources over 26 square degrees, source counts to 150 uJy, and a FIR-radio correlation for 394 H-ATLAS matches — and the data release looks genuinely useful. The strongest part is the empirically measured ATCA primary beam out to 3 FOV, which is a real advance over the truncated Miriad/ATCA polynomial and is validated against sub-band spectral fits.\n\nThe analysis uses standard tools (BLOBCAT, AEGEAN, PySE) and is well cross-checked: three extractors, comparisons to NVSS and RACS, and the flux densities look consistent. The catalog is released as additional material. That is reproducible work and should count as such.\n\nThe soft spot is the completeness/reliability claim. The 95% numbers come from injections in one representative field (s2242-3241), chosen by pixel-histogram similarity, and from negative-map reliability. Two things bother me. First, field-to-field noise varies by about 35%, so a single-field simulation is a stretch unless it is validated on at least a couple of other fields. Second — and more important — the simulation measures BLOBCAT's ability to recover injected sources on maps, but the published catalog went through a manual visual inspection that pruned 2649 SNR>=4.5 detections to 2294. Completeness is defined as N_det/N_inj for the extractor; the human pruning is not in the loop. If any of the 355 removed objects are real sources, the true completeness of the final catalog is lower. The paper gives no breakdown. So I would treat the 95% claim as 'extractor completeness on one representative field,' not 'final catalog completeness across 27 fields.'\n\nThe mixed use of Mancuso et al. (2017) for both injecting the simulation and comparing the final counts is a mild circularity, not a fatal one — the correction isn't fit to the model — but it should be flagged. There are also small internal inconsistencies (2697 vs 2649 BLOBCAT detections at SNR>=4.5; 2297 vs 2294 sources in the counts; 414 vs 394 in the FIR-radio sample; 497.5 vs 467.7 uJy for 95% completeness). Minor, but worth cleaning up.\n\nWho is this for: anyone working on sub-mJy radio source populations, CMB foregrounds, or FIR-radio correlation. It deserves a serious referee. I would send it to review, with a request to either include the visual inspection in the completeness estimate or limit the claim to the extractor stage, and to justify the single-field extrapolation.","headline":"A useful new 2.1 GHz catalog with a genuinely measured ATCA primary beam to 3 FOV; the 95% completeness claim is provisional until the visual-inspection step is folded into the simulations.","tokens_in":30256,"tokens_out":2916,"would_cite":true,"duration_ms":29995,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SHORES builds a 2294-source radio catalog at 2.1 GHz and pushes source counts to 150 microjansky.","keywords":["Extragalactic radio sources","Radio source catalogs","Radio interferometry","Surveys","Source counts","FIR-radio correlation","ATCA","H-ATLAS"],"falsifier":"Run the same completeness and reliability simulations on a different SHORES shallow field, or on a stack of several fields, injecting sources drawn from a source-count model other than Mancuso et al. (2017); if the recovered 95% completeness flux rises above 0.5 mJy or the false-detection rate at SNR 4.5 exceeds 5%, the quoted catalog statistics would not transfer to the whole survey.","tokens_in":29210,"feed_emoji":"📡","tokens_out":8125,"duration_ms":74938,"temperature":0.7,"pith_summary":"The paper presents SHORES, a 2.1 GHz ATCA survey of 27 shallow fields centered on candidate lensed galaxies inside the Herschel-ATLAS Southern Galactic Field, and argues that serendipitous observations of targets chosen for other reasons can be turned into a statistically useful survey of the faint radio sky. It claims a catalog of 2294 sources detected at signal-to-noise ratio $\\ge 4.5$ that is 95% complete above 0.5 mJy and 95% reliable, with Euclidean source counts measured down to 150 microjansky that agree with earlier work and with the Mancuso et al. (2017) model. The survey covers about 26 square degrees with sensitivity that peaks at $\\lesssim 33\\,\\mu$Jy rms toward each pointing center, and reaches a resolution of 3.2 by 7.2 arcsec, at which 81% of sources are unresolved. A measured correction for the ATCA primary beam lets the authors use each field out to three times the nominal field of view, increasing the bright-source statistics. The overlap with H-ATLAS supplies FIR counterparts for 457 sources and yields the FIR-radio correlation for 394, extending the earlier lensed-sample analysis to ordinary sub-mJy sources.","feed_headline":"Radio survey counts 2294 sources down to 150 microjansky","feed_subtitle":"27 ATCA pencil-beam fields in H-ATLAS give a 95% complete catalog and trace the sub-mJy radio sky at 2.1 GHz.","key_machinery":"The central machinery is a multiple pencil-beam survey geometry combined with three tools: a directly measured ATCA primary-beam profile (an 8th-order polynomial fit to PKSJ0537-441 observations across the four 512 MHz sub-bands) that extends usable imaging to about three FWHM from each pointing; BLOBCAT source extraction with corrections for peak bias, clean bias, and bandwidth smearing, cross-checked against AEGEAN and PySE; and simulated source injections on the representative field s2242-3241 that define the effective area, completeness, and the SNR $\\ge 4.5$ reliability threshold. Together these determine which faint sources are real and how the surveyed area depends on flux density, which is what the source counts and FIR-radio correlation rest on.","core_discovery":"On its own terms, the discovery is that the sub-mJy radio sky at ~2 GHz can be characterized from a set of pointed, serendipitous pencil-beam fields rather than from a dedicated mosaic. By imaging each pointing to three times the ATCA primary-beam FWHM and applying a custom-measured beam correction, the paper obtains a 26-square-degree survey whose effective area grows with flux density; 2294 sources are cataloged at SNR $\\ge 4.5$, 81% of them unresolved at 3.2 x 7.2 arcsec. Completeness and reliability, set by simulations on the representative field s2242-3241, reach 95% above 0.5 mJy and at SNR $\\ge 4.5$, respectively. The Euclidean source counts at 2.1 GHz extend to 150 microjansky and agree with previous determinations and the Mancuso et al. (2017) model. Cross-matching with H-ATLAS gives 457 counterparts, 394 with enough FIR data for the $q_{\\rm FIR}$ FIR-radio correlation, and 20 of the 27 central lensed candidates are detected, doubling the number with radio counterparts from the pilot campaign.","pith_inferences":["If the primary-beam profile is stable across epochs, a similar correction could be applied to archival ATCA observations, turning many historical pointed observations into wide-field source-count measurements.","The deep SHORES fields, reaching about 8 microjansky, should test whether the Mancuso et al. (2017) model continues to hold below the 150 microjansky limit of this paper.","The polarization calibration described but not analyzed here could provide a direct test of AGN versus star-formation origin for the sub-mJy sources, with relevance to CMB foreground studies."],"forward_implications":["Sub-mJy 2.1 GHz source counts are now measured to 150 microjansky over about 26 square degrees, reducing cosmic variance by a factor of $\\sqrt{27}$ relative to a single pencil-beam field.","The 2294-source catalog at 3.2 by 7.2 arcsec resolution, with 81% of sources unresolved, provides a faint radio-source sample for multi-wavelength follow-up and future SKA-era surveys.","The FIR-radio correlation is extended to ordinary sub-mJy sources, not just lensed candidates, using H-ATLAS counterparts and photometric redshifts for 394 sources.","Twenty of the 27 central lensed candidates now have radio detections, roughly doubling the number from the pilot campaign and supporting a star-formation origin for their radio emission.","The measured primary beam correction makes data usable to three times the nominal field of view, improving bright-source statistics and reducing sampling variance in the counts."],"supporting_citations":[{"why":"Supplies the 2.1 GHz source-count model used to inject simulated sources and as the reference curve for the measured counts.","marker":"Mancuso et al. (2017)"},{"why":"Pilot ATCA campaign on the same lensed candidates; SHORES re-observes these fields and roughly doubles the number with radio detections.","marker":"Giulietti et al. (2022)"},{"why":"BLOBCAT is the source extraction tool whose SNR > 4.5 detections define the SHORES catalog and whose flux corrections are adopted.","marker":"Hales et al. (2012)"},{"why":"Selected the candidate lensed galaxies that determine the pointing centers of the SHORES fields.","marker":"Negrello & Herschel-Atlas Team (2017)"},{"why":"H-ATLAS supplies the FIR photometry and photometric redshifts used for the FIR-radio correlation.","marker":"Eales et al. (2010)"},{"why":"AEGEAN is one of the independent detection methods used to validate the BLOBCAT detections.","marker":"Hancock et al. (2018)"},{"why":"Miriad is the calibration and self-calibration package used for all SHORES data reduction.","marker":"Sault et al. (1995)"},{"why":"WSCLEAN is the wide-field imager used to produce the final Stokes I maps.","marker":"Offringa et al. (2014)"},{"why":"NVSS provides the 1.4 GHz flux comparison that checks the SHORES flux scale.","marker":"Condon et al. (1998)"},{"why":"RACS-mid provides a similar-frequency flux comparison confirming the SHORES flux densities.","marker":"McConnell et al. (2020)"}],"fun_headline_variants":["Pencil-beam survey maps sub-mJy radio sky at 2.1 GHz","2294 radio sources from 27 H-ATLAS pencil beams","SHORES survey: sub-mJy radio counts from 26 sq deg","95% complete catalog: 2294 sources at 2.1 GHz","Radio pencil beams reach 150 microjansky in H-ATLAS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulations that set the 95% completeness and reliability assume that the single field s2242-3241 represents the noise and source population of all 27 SHORES fields, and that the injected sources follow the Mancuso et al. (2017) model; the visual inspection step that builds the final catalog is not included in those simulations.","fun_headline_variants_meta":{"raw":{"variants":["Pencil-beam survey maps sub-mJy radio sky at 2.1 GHz","2294 radio sources from 27 H-ATLAS pencil beams","SHORES survey: sub-mJy radio counts from 26 sq deg","95% complete catalog: 2294 sources at 2.1 GHz","Radio pencil beams reach 150 microjansky in H-ATLAS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000698,"raw_usage":{"total_tokens":3249,"prompt_tokens":1136,"completion_tokens":2113,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":752,"completion_tokens_details":{"reasoning_tokens":2015}},"tokens_in":752,"tokens_out":2113,"duration_ms":14573,"temperature":1.0,"reasoning_tokens":2015,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:46:21.044041+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same completeness and reliability simulations on a different SHORES shallow field, or on a stack of several fields, injecting sources drawn from a source-count model other than Mancuso et al. (2017); if the recovered 95% completeness flux rises above 0.5 mJy or the false-detection rate at SNR 4.5 exceeds 5%, the quoted catalog statistics would not transfer to the whole survey.","supporting_citations":[{"cited_title":"2022, , 511, 1408, 10.1093/mnras/stac145","cited_arxiv_id":null,"evidence_quote":"Pilot ATCA campaign on the same lensed candidates; SHORES re-observes these fields and roughly doubles the number with radio detections."},{"cited_title":"2017, Submillimeter Array Newsletter, 24, 10","cited_arxiv_id":null,"evidence_quote":"Selected the candidate lensed galaxies that determine the pointing centers of the SHORES fields."}],"review_version":1}