{"id":"0ee861ee-339f-4a32-aedd-a5d6dfe92354","arxiv_id":"2505.11453","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Deriving radio luminosity functions by statistically assigning redshifts from similar sources reproduces published results out to z≈0.5 and broadly out to higher redshifts.","lead":"Radio surveys detect millions of galaxies too faint for telescopes to measure their distances directly. This paper tests a simple statistical shortcut that estimates distances from radio brightness and compares the resulting galaxy census with earlier measurements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"High-z validation is semi-circular: GEM III redshifts are drawn from Smolčić et al. (2017), the same work used as the LF benchmark, so the agreement there does not independently test the method.","rationale":"The reader's weakest assumption correctly identifies the representativeness of the GEM I redshift template as the load-bearing premise, and the strongest claim centers on agreement with published LFs. My stress-test adds specificity about why the high-redshift comparison cannot carry that weight: the Smolčić et al. (2017) redshift distribution is used as the input prior for GEM III and then compared against Smolčić et al. (2017) as the benchmark, making the agreement partly self-confirming. The low-redshift Mauch & Sadler (2007) check is genuine and supports the method for z<=0.5 for GEM I+II, but GEM III dominates the sample size and controls the high-z LF. The post hoc AGN relabelling at L > 1e23.5 further weakens the separate SFG/AGN claim. I agree with the CONDITIONAL verdict: the method is promising and the low-z validation is real, but the headline 'match well with measured LFs' needs an independent high-z test and a pre-specified classification rule before it is fully supported. I found no internal inconsistency in the 1/Vmax implementation or the cosmology/K-correction choices worth flagging; the issue is external validation, not arithmetic. My proposed concrete test is a clean way to settle whether the high-z agreement survives an independent prior.","tokens_in":23424,"tokens_out":3024,"duration_ms":25472,"concrete_test":"Perform a fully external high-redshift validation using a redshift distribution and an RLF independent of Smolčić et al. (2017). For example, draw GEM III redshifts from the VLA-COSMOS photometric redshift distribution of Novak et al. (2017) or an equivalent independent COSMOS-based sample, keep the AGN/SFG classification fixed exactly as defined from GEM I (BPT-based plus the L888 > 1e23.5 radio-loudness criterion applied identically in all bins), and recompute Figure 15. If the derived GEM III LF agrees with the independent benchmark to within the plotted Poisson errors across all z>0.5 bins, the high-z claim is independently supported. If the LF shifts by more than ~0.5 dex in any bin, the current agreement is an artifact of using the benchmark as the prior.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that the statistical redshift assignment reproduces measured RLFs rests on two comparisons. The low-z check (GEM I+II against Mauch & Sadler 2007, z<=0.5) is a genuine external validation: GEM I redshifts are spectroscopic, GEM II redshifts are drawn from GEM I templates by (S888, mr) bin, and the benchmark is an independent local LF. That check works and is not the problem. The high-z check is not independent. In §6.1 the GEM III sources (28,821 of 39,812, about 72% of the sample) are assigned redshifts by sampling the Smolčić et al. (2017) redshift distribution, and in §7/Figure 15 the resulting LF is compared against the Smolčić et al. (2017) LF. Luminosity is computed from exactly those redshifts, so the comparison largely verifies that the 1/Vmax machinery and binning were coded consistently, not that the statistical assignment is physically representative. The paper's own §7.1 strengthens this concern: even a uniform distribution in 0.5<z<6 reproduces the Smolčić et al. (2017) LF to within ~0.6 dex, showing the high-z bins are insensitive to the input redshift distribution. The AGN/SFG split is also adjusted post hoc in §6.1 (all sources with L888 > 1e23.5 are relabelled AGN), adding further freedom when comparing against the benchmark. The load-bearing condition is that GEM I's (S888, mr) redshift distribution is representative of GEM II, and that some external high-z distribution is appropriate for GEM III; the first is supported only by the Mauch & Sadler agreement, the second is untested because the high-z benchmark was used to build the input.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses EMU early-science 887.5 MHz data in the GAMA G23 field to construct three samples: GEM I (6,425 radio sources with spectroscopy and r-band magnitudes), GEM II (4,566 with r-band magnitudes but no spectroscopy), and GEM III (28,821 with no optical counterpart). The authors assign statistical redshifts to GEM II by fitting Gaussians to GEM I redshift distributions in (log S888, mr) bins, and to GEM III first from GEM I flux-only templates and then, in §6.1, from a 'realistic' high-redshift distribution taken from Smolčić et al. (2017) or from a uniform distribution in 0.5<z<6. AGN/SFG classifications are transferred from GEM I via flux-bin probability mass functions and subsequently modified with a radio-luminosity threshold. RLFs are computed with the 1/Vmax method including completeness and area corrections, and are compared with Mauch & Sadler (2007) at low redshift and Smolčić et al. (2017) at high redshift. The central claim is that the statistical redshift assignment reproduces measured RLFs, including separate SFG and AGN LFs.","tokens_in":23805,"tokens_out":5727,"duration_ms":57844,"significance":"If the method is valid, it would offer a practical route to RLF measurements for future wide-area radio surveys where spectroscopic completeness is low and multiwavelength photometry is sparse. The low-redshift comparison against Mauch & Sadler (2007) is a genuine external test and works, and the multiple realisations provide a sensible treatment of statistical uncertainties. The uniform-redshift experiment in §7.1 is also a useful robustness check. However, the high-redshift validation is largely circular because the input redshift distribution and the benchmark LF come from the same Smolčić et al. (2017) work, and the separate SFG/AGN comparison is partly adjusted post hoc with a luminosity threshold. These issues need to be addressed before the central claim can be accepted.","major_comments":[{"comment":"The high-redshift comparison is not independent: the GEM III sources (28,821 of 39,812) are assigned redshifts by sampling the Smolčić et al. (2017) redshift distribution (§6.1, Fig. 13), and the resulting LFs are then compared with Smolčić et al. (2017) LFs (§7, Fig. 15). Because L888 is computed from exactly those assigned redshifts, the agreement largely verifies that the 1/Vmax machinery and binning are self-consistent, rather than validating that the statistical redshift assignment is representative. Please add an external high-redshift test, for example using sources with secure spectroscopic or reliable photometric redshifts from a deep survey such as VLA-COSMOS, or explicitly present the high-redshift comparison as a consistency check rather than as validation.","section":"§6.1 and §7, Fig. 15"},{"comment":"The AGN/SFG separation is adjusted post hoc: after the BPT-based PMF transfer, all sources with L888 > 1e23.5 W Hz–1 are relabelled as AGN in §6.1. The subsequent agreement of the separate SFG and AGN LFs with Mauch & Sadler (2007) in Fig. 14 is therefore partly built in by this criterion rather than being a prediction of the method. Please show the SFG and AGN LFs both before and after this relabelling and quantify the fraction of sources moved, so the reader can assess how much of the improvement comes from the redshift modelling and how much from the luminosity threshold.","section":"§6.1, Fig. 14"},{"comment":"The uniform-distribution experiment, while useful, actually weakens the high-redshift validation. Even with a flat input distribution in 0.5<z<6, the derived LFs agree with Smolčić et al. (2017) to within about 0.6 dex at the highest redshift bin. This indicates that the high-redshift RLF bins are insensitive to the input redshift distribution, so the agreement in Fig. 15 does not establish that the Smolčić-based assignment is physically realistic. The text should state this limitation explicitly when claiming that the LFs match well with measured LFs.","section":"§7.1"},{"comment":"The representativeness of the GEM I templates is only indirectly tested. The paper assumes that radio sources in the same (log S888, mr) bin share the GEM I redshift distribution, but GEM I is defined by the GAMA spectroscopic limit (mr<19.8), and no direct comparison is made between modelled GEM II redshifts and independent spectroscopic redshifts outside GEM I. The low-redshift agreement with Mauch & Sadler (2007) is promising, but a dedicated test, for example using withheld spectroscopic redshifts or an external overlapping survey, would substantially strengthen the method.","section":"§3.1"}],"minor_comments":[{"comment":"There is a typo in the sentence 'in the absence of redshfits' near the end of §7.1; it should read 'redshifts'.","section":"§7.1"},{"comment":"The citation 'Hopkins et al., submitted' appears in the text but is not included in the reference list; please add it or replace it with a published reference.","section":"§2.2"},{"comment":"The description in §3.1 of how mr and S888 vary across the panels of Figures 3 and 4 is confusing; please label the axes of the individual histograms or add a schematic so the bin layout is immediately clear to the reader.","section":"Figures 3 and 4"}],"recommendation":"major_revision","confidential_remarks":"The paper has a genuine external low-redshift test and a clear presentation, but the high-redshift and SFG/AGN claims need substantial revision: the high-redshift comparison is semi-circular, and the AGN/SFG agreement is partly produced by a post-hoc luminosity threshold. If the authors provide an independent high-redshift validation or explicitly narrow the claims, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The stress-test note is right: the low-redshift half of this paper is a legitimate external test that works, and the high-redshift half is not an independent validation. GEM I+II, with redshifts assigned from GAMA spectroscopic priors binned by (S888, mr), reproduce Mauch & Sadler (2007) out to z≈0.5. GEM III redshifts, by contrast, are drawn from Smolčić et al. (2017) and then compared with Smolčić et al. (2017) LFs. Agreement there mostly checks that the 1/Vmax code is consistent, not that the redshift assignment is right. The paper's own Section 7.1 makes this clear by showing a uniform 0.5<z<6 distribution also lands within about 0.6 dex of Smolčić.\n\nWhat is genuinely useful: the method is simple, cheap, clearly described, and aimed at a real problem, namely that most EMU sources have no spectroscopy and no optical counterpart. The GEM I template construction, the Gaussian fits per bin, and the multiple realizations are transparent. The authors do not oversell individual redshifts; they position the work as population statistics, which is the right framing. The citation pattern is fine—they engage the relevant statistical photo-z and clustering-z literature. The faint-end upturn at z<0.1 is an interesting byproduct, though it should be treated cautiously because it comes from repeated sampling of assigned redshifts.\n\nSoft spots in proportion. The high-z circularity is the main one. Since 72 percent of the sample is GEM III, the headline claim that the LFs match well rests largely on that comparison. The AGN/SFG split is also adjusted post hoc: sources with L888 > 10^23.5 W/Hz are relabeled AGN after seeing that the BPT-based classification underestimates the AGN LF. That may be physically justified, but it was not pre-specified, so the comparison with Smolčić is not clean. Agreement is assessed visually; there is no goodness-of-fit statistic. The quoted error bars are Poisson plus realization scatter; they do not include the systematic uncertainty in the redshift prior, which is the dominant term for GEM III. The representativeness assumption for GEM II is supported only by the low-z Mauch & Sadler agreement; there is no independent check for optically blank sources.\n\nBottom line: this is a promising proof-of-concept, not a validated measurement at z>0.5. The right reader is someone building population statistics for EMU or LoTSS-scale surveys without full photometric coverage. It deserves a serious referee, but the referee should ask for an independent high-z test, for example a subsample with photometric or spectroscopic redshifts not used to build the prior, and a pre-specified AGN classification. With those, the method could be genuinely useful. My recommendation: send to peer review, conditional on revision.","headline":"Low-z validation is genuine; high-z agreement is largely built into the input redshift distribution, so treat this as a promising proof-of-concept rather than a measured constraint above z≈0.5.","tokens_in":24483,"tokens_out":3526,"would_cite":true,"duration_ms":32967,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A statistical redshift-assignment scheme using only radio flux and r-band magnitude reconstructs the 888 MHz radio luminosity functions of star-forming galaxies and AGN, matching published surveys from z≈0 to z≈5.5.","keywords":["radio luminosity functions","statistical redshift estimation","radio AGN","star-forming galaxies","EMU early science","GAMA","1/Vmax method","radio continuum surveys"],"falsifier":"A reader could test the method by taking a complete radio survey with full spectroscopic redshifts, hiding a random subset of those redshifts, running the same bin-and-draw procedure, and comparing the reconstructed luminosity functions with the true ones; a significant mismatch in any redshift bin, particularly above z≈0.5, would show the statistical assignments are not representative. The same test can be run on the optically blank sample once follow-up spectroscopy reaches those sources.","tokens_in":23207,"feed_emoji":"📡","tokens_out":8720,"duration_ms":82915,"temperature":0.7,"pith_summary":"This paper argues that the radio luminosity function of a large radio-selected sample can be measured without individual redshifts for most sources. It takes a radio catalogue of 39,812 sources, of which only 6,425 have spectroscopic redshifts, and assigns the missing redshifts statistically from the redshift distributions of spectroscopically identified sources in the same bins of radio flux density and r-band magnitude; sources with no optical counterpart are assigned redshifts from a high-redshift distribution. Repeating the assignment one hundred times and building 1/Vmax luminosity functions, the paper finds that the total, star-forming, and AGN luminosity functions follow published 1.4 GHz luminosity functions from z≈0 to z≈0.5 for sources with optical counterparts, and out to z≈5.5 once the optically blank sources are redistributed. The point of the exercise is that population-level statistics, not individual redshifts, may be enough to characterise how the radio source population evolves.","feed_headline":"Statistical redshifts reproduce radio luminosity functions","feed_subtitle":"Sparse spectroscopy still recovers star-forming and AGN radio counts from z=0 to z≈5.5.","key_machinery":"The load-bearing device is distributional redshift assignment. Radio sources are divided into small bins of observed radio flux density and, when available, r-band magnitude; for each bin the redshift distribution of the spectroscopically complete subsample is approximated by a single Gaussian, and sources missing redshifts are assigned redshifts drawn at random from that fitted Gaussian. Two refinements carry the high-redshift side: optically blank sources are redistributed using a redshift distribution inferred from a published 1.4 GHz luminosity function extending to $z\\approx 6$, and the statistical AGN/SFG classification from BPT-diagram fractions is supplemented by a radio-luminosity cut. One hundred Monte Carlo realisations of the assignments yield the reported luminosity functions and their scatter.","core_discovery":"The central discovery is that statistical redshifts drawn from coarsely binned empirical distributions reproduce the measured radio luminosity functions. For the 6,425 radio sources with both an r-band magnitude and a spectrum, the authors bin them in ($\\log S_{888\\,\\mathrm{MHz}}$, $m_r$) space, fit a single Gaussian to each bin's redshift distribution, and draw random redshifts for the 4,566 sources with only $m_r$; the 28,821 sources with no optical counterpart are treated separately, first with the same flux-only bins and then with a higher-redshift prior. AGN versus star-forming classification is likewise assigned from the spectroscopic subsample through flux-binned probability mass functions, with a radio-luminosity threshold of $L_{888\\,\\mathrm{MHz}} > 10^{23.5}\\,\\mathrm{W\\,Hz^{-1}}$ used to label strong radio sources as AGN. The resulting 888 MHz luminosity functions for the full 39,812-source sample follow the reference low-redshift functions to $z=0.5$ and the high-redshift function to $z \\approx 5.5$; a faint-end upturn in the lowest redshift bin suggests a population of low-luminosity, low-mass sources not captured by the reference fit.","pith_inferences":["[editorial inference] Because the high-redshift prior for optically blank sources is drawn from the same published luminosity function later used as the high-redshift comparison, the agreement in that regime is a weaker test than it would be with an independent redshift prior; clustering-based or SED-based redshifts could supply that independent check.","[editorial inference] The same bin-and-draw machinery could be applied to other wavebands where spectroscopic completeness is low, such as far-infrared or X-ray surveys, with the same caveat that the training sample's redshift distribution must match the target population.","[editorial inference] The faint-end upturn at $L_{888\\,\\mathrm{MHz}}\\approx10^{20}\\,\\mathrm{W\\,Hz^{-1}}$ predicts an abundant population of low-mass, star-forming dwarf galaxies; deeper radio surveys with complete optical identifications could confirm or rule out this excess.","[editorial inference] The method's success is tied to the optical magnitude limit selecting a nearly complete low-redshift population; in surveys with a different depth or flux limit the bin-and-draw approach would need retesting before it can be trusted."],"forward_implications":["The full 39,812-source radio sample can be assigned redshifts statistically and its 888 MHz luminosity functions match published measurements, so the radio luminosity function does not require complete spectroscopy.","Separate star-forming and AGN luminosity functions are recovered once radio-loud sources are classified as AGN by luminosity rather than by optical line ratios alone.","Random assignments from flux and magnitude bins smooth over large-scale-structure features in the redshift distribution, so the recovered luminosity functions trace the broad population rather than individual structures.","The faint-end upturn implies a previously unpredicted population of low-luminosity radio sources at low redshift, corresponding to star formation rates near $0.07\\,M_\\odot\\,\\mathrm{yr^{-1}}$.","The approach scales to future wide radio surveys where most sources will lack deep multiwavelength photometry and spectroscopy."],"supporting_citations":[{"why":"Provides the low-redshift 1.4 GHz radio luminosity function and the parametric SFG/AGN fits that the optically matched sample's luminosity functions are compared against.","marker":"Mauch & Sadler (2007)"},{"why":"Supplies the high-redshift luminosity function and the redshift distribution used to redistribute the optically blank sources to $z\\gtrsim0.5$.","marker":"Smolčić et al. (2017)"},{"why":"Provides the EMU early science radio catalogue, including the 888 MHz flux densities, survey area, and flux-dependent completeness factors.","marker":"Gürkan et al. (2022)"},{"why":"Provides the GAMA DR4 spectroscopic redshifts and r-band magnitudes that define the spectroscopically complete training sample.","marker":"Driver et al. (2022)"},{"why":"The 1/Vmax method is the volume-correction technique used to construct the luminosity functions.","marker":"Schmidt (1968)"},{"why":"The BPT emission-line diagram is the classification scheme used to label the spectroscopically identified sources as star-forming or AGN.","marker":"Baldwin et al. (1981)"},{"why":"The theoretical maximum starburst line in the BPT diagram is used to separate star-forming galaxies from AGN.","marker":"Kewley et al. (2001)"},{"why":"The semi-empirical demarcation in the BPT diagram is used to distinguish star-forming galaxies, composites, and AGN.","marker":"Kauffmann et al. (2003)"},{"why":"The Saunders functional form is the star-forming luminosity function template used in the reference fits.","marker":"Saunders et al. (1990)"},{"why":"Provides the radio luminosity to star formation rate conversion used to argue that sources above $10^{23.5}\\,\\mathrm{W\\,Hz^{-1}}$ are AGN.","marker":"Hopkins et al. (2003)"}],"fun_headline_variants":["Coarse redshift bins reproduce radio luminosity functions","Statistical z's match radio counts from z=0 to 5.5","No spectra? Statistical redshifts still map radio galaxies","Sparse spectroscopy still yields accurate RLFs","Redshift estimates replicate radio luminosity functions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands on the assumption that radio sources without spectroscopy in a given radio flux density and r-band magnitude bin, and radio sources with no optical counterpart in a given radio flux bin, have the same redshift distribution as the spectroscopically measured sources in that bin; nothing else tests this representativeness except the low-redshift agreement with a published luminosity function.","fun_headline_variants_meta":{"raw":{"variants":["Coarse redshift bins reproduce radio luminosity functions","Statistical z's match radio counts from z=0 to 5.5","No spectra? Statistical redshifts still map radio galaxies","Sparse spectroscopy still yields accurate RLFs","Redshift estimates replicate radio luminosity functions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000449,"raw_usage":{"total_tokens":2301,"prompt_tokens":1015,"completion_tokens":1286,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":1213}},"tokens_in":631,"tokens_out":1286,"duration_ms":9889,"temperature":1.0,"reasoning_tokens":1213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:52:42.193635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could test the method by taking a complete radio survey with full spectroscopic redshifts, hiding a random subset of those redshifts, running the same bin-and-draw procedure, and comparing the reconstructed luminosity functions with the true ones; a significant mismatch in any redshift bin, particularly above z≈0.5, would show the statistical assignments are not representative. The same test can be run on the optically blank sample once follow-up spectroscopy reaches those sources.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the low-redshift 1.4 GHz radio luminosity function and the parametric SFG/AGN fits that the optically matched sample's luminosity functions are compared against."}],"review_version":1}