{"id":"56245fc5-8c93-4f57-a936-6d3cf3f54c07","arxiv_id":"2608.10347","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A machine-learning-selected sample of 11,293 ASKAP radio galaxies yields a completeness-corrected local star formation rate density of (1.4 +/- 0.5) x 10^-2 solar masses per year per cubic megaparsec, consistent with earlier measurements.","lead":"This paper uses a machine-learning classifier to pick out star-forming galaxies from radio and infrared survey data, then measures how fast the local Universe makes new stars. It is a test of whether this cheaper, automated route can reproduce star formation history results that usually need spectroscopy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Completeness correction C=0.11 is the load-bearing step: it rescales the local SFRD by ~9, yet carries no uncertainty and depends on extrapolating Matthews+21 to faint luminosities.","rationale":"The reader's weakest assumption is the same as mine: the completeness correction. I agree because without C=0.11 the local SFRD is 1.6e-3, an order of magnitude below the comparison literature; the corrected value's agreement with Mauch+07, Upjohn+19, and others is entirely produced by the 1/C multiplication. The correction has no uncertainty and is sensitive to choices that the paper does not vary. The ML classification itself is tested on held-out data (F1=0.93), and the Milliquas cross-match gives a low AGN contamination for z<0.1; those are genuine independent supports. The IRRC consistency claim in Section 5.3 is also weaker than stated (slopes 0.84-0.94 versus Molnar+21 m=1.114), but the completeness normalization is more load-bearing because it scales every SFR in the bin and determines whether the central number is credible. The proposed check is cheap: use a second local LF and vary the integration limits. If C remains near 0.11 with a modest uncertainty, conditional acceptance is justified; if not, the quantitative headline should be treated as a feasibility demonstration pending deeper data. The reader's CONDITIONAL verdict should therefore stand unchanged.","tokens_in":25482,"tokens_out":7038,"duration_ms":67593,"concrete_test":"Recompute C in Equation 8 using an independent local 1.4-GHz luminosity function (e.g., Mauch & Sadler 2007) and propagate the Matthews+21 LF uncertainties while varying the lower integration bound over log L = 18, 19, and 20. If C moves outside roughly 0.05-0.2, or if the resulting SFRD uncertainty exceeds the stated +/-0.5e-2, the headline consistency claim is not robust and the error budget in Table 5 must be revised to include the completeness systematic.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline value is produced by dividing the uncorrected z<0.1 SFRD (1.6e-3) by the completeness C=0.11 (Section 6), a factor of about nine. The paper itself cautions (Section 6) that large completeness corrections amplify small errors, but C is used as a point value with no propagated uncertainty. Equation 8 compares the 1/Vmax-corrected LF of the depth-matched sample to the Matthews+21 local LF and integrates from log L1.4 = 18 to 22.5. Three choices are unpropagated and could each change C materially: (1) the Matthews+21 LF is extrapolated well below the luminosity range where a ~10^4-galaxy NVSS sample constrains it; (2) the lower integration bound is fixed at 18 even though the sample and Matthews+21 LFs diverge near log L=22.5, so the ratio depends on the faint-end shape; (3) the depth-matching cut is imposed using q=2.54 from Molnar+21, so the missing population being corrected is partly defined by applying that IRRC before the IRRC is re-fit in Section 5.3. If the true faint-end LF normalization differs from Matthews+21 at the factor-of-two level, the corrected SFRD changes by roughly a factor of two, far outside the quoted +/-0.5e-2, which includes only SFR and Vmax uncertainties (Table 5). Thus the central consistency claim rests on an unvalidated normalization.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses a gradient-boosted decision tree to classify 389,392 RACS-mid/WISE cross-matched sources into galaxies and quasars, using labels from the WISE-PS1-STRM catalogue. It then depth-matches the resulting galaxy sample to the RACS-mid flux limit, fits the infrared-radio correlation (IRRC) for z<0.1, derives a modified 1.4-GHz SFR calibration (Eq. 7), and computes a 1/Vmax-weighted star formation rate density (SFRD) for z<0.1. After a completeness correction based on a comparison of the sample's luminosity function with the Matthews et al. (2021) local luminosity function (C=0.11), the paper reports a completeness-corrected local SFRD of (1.4±0.5)×10^-2 M_sun yr^-1 Mpc^-3, which is consistent with previous measurements. The paper argues this demonstrates that supervised learning can efficiently identify large, uncontaminated star-forming galaxy populations for SFRD studies.","tokens_in":25755,"tokens_out":7525,"duration_ms":62808,"significance":"If the completeness correction and SFR calibration are validated, the paper provides a novel and efficient path to measuring the local SFRD with machine-learning-selected radio galaxies, and it includes a careful ML pipeline with an independent AGN contamination check using Milliquas. The final value agrees with the literature, which is a useful consistency check. However, the headline number is dominated by a factor-of-nine completeness correction that is presented without a propagated uncertainty, and the SFR calibration is fit to the same sample to which it is applied. These issues must be addressed before the central claim can be considered robust.","major_comments":[{"comment":"The completeness correction C=0.11 is the dominant operation in the paper, scaling the SFRD by ~9, yet it is quoted as a point value with no uncertainty. The integration in Eq. (8) over 18 < log L1.4 < 22.5 relies on extrapolating the Matthews+21 luminosity function below the luminosity range where it is constrained, and the ratio is sensitive to the shape of both LFs near the upper limit. A factor-of-two change in the faint-end normalization changes the headline SFRD by a factor of two, far outside the stated ±0.5×10^-2, which only includes SFR and Vmax uncertainties. The text itself acknowledges that large completeness corrections amplify small errors; this unpropagated systematic should dominate the error budget and must be quantified before the central claim can be accepted.","section":"Section 6, Eq. (8), Fig. 7"},{"comment":"The modified 1.4-GHz SFR calibration in Eq. (7) is fit to the same z<0.1 depth-matched sample whose SFRs it is then used to compute for the SFRD in Section 7. This does not reduce the SFRD to the fitted parameters by construction, but it removes the independence of the SFR calibration from the measurement: any sample-specific bias (e.g., in photometric redshifts or in the depth-matching selection) is absorbed into the calibration and is not represented in the quoted uncertainties. In addition, the depth-matching step in Section 5.2 uses q_TIR=2.54 from Molnár+21 to decide which sources to keep, and the subsequent IRRC fit is conditioned on that choice; the statement in Section 7 that the analysis is not sensitive to the choice of q_TIR is based on comparing the Eq. (6) and Eq. (7) prescriptions on the same depth-matched sample and does not directly test this selection effect. I recommend validating Eq. (7) with an independent SFR indicator (e.g., H-alpha or UV for a subset) or using a calibration derived from an external sample.","section":"Sections 5.2, 5.3, Eq. (7)"},{"comment":"Equation (5) as written is dimensionally inconsistent. It defines f_scale = 10^-10 M_sun yr^-1 L_sun^-1 and then adds it directly to the intercept c in the logarithmic expression log(SFR) = (m+1) log L1.4 + (c + f_scale). However, the q_TIR in Eq. (4) is defined using L_TIR in watts, while the SFR calibration requires L_TIR in solar luminosities; the conversion between these units (a factor of about 3.83×10^26) is missing from Eq. (5). The intercepts in Eqs. (6) and (7) can only be reproduced if an additional constant of about -24 is included (i.e., c - 24.0 for the values in Table 4). The authors should correct Eq. (5) to show the full unit conversion explicitly, otherwise the derivation of the central SFR calibration is not reproducible as written.","section":"Section 5.3, Eq. (5)"},{"comment":"The comparison between the uncorrected SFRD from Eq. (7) (1.6×10^-3) and that from Eq. (6) (1.4×10^-3) is used to argue insensitivity to the choice of q_TIR or the SFR prescription, but both numbers are derived from the same depth-matched sample and are both uncorrected for completeness. This comparison does not validate the completeness correction, which is the step that actually brings the measurement into agreement with the literature. The paper should clarify that the consistency with previous SFRD measurements rests entirely on the unvalidated factor C=0.11, and should discuss what independent evidence supports that normalization.","section":"Section 7, Table 5"}],"minor_comments":[{"comment":"The sentence 'the best-fit models to the data for z >∼ 1 are unlikely to constrain the true behaviour of the IRRC' appears to contain a typo: the sample is limited to z<0.2 after depth-matching, so the statement should refer to z>~0.1 or z>~0.2, not z>1.","section":"Section 5.2"},{"comment":"The paper switches between two parameterizations of the IRRC: in Figure 2 the fit is log L1.4 = m log LTIR + c, while in Table 4 and Figure 6 the fit is q_TIR = m log L1.4 + c. This is confusing; each caption should explicitly state the fitting form used.","section":"Figure 2, Table 4"},{"comment":"The uncertainty columns σ_SFR and σ_total should clarify that neither includes the uncertainty in the completeness correction C; currently a reader might mistake σ_total for the total uncertainty on the corrected SFRD.","section":"Table 5"},{"comment":"The subscripts 1 and 2 in Eq. (8) are defined only in the following sentence; the equation would be clearer if the definitions appeared directly with the equation.","section":"Section 6, Eq. (8)"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of PASA and the ML pipeline is generally well executed. The main issue is the unpropagated completeness correction, which is the load-bearing step for the headline value. I would encourage the editor to request a robustness analysis of C (varying integration limits and the external LF) rather than new observations, and to require the authors to fix the dimensional inconsistency in Eq. (5)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is an honest proof-of-concept, not a new physical result. The authors show that an XGBoost classifier trained on RACS-mid and WISE photometry can isolate a star-forming galaxy sample big enough to estimate the local SFRD, and their completeness-corrected value lands where previous measurements sit. That is genuinely useful as a dry run for EMU, and the pipeline work is careful: proper cross-matching, resampling tests, hyperparameter tuning, SHAP analysis, and a Milliquas cross-match to bound AGN contamination. The paper is transparent about most of its limitations.\n\nThe soft spot is the completeness correction. C=0.11 rescales the uncorrected SFRD by a factor of about nine, and it is used as a point value with no propagated uncertainty. The correction is the ratio of the sample's 1/Vmax LF to the Matthews+21 local LF, integrated from log L=18 to 22.5. That is reasonable in spirit, but it depends on extrapolating Matthews+21 well below the luminosity range where their NVSS sample constrains the LF. If the faint-end normalization is off by a factor of two, the headline SFRD moves by a factor of two, far outside the quoted ±0.5e-2. The stress-test note is right: this is the load-bearing step.\n\nSecond, Equation 7 is the SFR calibration fit to the same z<0.1 depth-matched sample it is applied to. Not fatal, because the SFRD also depends on 1/Vmax weights and the external Matthews+21 LF, but it means the final number is not an independent test of the calibration. Relatedly, the text in Section 5.2 says the fitted L1.4–LTIR slopes (0.84–0.94) are consistent with Molnár+21's m=1.114; with the quoted errors, that is not true. The q_TIR fits in Table 4 do agree within ~1σ, but the statement as written is wrong and should be fixed.\n\nNone of this kills the central claim. The uncorrected SFRD is low, as expected from a shallow survey, and the corrected value is in the right neighborhood. For the stated purpose—demonstrating that ML-selected samples can feed SFRD measurements—the paper works. It would be stronger with uncertainty propagated through C, with the derived sample and code released, and with the consistency claims corrected. This is a paper for people building SFRD pipelines for EMU and similar surveys, not for cosmology at large. I'd send it to a referee and expect a solid published PASA paper after revision. I wouldn't quote the headline number without testing its sensitivity to the LF baseline, but I'd cite the method.","headline":"A useful proof-of-concept: ML-selected ASKAP galaxies give a plausible local SFRD, but the headline number rests on a factor-of-nine completeness correction with no propagated uncertainty.","tokens_in":26372,"tokens_out":4137,"would_cite":true,"duration_ms":35034,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Supervised learning that separates 336,674 galaxies from 52,718 quasars in ASKAP RACS-mid data yields a local star formation rate density consistent with previous measurements.","keywords":["star formation rate density","ASKAP RACS-mid","machine learning classification","XGBoost","galaxy-quasar separation","infrared-radio correlation","photometric redshifts","radio continuum galaxies"],"falsifier":"A deep 1.4 GHz survey (for example EMU's final catalogue) that counts $z<0.1$ star-forming galaxies down to $\\log(L_{\\rm 1.4\\,GHz})\\sim18$ would settle it: if the ratio of the machine-selected sample's counts to the deep survey's counts is not close to 0.11 in the range $18<\\log(L_{\\rm 1.4\\,GHz})\\le22.5$, the completeness correction and the headline SFRD would shift beyond the quoted $\\pm0.5\\times10^{-2}$. Comparing the SFRD computed with spectroscopic redshifts from the 4MOST Hemisphere Survey against the photometric-redshift version would isolate the redshift-error contribution to the correction.","tokens_in":25202,"feed_emoji":"📡","tokens_out":13561,"duration_ms":88815,"temperature":0.7,"pith_summary":"This paper tests whether supervised machine learning can replace traditional colour- or spectroscopic-based selection in measuring the cosmic star formation rate density (SFRD). Using a gradient-boosted decision tree trained on 389,392 RACS-mid radio sources labelled by the WISE-PS1-STRM catalogue, the authors separate 336,674 galaxies from 52,718 quasars and then estimate star formation rates from 1.4 GHz radio luminosity through a calibration fitted to the infrared–radio correlation. After depth-matching the radio and infrared catalogues and applying a completeness correction derived from the local 1.4 GHz luminosity function, they obtain a local ($z<0.1$) SFRD of $(1.4\\pm0.5)\\times10^{-2}\\,M_\\odot\\,\\mathrm{yr}^{-1}\\,\\mathrm{Mpc}^{-3}$ from 11,293 galaxies. This value agrees with earlier measurements, and the agreement is the evidence that machine-learning-selected radio galaxies form an uncontaminated population suitable for SFRD studies. The pipeline uses public all-sky catalogues and can be applied directly to deeper radio surveys.","feed_headline":"Machine learning recovers the local star formation rate density","feed_subtitle":"ASKAP radio plus WISE colours classify 389,000 sources into a local SFRD consistent with earlier measurements.","key_machinery":"The argument runs on three linked components. The first is a gradient-boosted decision tree (XGBoost) trained on 105 RACS-mid and WISE features, with random oversampling of the minority quasar class; this supplies the galaxy–quasar separation and the $z<0.1$ star-forming galaxy sample. The second is the infrared–radio correlation (IRRC), the observed relation between 1.4 GHz radio luminosity and total infrared luminosity, which carries the star-formation calibration: the authors fit $\\log L_{\\rm 1.4\\,GHz}$ versus $\\log L_{\\rm TIR}$ for the depth-matched $z<0.1$ sample and insert the best-fit slope and intercept into the Molnár et al. (2021) prescription, obtaining $\\log(\\mathrm{SFR}/M_\\odot\\,\\mathrm{yr}^{-1}) = (0.745\\pm0.088)\\log(L_{\\rm 1.4\\,GHz}/{\\rm W\\,Hz^{-1}})+(-15.8\\pm1.9)$. The third is the completeness correction: the sample's 1.4 GHz luminosity function is divided by the Matthews et al. (2021) local luminosity function over $18<\\log(L_{\\rm 1.4\\,GHz})\\le22.5$, giving $C=0.11$, and the observed SFRD is divided by this factor. The $1/V_{\\max}$ weighting and a 64 per cent sky-coverage correction convert the summed star formation rates into a density.","core_discovery":"The paper's central claim is that a supervised-learning galaxy–quasar classifier, applied to ASKAP RACS-mid radio data cross-matched with WISE, can select a large star-forming galaxy sample whose radio-inferred star formation rates reproduce the local SFRD. The optimised gradient-boosted model reaches a weighted F1 score of 0.93 and an accuracy of 0.94 on the held-out test set, classifies 336,674 of 389,392 sources as galaxies, and yields a $z<0.1$ depth-matched sample of 11,293 galaxies with an AGN contamination fraction of about 3 per cent. From these galaxies a modified 1.4 GHz SFR prescription is fitted using the infrared–radio correlation, and after a completeness correction the local SFRD is $(1.4\\pm0.5)\\times10^{-2}\\,M_\\odot\\,\\mathrm{yr}^{-1}\\,\\mathrm{Mpc}^{-3}$, consistent with previous measurements from radio, UV, IR, and multi-wavelength surveys. The authors argue that this consistency demonstrates that machine learning can identify sufficiently pure and complete galaxy populations to probe cosmic star formation history.","pith_inferences":["Editorial inference: the true local SFRD could be pinned down more sharply by measuring the faint-end 1.4 GHz luminosity function with a deeper survey, bypassing the factor-of-nine completeness extrapolation.","Editorial inference: the same trained classifier could be applied to the full RACS-mid catalogue of roughly 3.1 million sources once photometric redshifts cover all galaxies, reducing cosmic variance and shrinking error bars.","Editorial inference: because the classifier relies heavily on WISE W1/W2 colours, its galaxy–quasar boundary inherits the AGN selection limits of those infrared colours; adding optical or radio morphology features could reduce the high-redshift misclassifications reported in the appendix.","Editorial inference: the agreement with previous SFRD values validates the combination of machine-learning selection and the assumed luminosity-function baseline, not the selection alone; an independent spectroscopic $z<0.1$ sample would separate the two."],"forward_implications":["If the central claim is correct, a supervised-learning galaxy–quasar classifier trained on public radio and infrared catalogues yields a $z<0.1$ star-forming sample with an AGN contamination fraction of about 3 per cent, so AGN contamination is not inflating the local SFRD.","The completeness-corrected local SFRD of $(1.4\\pm0.5)\\times10^{-2}\\,M_\\odot\\,\\mathrm{yr}^{-1}\\,\\mathrm{Mpc}^{-3}$ is consistent with previous estimates from Mauch & Sadler (2007), Upjohn et al. (2019), Hopkins & Beacom (2006), and Behroozi et al. (2013), indicating no large systematic offset from the machine-learning selection.","The SFRD drops sharply beyond $z\\sim0.25$ because RACS is a shallow survey; applying the same pipeline to deeper 1.4 GHz surveys such as EMU should extend SFRD measurements to higher redshifts.","The choice of $q_{\\rm TIR}$ used for depth-matching changes the low-redshift SFRD by about 12 per cent, within the quoted uncertainties, so the headline result does not depend strongly on that assumption.","Classification errors concentrate at $z\\gtrsim0.5$, where WISE colours near $W1-W2=0.8$ are ambiguous, so the model's misclassifications affect redshifts well outside the $z<0.1$ measurement."],"supporting_citations":[{"why":"Supplies the WISE-PS1-STRM galaxy/quasar labels and photometric redshifts that train the classifier and provide the distance information.","marker":"Beck et al. (2022)"},{"why":"Provides the depth-matching procedure, the $q_{\\rm TIR}$ value, and the modified 1.4 GHz SFR calibration that the paper adapts to its sample.","marker":"Molnár et al. (2021)"},{"why":"Provides the spectroscopically complete local 1.4 GHz luminosity function used as the completeness baseline for the $C=0.11$ correction.","marker":"Matthews et al. (2021)"},{"why":"Releases the RACS-mid catalogue whose 1.4 GHz flux densities and 1.6 mJy completeness limit are the radio data and depth-matching threshold.","marker":"Duchesne et al. (2024)"},{"why":"Defines the WISE passbands and zero-magnitude flux densities used to convert WISE photometry to luminosities.","marker":"Wright et al. (2010)"},{"why":"Introduces the XGBoost algorithm that powers the galaxy/quasar classifier.","marker":"Chen & Guestrin (2016)"},{"why":"Provides the W3-PAH-to-TIR luminosity calibration used to compute total infrared luminosities.","marker":"Cluver et al. (2017)"},{"why":"The $1/V_{\\max}$ method used to correct the flux-limited RACS-mid sample and compute the SFRD and luminosity function.","marker":"Schmidt (1968)"},{"why":"Provides a local radio SFRD measurement that the completeness-corrected result is compared against.","marker":"Mauch & Sadler (2007)"},{"why":"Provides an analytical fit to compiled SFRD measurements used as a consistency check.","marker":"Hopkins & Beacom (2006)"}],"fun_headline_variants":["AI selects 336k galaxies from ASKAP for star formation census","Machine learning extracts local star formation density from ASKAP","ASKAP radio + machine learning measures local star formation rate","389k sources, one AI classifier: local SFRD from ASKAP","Gradient boosting classifies ASKAP radio to pin down SFRD"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the completeness correction is right: comparing how many radio-emitting galaxies the sample finds at each brightness with how many a complete spectroscopic catalogue finds implies that the sample captures 11 per cent of the total $z<0.1$ star-forming luminosity, and that the missing 89 per cent would, on average, look exactly like the detected galaxies.","fun_headline_variants_meta":{"raw":{"variants":["AI selects 336k galaxies from ASKAP for star formation census","Machine learning extracts local star formation density from ASKAP","ASKAP radio + machine learning measures local star formation rate","389k sources, one AI classifier: local SFRD from ASKAP","Gradient boosting classifies ASKAP radio to pin down SFRD"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001306,"raw_usage":{"total_tokens":5439,"prompt_tokens":1176,"completion_tokens":4263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":792,"completion_tokens_details":{"reasoning_tokens":4170}},"tokens_in":792,"tokens_out":4263,"duration_ms":24792,"temperature":1.0,"reasoning_tokens":4170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:23:17.808534+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A deep 1.4 GHz survey (for example EMU's final catalogue) that counts $z<0.1$ star-forming galaxies down to $\\log(L_{\\rm 1.4\\,GHz})\\sim18$ would settle it: if the ratio of the machine-selected sample's counts to the deep survey's counts is not close to 0.11 in the range $18<\\log(L_{\\rm 1.4\\,GHz})\\le22.5$, the completeness correction and the headline SFRD would shift beyond the quoted $\\pm0.5\\times10^{-2}$. Comparing the SFRD computed with spectroscopic redshifts from the 4MOST Hemisphere Survey against the photometric-redshift version would isolate the redshift-error contribution to the correction.","supporting_citations":[],"review_version":1}