{"id":"6591f168-e7c8-4511-8cef-1e0608312a1d","arxiv_id":"2412.08173","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A 270 square degree CRAFTS/FAST dataset is calibrated and validated, with noise matching theory within 5% and source fluxes matching NVSS and HI-MaNGA within 8.3% and 16.7%.","lead":"This paper builds and tests the data processing pipeline for the FAST/CRAFTS radio survey, calibrating 70 hours of drift-scan data covering 270 square degrees. The resulting maps have noise within a few percent of theoretical expectations, an encouraging early step for using FAST in hydrogen intensity mapping.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Flux validation is partially circular: the per-day correction factor fitted to NVSS sources (Sec. 3.7) is used to rescale the data, and the same sources then define the reported 8.3%/6.6% flux errors.","rationale":"The reader's weakest_assumption correctly identifies the circular use of NVSS sources for both calibration and validation. This is the most load-bearing concern because the central claim — that the calibrated data are of sufficient quality for HI intensity mapping — hinges on the accuracy of the absolute flux scale. The per-day c_f fitted to NVSS sources removes any day-dependent mean offset by construction, so the reported scatter (8.3% TOD, 6.6% map) reflects random errors after that fit, not the residual systematic error in the flux calibration. If c_f were a function of flux, time, or beam position, the single scalar would leave systematic residuals that the validation scatter would hide. The paper's own observation of larger errors for outer-circle beams (9.3% vs 6.7%) supports the possibility of beam-dependent residuals. The internal inconsistency in the reported map error (6.6% in Sec. 5.1.2 and the abstract versus 11.6% in Sec. 6) further underscores that the map-level validation is fragile and should be independently checked. The proposed split-half cross-validation and an independent-catalog check would settle whether the flux errors are truly at the claimed level. Given that the reader already issued a CONDITIONAL verdict with this concern explicitly flagged, no verdict change is needed; the conditionality is appropriate and should be resolved by the requested tests.","tokens_in":49413,"tokens_out":4876,"duration_ms":52867,"concrete_test":"Split the 447 NVSS sources by beam group (central/inner/outer) and by flux quartiles; fit c_f on half of each subgroup and validate on the held-out half. If the validation relative error increases significantly beyond the reported 8.3%/6.6%, or if the best-fit c_f values differ between subgroups beyond ~5%, the single-scalar correction is insufficient. Additionally, re-derive the map-level flux error using an independent catalog (e.g., FIRST or TGSS) without fitting c_f to those sources, to quantify the non-circular absolute calibration accuracy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that the per-day flux correction factor c_f (Sec. 3.7, Eq. 24), obtained by least-squares fitting to NVSS-selected point sources, is a single scalar that fully absorbs all day-dependent calibration errors (noise-diode overflow, T_ND and eta amplitude variations) without biasing the subsequent validation. The same 447 sources used to fit c_f are then used to report the 8.3% (TOD) and 6.6% (map) relative flux errors, so those errors measure only the scatter around the fitted line, not the accuracy of the absolute flux scale. If c_f varies with sky position (e.g., between central and outer beams), time, or source brightness, the single-scalar correction is incomplete and the reported errors underestimate the true calibration uncertainty. This directly affects the HI integral flux validation (16.7%) and the feasibility claim for HI intensity mapping, which relies on accurate absolute calibration. The paper itself reports larger errors for outer-circle beams (9.3% vs 6.7%), hinting at beam-dependent residuals that a per-day scalar cannot capture.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the calibration and data-processing pipeline for 70 hours of CRAFTS drift-scan observations with the FAST L-band 19-beam receiver, covering 270 deg^2. The pipeline combines pulsar-backend noise-diode calibration with spectrum-backend data, applies RFI flagging, temporal drift and baseline corrections, a per-day flux correction factor fitted to NVSS point sources, map-making, and standing-wave removal. Validation includes comparison of TOD and map noise levels with radiometer predictions, PCA foreground removal tests, continuum flux comparison with 447 NVSS sources (8.3% TOD, 6.6% map), and HI integral flux comparison with 90 HI-MaNGA galaxies (16.7%). The paper concludes that the calibrated data are suitable for further HI intensity mapping and galaxy studies.","tokens_in":49655,"tokens_out":5100,"duration_ms":53795,"significance":"If the validation holds, this is a useful methods paper for the HI intensity mapping community: it is the first detailed calibration and pipeline description for CRAFTS data, it quantifies systematics such as the high-cadence noise-injection overflow, and it provides an end-to-end product (calibrated TOD and maps) that can serve future cosmological analyses. The noise-level comparison with the theoretical radiometer expectation is carefully done, and the HI-MaNGA comparison is strengthened by re-processing the GBT spectra with the same baseline and integration choices as the CRAFTS data. However, the continuum flux validation is partially circular because the per-day correction factor is fitted to the same NVSS sources that are later used to quote the flux errors, so the reported 8.3%/6.6% values measure scatter around a fitted line rather than absolute flux accuracy. This limits the strength of one of the paper's headline claims.","major_comments":[{"comment":"The continuum flux validation is partially circular. The per-day correction factor c_f is obtained by least-squares fitting to NVSS-selected point sources (Sec. 3.7, Eq. 24), and the same 447 sources are then used in Sec. 5.1.1 to report the 8.3% (TOD) and 6.6% (map) relative flux errors. The mean agreement is therefore guaranteed by construction; only the scatter around the fitted line has evidential value. The manuscript should explicitly state this limitation, report the uncorrected flux residuals before applying c_f, and, if possible, test the stability of c_f across beams, time, and source flux. An independent absolute flux check (e.g., a calibrator observation or a comparison at a different frequency band with a different catalog) is needed to support the claim of accurate absolute calibration.","section":"Sec. 3.7, Eq. (24); Sec. 5.1.1; Fig. 25"},{"comment":"There is a numerical inconsistency in the reported map-level continuum flux error. Sec. 5.1.2 and the abstract report a relative flux error of 6.6% for the map, while Sec. 6 states 'we measure the flux of 447 continuum point sources near 1400 MHz. Compared with the NVSS catalog, our results yield a relative flux error of ~8.3% for TOD and ~11.6% for the map.' The 11.6% figure appears nowhere else and contradicts the 6.6% in Sec. 5.1.2 and the abstract. This needs to be corrected, since the map flux error is one of the headline validation numbers.","section":"Sec. 6 (Summary)"},{"comment":"The reported HI integral flux error of 16.7% is computed after excluding three galaxies with xi_mask > 0.1 (Sec. 5.2), but the abstract and Sec. 6 report this value without mentioning the exclusion. Since the exclusion criterion is a quality cut that improves the apparent agreement, the manuscript should either report the error with and without the three excluded galaxies, or at minimum state the exclusion explicitly in the abstract and summary.","section":"Sec. 5.2; Table A1"},{"comment":"The definition of the relative flux error in Eq. (35), deltaS = (S_CRAFTS - S_NVSS)/sqrt(S_CRAFTS * S_NVSS), is not applicable when S_CRAFTS is negative, because the denominator is imaginary. Table A1 contains such cases: for source 61 (HI-MaNGA 8604-9102), F_HI,FAST = -0.64 Jy km/s, yet a finite value deltaS = -226.49 is listed. The manuscript does not explain how negative flux measurements are handled in the reported 16.7% median error. This needs to be clarified, either by modifying the error definition, by excluding non-detections from this statistic, or by stating the convention used.","section":"Eq. (35); Table A1"},{"comment":"The paper itself provides evidence that a single per-day scalar correction factor c_f cannot fully absorb the calibration errors: the relative flux errors are 6.7% (central beam), 7.2% (inner circle), and 9.3% (outer circle). These beam-dependent residuals are expected if the effective beam shape or the calibration parameters vary across the 19-beam receiver, but the manuscript presents c_f as a day-dependent scalar and does not discuss this as a limitation of the absolute flux scale. The authors should quantify the beam-to-beam variation in c_f and either adopt a beam-dependent correction or state explicitly that the quoted flux errors do not include this systematic component.","section":"Sec. 5.1.1; Sec. 3.7"}],"minor_comments":[{"comment":"The reference to 'Fig. 3.5' in the paragraph following Eq. (24) appears to be a typo; the correct reference should be Fig. 15.","section":"Sec. 3.7"},{"comment":"The text states 'the flux correction factor (e.g. fc=1.26 in Fig. 15)', but the variable is defined as c_f in Sec. 3.7. The symbol should be c_f for consistency.","section":"Sec. 4.2"},{"comment":"There is a typo: 'CRFATS' should be 'CRAFTS' in the sentence beginning 'Nevertheless, the comparable results given by FAST confirm...'.","section":"Sec. 5.2"},{"comment":"In the description of the DAOStarFinder settings, 'FWHM = add sqrt(FWHM^2_beam + FWHM^2_kernel)' contains the stray word 'add'; the intended expression is likely the square root of the sum of squares.","section":"Sec. 5.1.2"},{"comment":"The abstract states that the noise level is 'consistent with the theoretical predictions within 5% at RFI-free channels.' This statement applies cleanly to the map (Sec. 4.4, right panel), but for the TOD the observed noise at 1050-1150 MHz is 6.2 mJy, which is approximately 10-15% above the theoretical level at that band, as acknowledged in Sec. 4.4. The wording should be adjusted to avoid overstating agreement for the TOD.","section":"Sec. 4.4; Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is in scope for a methods-focused journal in radio astronomy / cosmology. The central pipeline description is valuable and the noise validation is careful. The main risk is the circularity of the continuum flux validation and the internal contradiction in the reported map flux error (6.6% vs 11.6%). Both are fixable with additional analysis and clearer presentation, so I am recommending major revision rather than rejection. I would also ask the editor to ensure the authors address the treatment of negative HI flux measurements in Table A1, as the current error definition appears mathematically undefined for those entries."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new content here is the CRAFTS-specific calibration work: using the pulsar backend's high-cadence noise injection to calibrate the spectrum backend, identifying and correcting the noise-overflow synchronization effect, and validating a 270 deg2 field with 447 NVSS-matched continuum sources and 90 HI-MaNGA galaxies. The pipeline is an adaptation of fpipe/tlpipe, but the application to CRAFTS is not trivial, and the validation results are concrete and new. The noise levels (5.7 mJy TOD, 1.6 mJy map) matching theory within 5% at clean channels is a solid, reproducible check. The PCA residual test showing thermal-noise-dominated maps after 30 modes is also encouraging for the stated goal of intensity mapping. The paper is well-organized and unusually candid about known systematics, including the visual inspection step and the unknown cause of some bad data.\n\nThe main soft spot is the flux validation circularity. The per-day correction factor cf is fitted by least squares to NVSS-selected point sources (Sec. 3.7, Eq. 24), then the same survey's sources are used to report the 8.3% and 6.6% flux errors. Those errors measure scatter around the fitted line, not absolute flux accuracy. The paper partially acknowledges this by saying the correction absorbs TND and eta amplitude variations, but the presentation could mislead a reader into thinking the NVSS comparison independently validates the absolute flux scale. The beam-dependent scatter (9.3% vs 6.7%) hints that a single scalar per day does not capture everything, so the true accuracy could be worse for science cases that care about absolute calibration.\n\nOther issues are minor but worth fixing: the three HI galaxies with xi_mask > 0.1 are excluded before reporting the 16.7% error, which is defensible but should be stated more prominently as a post-hoc cut. Also, Sec. 6 quotes an 11.6% map flux error while Sec. 5.1.2 reports 6.6%; the paper needs to reconcile or clearly distinguish those numbers. Reproducibility is limited by no code/data release and the visual inspection step, but the request-based data availability is at least a start.\n\nOverall, this is a solid pipeline paper that deserves serious refereeing. The central feasibility claim—that CRAFTS data are usable for HI intensity mapping—is plausible and supported by the noise-level and PCA tests. The circular flux validation is a real limitation but not a fatal one; the paper should either refit cf on a subset and validate on a holdout set, or at least present the flux errors as precision around the calibrated scale rather than absolute accuracy. I would bring it to a reading group and cite it for the calibration details, but I would not use the 6.6% flux error as an absolute calibration number without further cross-checks. Recommend: send to peer review, with requests for clarification on the flux validation and the inconsistent map error.","headline":"CRAFTS pipeline paper is a careful, useful validation study for FAST HI intensity mapping, with the main caveat that the flux calibration error is partly circular because it is fitted to the same NVSS sources used for validation.","tokens_in":50323,"tokens_out":737,"would_cite":true,"duration_ms":11063,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that calibrated CRAFTS drift-scan data from FAST reach the noise and flux accuracy needed for HI intensity mapping.","keywords":["HI intensity mapping","CRAFTS","FAST","21 cm cosmology","radio calibration","drift scan","radio frequency interference","foreground removal"],"falsifier":"Re-compute the continuum flux comparison with the 447 sources divided by sky position, by flux, and by which of the 19 beams detected them: if the per-day correction factor actually varies across those splits, the residuals relative to NVSS will show systematic trends instead of scatter around zero, and the 8.3% and 6.6% errors will grow.","tokens_in":49170,"feed_emoji":"📡","tokens_out":8096,"duration_ms":73837,"temperature":0.7,"pith_summary":"This paper tries to establish that drift-scan data from the Commensal Radio Astronomy FAST Survey (CRAFTS) can be calibrated well enough to serve as a cosmological 21-centimetre intensity-mapping dataset. Using 70 hours of L-band 19-beam observations covering about 270 square degrees, the authors build a nine-step pipeline that combines a pulsar backend's fast noise-diode switching with a spectrum backend's fine frequency resolution. The maturity of the data is tested three ways: measured noise levels match radiometer predictions within about 5 percent in clean frequency channels; 447 bright continuum sources match NVSS fluxes within 8.3 percent on time-ordered data and 6.6 percent on maps; and 90 low-redshift galaxies match HI-MaNGA integrated fluxes within 16.7 percent. The paper also shows that after removing 30 principal-component foreground modes, the residual map behaves like thermal noise at about 1.6 mJy. If these numbers hold, CRAFTS is a viable path to large-scale HI structure measurements.","feed_headline":"FAST drift-scan data hit theoretical noise within 5 percent","feed_subtitle":"A 270 deg2 patch is now calibrated and vetted for HI intensity mapping, with source fluxes matching catalogs to ~8%.","key_machinery":"The load-bearing mechanism is the pairing of two backends in calibration: the pulsar backend resolves the 198.6 microsecond noise-diode on/off cycle, giving a per-0.2-second measure of gain and noise-diode response, while the spectrum backend supplies 7.6 kHz spectral resolution that is later rebinned to 30 kHz. The ratio of spectrum-backend power to pulsar-backend power, together with the noise-on minus noise-off difference, factors the response into bandpass and temporal drift components; the absolute scale comes from the measured noise-diode temperature $T_{\\rm ND}(\\nu)$ and aperture efficiency $\\eta(\\theta_{\\rm ZA},\\nu)$. A per-day least-squares flux correction $c_f$ fitted to isolated NVSS sources (Eq. 24) absorbs residual scale errors, including a known noise-injection overflow effect in early data.","core_discovery":"The central claim is that the calibrated CRAFTS data product is of sufficient quality for HI intensity mapping: the calibrated time-ordered data reach a noise level of $\\sim 5.7$ mJy and the map $\\sim 1.6$ mJy per beam, within 5% of theoretical predictions at RFI-free channels, while continuum source fluxes agree with NVSS at 8.3% (time-ordered data) and 6.6% (map level) and HI integral fluxes agree with HI-MaNGA at 16.7%. The pipeline achieves this by using the pulsar backend's high-cadence noise injection to calibrate bandpass and temporal drift, applying absolute flux calibration from noise-diode temperature and aperture efficiency, and then correcting residual day-to-day scale errors with a per-day flux correction factor fitted to NVSS-selected point sources. This is presented as the first systematic feasibility assessment for cosmological HI detection with CRAFTS.","pith_inferences":["If the per-day flux correction factor varies across the field, the quoted flux errors are optimistic: splitting the 447 calibrators by sky position or flux and re-fitting would show residual trends, and this can be done with data already in hand.","The paper's own account of the noise-overflow effect implies data taken before winter 2021 carry a roughly 30% scale error that is only removed per-day; co-adding those days with later data without modelling day-to-day discontinuities would bias any stacked HI auto-spectrum.","The standing-wave removal is confined to a narrow delay-space peak around $k_\\parallel \\sim 2\\,h\\,{\\rm Mpc}^{-1}$ at $z\\sim0.07$; injecting simulated standing waves through the same pipeline would turn this claimed confinement into a quantified transfer function for the auto-spectrum.","A natural end-to-end test that needs no new observations is cross-correlating the cleaned 270 deg2 map with an optical galaxy catalog in the footprint; given the stated noise level, a detected cross-signal would validate the whole calibration chain for cosmology."],"forward_implications":["The 270 deg2 calibrated cube gives a working testbed for HI power-spectrum estimation at redshift $0<z<0.07$ and $0.23<z<0.35$, the frequency bands left after masking the strong RFI band.","The stated flux agreement (8.3% on time-ordered data, 6.6% on maps, 16.7% on HI integrals) defines the current calibration floor that any cosmological analysis with CRAFTS must fold into its error budget.","The PCA result that residual maps become thermal-noise-dominated at $\\sim1.6$ mJy after removing 30 modes means foreground subtraction is not the immediate limiting step; the next limit is scheduled observing time.","The larger point-source errors for outer-ring beams (9.3% versus 6.7% for the central beam) point to the beam model as the next improvement, since a Gaussian profile was assumed for flux fitting.","Residual beam and day stripes at $\\sim0.1$ K, reduced from several kelvins by temporal baseline subtraction, can be further suppressed by repeated scans, which the survey already plans."],"supporting_citations":[{"why":"Provides the fpipe pipeline and pilot-survey calibration methods that this CRAFTS pipeline adapts, plus the previous 60 deg2 dataset.","marker":"Li et al. (2023)"},{"why":"Supplies the 19-beam receiver parameters, aperture efficiency fits, and system temperatures used for absolute flux calibration and noise predictions.","marker":"Jiang et al. (2020)"},{"why":"Is the NVSS catalog used to select and measure 447 continuum sources and to fit the per-day flux correction.","marker":"Condon et al. (1998)"},{"why":"Provides the HI-MaNGA galaxy catalog and the 1.2 W20 integration convention used for the 90-galaxy HI flux comparison.","marker":"Masters et al. (2019)"},{"why":"Supplies the HI-MaNGA sample definitions and data products used for the low-redshift HI galaxy comparison.","marker":"Stark et al. (2021)"},{"why":"Contributes the SumThreshold and SIR RFI flagging algorithms implemented in the pipeline.","marker":"Zuo et al. (2021)"},{"why":"Is the source of the PCA foreground-subtraction method used to test residual map noise.","marker":"Alonso et al. (2015)"},{"why":"Provides the four-channel difference method for estimating noise level from real data.","marker":"Wang et al. (2021a)"},{"why":"Supplies the Zernike beam model used to justify the Gaussian beam profile in flux fitting.","marker":"Zhao et al. (2024)"}],"fun_headline_variants":["FAST's CRAFTS data passes validation for HI mapping","CRAFTS calibration achieves noise within 5% of theory","270 sq deg HI survey hits 5% noise bound","FAST pipeline validates HI fluxes to 8% vs NVSS","First CRAFTS data product ready for HI cosmology tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a single per-day flux correction factor, fitted to bright isolated NVSS sources, absorbs all day-dependent calibration errors (noise overflow, noise-diode and efficiency amplitude variations) without overfitting or biasing the validation.","fun_headline_variants_meta":{"raw":{"variants":["FAST's CRAFTS data passes validation for HI mapping","CRAFTS calibration achieves noise within 5% of theory","270 sq deg HI survey hits 5% noise bound","FAST pipeline validates HI fluxes to 8% vs NVSS","First CRAFTS data product ready for HI cosmology tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00087,"raw_usage":{"total_tokens":3846,"prompt_tokens":1098,"completion_tokens":2748,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":714,"completion_tokens_details":{"reasoning_tokens":2663}},"tokens_in":714,"tokens_out":2748,"duration_ms":21151,"temperature":1.0,"reasoning_tokens":2663,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:07:20.952265+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-compute the continuum flux comparison with the 447 sources divided by sky position, by flux, and by which of the 19 beams detected them: if the per-day correction factor actually varies across those splits, the residuals relative to NVSS will show systematic trends instead of scatter around zero, and the 8.3% and 6.6% errors will grow.","supporting_citations":[],"review_version":1}