{"id":"a4fc1d6a-35c7-4747-bff0-77851733b880","arxiv_id":"2412.18767","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"An all-sky automated Fermipy search finds 1379 candidate gamma-ray sources above 4 sigma, with inconsistent headline counts and no trial-correction estimate.","lead":"This paper uses the Fermipy analysis package to automatically scan the entire gamma-ray sky in 72 overlapping regions, finding 1379 candidate new sources (above a 4 sigma threshold) in 15.41 years of Fermi-LAT data. If the catalog is reliable, it would add hundreds of faint gamma-ray candidates for follow-up, but the paper's own numbers are internally inconsistent and the false-positive rate is uncontrolled.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 1379 new sources rely on a raw 4-sigma threshold with no trial correction; background fluctuations over the all-sky search can plausibly account for a large fraction of the faint 4–5 sigma population.","rationale":"The reader's weakest assumption correctly identifies the absence of trial correction at a low 4-sigma threshold as the core flaw. This is the most load-bearing concern because the entire delta of the paper is the claimed discovery of 1379 new sources, most of which are in the 4–5 sigma range; if even a moderate fraction are background fluctuations, the catalog is not a reliable extension of 4FGL-DR4. The paper's own robustness check only compares flux normalization for known sources and thus cannot detect systematic over-detection of faint candidates. The internal count inconsistency strengthens the case that the reported numbers have not been carefully validated. My proposed Monte Carlo background-only test would directly quantify the false-positive rate and settle whether the concern lands; until such a test is performed, the central claim remains unsupported. I therefore agree with the reader's REJECT verdict and see no reason to adjust it.","tokens_in":14381,"tokens_out":3199,"duration_ms":31533,"concrete_test":"Run the identical Fermipy pipeline on 72 ROIs using simulated all-sky data containing only the Galactic and isotropic diffuse models (no point sources), with the same exposure, energy selection, and sqrt_ts_threshold=4.0. Count how many candidates survive the full two-step threshold across the sky. If the mean number of false positives exceeds about 50, or is comparable to the 886 claimed faint sources, the 4-sigma catalog is dominated by background fluctuations. Separately, reconcile the counts in the abstract (1379), Section 4.1 (1451), and the 886+508 breakdown by inspecting the released 4FGL-Xiang.fits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on treating every candidate with TS >= 16 (Section 3.2, sqrt_ts_threshold=4.0) as a real gamma-ray source. The search spans 72 ROIs of 30x30 degrees at 0.1-degree pixel resolution, i.e., millions of spatial pixels. At TS=16, the single-trial p-value is about 5.7e-5, so the expected number of background fluctuations with TS>=16 across the whole sky is of order hundreds, even after accounting for PSF smoothing and the second fit that drops some low-TS sources. The paper provides no Monte Carlo validation, no trial correction, no false-discovery-rate control, and no direct test of the false-positive rate. The robustness check in Section 4.2 validates flux normalization of known 4FGL-DR4 sources, but it does not validate detection efficiency or the reliability of new faint sources. The source count is also internally inconsistent: the abstract and Section 4.4 say 1379, Section 4.1 says 1451, and the stated breakdown (886 between 4 and 5 sigma, 508 above 5 sigma) sums to 1394. Because the faint 4–5 sigma population is where the noise contamination is largest, the headline number of 1379 and all derived extension, spectral, and variability results are not supported as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an automated all-sky search for new gamma-ray sources using Fermipy, dividing the sky into 72 30°x30° regions aligned with the Galactic diffuse emission template, and analyzing 15.41 years of Fermi-LAT data above 500 MeV. The authors report identifying 1379 new sources at >4σ, of which 497 are >5σ, and then characterize 21 extended sources, 23 sources with spectral curvature above 10 GeV, and 44 variable sources above 1 GeV. The method is partially validated by comparing fluxes of known 4FGL-DR4 sources, and the resulting 4FGL-Xiang catalog is released through China-VO.","tokens_in":14617,"tokens_out":3891,"duration_ms":34504,"significance":"If the catalog were reliable, this would be a substantial extension of the 4FGL-DR4 source population and a potentially useful resource for the community. The paper demonstrates a practical parallelized pipeline and makes data products publicly available, which is commendable. However, the statistical validation is insufficient to support the headline number of new sources: the absence of trial correction over the all-sky search means that a large fraction of the faint 4–5σ candidates could be background fluctuations. The internal inconsistencies in the reported source counts further undermine confidence in the catalog. The work would be significant only after the detection threshold is properly calibrated and the counts reconciled.","major_comments":[{"comment":"The detection threshold is set to sqrt_ts_threshold=4.0 (TS=16) with no correction for the large number of spatial trials. With 72 ROIs of 30°x30° at 0.1° binning, the search effectively involves millions of spatial pixels; even after PSF smoothing, the expected number of noise fluctuations with TS≥16 is of order hundreds. The claimed 1379 new sources are therefore not supported as genuine detections without Monte Carlo validation or false-discovery-rate control. This is the central claim of the paper and is load-bearing.","section":"Section 3.2"},{"comment":"The reported number of new sources is internally inconsistent: the abstract and Section 4.4 state 1379, Section 4.1 states 1,451, and the breakdown in Section 4.1 of 886 faint (4–5σ) plus 508 significant (>5σ) sources sums to 1394. The catalog count must be reconciled and the cause of the discrepancies explained, since the headline number is the primary result of the paper.","section":"Abstract, Section 4.1, Section 4.4"},{"comment":"The robustness test only compares flux normalization and significance of known 4FGL-DR4 sources within 5° of ROI centers. It does not test the false-positive rate for new sources, nor does it measure detection efficiency as a function of flux, position, or background level. Consequently it cannot validate the reliability of the 4–5σ population, which is the most noise-contaminated part of the catalog.","section":"Section 4.2"},{"comment":"The duplicate-removal procedure is not reproducible: the manuscript states that duplicates were removed 'utilizing the Excel functionality in WPS' without specifying the matching radius, whether 1σ or 2σ error circles were used, or how sources near ROI boundaries were assigned. This directly affects the final source count and should be specified precisely.","section":"Section 3.2"},{"comment":"The extension analysis uses TS_ext>16 without accounting for the number of tested spatial models and the range of tested radii (0.1°–3.0° in 0.01° increments). This introduces a multiple-testing problem for each of the 1379 sources, which can produce spurious extended-source claims, especially for faint sources. A Monte Carlo or other trial-correction scheme is needed.","section":"Section 3.3"}],"minor_comments":[{"comment":"The abstract contains a typo: 'Fermi p y' should be 'Fermipy'.","section":"Abstract"},{"comment":"The word 'applyed' in Section 2 should be 'applied'.","section":"Introduction"},{"comment":"The text uses 'IOR' where 'ROI' is meant; this appears in the first sentence of Section 4.2.","section":"Section 4.2"},{"comment":"The caption states the map size as 180°×360°, while the text in Section 3.1 says the sky is 360°×180°. These should be consistent.","section":"Figure 1 caption"},{"comment":"The caption reads 'Analysis results of robustness verification', but the table lists column descriptions and units; this caption should be changed to something like 'Description of catalog columns'.","section":"Table 5 caption"},{"comment":"The reference list includes Strong (1977), but this work does not appear to be cited in the text.","section":"References"}],"recommendation":"reject","confidential_remarks":"The statistical flaw in the detection threshold is not a local presentation issue but affects the core result. Without trial correction, the catalog cannot be considered reliable, and the reported counts are also internally inconsistent. A proper re-analysis with Monte Carlo or FDR control would be needed, and it is likely to change the number of real sources substantially. In its current form, the manuscript is not suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague: the paper does something genuinely useful: it runs Fermipy's find_sources over the whole sky on 15.41 years of Fermi-LAT data, using the Galactic diffuse map to tile 72 ROIs and parallelize. That all-sky tiling plus the longer dataset is a legitimate new application. The sanity check against 4FGL-DR4 for 556 known sources shows their flux and significance measurements are consistent, which is real evidence that the pipeline is not broken. They also make the catalog and TS maps available at China-VO.\n\nThe soft spots are load-bearing. The headline 1379 new sources comes from a raw sqrt_ts_threshold=4.0 with no trial correction. Fermipy's find_sources scans millions of pixels; even after PSF smoothing there are tens of thousands of independent trials, so a 4-sigma single-trial cut will produce some number of background fluctuations. The paper does not quantify that number, provides no Monte Carlo or false-discovery control, and the robustness check only tests flux normalization of bright known sources, not detection efficiency for faint candidates at ROI edges. The count is also internally inconsistent: abstract and summary say 1379, Section 4.1 says 1451, and the breakdown in Section 4.1 (886 + 508) sums to 1394. That kind of discrepancy undermines confidence in the bookkeeping. The de-duplication step is described as done in WPS Excel, with no stated matching radius or treatment of overlapping ROIs.\n\nThe extension, spectral curvature, and variability sub-analyses depend on the same unreliable source list, so they inherit the problem. There is likely a real population of new sources in here—15.41 years of data should reveal more than 4FGL-DR4, especially above 500 MeV—but as presented the paper cannot distinguish genuine detections from the tail of the noise distribution.\n\nMy take: this is a candidate list in need of a statistical upgrade, not a finished catalog. A good referee would ask for trial correction or injection tests, a single consistent count, and a documented de-duplication. The all-sky approach and data release are worth engaging with. Send it to review; it is fixable, and the fixed version would be citable.","headline":"A useful all-sky candidate search undermined by an uncorrected 4-sigma threshold and inconsistent counting; the catalog is worth reviewing after revision.","tokens_in":15198,"tokens_out":2888,"would_cite":false,"duration_ms":27101,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automated all-sky search of 15.41 years of Fermi-LAT data identifies 1379 new gamma-ray sources, 497 of them above 5 sigma.","keywords":["gamma-ray sources","Fermi-LAT","all-sky survey","source detection","binned likelihood","TS map","4FGL-Xiang catalog","Fermipy"],"falsifier":"Run the same 72-region pipeline on scrambled photon maps generated from the best-fit background models and count how many fitted candidates still reach TS at least 16; if the false-positive rate is a few percent or higher, a substantial part of the 1379 new sources, particularly the 886 faint ones between 4 and 5 sigma, would be statistical fluctuations rather than genuine emitters.","tokens_in":14130,"feed_emoji":"🔭","tokens_out":9739,"duration_ms":91935,"temperature":0.7,"pith_summary":"This paper claims to have built an efficient, automated all-sky search for GeV gamma-ray sources and to have applied it for the first time to 15.41 years of Fermi-LAT data. The search divides the sky into 72 overlapping $30^\\circ\\times30^\\circ$ regions, uses Fermipy's find_sources() routine to seed and fit candidates with $\\mathrm{TS}\\ge16$ ($4\\sigma$), and collects those not present in 4FGL-DR4 into a new catalog named 4FGL-Xiang. The paper reports 1379 new sources above 500 MeV, with 497 claimed above $5\\sigma$; among them are 21 spatially extended sources, 23 sources with spectral curvature above 10 GeV, and 44 variable sources above 1 GeV. If the detections hold, the catalog substantially extends the census of known GeV emitters and shows that automated pipelines can replace laborious visual TS-map inspection for future all-sky searches.","feed_headline":"1379 new gamma-ray sources found in 15 years of Fermi data","feed_subtitle":"Automated all-sky re-analysis of Fermi-LAT data adds 1379 sources and flags 44 variables for follow-up","key_machinery":"The load-bearing machinery is Fermipy's GTAnalysis pipeline, in particular the find_sources() method, which builds a TS map over each region, places a point-source model at every pixel exceeding a $\\sqrt{\\mathrm{TS}}=4.0$ threshold, and fits the resulting model jointly; localize() then refines each candidate position. The all-sky search becomes tractable by tiling the sky into 72 $30^\\circ\\times30^\\circ$ regions derived from the Galactic diffuse emission template, running binned likelihood fits in parallel with all 4FGL-DR4 sources included within $30^\\circ$ of each region center, and then removing duplicate candidates found in the overlapping region boundaries. The follow-up classification uses standard 4FGL statistics: $\\mathrm{TS}_{\\rm ext}>16$ for spatial extension, $\\mathrm{TS}_{\\rm cur}>16$ for spectral curvature, and $\\mathrm{TS}_{\\rm var}>21.67$ for variability at 99% confidence.","core_discovery":"On the paper's own terms, the central discovery is that an automated Fermipy pipeline can recover the known 4FGL-DR4 source population and simultaneously reveal thousands of additional gamma-ray candidates without per-candidate visual scanning. The pipeline finds 1379 new sources whose $2\\sigma$ error regions do not overlap any 4FGL-DR4 source, and it provides their positions, spectra, significance levels in three energy bands, and variability information. A cross-check on 556 known sources re-analyzed over the same 14-year interval used by 4FGL-DR4 returns significance ratios and energy fluxes consistent with the published catalog, which the paper takes as evidence that the new detections are genuine rather than artifacts of the multi-region fitting scheme.","pith_inferences":["A Monte Carlo or scramble test would likely show that some fraction of the 886 sources at 4-5 sigma are spurious, since no trials factor is applied; the paper's true new-source count may therefore be closer to the 5-sigma subset than to 1379.","Cross-matching 4FGL-Xiang with radio, X-ray, and neutrino catalogs could reveal whether the faint candidates form a new population of extragalactic emitters or are mostly residual Galactic diffuse fluctuations.","The paper's internal count of 1451 sources above 4 sigma in one section versus 1379 in the summary, and 497 versus 508 above 5 sigma, should be reconciled by the authors before the catalog is used for population statistics.","The same automated tiling pipeline could be run on 100 MeV data with the appropriate foreground model, or on future survey data, to test whether the 500 MeV energy threshold is what keeps the false-positive rate acceptable."],"forward_implications":["The 4FGL-Xiang catalog enlarges the gamma-ray source population by 1379 candidates, giving counterpart-search programs a new target list across the whole sky.","The three-band significance columns (Sig0.5, Sig1, Sig10) let researchers pick out hard-spectrum candidates without re-running the likelihood analysis.","The 21 extended candidates and 23 curved-spectrum candidates above 10 GeV are ready-made samples for follow-up spectral and morphological studies.","The 44 variable sources above 1 GeV provide a sample for studying flare mechanisms and searching for quasi-periodic behavior.","The method itself implies that future Fermi-LAT catalog increments can be produced automatically and quickly as more data accumulate, instead of through manual TS-map inspection."],"supporting_citations":[{"why":"Supplies the 4FGL-DR4 input source model and the known-source comparison set used for robustness checks.","marker":"Ballet et al. 2024"},{"why":"Introduces Fermipy and the GTAnalysis functions (find_sources, localize, extension, curvature, lightcurve) that carry the entire search.","marker":"Wood 2017"},{"why":"The 4FGL catalog paper that defines the significance, TS, and point-spread-function conventions the analysis adopts.","marker":"Abdollahi et al. 2020"},{"why":"Defines TScur and the TSvar threshold used to classify spectral curvature and variability.","marker":"Nolan et al. 2012"},{"why":"Defines TSext = 2 log(L_ext/L_ps), the test statistic used to identify spatially extended sources.","marker":"Lande et al. 2012"},{"why":"Provides the TSext > 16 threshold adopted for significant spatial extension.","marker":"Acero et al. 2016"},{"why":"The scipy.stats package used to compute the TSvar = 21.67 threshold at 99% confidence.","marker":"Virtanen et al. 2020"}],"fun_headline_variants":["Automated Fermipy pipeline uncovers 1379 new gamma-ray sources","All-sky Fermi reanalysis adds 1379 gamma-ray sources","1379 new gamma-ray sources from 15 years of Fermi data","Fermi survey reveals 1379 new gamma-ray sources","New Fermipy analysis finds 1379 gamma-ray sources"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The catalog rests on the assumption that a candidate with likelihood TS at least 16, selected without any trial correction over the all-sky grid and without Monte Carlo validation, is a real source rather than a background fluctuation.","fun_headline_variants_meta":{"raw":{"variants":["Automated Fermipy pipeline uncovers 1379 new gamma-ray sources","All-sky Fermi reanalysis adds 1379 gamma-ray sources","1379 new gamma-ray sources from 15 years of Fermi data","Fermi survey reveals 1379 new gamma-ray sources","New Fermipy analysis finds 1379 gamma-ray sources"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000905,"raw_usage":{"total_tokens":3856,"prompt_tokens":872,"completion_tokens":2984,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":2895}},"tokens_in":488,"tokens_out":2984,"duration_ms":20352,"temperature":1.0,"reasoning_tokens":2895,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:29:05.525874+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 72-region pipeline on scrambled photon maps generated from the best-fit background models and count how many fitted candidates still reach TS at least 16; if the false-positive rate is a few percent or higher, a substantial part of the 1379 new sources, particularly the 886 faint ones between 4 and 5 sigma, would be statistical fluctuations rather than genuine emitters.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces Fermipy and the GTAnalysis functions (find_sources, localize, extension, curvature, lightcurve) that carry the entire search."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The 4FGL catalog paper that defines the significance, TS, and point-spread-function conventions the analysis adopts."},{"cited_title":"L., Abdo, A","cited_arxiv_id":null,"evidence_quote":"Defines TScur and the TSvar threshold used to classify spectral curvature and variability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines TSext = 2 log(L_ext/L_ps), the test statistic used to identify spatially extended sources."},{"cited_title":"E., et al., 2020, NatMe, 17, 261 2","cited_arxiv_id":null,"evidence_quote":"The scipy.stats package used to compute the TSvar = 21.67 threshold at 99% confidence."}],"review_version":1}