{"id":"090029e0-291a-4c10-a12a-1086f830267b","arxiv_id":"2510.23945","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Long gamma-ray burst formation exceeds the cosmic star-formation rate at low redshift in the spectroscopic sample, and the excess shrinks when machine-learned redshifts are included.","lead":"Using a larger sample of long gamma-ray bursts that includes machine-learned redshifts, the authors measure how burst luminosity evolves with cosmic distance and compare the burst birth rate with the star-formation rate. They find a low-redshift excess in spectroscopic data, but adding machine-learned redshifts makes the excess smaller.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ML-redshift pooling is not selection-corrected; the reduced low-z excess in the full catalog may be an artifact of the ML sample's z-dependent incompleteness.","rationale":"The reader's weakest assumption is the same one I identify: pooling ML redshifts with spectroscopic redshifts without modeling the selection that produces the ML sample. I agree with the conditional verdict. The paper is otherwise a careful nonparametric analysis: the spectroscopic-only sample reproduces earlier low-z excess, and the authors honestly report the distribution mismatch and run both catalogs. However, the headline new result—that ML redshifts dilute the excess—is not protected against the selection effect they themselves document. The EP method's associated-set construction assumes truncation only in luminosity (via L_min(z)); it has no term for the probability that a burst enters the ML sample as a function of z, NH, peak flux, etc. Because the ML training set has different NH/peak-flux distributions for GRBs with and without z, the 'full' sample is not a random augmentation. A simple injection test on the spectroscopic sample would settle whether the observed dilution is quantitatively consistent with the ML selection function. If it is, the full-catalog comparison should be presented as an upper limit or as requiring reweighting, not as evidence against the low-z excess. The authors' main claim about the spectroscopic sample stands; only the ML-based comparison is weakened.","tokens_in":10246,"tokens_out":5526,"duration_ms":62215,"concrete_test":"Using the 304-burst spectroscopic sample, re-run the EP/C^- pipeline after artificially removing or reweighting sources according to the ML selection function implied by Fig. 1 (e.g., acceptance probability ~0 for z<1.5, rising to a peak at 1.5<z<3). If the low-z excess drops by the same factor as observed when ML sources are added, the full-catalog reduction is fully explained by selection. Alternatively, on the Dainotti/Narendra validation set, compare ML-predicted versus spectroscopic z as a function of z; if ML z are biased at low z or low-z GRBs are absent from the training set, the absence of low-z ML bursts is an artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The novel comparison is between the full (spectroscopic + ML) and spectroscopic-only rate histories, and it drives the claim that adding ML redshifts reduces the low-z excess. This requires that ML redshifts can be pooled with spectroscopic redshifts in the Efron–Petrosian/C^- estimator. The paper itself shows they cannot be pooled naively: the Anderson–Darling test gives p=0.001, and Fig. 1 shows the ML sample is concentrated at 1.5<z<3 with almost no z<1.5 sources. The authors attribute this to differences in the training-set variables (NH, peak flux) between GRBs with and without measured redshifts. Thus the ML subsample is selected on variables correlated with z and luminosity. The EP method corrects only for the flux truncation L_min(z); it does not correct for this z-dependent inclusion probability. If low-z ML bursts are absent because the training set lacks low-z examples or because predictor variables are biased, the dilution of the excess in the full catalog is a selection artifact. A second, related gap is that ML z are point estimates with no reported uncertainties; redshift errors propagate into ranks and associated sets in Eqs. (6), (10), and (11), and into K-corrections and luminosities, but no error propagation is given. The k values for the two catalogs (2.8 vs 3.7) are also within each other's large 1σ ranges, so the evolution correction does not robustly distinguish them.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes the cosmological evolution of long GRBs using a sample of Swift LGRBs with spectroscopic redshifts (444 after cuts) augmented by 251 ML-estimated redshifts from Dainotti et al. (2025), for a working sample of 499 bursts with fluxes and spectral fits. The authors apply the nonparametric Efron–Petrosian method to remove luminosity evolution and the Lynden-Bell C^- method to derive luminosity functions and formation rates for both the full (spectroscopic+ML) and the spectroscopic-only catalogs. They find that the spectroscopic sample shows a low-redshift excess of the formation rate relative to the Madau–Dickinson SFR, growing to about two orders of magnitude at the lowest redshifts, while the full catalog shows a smaller rise because the ML sample contributes few low-z bursts. The main claimed result is that the LGRB formation rate closely tracks the SFR for z>=1.5 in both catalogs, with the low-z excess confirming earlier work.","tokens_in":10657,"tokens_out":4969,"duration_ms":55828,"significance":"If the ML-augmented comparison were robust, this would be a valuable step forward in using larger, less incomplete GRB samples for population studies. The paper provides a detailed description of the data pipeline, flux limits, K-corrections, and spectral fits, and it explicitly documents the distributional difference between ML and spectroscopic redshifts (Anderson–Darling p=0.001), which is commendable transparency. The spectroscopic-only analysis reproduces and strengthens previously reported low-z LGRB excess with a larger sample. However, the novel full-catalog comparison is not yet convincing because the ML sample is not corrected for its z-dependent inclusion probability, and because the analysis lacks error propagation and uncertainty estimates on key outputs.","major_comments":[{"comment":"","section":"Section 2, Fig. 1, Section 5"},{"comment":"","section":"Section 4.2, 4.3, Tables 1-2, Fig. 7"},{"comment":"","section":"Section 3, Eqs. (6), (10), (11)"},{"comment":"","section":"Section 2, Fig. 2"}],"minor_comments":[{"comment":"","section":"Section 2"},{"comment":"","section":"Section 4.1"},{"comment":"","section":"Fig. 6 caption"},{"comment":"","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The spectroscopic-only analysis is a useful confirmation of the low-z LGRB excess, but the ML-augmented comparison, which is the paper's main new result, is not yet defensible because of the unmodeled ML selection function and the lack of uncertainties. The authors are transparent about the distributional mismatch, which is good, but they need to either model the ML selection or downgrade the full-catalog claims. I would not reject the paper; the underlying methods and data pipeline are sound for the spectroscopic sample, and the issues are fixable with additional analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on GRB rates, but the headline number is weaker than it looks. The paper does exactly what the title says: applies the standard Efron-Petrosian and Lynden-Bell C-minus machinery to a larger sample of LGRBs — 304 spectroscopic plus 251 ML-estimated redshifts — and compares the inferred formation rate to Madau-Dickinson. The spectroscopic-only sample shows the now-familiar low-redshift excess over SFR; adding ML redshifts dilutes that excess. That dilution is the new bit, and it's where the trouble starts.\n\nCredit where due: the pipeline is described carefully (K-corrections, flux limit, truncation), the nonparametric methods are appropriate, and the authors don't bury the fact that the ML and spectroscopic redshift distributions differ (Anderson-Darling p=0.001). They run both catalogs separately and tell you what changes. That's honest.\n\nSoft spots, in order. First, the ML sample is selected on variables other than flux — column density and peak flux in the training set differ from the spectroscopic sample — and the EP method only corrects the flux truncation L_min(z). If the training set lacks low-z examples, the low-z ML deficiency is a selection artifact, and the reduced excess in the full catalog is not a clean astrophysical statement. The paper acknowledges the distribution mismatch but doesn't correct for the inclusion probability. That's the load-bearing weakness, and the stress-test note got it right. Second, no uncertainties on the fit parameters in Tables 1 and 2, and no error bands on the rate curves in Figure 7. The k values for the two catalogs overlap within their large 1σ ranges (2.8 vs 3.7), so the evolution correction doesn't robustly distinguish them. Third, ML redshifts are point estimates; rank-based methods propagate those errors into the associated sets and rates, and that's not addressed.\n\nThe low-z excess in the spectroscopic sample confirms earlier work — not new, but useful. The ML dilution claim is the candidate new result, and it's not secure. I'd send it to a referee: the data compilation and pipeline deserve scrutiny, and the ML-selection issue is a real, fixable problem if the authors model the inclusion probability rather than just noting it.","headline":"Competent EP/C-minus analysis of a larger GRB sample; the low-z excess in the spectroscopic sample is solid, but the ML-driven dilution is not selection-corrected and should be treated as tentative.","tokens_in":11107,"tokens_out":2700,"would_cite":true,"duration_ms":30795,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.70.Rz"],"model":"deepseek-v4-flash","headline":"This paper claims that long gamma-ray bursts track the cosmic star formation rate above redshift 1.5, but exceed it by up to two orders of magnitude at low redshift in the spectroscopic sample, with machine-learning-estimated redshifts dilu","keywords":["gamma-ray bursts","long GRBs","star formation rate","luminosity evolution","Efron-Petrosian method","Lynden-Bell C- method","redshift estimation","machine learning"],"falsifier":"Obtain spectroscopic redshifts for a large, complete sample of satellite-detected long GRBs at z<1 (for instance from a future wide-field mission with immediate IR follow-up) and compare their space density directly to the star formation rate at those redshifts; a measured deviation of two orders of magnitude would confirm the claim, while a rate consistent with the SFR would refute it. Alternatively, show that the ML redshift predictor systematically assigns low-z sources to z>1.5 because of its training-set selection, which would invalidate the combined-catalog comparison.","tokens_in":10181,"feed_emoji":"💥","tokens_out":5628,"duration_ms":54125,"temperature":0.7,"pith_summary":"The paper asks whether long gamma-ray bursts (LGRBs) really form only from collapsing massive stars, tracking the cosmic star formation rate (SFR), or whether an excess of low-redshift LGRBs points to an additional progenitor channel. Using the largest sample to date—695 bursts, merging spectroscopic redshifts with 251 machine-learning-estimated redshifts—and applying non-parametric Efron-Petrosian and Lynden-Bell C- methods to correct for selection bias, the authors find that the LGRB formation rate closely tracks the SFR for z≥1.5. Below z=1.5 the spectroscopically confirmed sample shows a deviation that grows to two orders of magnitude at the lowest redshifts, confirming earlier claims. Adding ML-estimated redshifts, which cluster at 1.5<z<3, reduces the excess to about a factor of ten, because few low-redshift bursts are in the ML sample. The paper concludes that the low-redshift excess is real in the spectroscopic data and may reflect a population of long GRBs produced by compact mergers, with consequences for gravitational-wave rates.","feed_headline":"Long GRB formation rate outpaces star formation at low redshift","feed_subtitle":"Spectroscopic sample shows a 100-fold excess below z=1.5; ML redshifts dilute but do not erase it.","key_machinery":"The machinery is the Efron-Petrosian (EP) method, a non-parametric test for correlation between luminosity and redshift under one-sided truncation, which the authors use to measure luminosity evolution via a broken power-law g(Z) that renders a transformed luminosity L0 independent of redshift. Once the evolution is removed, the Lynden-Bell C- method reconstructs the local luminosity function and the cumulative formation rate from the rank structure of the truncated sample. The flux limit and averaged K-correction define the truncation boundary that these methods require.","core_discovery":"The central claim, stated in Section 5, is that the formation rate of long GRBs closely tracks the cosmic star formation rate for z≥1.5 in both the combined catalog and the spectroscopic-only catalog. In the spectroscopic catalog, the density rate deviates increasingly from the SFR below z<1.5, reaching two orders of magnitude at the lowest redshift—confirming earlier results. When machine-learning-estimated redshifts are added, the concordance with SFR persists until z=1, where the formation rate breaks and increases by a factor of ten; the smaller rise is attributed to the absence of low-redshift bursts in the ML sample. The paper interprets the low-z excess as evidence that a significant","pith_inferences":["If the low-z excess is treated as astrophysical, a direct cross-check is to compare the implied local rate of compact mergers from LGRBs with the LIGO/Virgo binary neutron star merger rate; an inconsistency would force a rethink of the progenitor interpretation.","The Anderson-Darling test reported (p=0.001) shows the ML and spectroscopic redshift distributions are not drawn from the same parent population. A joint-likelihood analysis that models both selection functions simultaneously would be a natural next step, rather than pooling the samples as done here.","The same EP + C- pipeline could be applied to short GRBs with increasing redshift samples to test whether their formation rate follows a delayed SFR, providing an independent handle on merger delay-time distributions.","If the ML redshift distribution's concentration at 1.5<z<3 reflects training-set bias rather than a true dearth of low-z bursts, then the combined-catalog reduction is an artifact; this could be tested by constructing an ML estimator trained on a redshift-complete sample and checking where the predicted redshifts fall."],"forward_implications":["If the low-redshift excess is real, long GRBs cannot be exclusively collapsars at low z; a merger channel would increase the predicted rate of gravitational-wave sources and could explain the kilonova associations seen in GRB 211211A and GRB 230307A.","The formation rate tracking the SFR at z≥1.5 supports the collapsar model for the high-redshift LGRB population, where lower metallicity favors envelope retention and jet production.","The dilution of the excess when ML redshifts are added implies that redshift-estimation methods with selection functions concentrated at intermediate z can mask genuine low-z features; future population studies should treat ML and spectroscopic samples carefully.","The broken power-law luminosity evolution index (k≈2.8 for the full catalog, 3.7 for the spectroscopic) is consistent with earlier estimates, validating the non-parametric approach on a larger sample."],"fun_headline_variants":["Long GRB excess at low redshift survives larger sample","Low-z GRB excess weakens with ML redshifts but persists","Long GRBs mirror star formation until z~1, then overproduce","Spectroscopic long GRB rate deviates 100-fold at low z"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's combined-catalog conclusion that the low-redshift excess is reduced rests on treating machine-learning-estimated redshifts as valid measurements on par with spectroscopic ones; if those estimates are biased toward intermediate redshifts by their training set, the reduction is an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Long GRB excess at low redshift survives larger sample","Low-z GRB excess weakens with ML redshifts but persists","Long GRBs mirror star formation until z~1, then overproduce","Spectroscopic long GRB rate deviates 100-fold at low z"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2806,"prompt_tokens":856,"completion_tokens":1950,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":1876}},"tokens_in":600,"tokens_out":1950,"duration_ms":13259,"temperature":1.0,"reasoning_tokens":1876,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T07:49:58.216944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain spectroscopic redshifts for a large, complete sample of satellite-detected long GRBs at z<1 (for instance from a future wide-field mission with immediate IR follow-up) and compare their space density directly to the star formation rate at those redshifts; a measured deviation of two orders of magnitude would confirm the claim, while a rate consistent with the SFR would refute it. Alternatively, show that the ML redshift predictor systematically assigns low-z sources to z>1.5 because of its training-set selection, which would invalidate the combined-catalog comparison.","supporting_citations":[],"review_version":1}