{"id":"9c311a4a-0556-4fd8-8d12-3da7d6eb81bb","arxiv_id":"2601.03861","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A flux-complete hierarchical Bayesian analysis finds short gamma-ray burst delays below ~1 Gyr and attributes the multi-Gyr delays of earlier studies to mishandled selection effects.","lead":"This paper re-analyzes short gamma-ray bursts — bright flashes thought to come from merging neutron stars — with a careful model of which bursts telescopes could actually detect, and finds the mergers probably happen within about a billion years of their parent stars forming, not after several billion years as earlier work claimed. The authors argue the earlier long-delay estimates were artifacts of including bursts whose detectability was modeled incorrectly.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim hinges on unverified GBM completeness threshold p_lim=3.5; if wrong, short-delay finding and bias diagnosis both fail.","rationale":"The reader and I identify the same weakest assumption: the GBM completeness threshold p_lim=3.5 is adopted from S23 and not independently verified. This is load-bearing because both the headline inference and the explanation of previous long delays depend on it. The paper's own injection validation is limited: it shows that a step-function approximation at the assumed threshold recovers the truth in a simulated regime that is not the inferred one. The W15* reproduction and four-model consistency are genuine strengths, and the paper is transparent about prior sensitivity in Appendix A.3; these do not remove the need to verify the threshold. Since the reader already marks the verdict CONDITIONAL, my stress test does not move it.","tokens_in":20966,"tokens_out":3979,"duration_ms":36139,"concrete_test":"Re-run all four model combinations using (a) the S23 smooth detection efficiency p_det,GBM instead of the step function in Eq. 15, and (b) the step function at p_lim,GBM = 4.5 and 5.5 cm⁻² s⁻¹, rebuilding the observer-frame and rest-frame samples with the corresponding flux cuts. If the 90% credible interval for ⟨τ_d⟩ shifts so that its lower bound exceeds 1 Gyr, or its median moves by more than a factor of 2 from the values in Table 2, the short-delay conclusion is not robust to the completeness assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim — that SGRB delay times are short (⟨τ_d⟩ ≈ 10–800 Myr) and that previous multi-Gyr delays are selection artifacts — rests on the completeness threshold p_lim,GBM = 3.5 cm⁻² s⁻¹ (§2.4.1, Eq. 15). This value is adopted from the authors' previous paper (S23) and implemented as a hard Heaviside cut in the selection function. The denominator D(λ') in Eq. 3 is the fraction of the population passing this cut; because the posterior in Eq. 1 divides by D(λ'), an incorrect threshold directly biases the DTD hyperparameters. The injection test in §4.1 does not fully validate this assumption: it uses the S23 smooth efficiency only to generate the simulated sample, then infers with the step function at p_lim=3.5, and the simulated truth (ατ=1, τmin=0.1 Myr, ⟨τd⟩=2.78 Gyr) lies far from the inferred regime (ατ≈3.3, τmin≈80 Myr, ⟨τd⟩≈160 Myr). Thus the magnitude of selection bias in the actual regime is not demonstrated. If the true completeness limit is higher than 3.5, the observer-frame sample is flux-incomplete in precisely the manner the paper attributes to W15; if lower, the bias correction applied to earlier samples is overstated. The rest-frame sample's redshift completeness is asserted (§2.4.2) rather than modeled, so it cannot independently rescue the inference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a hierarchical Bayesian inference of the short gamma-ray burst (SGRB) population, jointly modeling the delay-time distribution (DTD) and luminosity function. The analysis uses two DTD models (power-law with minimum delay, log-normal), two luminosity models (empirical broken power law, ELF; quasi-universal structured jet, QUSJ), and two samples: a 210-event Fermi/GBM observer-frame sample and an 18-event rest-frame sample. The authors report short average delays, 10 ≲ ⟨τ_d⟩/Myr ≲ 800, and a minimum delay ≲ 350 Myr, in contrast with earlier studies that found multi-Gyr delays. They reproduce the earlier Wanderman & Piran (2015) constraints using the same framework with the earlier sample and selection model (W15*), and perform an injection test in §4.1 to argue that the longer delays in previous studies are due to incorrect treatment of selection effects, specifically using flux thresholds below the completeness limit.","tokens_in":21368,"tokens_out":8159,"duration_ms":75295,"significance":"If correct, the paper would resolve a major discrepancy in SGRB population studies: it would show that SGRB rate evolution is consistent with short-delay BNS merger channels and that previous multi-Gyr delay inferences are selection artifacts. The paper's strengths are the W15* reproduction, which anchors the new framework to an external published result, the consistency of the central finding across four model combinations, and the injection test that directly demonstrates the bias direction from incomplete flux samples. The analysis builds on the publicly available grbpop code, which is a reproducibility asset. The main risk is that the central claim and the bias diagnosis both depend on the completeness threshold p_lim,GBM = 3.5 cm⁻² s⁻¹ adopted from the authors' earlier S23 work, and on the assertion that the rest-frame sample is redshift-complete.","major_comments":[{"comment":"The injection test does not validate the bias in the parameter regime actually inferred. The simulated truth is α_τ=1, τ_min=0.1 Myr, giving ⟨τ_d⟩=2.78 Gyr, whereas the real posterior is α_τ≈3.3, τ_min≈0.08 Gyr, ⟨τ_d⟩≈160 Myr. The test therefore shows that, for a long-delay population, using an incomplete flux threshold biases the inferred delay even longer; it does not demonstrate the magnitude of the selection bias for the short-delay population that the paper claims to measure. Please rerun the injection with truth parameters drawn from the inferred posterior regime, and also report the sensitivity of the real inference to p_lim values around 3.5 (e.g., 3.0, 4.0, 5.0 cm⁻² s⁻¹). Without this, the quantitative claim '10–800 Myr' and the causal diagnosis for W15 are not fully supported.","section":"§4.1, Fig. 5"},{"comment":"The entire analysis, and especially the bias diagnosis, rests on p_lim,GBM = 3.5 cm⁻² s⁻¹, which is adopted from S23 and modeled as a hard Heaviside cut. The denominator D(λ') in Eq. (3) divides the per-event likelihood by the fraction of the population passing this cut, so an incorrect completeness threshold directly biases the DTD hyperparameters. The injection test validates this threshold only against S23's smooth efficiency at E_p,obs = 100 keV and at a single threshold value. If the true completeness limit is higher than 3.5, the observer-frame sample is flux-incomplete in precisely the manner attributed to W15; if it is lower, the bias correction applied to earlier samples may be overstated. The paper needs an independent or otherwise stronger justification of p_lim=3.5, or a sensitivity analysis showing that the short-delay conclusion and the W15-bias diagnosis are robust across","section":"§2.4.1, Eq. (15)"},{"comment":"The rest-frame sample is asserted to maximize redshift completeness, but the selection function does not include the probability that a redshift is measured. The sample contains 18 events and only 16 with measured redshift; the cuts that produce the sample may correlate with luminosity and redshift. If redshift measurement probability depends on L or z, the likelihood in Eqs. (1)–(3) is biased. The paper should either model the redshift-measurement efficiency explicitly or demonstrate by simulation that the stated cuts leave the redshift distribution unbiased. This is load-bearing because the rest-frame sample is one of the two anchors of the redshift-evolution inference.","section":"§2.4.2, Eqs. (15)–(16)"},{"comment":"For the log-normal DTD, σ_τ is essentially unconstrained: ELF gives σ_τ = 0.30 +2.16 −0.28 and QUSJ gives σ_τ = 0.14 +1.59 −0.13 (90% intervals). Since the mean delay of a log-normal is exp(μ + σ²/2), large σ admits substantially longer delays. The abstract's claim 'Regardless of the chosen parametrization' is therefore weaker for the log-normal case than for the power-law case. The paper should either report the posterior on ⟨τ_d⟩ directly (which is the meaningful quantity) or qualify the log-normal conclusion. This is not a fatal flaw, but it is important for the robustness claim and should be addressed.","section":"Table 2, log-normal DTD rows"}],"minor_comments":[{"comment":"The key words include 'gamma-ray burst: individual: GRB 221009A', but GRB 221009A is not discussed anywhere in the paper and is not an SGRB. Remove or replace.","section":"Abstract / Keywords"},{"comment":"The log-normal DTD formula has a typo: the exponent appears as '−1/2 (ln τ_d − ln μ_τ)^2 / 2σ_τ^2', which is dimensionally inconsistent. It should be exp[−(ln τ_d − ln μ_τ)^2 / (2σ_τ^2)].","section":"Eq. (8)"},{"comment":"The prior for R_0 is listed as 'Uniform-in-log, σ_c ∈ [1,10^5]' in Table 1; the parameter name should be R_0. In Table A.1 the units for R_0 are given as 'Gpc⁻³ s⁻¹'; should be 'Gpc⁻³ yr⁻¹'.","section":"Table 1 / Table A.1"},{"comment":"The text says the W15* local rate density is 'one order of magnitude higher' than W15, but Table A.2 gives R_0 = 13.8 vs 7.7 Gpc⁻³ yr⁻¹ (factor ~1.8). Please correct or rephrase.","section":"Appendix A.3 / Table A.2"},{"comment":"The statement that the detection efficiency is '≳80% above the chosen threshold' is used to justify treating the sample as complete. 80% is not 100%; the hard-threshold approximation should be acknowledged more explicitly, especially because it is the basis of Eq. (15).","section":"§4.1, Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is interesting and likely publishable after revision. The main technical risk is the dependence of both the main result and the bias diagnosis on the S23 completeness threshold, plus the injection test's mismatch with the inferred regime. I would ask the authors to add a sensitivity analysis on p_lim and an injection test with truth parameters in the actual posterior regime; these are within the scope of the current manuscript. The log-normal σ_τ constraint should also be reported honestly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a serious paper and the short-delay conclusion looks defensible, but the load-bearing assumption is the GBM flux-completeness threshold, which is borrowed from the authors' own earlier work rather than independently established here.\n\nThe paper's real contribution is the package: a flux-complete GBM sample run through an established hierarchical Bayesian pipeline (grbpop), with a Poisson observed-number term added; four model combinations (two luminosity functions × two DTDs) all give average delays between ~10 and ~800 Myr; a reproduction of W15 within the same framework (W15*) recovers the old multi-Gyr delays; and a simulation experiment shows that using a flux threshold below the completeness limit biases inferred delays toward longer values. That last piece is the most valuable part of the paper: it gives a mechanism, not just a different answer.\n\nI think the reader's conditional verdict is about right. The W15* reproduction anchors the method against an external result, and the bias direction in the injection test is clean. The central claim does depend on p_lim,GBM = 3.5 cm^-2 s^-1 from S23, and the selection function is a hard Heaviside cut. If the true completeness limit is higher, this sample is biased in exactly the way they attribute to W15. That is a legitimate concern, though I'd call it a verification gap rather than a flaw: S23 is published and the threshold is a concrete, testable number. The hard-step approximation is also less severe than the stress-test implies, because the injection test shows the bias appears as soon as the assumed threshold drops below the true one.\n\nLesser soft spots: the rest-frame sample's redshift completeness is asserted rather than modeled; the log-normal σ_τ is barely constrained; the faint end of the ELF is unconstrained; and the W15* reproduction is prior-sensitive, which the authors note in the appendix. The injection test's truth (α_τ=1) does not match the inferred regime (α_τ≈3), so the quantitative magnitude of the bias in the actual regime is not directly demonstrated. None of these kill the argument, but they keep it from being airtight.\n\nWho this is for: anyone working on SGRB rate evolution, BNS merger rates, or r-process enrichment. It deserves a serious referee. I'd send it to review and let the referees push on p_lim and the completeness modeling.","headline":"A careful HBM reanalysis that plausibly corrects W15's multi-Gyr SGRB delays to sub-Gyr, but the whole edifice leans on a completeness threshold taken from the authors' prior paper.","tokens_in":21931,"tokens_out":2320,"would_cite":true,"duration_ms":22392,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.70.Rz"],"model":"deepseek-v4-flash","headline":"This paper argues that short gamma-ray bursts occur just 10–800 million years after their parent stars form, not billions of years later — and that the earlier long delays came from mishandled selection effects.","keywords":["short gamma-ray bursts","delay time distribution","binary neutron star mergers","selection effects","hierarchical Bayesian inference","cosmic star formation history","luminosity function","flux completeness"],"falsifier":"Re-run the paper's own injection test (Section 4.1) with the delay parameters it actually infers — a power law with α_τ ≈ 3 and τ_min ≈ 0.08 Gyr, giving an average delay around 160 Myr — and check whether lowering the assumed flux threshold below the completeness limit still pushes the recovered average delay upward; if it does not, the proposed bias mechanism is not operative in the regime the paper concludes. A direct measurement of Fermi/GBM's short-burst detection efficiency near 3.5 cm⁻² s⁻¹ would settle whether the completeness premise itself holds.","tokens_in":20763,"feed_emoji":"⏱️","tokens_out":15886,"duration_ms":140077,"temperature":0.7,"pith_summary":"The paper sets out to measure how long after star formation a short gamma-ray burst actually happens, under the established idea that these bursts are neutron-star mergers. Applying a hierarchical Bayesian population model with an explicit treatment of detection selection to a flux-complete sample of 210 Fermi/GBM bursts and an 18-event, nearly redshift-complete subsample, the authors find average delay times of roughly 10–800 Myr and a minimum delay below about 350 Myr — consistent across two delay-time models (power-law and log-normal) and two luminosity models (an empirical broken power law and a structured-jet model). That contradicts earlier population studies that inferred multi-gigayear delays. The paper demonstrates the likely cause: when a sample's flux threshold sits below the true completeness limit, the inferred delay is systematically pushed upward; a simulation shows exactly this bias, and re-running the earlier analysis within the same framework reproduces its long delays. If correct, the result means neutron-star binaries merge quickly, SGRB rates track the cosmic star formation history with little lag, and fast merger channels dominate.","feed_headline":"Short gamma-ray bursts merge within 800 Myr, not billions of years","feed_subtitle":"A selection-corrected analysis of 210 bursts blames earlier multi-gigayear delays on flux-cut bias.","key_machinery":"The load-bearing element is the selection term in the hierarchical Bayesian likelihood — the denominator D(λ'_pop) = ∫ P_det(λ_src) P_pop(λ_src) dλ_src, which encodes the probability that a source with given properties would be detected and included. The paper models detection as a hard Heaviside threshold on the 64-ms peak photon flux, P_det = Θ(p − p_lim), and takes the flux-completeness limit of Fermi/GBM to be p_lim,GBM = 3.5 cm⁻² s⁻¹: above this flux the sample is assumed genuinely complete. That threshold carries the whole argument — it is the difference between this analysis and earlier ones, which used effective thresholds low enough to admit flux-incomplete samples. The supporting m","core_discovery":"On the paper's own terms, the discovery is that short gamma-ray bursts merge shortly after their progenitors form: a hierarchical Bayesian reanalysis of a flux-complete 210-burst Fermi/GBM sample and a nearly redshift-complete 18-event subsample, run with two delay models and two luminosity models, yields average delays of about 10–800 Myr, minimum delays below ~350 Myr, and steep power-law indices α_τ ≈ 3. It further establishes that earlier multi-gigayear estimates were likely an artifact of selection effects: in simulation, re-inferring a population with known delays using flux thresholds below the completeness limit shifts recovered delays upward, and the same framework on the earlier st","pith_inferences":["A directly testable corollary the paper leaves implicit: a truly flux- and redshift-complete SGRB sample should show a redshift distribution that peaks near the star-formation peak and decays less steeply than the long-delay prediction — a larger complete redshift sample would settle this observationally.","The bias mechanism plausibly generalizes beyond SGRBs to any transient population inferred from a flux-limited catalog without a verified completeness threshold, including long gamma-ray bursts and other compact-merger counterparts.","The injection test's simulated population has a long true average delay (2.78 Gyr, α_τ = 1), not the short-delay regime the paper infers; repeating the test at the inferred parameters (α_τ ≈ 3, τ_min ≈ 0.08 Gyr) would confirm the size of the bias in the actual regime.","As gravitational-wave catalogs grow, the delay-time distribution inferred here from SGRBs can be cross-checked against the merger-delay distribution reconstructed directly from detected BNS systems, providing a largely independent test of the short-delay claim."],"forward_implications":["SGRB rate evolution should closely track the cosmic star formation history with little time lag, so the high-redshift SGRB rate is higher than earlier long-delay models implied.","A steep power-law index α_τ ≈ 3 replaces the simple α_τ = 1 expectation for binaries shrinking purely by gravitational radiation, implying fast merger channels — short initial orbital separations, common-envelope and mass-transfer evolution, and neutron-star natal kicks.","The inferred short delays bring SGRB population constraints in line with Galactic r-process element studies that require fast mergers (minimum delays ≲ 40 Myr and steep indices).","For most model combinations, the inferred local SGRB rate falls within the binary-neutron-star merger rate measured by gravitational-wave observatories, reinforcing BNS mergers as the dominant SGRB progenitor channel.","Future population studies must either build flux-complete samples or model the full detection efficiency; analyses that ignore completeness will systematically overestimate delay times."],"fun_headline_variants":["Short gamma-ray bursts merge under 800 Myr","New analysis cuts SGRB merger delays to ~800 Myr","Selection bias caused overlong SGRB delay estimates","SGRBs: short delay times after correcting flux bias","Gamma-ray burst delays under a billion years"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The result — and the diagnosis that earlier studies were biased — rests on the assumption that Fermi/GBM's detection efficiency is genuinely complete above a 64-ms peak photon flux of 3.5 cm⁻² s⁻¹, so that the hard threshold used in the selection model is the true completeness limit; if real efficiency falls below that even above the threshold, this analysis carries the same bias it attributes to earlier work.","fun_headline_variants_meta":{"raw":{"variants":["Short gamma-ray bursts merge under 800 Myr","New analysis cuts SGRB merger delays to ~800 Myr","Selection bias caused overlong SGRB delay estimates","SGRBs: short delay times after correcting flux bias","Gamma-ray burst delays under a billion years"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3132,"prompt_tokens":803,"completion_tokens":2329,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":2251}},"tokens_in":547,"tokens_out":2329,"duration_ms":15874,"temperature":1.0,"reasoning_tokens":2251,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T12:12:00.672446+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the paper's own injection test (Section 4.1) with the delay parameters it actually infers — a power law with α_τ ≈ 3 and τ_min ≈ 0.08 Gyr, giving an average delay around 160 Myr — and check whether lowering the assumed flux threshold below the completeness limit still pushes the recovered average delay upward; if it does not, the proposed bias mechanism is not operative in the regime the paper concludes. A direct measurement of Fermi/GBM's short-burst detection efficiency near 3.5 cm⁻² s⁻¹ would settle whether the completeness premise itself holds.","supporting_citations":[],"review_version":1}