{"id":"97885772-82ec-49c0-9499-70d8b78c055a","arxiv_id":"2507.16192","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Simulations show that pseudo-redshifts from three popular GRB correlations badly distort the true redshift distribution and GRB rate, so the correlations cannot be trusted for population-level distance measurements.","lead":"This paper tests whether empirical correlations of gamma-ray bursts can serve as distance indicators for population studies. Using synthetic catalogs of 4000 bursts, it finds that pseudo-redshifts from the Yonetoku, 3D Dainotti, and L-T-E relations produce redshift and rate distributions that differ strongly from the true input, even when selection effects are ignored.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Population-level claim is overbroad: only per-object point inversion and one Gaussian-likelihood variant are tested, leaving untested the hierarchical forward-model route that the paper itself recommends.","rationale":"The reader's weakest-assumption diagnosis is exactly the point I consider load-bearing. Sections 2-3 and Table 1 establish, under the paper's stated generative assumptions, that equating each burst's zi with the root of the parametric-curve/plane intersection produces a redshift distribution and rate that differ from the injected ones. That is a valid negative result for point-inversion, and the KS p-values are convincing. But the abstract generalizes this to 'population level' without testing the standard alternative: a hierarchical forward model that treats z as a latent variable, computes the per-burst likelihood of the observed quantities as a function of z, and marginalizes over z with a population prior. Such a model does not need individual zi to be unbiased; it can in principle recover Psi even when the mode of the posterior for z is shifted, because the full likelihood retains information about the population through the joint distribution of observables over many bursts. The paper's Section 4 likelihood variant is not this: it fixes mu and Sigma to the true intrinsic distribution and then picks the z that maximizes a Gaussian likelihood, which is a point estimate conditioned on the answer. Varying the luminosity function would also matter, since the paper's invariance claim applies only to the geometry of the parametric curve, not to the hierarchical likelihood. The blank Gitlab link is a reproducibility problem, but it is secondary. I recommend no change to the reader's conditional verdict: the broad claim should not be accepted as proven until the hierarchical forward-model test is performed; if the test recovers the injected Psi, the claim must be narrowed, and if it fails, the conditional verdict can stand.","tokens_in":12170,"tokens_out":4518,"duration_ms":51942,"concrete_test":"Implement a hierarchical Bayesian fit on the same synthetic catalog: latent z_i for each burst, likelihood from the Yonetoku relation (Eq. 1) with log-normal scatter sigma=0.25, and population prior Psi(z) proportional to (1+z)^a / (1+(1+z)/B)^c with unknown (a,B,c), plus a log-normal luminosity function with unknown mean and variance. Sample the posterior with HMC/NUTS and check whether the 90% credible region for Psi(z) contains the true injected Psi for N=4000 and c=5.6; repeat for the 3D Dainotti and L-T-E relations. If the posterior covers truth, the paper's population-level claim is falsified by its own generative model; if it does not, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the Abstract, that empirical GRB correlations cannot serve as reliable distance indicators at the population level, is stronger than what Sections 2-3 actually demonstrate. The simulations infer each redshift by root-finding the intersection of a parametric curve with the best-fit plane, then compare the histogram of zi to zg. This tests only point-estimate inversion. Section 4 adds a multivariate Gaussian likelihood using the true intrinsic mean and covariance, but it still assigns one zi per burst and conditions on the very population parameters that a population study would need to infer; it does not marginalize over latent redshifts with a population prior. A standard hierarchical forward model, with latent z_i, per-burst likelihoods from the same correlations, and a flexible prior Psi(z; a,B,c), could in principle recover Psi even when individual pseudo-redshifts are biased, because the likelihood for each burst retains information about z through the observed flux, fluence, Ep, and T_a. The paper's own discussion concedes that 'statistically robust, population-level methodologies' might improve inference, but it never tests one. Without such a test, the unequivocal population-level conclusion is an extrapolation. The narrow result, that direct inversion yields biased zi distributions, is internally consistent and well demonstrated; the broad result is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests whether empirical GRB correlations (Yonetoku, 3D Dainotti, L-T-E) can serve as distance indicators for population studies. It generates synthetic GRB catalogs from assumed intrinsic distributions and best-fit correlation parameters, infers a pseudo-redshift for each burst by intersecting the redshift-dependent parametric locus with the correlation (or, in Section 4, by maximizing a multivariate Gaussian likelihood built from the true intrinsic mean and covariance), and compares the inferred redshift distribution and rate Psi(z) with the truth using KS tests. The simulations give statistically significant mismatches for all three correlations and all tested widths of Psi(z). The paper concludes that pseudo-redshift estimates from these correlations cannot constrain Psi_GRB and that the correlations cannot serve as reliable distance indicators at either the individual or population level.","tokens_in":12457,"tokens_out":5844,"duration_ms":61268,"significance":"The controlled simulation is a strength: the test is self-consistent, since synthetic bursts are generated from the same relations being inverted, and the negative result is not an artifact of unknown ground truth. The paper also explicitly separates the assumption of no selection effects. If the narrow claim (direct pseudo-redshift inversion produces biased redshift and rate distributions) is accepted, it is a useful caution to pseudo-redshift population studies. The broad population-level claim, however, is not established because the only alternatives tested are point-estimate methods; a hierarchical forward model that marginalizes over latent redshifts is not tested, despite being recommended in the discussion. The paper's strength is therefore in the controlled demonstration of the failure of root-finding and likelihood-point-estimate approaches.","major_comments":[{"comment":"The Gaussian likelihood test conditions on the true intrinsic mean vector mu and covariance matrix Sigma. A population study does not know these; they must be inferred jointly with the latent redshifts. Because the paper conditions on truth and still assigns one z_i per burst, it does not test the Bayesian hierarchical forward model it recommends. The abstract's claim that the correlations 'cannot serve as reliable distance indicators ... at the population level' is an extrapolation beyond the evidence. To support the strong claim, add a hierarchical model test or narrow the conclusion to direct pseudo-redshift methods.","section":"Section 4, likelihood expression"},{"comment":"The claim that 'the specific form of the luminosity function does not affect the core results' is asserted but not demonstrated. The no-solution and two-solution fractions in Sections 3.2 and 3.3 and the low-redshift pile-up described in the Appendix depend on the joint distribution of (L_p, E_p,z, T_a*, L_X) and on the intrinsic scatter sigma_int. Since the conclusion is stated as holding 'regardless of the intrinsic distribution's characteristic width', the parameter coverage is too narrow: only the width parameter c of Eq. (8) is varied, not the luminosity function or the intrinsic scatter, so the generality of the population-level claim is not tested.","section":"Section 2, luminosity function and scatter"},{"comment":"The paper acknowledges that 'statistically robust, population-level methodologies' might improve inference but does not implement one. This is more than a cosmetic gap because the information content of the correlations is preserved in per-burst likelihoods even when point pseudo-redshifts are biased. A forward model with latent z_i and a flexible population prior could in principle recover Psi_GRB even when individual pseudo-redshifts are poor. Without such a test, the 'unequivocal' wording in the Abstract overstates what the simulations can establish; the central claim needs either a new test or a softening.","section":"Section 4, Discussion and Conclusion"}],"minor_comments":[{"comment":"The Gitlab repository link in Section 2 is empty (shown as '</>'); please provide a working URL so the code can be checked.","section":"Section 2, code availability"},{"comment":"There is a typo: 'observ ed plateau flux' should read 'observed plateau flux'.","section":"Section 2"},{"comment":"Equation (A1) defines g(z) as -log L_X + C0 + alpha log E_iso + beta log T_a*, while Eq. (A3) and the surrounding text use g(z_true) as log L_X - log L_plane, i.e., with the opposite sign. Please make the notation consistent.","section":"Appendix, Eq. (A1) and (A3)"},{"comment":"The p-values are reported as '0.00'; since KS p-values cannot be exactly zero, report the actual bound (for example, p < 10^-6).","section":"Table 1"},{"comment":"The phrase 'where P(z) denotes the unnormalized histogram' is unclear; please specify exactly how Psi_GRB,i is computed from the histogram of z_i.","section":"Section 3.1"},{"comment":"The caption has formatting errors ('d f /dzin', 'figures 4c and 4d , respectively'); please fix them.","section":"Figure 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's strongest point is the controlled simulation; the weakest is the breadth of the conclusion. I recommend revision rather than rejection because the fix is either to add a hierarchical forward-model test or to narrow the abstract and conclusions to direct pseudo-redshift inversion methods."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper: the narrow result is solid. In synthetic catalogs built from the Yonetoku, 3D Dainotti, and L-T-E correlations, inverting those same correlations to assign each burst a pseudo-redshift produces a distribution that is statistically incompatible with the true redshift distribution, across a range of assumed GRB rate shapes. The KS tests and rate ratios are decisive, and the appendix gives a clean geometric explanation of the low-redshift pile-up. This is a legitimate extension of the authors' earlier individual-level work to population-level claims.\n\nWhat they do well is make it a controlled self-consistency test: they generate bursts from the very relations they then try to invert, so the failure is not an artifact of unknown physics or selection effects. That is the right experimental design. They also test three different correlations, including the higher-dimensional planes, and show the inversion gets worse, not better, with more parameters. For anyone using pseudo-redshifts to infer rates or luminosity functions, this is a useful cautionary result.\n\nThe soft spot is the abstract. It says empirical GRB correlations alone cannot serve as reliable distance indicators at the population level. That is stronger than what they actually test. Their inversion is per-object point estimation (root-finding), plus one Gaussian-likelihood variant that uses the true intrinsic mean and covariance of the simulated population. They do not test a hierarchical forward model that marginalizes over latent redshifts with a population prior, which is precisely the approach they recommend in the discussion. Such a model could in principle recover the rate even when individual pseudo-redshifts are biased. Without testing that route, the unconditional population-level conclusion is an extrapolation. The reader's stress-test note is correct on this point.\n\nTwo smaller issues. The Gitlab link in the paper is blank, so the simulations are not publicly reproducible as shipped. And the luminosity function is fixed; the argument that varying it would not change the core result is plausible, but it is not demonstrated.\n\nBottom line: this deserves a serious referee and, if accepted, the authors should temper the abstract and either ship the code or add a hierarchical test. I'd bring it to a GRB or astro-statistics reading group, but the broad claim needs revision before I'd trust it at face value.","headline":"Solid controlled simulation showing direct pseudo-redshift inversion is biased, but the abstract's population-level conclusion overreaches what the tests actually cover.","tokens_in":12999,"tokens_out":2721,"would_cite":false,"duration_ms":31381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even in idealized selection-free simulations, empirical GRB correlations cannot recover the GRB rate.","keywords":["gamma-ray bursts","GRB rate","pseudo-redshift","Yonetoku relation","3D Dainotti relation","L-T-E correlation","redshift inference","population studies"],"falsifier":"Construct the same synthetic population and fit a full forward generative model that treats the correlation plane, intrinsic scatter, luminosity function, and selection function as latent variables and marginalizes over them; if the posterior for $\\Psi_{\\mathrm{GRB}}$ contains the injected value while the per-burst $z_i$ remain biased, then the blanket claim that the correlations cannot serve population studies is falsified.","tokens_in":11963,"feed_emoji":"🔭","tokens_out":13107,"duration_ms":118147,"temperature":0.7,"pith_summary":"Gamma-ray bursts with measured redshifts are rare, so several empirical correlations—between spectral peak energy and luminosity (the Yonetoku relation) and between X-ray plateau luminosity, prompt luminosity or energy, and plateau duration (the 3D Dainotti and $L$–$T$–$E$ relations)—have been used to assign pseudo-redshifts and then study the GRB population. This paper tests whether that population-level use can work, even under the idealized assumption that the correlations are intrinsic and free of detector selection biases. The authors generate a synthetic catalog of 4000 bursts with known true redshifts, invert each correlation to infer pseudo-redshifts, and compare the inferred redshift distribution and GRB rate $\\Psi_{\\mathrm{GRB}}$ with the input values. For all three correlations and every width of the intrinsic redshift distribution, the inferred distribution is statistically incompatible with the true one in a two-sample Kolmogorov-Smirnov test (p-value effectively zero), with rate errors at $z=0$ growing from a factor $\\sim 2$ for the Yonetoku relation to $\\sim 10$ and $\\sim 100$ for the 3D Dainotti and $L$–$T$–$E$ relations. The paper concludes that empirical GRB correlations alone cannot serve as reliable distance indicators, for individual bursts or for population studies.","feed_headline":"GRB pseudo-redshift correlations cannot recover the true burst rate","feed_subtitle":"Simulated bursts show Yonetoku, 3D Dainotti, and L-T-E pseudo-redshifts distort the redshift distribution and the inferred GRB rate.","key_machinery":"The central object is the parametric redshift locus $Y(z; E_{p,o}, f_\\gamma) = (L_{p,z}(z), E_{p,z}(z))$ for the Yonetoku relation, together with the analogous $D(z; T^*_{a,o}, f_x, f_\\gamma)$ and $L(z; S_\\gamma, T^*_{a,o}, f_x)$ curves for the 3D Dainotti and $L$–$T$–$E$ correlations. Each observer-frame observation defines one such curve, and the inferred pseudo-redshift is the point where that curve intersects the best-fit correlation plane. The appendix shows these curves rise steeply at low redshift, so their roots concentrate at $z\\lesssim1$, especially for bursts whose true X-ray luminosity lies above the plane. This intersection-plus-Kolmogorov-Smirnov machinery carries the negative result: the distribution of intersections is compared with the injected true distribution, and that comparison fails for every tested width.","core_discovery":"The central claim is that no empirical GRB correlation can recover the redshift distribution needed for population studies, even when the correlation is assumed intrinsic and no selection effects are included. In the mock catalog, each inferred pseudo-redshift $z_i$ is set by tracing the parametric curve $Y(z; E_{p,o}, f_\\gamma)$ (or the analogous $D$ and $L$ curves for the 3D Dainotti and $L$–$T$–$E$ planes) and taking its intersection with the best-fit correlation plane. For the Yonetoku relation the Kolmogorov-Smirnov statistic against the true redshifts is 0.13–0.20 across the tested widths; for 3D Dainotti it is 0.23–0.26 and for $L$–$T$–$E$ it is 0.26–0.30, all with p-values effectively zero. Many bursts yield no solution or two solutions (about 22% and 6% for 3D Dainotti; 19% and 2.5% for $L$–$T$–$E$), and inferred redshifts pile up below $z\\sim1$ because the plane-intersection curves rise steeply at low redshift. The paper concludes that higher-dimensional correlations are not better distance indicators than the simpler Yonetoku relation, and that a flux-limited subsample, equivalent to a narrower intrinsic distribution, does not fix the problem.","pith_inferences":["The paper tests direct root-finding and one Gaussian-likelihood variant that uses the true intrinsic mean and covariance, but it does not test a full Bayesian hierarchical forward model; such a model could in principle recover $\\Psi_{\\mathrm{GRB}}$ even when individual $z_i$ are biased, so the blanket cannot-serve conclusion is strongest against inversion-based estimators rather than against every","Varying the intrinsic scatter of the correlations while keeping the plane fixed would isolate whether the failure is driven by scatter or by the functional form; the current design holds scatter fixed and varies only the width of the redshift distribution.","The appendix's asymmetry between bursts with $L_X > L_{\\rm plane}$ and $L_X < L_{\\rm plane}$ suggests a possible analytic correction for the low-redshift pile-up, but the paper does not pursue one.","If these results hold for real data, existing pseudo-redshift catalogs contain systematically biased redshift distributions, and re-fitting published population constraints with a forward model would provide a direct observational check."],"forward_implications":["Any published GRB rate or luminosity-function result built on pseudo-redshifts from the Yonetoku, 3D Dainotti, or $L$–$T$–$E$ correlations rests on inferred redshift distributions that this paper shows are incompatible with the true ones.","The failure is not a selection-effect artifact: the paper's idealized mock catalog includes no detector selection, so the problem lies in the geometry of the inversion itself.","Moving from the two-parameter Yonetoku relation to the three-dimensional fundamental planes makes population inference worse, not better, because of their larger scatter and steep low-redshift plane-intersection curves.","Population-level GRB studies should therefore move to forward modeling that fits the full data, including spectral shapes, light curves, and selection functions, rather than plugging empirical correlations into inverse redshift estimates."],"supporting_citations":[{"why":"It supplies the best-fit Yonetoku correlation parameters ($a_Y=0.625$, $b_Y=-30.22$) used to define the plane for pseudo-redshift extraction.","marker":"D. Yonetoku et al. 2010"},{"why":"It supplies the 3D Dainotti fundamental-plane parameters ($C_0=15.75$, $\\alpha=0.67$, $\\beta=-0.77$) and the dispersion used to simulate and invert that relation.","marker":"M. G. Dainotti et al. 2016"},{"why":"It supplies the $L$–$T$–$E$ fundamental-plane parameters and dispersion used for the corresponding mock catalog and inversion.","marker":"C. Deng et al. 2023"},{"why":"It provides the parametrized GRB redshift distribution whose characteristic width $c$ is varied across the simulations.","marker":"P. Madau & M. Dickinson 2014"},{"why":"It provides the intrinsic dispersion $\\sigma_{\\log E_{p,z}}=0.25$ added when generating mock bursts along the Yonetoku relation.","marker":"G. Ghirlanda et al. 2005"},{"why":"It supplies the Planck 2018 flat-$\\Lambda$CDM cosmology ($H_0$, $\\Omega_M$) used in the luminosity-distance calculation.","marker":"N. Aghanim et al. 2020"},{"why":"It provides the Fermi limiting flux used in the flux-limited test showing that selection does not rescue the correlations.","marker":"A. Goldstein et al. 2017"},{"why":"It is the preceding paper whose analytic results on individual-redshift failures of the Amati and Yonetoku relations this work extends to population-level inference.","marker":"E. S. Yorgancioglu et al. 2025"},{"why":"It is the source of the two-sided Kolmogorov-Smirnov test used to compare the true and inferred redshift distributions.","marker":"J. L. Hodges 1958"},{"why":"It supplies the k-correction formula applied to Swift/BAT and XRT bandpasses for the 3D Dainotti and $L$–$T$–$E$ simulations.","marker":"J. S. Bloom et al. 2001"}],"fun_headline_variants":["GRB correlations fail to recover true burst rate","Pseudo-redshifts distort GRB population studies","No empirical GRB correlation can judge burst rate","Yonetoku, Dainotti, L-T-E all fail as distance rulers","Idealized GRB correlations still miss the redshift"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion assumes that a population study must recover the distribution of individual pseudo-redshifts $z_i$, and the paper tests direct root-finding plus one Gaussian-likelihood variant but not a full Bayesian hierarchical forward model that could in principle recover $\\Psi_{\\mathrm{GRB}}$ even when individual inferred redshifts are biased.","fun_headline_variants_meta":{"raw":{"variants":["GRB correlations fail to recover true burst rate","Pseudo-redshifts distort GRB population studies","No empirical GRB correlation can judge burst rate","Yonetoku, Dainotti, L-T-E all fail as distance rulers","Idealized GRB correlations still miss the redshift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1332,"prompt_tokens":999,"completion_tokens":333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":253}},"tokens_in":615,"tokens_out":333,"duration_ms":4038,"temperature":1.0,"reasoning_tokens":253,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:16:21.818847+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct the same synthetic population and fit a full forward generative model that treats the correlation plane, intrinsic scatter, luminosity function, and selection function as latent variables and marginalizes over them; if the posterior for $\\Psi_{\\mathrm{GRB}}$ contains the injected value while the per-burst $z_i$ remain biased, then the blanket claim that the correlations cannot serve population studies is falsified.","supporting_citations":[{"cited_title":"2010, Publications of the Astronomical Society of Japan, 62, 1495, doi: 10.1093/pasj/62.6.1495","cited_arxiv_id":null,"evidence_quote":"It supplies the best-fit Yonetoku correlation parameters ($a_Y=0.625$, $b_Y=-30.22$) used to define the plane for pseudo-redshift extraction."},{"cited_title":"S., Du, Y.-F., Yi, S.-X., et al","cited_arxiv_id":null,"evidence_quote":"It is the preceding paper whose analytic results on individual-redshift failures of the Amati and Yonetoku relations this work extends to population-level inference."}],"review_version":1}