{"id":"e538d913-94ff-4027-9c57-67806095a20e","arxiv_id":"2506.20374","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Peak flux selection biases the best-fit Ep,i-Liso correlation of long GRBs, and a small set of low-luminosity high-flux bursts may be non-collapsars.","lead":"Using simulated long gamma-ray burst (GRB) populations, this paper shows that the observed correlation between peak energy and luminosity, a key tool for cosmology, depends on the detector's peak flux threshold. It also suggests that a handful of low-luminosity, high-flux bursts may have non-collapsar origins, which could refine how GRBs are classified.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulation imports the observed (Ep,i,Liso) distribution as intrinsic, making the claim that P-selection changes the correlation and that Ep,i–Liso is intrinsic potentially circular; the no-correlation test is not shown.","rationale":"The paper's central claim has two parts: (1) the observed (Ep,i,Liso) distribution and the best-fit Ep,i–Liso relation depend on peak flux P, so selection effects must be modeled; (2) the Ep,i–Liso connection is intrinsic, as shown by the failure of simulations without it to reproduce the observed P distribution. The first part is a useful practical warning and is consistent with previous work, but it is not a completely surprising result because P is constructed from Liso and z. The second part rests on the simulation's input: Section 2.1 step 3 uses the observed Swift (Ep,i,Liso) distribution as the parent distribution for mock Ep,i. That observed distribution is itself the outcome of the P≥0.5 selection and the instrument's trigger and redshift-measurement probabilities. If the observed correlation is partly produced by selection, then feeding it back into the simulation and finding that the P distribution can be reproduced is not evidence for an intrinsic correlation. The paper does not display the no-correlation simulation or its KS result, so the key negative result is not verifiable from the manuscript. This is the weakest load-bearing assumption, matching the reader's analysis. The proposed test directly addresses it by comparing the P distribution under a no-correlation intrinsic model and by calibrating the pipeline on a known zero-correlation population. The subgroup identification using four bursts is a minor, clearly labeled speculation and does not affect the main concern. Given these issues, the reader's CONDITIONAL verdict remains appropriate: the practical selection-effect warning can be accepted, but the intrinsic-correlation claim should be conditional on the missing test and on using a selection-corrected intrinsic distribution. No change to the verdict is needed from this stress-test pass.","tokens_in":10819,"tokens_out":9026,"duration_ms":89144,"concrete_test":"Run the Section 2.1 pipeline with the same z and Liso draws, but sample mock Ep,i independently of Liso from the observed marginal Ep,i distribution, and apply the same Swift selection and 2D KS test to the resulting P distribution. If the no-correlation P CDF is not rejected at the 5% level, the Section 3.1 claim that an Ep,i–Liso correlation is mandatory is refuted; if it is rejected, repeat the test on a calibration sample with a known zero intrinsic correlation to check that the pipeline does not produce a false positive. Either outcome settles whether the observed (Ep,i,Liso) input distribution is driving the conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 step 3 draws mock Ep,i from the observed Swift (Ep,i,Liso) distribution of 259 bursts with P≥0.5 ph cm−2 s−1. That distribution is already shaped by the same P-selection the paper studies. When a mock Liso is drawn and Ep,i is then drawn from this observed conditional distribution, the simulated joint distribution is forced to contain whatever correlation—intrinsic or selection-induced—exists in the observed sample. Because P is computed from Liso, z, and the spectral parameters, different log P bins correspond largely to different Liso (and z) ranges; the differing best-fit slopes in Table 1 are therefore expected and do not independently demonstrate a selection bias on the intrinsic relation. The distinctive Section 3.1 claim—that the observed P distribution can only be reproduced if an Ep,i–Liso correlation is included—is asserted without showing the no-correlation model, its P CDF, or the KS statistic. If the no-correlation variant draws Ep,i independently from the observed marginal distribution, the failure could be caused by the marginal distribution being incompatible with the assumed luminosity function, not by the absence of an intrinsic correlation. Conversely, if the observed (Ep,i,Liso) correlation is partly selection-induced, importing it as the intrinsic distribution would trivially reproduce the observed P distribution. Thus the evidential value of the 'only if' statement for an intrinsic correlation is unclear. None of this removes the practical conclusion that P-limited samples must be handled carefully in Ep,i–Liso fits, but the stronger claim about the intrinsic nature of the correlation rests on an unvalidated assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper builds a simulated population of 10,000 long GRBs from the Swift/BAT world model of Lan et al. (2021) in order to study how peak-flux (P) selection affects the observed Ep,i--Liso distribution and the best-fit Ep,i--Liso correlation. Redshifts and isotropic luminosities are drawn from a cosmic star-formation-rate model, a triple power-law luminosity function, and a trigger/redshift-measurement probability function; rest-frame peak energies are drawn from the observed Swift 2D (Ep,i, Liso) distribution; spectral indices are drawn from the observed (alpha, Ep,o) distribution; and P is computed self-consistently from the spectral model. The mock sample passes 1D and 2D KS tests against the 259-burst Swift sample with P >= 0.5 ph cm^-2 s^-1. The authors find that the best-fit Ep,i--Liso slope depends on the P range considered (Table 1), that the observed P distribution can be reproduced only if an Ep,i--Liso correlation is included in the simulation, and that a small group of low-Ep,i, low-Liso, high-P bursts contains objects such as GRB 060614 and GRB 191019A that may have non-collapsar origins. The paper concludes that peak-flux selection must be taken into account in applications of the Ep,i--Liso relation.","tokens_in":11203,"tokens_out":6072,"duration_ms":73805,"significance":"The practical message of the paper--that selection on peak flux can change the apparent Ep,i--Liso relation and that standard-candle applications need to model this selection--is important and timely for the growing use of GRB correlations in cosmology. The simulation pipeline is detailed, the construction is explicitly described, and the use of KS tests on both 1D and 2D distributions is a good validation practice that goes beyond what is often done in this literature. If the central claims hold, the paper would strengthen the case for correcting selection effects before fitting the Ep,i--Liso relation. However, the evidence for the strongest claim--that the Ep,i--Liso correlation is an intrinsic property required to reproduce the P distribution--is weakened by a circular construction: the intrinsic 2D distribution is imported from the very observed sample whose selection effects are being studied. The paper is therefore best viewed as a useful forward-modeling exercise whose conclusions need a more independent test before they can be fully accepted.","major_comments":[{"comment":"The claim that the observed P distribution can be reproduced only if an Ep,i--Liso correlation is included in the simulation is asserted but never demonstrated. The paper states that the mock P distribution cannot be reproduced 'well' without the correlation, but it does not show the no-correlation variant: no P CDF, no KS statistic, and no description of what exactly was tried. This is a load-bearing point, since it is used to conclude that the Ep,i--Liso relation is a crucial intrinsic property of LGRBs. The authors should show the full comparison between the observed P distribution and the mock P distributions with and without an intrinsic Ep,i--Liso correlation, including the quantitative KS results.","section":"Section 3.1 and Section 4"},{"comment":"The mock Ep,i values are drawn from the observed Swift 2D (Ep,i, Liso) distribution of the same P >= 0.5 ph cm^-2 s^-1 sample whose selection effects are under study. If that observed 2D distribution is already distorted by the P selection, then the simulation imports the very effect it claims to discover. Because P is computed directly from Liso, z, and the spectral parameters, splitting the mock sample by P naturally selects different Liso and z ranges, and the different best-fit slopes in Table 1 are a necessary consequence of coupling Ep,i to Liso. This does not independently demonstrate that the intrinsic relation changes with P. A more convincing test would generate the intrinsic (Ep,i, Liso) distribution from an assumed functional form (e.g., a bivariate lognormal or a physically motivated correlation), apply the same selection machinery, and show that the observed P and 2D distributions are recovered only when the intrinsic correlation is present, while a no-correlation intrinsic model fails even after applying the same trigger and redshift-selection probabilities.","section":"Section 2.1, step 3"},{"comment":"The proposed identification of a distinct subgroup in the low-Ep,i, low-Liso, high-P region is based on four bursts selected by cuts that were chosen after inspecting the same diagram. The paper notes that two of the four (GRB 060614 and GRB 191019A) have been proposed to have non-collapsar origins, but it does not quantify the expected number of non-collapsar contaminants among LGRBs selected this way, nor does it apply the same cuts to a simulated collapsar-only population to estimate a false-positive rate. As a classifier, the proposed region needs a completeness and purity estimate; as evidence for a separate population, the present argument is anecdotal and should be either strengthened or explicitly labeled as a speculative illustration.","section":"Section 3.3"},{"comment":"The reported differences in best-fit slopes across P bins are not accompanied by a statistical test of whether the slopes differ beyond what is expected from the different dynamic ranges and sample sizes. The 1-sigma uncertainties in Table 1 do not account for the strong covariance between the fitted slope a and intercept b, nor for the fact that low-P and high-P bins occupy very different regions of the (Ep,i, Liso) plane. Reporting the Liso range, sample size, and confidence contours (or a fit with a single relation plus a selection term) would make the claim that the P dependence is significant and not merely a boundary effect more convincing.","section":"Table 1 and Section 3.2"}],"minor_comments":[{"comment":"The word 'instrinsic' should be corrected to 'intrinsic'.","section":"Section 4"},{"comment":"The selection probability function theta(P(L,z)) is used in the CDFs but is not explicitly defined in the text; a brief definition or a pointer to the corresponding Lan et al. (2021) equations would improve reproducibility.","section":"Equations (1)-(2)"},{"comment":"The caption states that the figure shows cumulative distributions of six quantities, but the text reproduction shows only one panel; if the published figure is multi-panel, the caption should clearly identify each panel, and if not, the caption should be revised.","section":"Figure 3"},{"comment":"The 4% of mock events with P below Plim are retained in the main analysis and then removed in a robustness check. It would be cleaner to exclude them from the beginning or to model the detection probability explicitly, rather than relying on a post-hoc removal.","section":"Section 3.4"},{"comment":"The KS acceptance criterion at the 5% level is applied to many 1D and 2D distributions; reporting the number of tests performed and the resulting multiple-testing caveat would be useful.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for astro-ph.HE and addresses a question of real practical importance. The main concern is circularity: the intrinsic (Ep,i, Liso) distribution is taken from the observed sample whose selection effects are the subject of the paper, and the critical no-correlation baseline is not shown. This is fixable with additional simulations and quantitative comparisons, so I am not recommending rejection, but the claims as written outrun the evidence. The classification suggestion in Section 3.3 would also benefit from a proper false-positive estimate before being highlighted in the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper makes one point I think is right and one that doesn't yet land. The right point: if you fit Ep,i–Liso on Swift-like samples, applying different peak-flux cuts gives you different best-fit slopes. Their Table 1 shows that cleanly in a mock, and it's a useful warning for anyone using this relation as a standard candle. The second point—that the Ep,i–Liso connection is a crucial intrinsic property because the mock can't reproduce the observed P distribution without it—is asserted but not demonstrated. Section 3.1 says the no-correlation version fails, but no CDF, no KS statistic, no figure. That's the load-bearing support for the \"intrinsic\" language, and it's missing.\n\nThe bigger problem is circularity. In step 3 of Section 2.1, mock Ep,i is drawn from the observed Swift (Ep,i, Liso) distribution of 259 bursts with P≥0.5. That observed distribution is already shaped by the very selection the paper studies. So when they bin the mock by P and see different slopes, they are re-discovering the selection effect they imported. It's a valid demonstration that P cuts matter for the observed correlation, but it doesn't tell you about the intrinsic relation. The no-correlation test, if shown properly with Ep,i drawn independent of Liso from a plausible marginal, could fix this—but it has to be shown, and the intrinsic distribution has to be constructed without using the biased sample.\n\nWhat the paper does well: the simulation is carefully built on the Lan et al. (2021) world model, the KS checks are standard, and the authors are transparent about mock artifacts like the 4% faint-P excess and the Ep,o fitting jump. The four-burst subgroup (060614, 161219B, 191019A, 230328B) is a suggestive post-hoc selection; two of the four are known non-collapsars, which is interesting but a sample of four can't carry much weight. The references to Butler et al., Shahmoradi, and Palmerio & Daigne are appropriate; this is a quantitative extension of known selection concerns, not a wholly new effect.\n\nWho's this for: anyone fitting spectral–luminosity correlations on Swift GRBs, and people building population models. It deserves peer review, but the referee should ask for the no-correlation baseline, a discussion of how the imported 2D distribution limits the intrinsic claim, and a softer framing of the subgroup result.","headline":"Useful caution about peak-flux cuts in Ep,i–Liso fits, but the claim that the correlation is intrinsic is not yet supported by the simulation as constructed.","tokens_in":11736,"tokens_out":2627,"would_cite":true,"duration_ms":28441,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that peak-flux selection significantly changes the observed ($E_{\\rm p,i}$, $L_{\\rm iso}$) distribution of long gamma-ray bursts, so the correlation's use as a standard candle requires correcting for this selection.","keywords":["gamma-ray bursts","long GRBs","Ep,i-Liso correlation","peak flux selection","selection effects","GRB luminosity function","GRB classification","standard candles"],"falsifier":"One concrete check would be to take a large, redshift-complete sample of long GRBs with a well-understood peak-flux selection function, correct the observed ($E_{\\rm p,i}$, $L_{\\rm iso}$) distribution for that selection, and then re-fit the correlation in separate peak-flux bins. If the corrected distributions in different $P$ bins coincide, the claim that $P$ reshapes the ($E_{\\rm p,i}$, $L_{\\rm iso}$) plane is contradicted. A simpler version within current data is to restrict the mock comparison to $L_{\\rm iso}\\ge10^{50.5}$ erg s$^{-1}$ and test whether the predicted ordering of best-fit slopes across $\\log P$ bins survives when the selection function is varied.","tokens_in":10608,"feed_emoji":"💥","tokens_out":10153,"duration_ms":91930,"temperature":0.7,"pith_summary":"This paper argues that peak-flux selection—the observational filter that decides which long gamma-ray bursts (LGRBs) are detected and get redshifts—significantly shapes the observed joint distribution of rest-frame peak energy $E_{\\rm p,i}$ and isotropic luminosity $L_{\\rm iso}$. Because the best-fit $E_{\\rm p,i}$–$L_{\\rm iso}$ correlation is computed from that distribution, its slope and intercept change when samples are cut by peak flux, so the relation cannot be used as a standard candle or as a probe of GRB physics without modeling this selection. To show this, the authors simulate a 10,000-burst population calibrated to Swift observations. They also find the observed peak-flux distribution is reproduced only when an intrinsic $E_{\\rm p,i}$–$L_{\\rm iso}$ connection is included, strengthening the case that the correlation is real, and they identify two bursts in the low-luminosity corner whose likely merger origins suggest the diagram plus peak flux can flag non-collapsar LGRBs.","feed_headline":"Peak-flux selection distorts the GRB luminosity correlation","feed_subtitle":"Brightness-cut samples yield different Ep,i-Liso slopes, so standard-candle use must correct for selection.","key_machinery":"The load-bearing instrument is a synthetic population of 10,000 LGRBs generated from the GRB world model of Lan et al. (2021): mock redshifts and luminosities are drawn from rate and luminosity functions that include trigger and redshift-measurement probabilities, mock rest-frame peak energies are sampled from the observed Swift ($E_{\\rm p,i}$, $L_{\\rm iso}$) distribution, mock spectral indices from the observed ($\\alpha$, $E_{\\rm p,o}$) distribution, and mock peak fluxes are then computed self-consistently from $z$, $L_{\\rm iso}$, $E_{\\rm p,o}$, and $\\alpha$. The sample is retained only if one- and two-dimensional Kolmogorov-Smirnov tests fail to reject agreement with the observed Swift sample at the 5% significance level. The argument proceeds by comparing best-fit $E_{\\rm p,i}$–$L_{\\rm iso}$ relations across peak-flux bins and redshift bins, and by checking whether the simulated $P$ distribution matches observations only when an intrinsic $E_{\\rm p,i}$–$L_{\\rm iso}$ dependence is present.","core_discovery":"On the paper's own terms, the central discovery is that the two-dimensional ($E_{\\rm p,i}$, $L_{\\rm iso}$) distribution of long GRBs is not a fixed population: it depends explicitly on the peak flux $P$. Fitting the $E_{\\rm p,i}$–$L_{\\rm iso}$ relation separately for different $\\log P$ ranges yields slopes from $0.27\\pm0.02$ to $0.37\\pm0.01$ across the full simulated sample, and the low- and high-$P$ populations barely overlap. The same simulation shows the observed $P$ distribution cannot be matched unless an intrinsic $E_{\\rm p,i}$–$L_{\\rm iso}$ dependence is built in, which the authors take as evidence that the correlation is a crucial physical property of LGRBs. The paper further claims that high-peak-flux bursts in the low-$E_{\\rm p,i}$, low-$L_{\\rm iso}$ corner are not the simple extrapolation of the high-peak-flux population, and that selecting $L_{\\rm iso}\\le10^{50}$ erg s$^{-1}$, $E_{\\rm p,i}\\le10^{2.5}$ keV, and $P\\ge10^{0.5}$ ph cm$^{-2}$ s$^{-1}$ picks out GRB 060614 and GRB 191019A, both plausibly of merger rather than massive-star origin.","pith_inferences":["Editorial extensions beyond the paper's claims: if the peak-flux dependence is real, the next necessary step is to invert the selection function and reconstruct the intrinsic two-dimensional distribution directly from data, rather than sampling the observed distribution as the simulation does.","A testable prediction follows for future samples with nearly complete redshift follow-up at lower $E_{\\rm p,i}$ and $L_{\\rm iso}$: the low-corner bursts should split into two populations, one with supernova associations and one without.","The same selection logic should apply to other GRB correlations such as $E_{\\rm p,i}$–$E_{\\rm iso}$ or multi-parameter relations, where peak-flux cuts may alter fitted slopes in an analogous way.","The paper's classification criterion is conditional on peak flux, which implies that proposed boundaries in the ($E_{\\rm p,i}$, $L_{\\rm iso}$) plane should be treated as functions of $P$, a concrete and checkable prediction."],"forward_implications":["Any flux-limited sample of LGRBs selects a specific slice of the ($E_{\\rm p,i}$, $L_{\\rm iso}$) plane, so the best-fit correlation from one sample should not be compared directly with another sample selected at a different peak-flux threshold.","Cosmological uses of the $E_{\\rm p,i}$–$L_{\\rm iso}$ relation, such as standardizing LGRB luminosities to build a Hubble diagram, need an explicit model of peak-flux selection; otherwise the derived distances can be biased.","The fact that the observed peak-flux distribution is reproduced only when an intrinsic $E_{\\rm p,i}$–$L_{\\rm iso}$ link is included supports the physical reality of the correlation rather than a purely selection-driven origin.","Selecting $L_{\\rm iso}\\le10^{50}$ erg s$^{-1}$, $E_{\\rm p,i}\\le10^{2.5}$ keV, and $P\\ge10^{0.5}$ ph cm$^{-2}$ s$^{-1}$ isolates bursts such as GRB 060614 and GRB 191019A, suggesting the diagram combined with peak flux can identify long GRBs with non-collapsar progenitors.","Earlier estimates of the $E_{\\rm p,i}$–$L_{\\rm iso}$ relation were probably dominated by luminous bursts with moderate-to-high peak flux, so the full population may have a wider and differently sloped correlation."],"supporting_citations":[{"why":"supplies the GRB world model—star-formation-based redshift distribution and triple-power-law luminosity function—from which mock z and Liso are drawn.","marker":"Lan et al. (2021)"},{"why":"independently argued for an intrinsic Ep,i-Liso connection and motivated the faint-end peak-flux cut used in sample construction.","marker":"Palmerio & Daigne (2021)"},{"why":"provides the reference best-fit Ep,i-Liso relations for total and complete (P > 2.6 ph cm^-2 s^-1) samples that the new fits are compared against.","marker":"Nava et al. (2012)"},{"why":"introduced the Ep,i-Liso correlation that is the object of the selection-effect study.","marker":"Yonetoku et al. (2004)"},{"why":"supplies the two-dimensional Kolmogorov-Smirnov test used to validate that the mock sample matches the observed Swift distributions.","marker":"Peacock (1983)"},{"why":"established that GRB 060614 has no supernova association, supporting its use as a non-collapsar example.","marker":"Gehrels et al. (2006)"},{"why":"identified GRB 191019A as a merger-origin burst in a dense stellar environment, supporting its use as a non-collapsar example.","marker":"Levan et al. (2023)"}],"fun_headline_variants":["Peak flux skews the GRB Ep,i-Liso relation","Brightness selection distorts GRB standard-candle relation","GRB luminosity correlation shifts with peak flux cut","Peak flux reveals possible merger-origin GRBs","Low-luminosity GRBs at high peak flux point to mergers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the observed Swift two-dimensional ($E_{\\rm p,i}$, $L_{\\rm iso}$) distribution used to seed the simulation is itself an unbiased picture of the intrinsic GRB population; if that distribution is already distorted by the same peak-flux selection under study, the simulated $P$-dependence could be a circular artifact.","fun_headline_variants_meta":{"raw":{"variants":["Peak flux skews the GRB Ep,i-Liso relation","Brightness selection distorts GRB standard-candle relation","GRB luminosity correlation shifts with peak flux cut","Peak flux reveals possible merger-origin GRBs","Low-luminosity GRBs at high peak flux point to mergers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1456,"prompt_tokens":1293,"completion_tokens":163,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":909,"completion_tokens_details":{"reasoning_tokens":81}},"tokens_in":909,"tokens_out":163,"duration_ms":2468,"temperature":1.0,"reasoning_tokens":81,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:49:16.600227+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete check would be to take a large, redshift-complete sample of long GRBs with a well-understood peak-flux selection function, correct the observed ($E_{\\rm p,i}$, $L_{\\rm iso}$) distribution for that selection, and then re-fit the correlation in separate peak-flux bins. If the corrected distributions in different $P$ bins coincide, the claim that $P$ reshapes the ($E_{\\rm p,i}$, $L_{\\rm iso}$) plane is contradicted. A simpler version within current data is to restrict the mock comparison to $L_{\\rm iso}\\ge10^{50.5}$ erg s$^{-1}$ and test whether the predicted ordering of best-fit slopes across $\\log P$ bins survives when the selection function is varied.","supporting_citations":[],"review_version":1}