{"id":"3714ef8e-d8cf-4295-9648-4375187a8609","arxiv_id":"2504.15685","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A combined model with both Lorentz-violation and energy-dependent intrinsic delays recovers simulated parameters across all mock datasets and fits multi-GeV and TeV GRB photons, yielding a subluminal LV scale of about 3 x 10^17 GeV.","lead":"This paper tests three models for gamma-ray burst photon time delays, where one model adds an energy-dependent intrinsic delay to a possible Lorentz-invariance-violation delay. Using simulations and real photon data, it argues the combined model best fits the data and implies a Lorentz-violation scale of about 3 x 10^17 GeV.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation-based model-discrimination claims rest on 1000 photons per GRB, not the 17-photon real sample; the paper's robustness claims are accordingly overstated.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the simulation uses 1000 high-energy photons per GRB while the real dataset contains only 17 photons. This is a genuine weakness because the paper's strongest quantitative claims about Model A/B failure are calculated under the dense-sample mock data. The real-data analysis in Sec. V is largely a reproduction of previous work by the same authors, and while it shows a nonzero aLV in Model C, the 'catastrophic' and 'conclusively demonstrated' language in the conclusions is not supported by the realistic sample size. My concern does not move the verdict because the reader already judged the paper CONDITIONAL; it reinforces that conditionality. I agree with the reader's assessment that the methods are sound and the framework may be useful, but the paper overstates the robustness of its model discrimination by relying on unrealistic photon counts. The proposed concrete test would directly quantify whether the reported significance levels survive at realistic sample sizes. No ad hominem is intended; the issue is solely with the match between the simulation design and the real data it is meant to emulate.","tokens_in":14592,"tokens_out":4566,"duration_ms":45645,"concrete_test":"Rerun the full 250-seed simulation suite with realistic event counts: for each of the 10 GRBs, draw the same number of high-energy photons as in the actual 17-photon sample (e.g., 1-3 photons per GRB plus the three TeV photons), using the observed energy and redshift distributions and the same injected parameters and noise model. Then recompute the posterior tensions of Models A and B on Sets A-C. If the >4.7 sigma biases collapse below ~1-2 sigma, the model-discrimination claims are an artifact of the 1000-photon mock data; if they persist, this concern is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Model C robustly identifies a subluminal LV scale E_LV ~ 3e17 GeV rests on two legs: recovery tests on mock datasets and a fit to 17 real photons. The recovery tests (Secs. II-IV) generate 1000 high-energy photons per GRB, while the real sample contains only 17 photons (14 Fermi-LAT multi-GeV photons, one 99.3 GeV Fermi photon, one MAGIC TeV photon, and one LHAASO multi-TeV photon). The paper's quantitative demonstrations of Model A/B failure (>4.7 sigma, 6.9 sigma, 11 sigma, 12 sigma in Tables I-III and Sec. IV) are all computed at that dense-sample limit. With 17 photons and sparse TeV coverage, posterior widths would scale roughly as sqrt(1000/17) ~ 7.7 larger if the information content per photon were similar, so these tensions would be expected to shrink dramatically. The real-data fit in Fig. 4 does show a nonzero aLV, but the 'catastrophic failure' language and the claim that energy-dependent intrinsic delays are 'conclusively demonstrated' are supported mainly by the unrealistic mock data. The authors themselves concede in Sec. V.D that 'further investigation is still required,' which is appropriately cautious. The headline conclusion is therefore conditional: the framework may be useful, but the quantitative model-discrimination claims have not been shown to survive at realistic photon counts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a Monte Carlo study of three competing time-delay models for gamma-ray burst (GRB) spectral lags in the context of Lorentz-invariance violation (LV): Model A (LV delay plus a constant intrinsic delay), Model B (energy-dependent intrinsic delay only), and Model C (LV delay plus a linear energy-dependent intrinsic delay). Mock datasets of ten GRBs, each containing 1000 high-energy photons drawn from a GRB 221009A-like spectrum, are generated under each model, and Bayesian parameter estimation (Eqs. (12)-(16)) is used to test parameter recovery (Sec. IV, Tables I-III). The authors find that Models A and B exhibit 4.7-12 sigma biases when applied to data generated under alternative mechanisms, while Model C recovers the injected parameters in all cases. The same machinery is then applied to a realistic sample of 17 photons (14 Fermi-LAT multi-GeV photons plus one 99.3 GeV, one MAGIC 1.07 TeV, and one LHAASO 12.2 TeV photon), with Model C reported to yield a subluminal LV scale E_LV about 3 x 10^17 GeV consistently across datasets. The paper concludes that energy-dependent intrinsic delays are essential and that Model C is the definitive framework for future LV searches.","tokens_in":14873,"tokens_out":39109,"duration_ms":345062,"significance":"The Bayesian machinery appears correctly implemented: the marginal likelihoods in Eqs. (12)-(15) properly integrate out the per-photon intrinsic delay with parameter-dependent normalization, and the 250-seed repetition of the injection-recovery test is a reasonable way to check the pipeline. The single-GRB degeneracy a_LV f(z) + alpha identified in Sec. V.A is correct and clearly explained, and the closure tests do establish that the four-parameter model family is identifiable from multi-GRB data at high statistics. However, the paper's central quantitative claims - that Models A and B fail 'catastrophically' at >5 sigma and that energy-dependent intrinsic delays are required by the data - are established only for mock datasets with about 10^4 photons, roughly 600 times the real sample, and the model discrimination is never quantified on the 17 observed photons (the AIC comparison is cited to the authors' own Ref. [33]). The headline E_LV of about 3 x 10^17 GeV is a fitted parameter under Model C, not a prediction, and the 'validation' is a closure test with injections taken from the authors' own earlier posteriors.","major_comments":[{"comment":"The mock datasets used for every quantitative discrimination claim contain 1000 high-energy photons per GRB (Sec. II: 'we randomly draw 1000 high-energy photons ... for each GRB'), i.e., 10^4 photons per set, while the realistic analysis of Sec. V uses 17 photons. The reported tensions - 11 sigma and 4.7 sigma in the Set A/B analyses, 6.9 sigma and 12 sigma in Set C, and the Sec. VI statement that Models A and B 'exhibit catastrophic failures ... exceeding about 5 sigma' - are all computed from the dense mock posteriors of Tables I-III, whose widths shrink roughly as N^{-1/2} in the information-dominated limit. These significances therefore do not transfer to the 17-photon sample, and the paper never performs an injection-recovery or model-selection test at realistic photon counts. The real-data evidence for model discrimination is presented only qualitatively in Sec. V.C (the AIC-based preference for Model C is cited to the authors' own Ref. [33] and not computed here), so the abstract/conclusion claims that intrinsic delays must be energy-dependent are not established at the actual sample's statistical weight. A realistic-count simulation (e.g., using the actual 17 observed energies, with roughly two photons per burst) or an explicit Bayes-factor/evidence computation on the 17-photon sample should be added.","section":"II, IV, VI"},{"comment":"The recovery tests are closure tests by construction. The injected parameters in Tables I-III are the posterior modes of the same authors' fits in Refs. [28, 32, 33] (as the captions state and Sec. VI confirms: parameters 'injected from real-data posterior distributions'), and the mock generator draws the common intrinsic delay independently for each photon (Sec. IV), exactly matching the likelihood's N(mu, upsilon^2) noise model in Eqs. (12)-(15). Such tests verify the internal consistency of the machinery and the identifiability of the model family at high statistics, but they cannot provide independent support for the physical content of Model C, because any misspecification of the noise model or of the Taylor truncation in Eq. (5) would be reproduced identically in generation and recovery. The concluding phrase 'To validate our framework' (Sec. VI) therefore overstates the evidential value; the tests should be labeled as internal-consistency/closure checks, with the independent case for Model C resting on the real-data analysis of Sec. V.","section":"IV, VI"},{"comment":"The manuscript never lists the observed arrival-time differences Delta_t_obs of the 17 real photons or the residuals of the best-fit Model C, so the central claim of consistency at E_LV about 3 x 10^17 GeV cannot be verified from the text. This is not a formality: for GRB 221009A, Eq. (8) with E_LV about 3 x 10^17 GeV assigns an LV delay of about 2.8 x 10^3 s to the 12.2 TeV photon, which the fit must nearly cancel with an intrinsic delay of the opposite sign (alpha E_h,s about -2.8 x 10^3 s); the predicted arrival time is thus the small difference of two large terms, and whether the 17-photon posterior actually reproduces the observed near-simultaneous TeV arrival depends on the joint (alpha, mu) correlations, which are not shown. I request a table of observed versus predicted delays (with 68% credible bands) for all 17 photons under each model, and a stress test of the linear extrapolation of alpha from about 0.1 TeV to about 14 TeV (e.g., adding a beta E^2 term and checking the evidence), since the TeV-photon 'consistency' rests entirely on that linear form.","section":"V, Fig. 4, Eq. (8)"}],"minor_comments":[{"comment":"Sec. VI states that the recovery achieves 'subpercent biases in E_LV and alpha estimates,' but Table III (Set C) shows a recovered a_LV that differs from the injected value by about 3% and a recovered alpha differing by about 7%; the wording should be 'percent-level' or the numbers should be recomputed.","section":"VI"},{"comment":"Sec. II justifies truncating the intrinsic-delay Taylor series at the linear term by noting that beta and gamma 'converge toward 0 through data fitting' in Ref. [32]; since this truncation defines Model C, the manuscript should briefly display that evidence (e.g., fitted beta, gamma and their uncertainties) or explicitly refer to it as an assumption inherited from Ref. [28].","section":"II"},{"comment":"Fig. 4 is difficult to parse in the compiled text because the numerical values and axis labels are interleaved in the captions and the panel structure is not self-evident; a single clean figure with separately labeled panels per model (posterior means and 68% credible intervals for a_LV, alpha, mu, sigma) would be much clearer.","section":"Fig. 4"},{"comment":"Eqs. (13)-(15) treat the intrinsic delay of every photon, including multiple photons from the same GRB, as an independent draw from N(mu, upsilon^2); the justification for this independence (as opposed to a per-GRB common offset) should be stated, since it affects the interpretation of the 14+3-photon constraints.","section":"III"},{"comment":"Sec. II's opening sentence ('Physics is an experimental science...') is rhetorical and unrelated to the mock-data construction; it could be removed. Also, please clarify whether the same spectral template (Eq. (9), fitted to GRB 221009A) is applied to bursts at other redshifts without a k-correction, and whether that is a deliberate choice.","section":"II"},{"comment":"For a Monte Carlo study of this kind, the mock-generation and fitting code and the list of the 17 observed photons (energies, arrival times, uncertainties) should be made available as supplemental material or a repository for reproducibility.","section":"II, IV"}],"recommendation":"major_revision","confidential_remarks":"The manuscript header indicates the work is published (Phys. Rev. D 111, 103015), and the paper functions largely as a companion to the authors' own Refs. [28, 32, 33]: the real-data fit (Fig. 4) and the AIC-based model preference invoked in Sec. V.D are taken from those papers, and the 'validation' injections are the same papers' posterior modes. The novel element is the Monte Carlo closure study; the citation pattern is heavily self-referential, which is a scope/novelty consideration for the editor. Given the dense-mock overclaims (major comment 1) and the missing residual verification (major comment 3), I would want the revision to include a realistic-count simulation and a residual table before the paper's conclusions are taken at face value. The framing of the conclusions as a 'definitive framework' is not supported by the demonstrated evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a Monte Carlo closure test for the authors' three GRB time-delay models. The simulations are careful and the Bayesian machinery is sound, but the model-discrimination headlines rest on mock data with 1000 photons per GRB, while the real sample has 17 photons. The quantitative failure claims for Models A and B do not survive a realistic photon count, so this is a useful methodological exercise rather than a definitive result.\n\nWhat's new: the systematic 3x3x250 simulation (three generative models, three fitted models, 250 noise seeds) is new, and it does show that Model C can recover its own parameters and that simpler models develop biases when fit to data generated under other assumptions. That is a legitimate extension of earlier work, and the likelihood functions and priors look correct. The authors wisely include a sentence in Sec. V.D admitting that further investigation is required, which is more cautious than the conclusions section.\n\nThe soft spots are real. First, the simulations draw 1000 high-energy photons per GRB from 2 GeV to 20 TeV, while the actual analysis uses 14 Fermi-LAT photons plus one 99.3 GeV Fermi photon, one MAGIC TeV photon, and one LHAASO multi-TeV photon. Posterior widths in real data are roughly an order of magnitude wider, so the 12-sigma and \"catastrophic failure\" numbers in Tables I-III are artifacts of the dense mock sample. With 17 photons the tensions shrink dramatically, even if the qualitative ranking of models remains. Second, the injected parameters in all three sets come from the authors' own posterior fits (Ref. [28]); the recovery test is a self-consistency closure check, not independent validation. It does not test the models against, say, a different functional form for intrinsic delays or a different LV dispersion. Third, the real-data constraint on E_LV is a fitted parameter with large uncertainty (for 14+3 photons, a_LV = 3.34 +/- 0.90 x 10^-18, about a 3.7-sigma hint), so calling this a \"robust identification\" of 3 x 10^17 GeV is generous. Finally, the AIC model comparison is only cited to their own previous paper, not shown here, so the reader cannot check the model-selection evidence.\n\nThe paper deserves a serious referee because the simulation framework could be useful for planning future LV analyses, e.g., for CTA. But the authors should be asked to rerun the model comparison at realistic photon counts, report actual AIC values, and soften the \"definitive\" and \"catastrophic\" language. As it stands, it's an incremental contribution within an established program, and the headline claims outrun the evidence.","headline":"A careful but self-referential Monte Carlo closure test that overstates model discrimination because the mock data are far denser than the real 17-photon sample.","tokens_in":15407,"tokens_out":4469,"would_cite":false,"duration_ms":39467,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GRB time delays need an energy-dependent intrinsic term before Lorentz-violation claims hold.","keywords":["Lorentz invariance violation","gamma-ray bursts","spectral lag","intrinsic time delay","Bayesian parameter estimation","TeV photon","GRB 221009A","subluminal LV scale"],"falsifier":"Re-analyze the 14+3 real photons with a model that lets $\\alpha$ vary per GRB or adds a quadratic intrinsic term; if the posterior for $a_{\\rm LV}$ then includes zero or the bound moves above the Planck scale, the claimed $3\\times 10^{17}$ GeV detection would not survive. Alternatively, a future bright GRB with many TeV photons whose Model-C fit forces $\\alpha$ consistent with zero would remove the need for the energy-dependent intrinsic delay.","tokens_in":14328,"feed_emoji":"⚡","tokens_out":5100,"duration_ms":41860,"temperature":0.7,"pith_summary":"This paper argues that previous claims of Lorentz-invariance violation from gamma-ray burst (GRB) photon lags rest on an incomplete model. It compares three delay models and shows, through Monte Carlo mock datasets and Bayesian fits, that only a model combining a Lorentz-violating (LV) propagation delay with an energy-dependent intrinsic emission delay (Model C) recovers the injected parameters in all simulated universes. Simpler alternatives—LV plus a constant intrinsic delay (Model A) or energy-dependent intrinsic delay without LV (Model B)—fail badly on datasets generated by the other mechanism, with tensions above 4.7σ. Applied to 14 Fermi-LAT multi-GeV photons plus three TeV photons, Model C yields a subluminal LV scale $E_{\\rm LV} \\simeq 3\\times 10^{17}$ GeV, consistent with earlier Fermi-LAT constraints. The practical stake: if right, future LV searches must jointly fit source physics and propagation effects or they will mistake astrophysics for quantum gravity.","feed_headline":"Simulations back one GRB delay model, pinning the LV scale near 3e17 GeV","feed_subtitle":"Model C, which adds an energy-dependent intrinsic delay, recovers parameters where simpler alternatives fail on mock and real data.","key_machinery":"The load-bearing object is the combined delay relation $\\Delta t_{\\rm obs}/(1+z) = a_{\\rm LV} K_1 + \\alpha E_{h,s} + \\mu$ (with scatter $\\upsilon$), which unifies the LV propagation term and a linear intrinsic emission delay. A Taylor-expansion argument motivates the intrinsic term: any analytic source-frame emission time can be expanded in powers of energy, and fits to real data drive the quadratic and cubic coefficients to zero, leaving the linear term $\\alpha$. The redshift function $f(z)$ in the single-burst form $C = a_{\\rm LV} f(z) + \\alpha$ shows that $a_{\\rm LV}$ and $\\alpha$ are completely degenerate for photons from one GRB; the degeneracy is broken only by combining multiple GRBs at different redshifts, which is what the multi-GRB likelihood does. On top of this, the mock-data machinery draws 1000 high-energy photons per burst from the GRB 221009A spectrum and repeats generation-plus-fit 250 times with different noise seeds to suppress sampling bias.","core_discovery":"The paper's central claim is that the observed arrival-time lags of high-energy GRB photons are best described by Model C: $\\Delta t_{\\rm obs}/(1+z) = a_{\\rm LV} K_1 + \\alpha E_{h,s} + \\Delta t_{\\rm in,c}$, where $a_{\\rm LV}=1/E_{\\rm LV}$ is the inverse LV scale, $K_1$ a redshift-dependent propagation factor, and $\\alpha E_{h,s}$ an intrinsic delay linear in source-frame energy. Using mock datasets generated under each of the three models and analyzed with the same Bayesian machinery, the paper reports that Model C recovers the injected $a_{\\rm LV}$ and $\\alpha$ within 1σ for all three datasets, while Models A and B show >4.7σ biases when applied to data generated under an alternative mechanism (up to 12σ for $\\alpha$ in some cases). On real data—14 Fermi-LAT photons from eight GRBs plus the 99.3 GeV, 1.07 TeV, and 12.2 TeV photons from GRB 221009A and GRB 190114C—Model C gives a consistent subluminal scale $E_{\\rm LV} \\simeq 3\\times 10^{17}$ GeV, with a negative $\\alpha$ indicating high-energy photons are emitted earlier in the source frame. The paper therefore positions Model C as the generalized framework for future LV searches.","pith_inferences":["If the source-frame intrinsic delay is truly linear in energy, the same Model-C machinery could be applied to other transients, such as fast radio bursts or AGN flares, where a similar degeneracy between propagation and emission effects exists.","The paper's mock GRBs contain 1000 high-energy photons each, whereas the real dataset has only 17 photons; a natural follow-up is to re-run the recovery tests with realistic sparse samples to see how the >4.7σ verdicts scale.","A testable prediction is that the inferred $\\alpha$ should remain stable as more GRBs are added; if $\\alpha$ drifts toward zero with a larger sample, the claimed LV scale would need upward revision."],"forward_implications":["If Model C is correct, earlier analyses that assumed a constant intrinsic delay (Model A) systematically overestimate the LV scale by absorbing energy-dependent source delays into $a_{\\rm LV}$.","Energy-dependent intrinsic delays must be included in all future LV fits; without them, fits to TeV photons and multi-GeV photons cannot be reconciled.","The LV scale $E_{\\rm LV} \\simeq 3 \\times 10^{17}$ GeV (subluminal, $n=1$) is consistent across Fermi-LAT, MAGIC, and LHAASO data, making it a cross-instrument constraint.","The negative $\\alpha$ under Model C implies that high-energy photons are emitted before low-energy photons in the GRB source frame, a physical claim about jet emission.","Next-generation TeV observatories such as LHAASO and CTA can use Model C for joint $(E_{\\rm LV},\\alpha,z)$ reconstruction with roughly a thousand GRBs."],"supporting_citations":[{"why":"Introduces the three competing models and the energy-dependent intrinsic delay framework that this paper tests.","marker":"[28]"},{"why":"Provides the Fermi-LAT-only Model C analysis and the $E_{\\rm LV} \\simeq 2.96\\times 10^{17}$ GeV result extended here.","marker":"[32]"},{"why":"Gives the 14+3 photon analysis and the Akaike-information-criterion model selection favoring Model C.","marker":"[33]"},{"why":"Earlier constant-intrinsic-delay Fermi-LAT analysis that Model C reproduces as a special case.","marker":"[22]"},{"why":"Companion earlier analysis of GRB 160509A with constant intrinsic delay, used as the prior baseline for Model A.","marker":"[23]"},{"why":"LHAASO detection of the GRB 221009A >10 TeV photon used in the combined real dataset.","marker":"[30]"},{"why":"MAGIC detection of the GRB 190114C TeV photon used in the combined real dataset.","marker":"[31]"},{"why":"LHAASO-WCDA spectrum of GRB 221009A used as the universal spectral template for generating mock photons.","marker":"[34]"}],"fun_headline_variants":["Simulations pick GRB model C, LV scale near 3e17 GeV","GRB lag analysis favors intrinsic delay, sets LV scale ~3e17 GeV","Model C with intrinsic delay best fits GRBs, LV scale 3e17 GeV","Monte Carlo GRB tests settle LV model, scale ~3e17 GeV","Energy-dependent delay key to GRB LV fit, scale 3e17 GeV"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline conclusions rest on mock datasets with 1000 high-energy photons per GRB, whereas the real data contain only 17 photons; if the simulated data are not representative of the sparse real observations, the claimed failures of Models A and B may be overstated.","fun_headline_variants_meta":{"raw":{"variants":["Simulations pick GRB model C, LV scale near 3e17 GeV","GRB lag analysis favors intrinsic delay, sets LV scale ~3e17 GeV","Model C with intrinsic delay best fits GRBs, LV scale 3e17 GeV","Monte Carlo GRB tests settle LV model, scale ~3e17 GeV","Energy-dependent delay key to GRB LV fit, scale 3e17 GeV"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1951,"prompt_tokens":1151,"completion_tokens":800,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":767,"completion_tokens_details":{"reasoning_tokens":693}},"tokens_in":767,"tokens_out":800,"duration_ms":6746,"temperature":1.0,"reasoning_tokens":693,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:19:44.359689+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the 14+3 real photons with a model that lets $\\alpha$ vary per GRB or adds a quadratic intrinsic term; if the posterior for $a_{\\rm LV}$ then includes zero or the bound moves above the Planck scale, the claimed $3\\times 10^{17}$ GeV detection would not survive. Alternatively, a future bright GRB with many TeV photons whose Model-C fit forces $\\alpha$ consistent with zero would remove the need for the energy-dependent intrinsic delay.","supporting_citations":[{"cited_title":"Lorentz Invariance Violation from Gamma-Ray Bursts","cited_arxiv_id":"2504.00918","evidence_quote":"Provides the Fermi-LAT-only Model C analysis and the $E_{\\rm LV} \\simeq 2.96\\times 10^{17}$ GeV result extended here."},{"cited_title":"Examining Lorentz invariance violation with three remarkable GRB photons","cited_arxiv_id":"2504.14295","evidence_quote":"Gives the 14+3 photon analysis and the Akaike-information-criterion model selection favoring Model C."}],"review_version":1}