{"id":"8aecf8d2-3ae8-4a8a-bda7-4926e665dc08","arxiv_id":"1908.09452","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"ZEBRA, a new zenith-dependent analysis framework, produces daily gamma-ray light curves for Crab and Mrk 421 consistent with the existing LiFF method, with the promise of shorter-timescale fluxes.","lead":"A new HAWC analysis framework called ZEBRA computes gamma-ray fluxes using a zenith-dependent detector response, and its daily light curves for the Crab and Mrk 421 agree with the previous HAWC method. If validated further, ZEBRA could let astronomers measure blazar flares on timescales shorter than a full nightly transit.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Daily-binned validation does not support ZEBRA's advertised short-timescale (sub-transit) capability, so the consistency claim is not yet established for the paper's main advantage.","rationale":"The reader's conditional verdict is appropriate. The paper presents a plausible implementation of ZEBRA and a reasonable visual/quantitative comparison on daily bins, but the central advertised advantage of arbitrary timescales is not validated by any sub-transit test. This is not an internal inconsistency; rather, the load-bearing assumption that daily agreement implies unbiased hour-scale fluxes is unsupported. The concrete simulation test would settle whether the zenith-dependent response and quality cuts are reliable at short timescales. Because the paper is an ICRC proceedings contribution and the missing demonstration is a scope-of-validation issue rather than a demonstrated error, keeping the conditional verdict is the right call.","tokens_in":4040,"tokens_out":3483,"duration_ms":37712,"concrete_test":"Run ZEBRA on simulated HAWC data with a known constant Crab-like source injected through the collaboration's Monte Carlo chain, then fit fluxes in 1-hour time windows. Construct the pull (flux_fit - flux_true)/sigma_fit for all windows; if the pull distribution deviates from N(0,1) or shows a trend with zenith angle, the daily-timescale consistency does not transfer to short timescales.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that ZEBRA and LiFF give consistent flux measurements, and its stated advantage is the capability to derive fluxes for arbitrary timescales, including hours or less (Sections 2.2 and 4). The only quantitative validation is daily-binned light curves for the Crab and Mrk 421 (Section 3, Figures 1 and 2). A daily bin averages over a full transit, so the source is observed across the full range of zenith angles and the convolution of detector response with exposure is tested only in aggregate. Systematic errors in the zenith-dependent PSF, in the background model, or in quality cuts could partially cancel or average down over a transit while becoming significant in hour-scale windows. The conclusion states \"We have proved that both methods LiFF and ZEBRA give consistent results\", but the evidence is limited to daily timescales; no simulation, sub-transit cross-check, or test statistic (e.g., chi-square per degree of freedom) is presented. The advertised future studies (Bayesian Blocks, X-ray/gamma-ray correlation, variability) depend precisely on this untested short-timescale regime. This is the load-bearing gap: if the zenith-dependent response is miscalibrated, the daily comparison could still appear consistent while sub-transit fluxes are biased.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ZEBRA, a new analysis framework for deriving gamma-ray light curves from HAWC data in which the detector response is computed as a function of zenith angle and convolved with the actual exposure of a given time window, in contrast to the previous LiFF method that assumes a full-transit exposure. Using the first 17 months of HAWC data, the authors derive daily-binned light curves for the Crab and Mrk 421 with ZEBRA and compare them with those obtained with LiFF. They report that the flux difference is less than 10%, consistent with systematic uncertainties, and conclude that the two methods give consistent results, thus validating ZEBRA for future studies of variability on short timescales.","tokens_in":4264,"tokens_out":2668,"duration_ms":28258,"significance":"If the consistency claim holds and ZEBRA indeed delivers unbiased fluxes on arbitrary timescales, the framework would enable HAWC to probe intra-night variability of blazars, a key capability for constraining emission models. The paper also strengthens the multi-messenger and variability program of HAWC by providing a more flexible analysis tool. However, the significance is conditional: the manuscript itself provides only a visual and qualitative comparison on daily timescales, so the load-bearing claim of short-timescale capability is not yet demonstrated.","major_comments":[{"comment":"The quantitative evidence for consistency is limited to the statement that the flux difference is 'less than %10' and 'consistent with the systematic uncertainties.' No test statistic is provided: the lower panels show (flux_ZEBRA - flux_LiFF)/sigma_ZEBRA, but the reader is not told the mean, standard deviation, or fraction of points exceeding 2 sigma or 3 sigma of this pull distribution. I request a quantitative comparison, e.g., the chi-square per degree of freedom of the two light curves, the mean and RMS of the pull, and a histogram or table of the daily differences, so that the 'consistent results' claim can be checked rather than taken on visual inspection.","section":"Section 3, Figures 1 and 2"},{"comment":"The advertised main advantage of ZEBRA is the capability to derive fluxes for arbitrary timescales, including windows of hours or less, yet the only validation presented is on daily bins. A daily bin averages over a full transit, so the zenith-dependent point spread function and background systematics are only tested in aggregate. This leaves open the possibility that systematic errors which average out over a transit become significant in hour-scale windows. To support the central claim of the paper's stated advantage, please provide either a simulation study of bias as a function of time-window length, a cross-check on sub-transit bins, or at least a comparison of ZEBRA on half-transit or two-hour windows with an independent estimator.","section":"Sections 2.2 and 4"}],"minor_comments":[{"comment":"In the Mrk 421 model description, 'an exponential cutoff at 5 eV' should almost certainly read '5 TeV'; as written, the cutoff energy is physically implausible and inconsistent with the HAWC energy range.","section":"Section 3"},{"comment":"The phrase 'There is an overall flux difference less than %10' should read 'less than 10%.' It would also be helpful to specify whether this refers to the mean absolute difference, the maximum, or a per-point property, and whether it is computed before or after applying the <0.5 coverage cut.","section":"Section 3"},{"comment":"The lower panels plot (flux_ZEBRA - flux_LiFF)/sigma_ZEBRA, but sigma_ZEBRA is not defined; state explicitly whether it is the statistical uncertainty of the ZEBRA flux only or includes systematic contributions.","section":"Section 3, Figures 1 and 2"},{"comment":"The paper states that the previous LiFF method 'assumes a minimum exposure on a full transit of the source,' but the discussion of the 'correction factor' in the previous work is vague; a more precise description of what the correction factor does and why it fails for short windows would help the reader evaluate the claimed advantage of ZEBRA.","section":"Section 1"},{"comment":"The comparison with HESS for the Crab at >1 TeV is mentioned, but the energy range of the ZEBRA and LiFF results (300 GeV to 100 TeV) is not explicitly restated in the comparison; clarify whether the HESS comparison is on the integrated flux or on a specific energy band.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"This is a short ICRC proceedings paper, so the authors may have been constrained by space, but the central validation claim is central to the paper's contribution and is currently supported only by visual inspection and a single percentage statement. The requested additional analyses are feasible and should be added in a revision. I would also note that the paper does not include the figures in the provided text, which makes it hard to assess the visual agreement; the final version should ensure figures are fully legible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is ZEBRA, a Zenith Band Response Analysis framework that convolves a zenith-dependent detector response with exposure information, letting HAWC in principle fit fluxes on arbitrarily short time windows. That is a real technical improvement over LiFF, which assumes a full transit. The paper compares ZEBRA against LiFF on the same 17-month Crab and Mrk 421 datasets and shows that the daily light curves agree to within 10% for the Crab, with the Mrk 421 comparison looking similar. For an ICRC proceedings, that is a reasonable and useful validation.\n\nWhat the paper does well: it uses the same selection criteria and spectral models as the earlier LiFF analysis, so the comparison is clean; the stated <10% difference is consistent with LiFF's systematic uncertainties; and the motivation for the new framework is clearly explained. The two figures show no obvious disagreements, and the claim that ZEBRA reproduces the previously published light curves is plausible.\n\nThe soft spot is exactly what the stress-test note says: the advertised sub-transit capability is never demonstrated. All validation is on daily bins. A daily bin averages over a full transit, so zenith-dependent errors in the PSF, background, or quality cuts could partially cancel, while still biasing hour-scale fluxes. The paper offers no simulation, no sub-transit cross-check, and no test statistic beyond an eyeball comparison and a percentage. The word \"proved\" in the conclusion is too strong for what is shown. This matters because the planned applications (Bayesian blocks, X-ray/gamma-ray correlations) depend on exactly that untested short-timescale regime.\n\nMinor points: the Mrk 421 cutoff is given as \"5 eV\" in Section 3, which is presumably a typo for TeV. No data or code artifacts are provided, which is common for conference proceedings but would be needed for a full journal paper.\n\nBottom line: a solid, incremental technical contribution for the HAWC collaboration. If this is submitted to a journal, it deserves a serious referee, but the referee should ask for a short-timescale validation—ideally on simulated data or on a flaring interval with sub-transit bins—before accepting the central capability claim. I would engage with it, but treat the arbitrary-timescale advantage as an assertion, not a demonstrated result.","headline":"ZEBRA is a sound new analysis tool for HAWC, but the paper proves consistency only on daily timescales, not the advertised sub-transit capability.","tokens_in":4810,"tokens_out":2077,"would_cite":false,"duration_ms":21380,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ZEBRA, a new HAWC analysis framework, reproduces the previously published gamma-ray light curves of Mrk 421 and the Crab, with overall flux differences under ten percent.","keywords":["HAWC","Mrk 421","Crab nebula","gamma-ray light curves","ZEBRA","LiFF","blazar variability","very-high-energy gamma rays"],"falsifier":"A day-by-day residual analysis between ZEBRA and LiFF on all 17 months that shows a substantial fraction of days with deviations above 2σ, or an hour-scale ZEBRA light curve around a Mrk 421 flare that disagrees with simultaneous observations, would falsify the claim.","tokens_in":3844,"feed_emoji":"🔭","tokens_out":6041,"duration_ms":56275,"temperature":0.7,"pith_summary":"This paper tries to establish that a new gamma-ray flux estimation framework, ZEBRA, gives the same answers as the established LiFF likelihood method when applied to HAWC data for Mrk 421 and the Crab nebula. Using the first 17 months of HAWC observations, the authors build daily light curves with ZEBRA and compare them day by day with the light curves previously reported by the HAWC collaboration. They find overall flux differences below ten percent, within the systematic uncertainties assigned to the earlier analysis, and no day-to-day variations larger than 2σ between the two methods. If the claim holds, ZEBRA can replace LiFF for these sources and the physics conclusions drawn from the original light curves remain unchanged, while gaining the ability to estimate fluxes on shorter time windows.","feed_headline":"New HAWC flux method reproduces Mrk 421 light curves","feed_subtitle":"Zenith-based ZEBRA analysis agrees with the previous likelihood method within 10 percent.","key_machinery":"The central object is ZEBRA (Zenith Band Response Analysis), a Monte-Carlo-based framework that characterizes the HAWC detector response as a function of zenith angle. It convolves that response with the actual exposure time at each zenith angle to estimate counts from a source over an arbitrary period, updating the point spread function per zenith band. This replaces LiFF's approach of computing detector response as a function of declination under a minimum full-transit exposure assumption. The zenith-resolved response is the mechanism that removes the need for an overall correction factor when short time windows are wanted.","core_discovery":"On its own terms, the paper's discovery is that a flux measurement is independent of the fitting framework: ZEBRA and LiFF produce consistent results for a flux measurement of the same source and data. Concretely, for both the Crab and Mrk 421, the ZEBRA and LiFF daily light curves agree in their high and low flux states, have similar average fluxes within uncertainties, and show no differences above 2σ. The paper states that the overall flux difference is less than 10%, consistent with the systematic uncertainties considered for LiFF in the 17-month analysis. This consistency is what allows the authors to conclude that future analyses with ZEBRA will recover the same physics, such as Bayesian-block variability structure and correlations with other wavelengths.","pith_inferences":["Beyond the paper, the same zenith-resolved machinery should apply to other HAWC-monitored blazars, so the consistency test could be repeated on Mrk 501 or a fainter source with little additional work.","The paper asserts, but does not demonstrate, that ZEBRA works on sub-transit time scales; a natural next test is to compute hour-scale fluxes around a known Mrk 421 flare and check their residuals against simultaneous observations.","If the 10% agreement band reflects a systematic floor, then future ZEBRA-based variability claims will need to show that flux changes exceed that floor rather than just statistical errors."],"forward_implications":["Future HAWC light curves for Mrk 421 and the Crab can be produced with ZEBRA, and the physics derived from them should match the earlier 2017 analysis.","Bayesian Blocks, X-ray/gamma-ray correlation studies, and variability analyses planned with ZEBRA can proceed on the strength of this consistency check.","The zenith-dependent response opens a route to flux measurements on time scales shorter than a full transit, where LiFF required a correction factor.","The small overall flux difference can be read as the approximate systematic floor for ZEBRA's daily flux points, comparable to the systematic uncertainties already quoted for LiFF."],"supporting_citations":[{"why":"Supplies the previously published HAWC light curves and the selection criteria, spectral models, and quality cuts that ZEBRA reproduces and compares against.","marker":"[Abeysekara et al.(2017)]"},{"why":"Describes the LiFF likelihood fitting framework that ZEBRA is compared to and that defines the baseline method for HAWC flux estimation.","marker":"[Younk et al.(2015)]"},{"why":"Provides independent very-high-energy Crab flux measurements used to check that ZEBRA's Crab flux is consistent with an external measurement.","marker":"[Aharonian et al.(2006)]"}],"fun_headline_variants":["New HAWC method matches prior Mrk 421 light curves","ZEBRA and LiFF agree on Mrk 421 flux within 10%","HAWC's ZEBRA reproduces Mrk 421 variability","Mrk 421 light curves consistent across HAWC methods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that agreement with LiFF on daily binned light curves guarantees unbiased fluxes on much shorter time windows, since ZEBRA's stated advantage of arbitrary timescales is asserted but not validated with hour-scale data.","fun_headline_variants_meta":{"raw":{"variants":["New HAWC method matches prior Mrk 421 light curves","ZEBRA and LiFF agree on Mrk 421 flux within 10%","HAWC's ZEBRA reproduces Mrk 421 variability","Mrk 421 light curves consistent across HAWC methods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1557,"prompt_tokens":889,"completion_tokens":668,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":592}},"tokens_in":505,"tokens_out":668,"duration_ms":6377,"temperature":1.0,"reasoning_tokens":592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:10:07.478968+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A day-by-day residual analysis between ZEBRA and LiFF on all 17 months that shows a substantial fraction of days with deviations above 2σ, or an hour-scale ZEBRA light curve around a Mrk 421 flare that disagrees with simultaneous observations, would falsify the claim.","supporting_citations":[{"cited_title":"U., Albert, A., Alfaro, R., et al.\\ 2017, , 841, 100","cited_arxiv_id":null,"evidence_quote":"Supplies the previously published HAWC light curves and the selection criteria, spectral models, and quality cuts that ZEBRA reproduces and compares against."},{"cited_title":"W., Lauer, R","cited_arxiv_id":null,"evidence_quote":"Describes the LiFF likelihood fitting framework that ZEBRA is compared to and that defines the baseline method for HAWC flux estimation."},{"cited_title":"G., Bazer-Bachi, A","cited_arxiv_id":null,"evidence_quote":"Provides independent very-high-energy Crab flux measurements used to check that ZEBRA's Crab flux is consistent with an external measurement."}],"review_version":1}