{"id":"853320e3-3db3-41ec-823f-cd4cc1392613","arxiv_id":"2607.27821","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A Benford's Law goodness-of-fit metric, applied to sliding windows of seismic amplitude, is reported to detect local earthquake onsets across two Indian seismic networks.","lead":"This paper tests whether the digit pattern known as Benford's Law can flag earthquakes in continuous seismic recordings. The authors report that a Benford goodness-of-fit measure spikes at earthquake onsets and detects 268 of 275 cataloged events in the Deccan Volcanic Province.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 3 defines φ = (1−χ)×100%, but χ is the χ² statistic from Eq. 2 (typically ~8), so φ would be negative; reported φ values near 90% imply an unstated normalization.","rationale":"The reader's weakest assumption correctly identifies the Goodness-of-Fit conversion in Eq. 3 as a key unstated dependency. My reading confirms that the equation is mathematically inconsistent with Eq. 2: a chi-square statistic typically near 8 cannot yield φ values above 90% under φ = (1 − χ) × 100%. All results — the temporal φ curves, the normalized deviations, the detection thresholds, and the headline 97.5% recall — inherit this inconsistency. Even if one treats Eq. 3 as a typographical shorthand for a p-value or reduced chi-square, the paper does not define the actual formula, making the method non-reproducible. This is an internal soundness problem rather than a matter of disagreeing with the community consensus on Benford's Law in seismology. The proposed test — recomputing φ for a representative window from the raw equations — settles whether the published numerical results can be reproduced from the text. If they cannot, the central claim is unsupported as written. I therefore agree with the reader's REJECT verdict; no adjustment is needed.","tokens_in":10315,"tokens_out":5647,"duration_ms":55155,"concrete_test":"Take one of the waveforms used in Figure 2 (e.g., Event 1 at station ABR) and recompute the φ(t) trace exactly as specified: for each sliding window, compute χ² from Eq. 2, then apply Eq. 3 literally. If the maximum φ is not close to 95% as plotted but instead is negative, the reported results rely on an unstated normalization. Alternatively, if the authors supply the actual formula, verify it against Eq. 3 and check whether the detection metrics remain unchanged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing flaw is the definition of the Goodness-of-Fit φ in Eq. 3. In Eq. 2, χ² is the chi-square statistic Σ (O_d − E_d)² / E_d, a sum of nonnegative terms. For a perfectly BL-distributed window, the expected value of χ² is 8 (df = 9 digits − 1), and typical values are on the order of 5–15. Eq. 3 then gives φ = (1 − χ) × 100%, which for χ ≈ 8 yields −700%, not the 90–100% values reported throughout the paper. Thus the equation as written cannot produce the φ(t) traces in Figures 2, 4, 5, or 6. Either the symbol χ in Eq. 3 denotes some other quantity (e.g., a p-value, reduced χ², or normalized misfit) that is never defined, or the figures were generated with a different formula. Since every downstream quantity — Δφ_N (Eq. 4), Δφ'_N (Eq. 7), the thresholds φ>90%, Δφ_N>0.5, Δφ'_N>0.5, and the 97.5% detection rate — is computed from φ, the central claim that BL provides a reliable detector is not reproducible from the stated methodology. This is an internal inconsistency, not merely a disagreement with consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Benford's Law (BL) based framework for local earthquake detection in continuous seismic waveforms. Using sliding windows, the authors compute a chi-square misfit between observed first-digit distributions and the BL prediction, convert it to a Goodness-of-Fit measure φ, and derive two normalized deviation metrics Δφ_N and Δφ'_N. They apply the method to two Indian datasets (Deccan Volcanic Province and NAMASTE/Himalaya), report that earthquake onsets show enhanced BL conformity relative to pre-event noise, claim detection of 268/275 cataloged events (97.5%) with no false positives, and argue that BL-based detection is parameter-light and training-free compared to STA/LTA and machine-learning pickers.","tokens_in":10671,"tokens_out":3169,"duration_ms":34464,"significance":"If the claims were supported, the paper would offer an appealing low-cost, training-free detector with potential use in noisy or data-sparse environments. The use of two contrasting tectonic datasets is a strength, and the idea that BL conformity changes at signal onset is worth investigating. However, the central methodological definition of φ is internally inconsistent, the performance evaluation is circular, and the 'no false positives' claim is not meaningful given the experimental design. These issues are load-bearing, so the contribution as presented cannot be accepted or reproduced from the stated methods.","major_comments":[{"comment":"Eq. (2) defines χ² as a sum of nonnegative terms with expected value ≈8 for 9 digits. Eq. (3) then defines φ = (1−χ)×100%, which for any χ>1 gives a negative percentage. The paper reports φ values near 90–100% throughout (e.g., Figures 2, 4, 5, 6), so the equation as written cannot produce those curves. Either χ in Eq. (3) denotes some unstated normalized quantity (e.g., reduced χ² or a p-value) or the figures were generated with a different formula. Because every downstream metric, threshold, and detection rate depends on φ, this internal inconsistency makes the central claim non-reproducible.","section":"§2.2, Eq. (3)"},{"comment":"The detection thresholds (φ>90%, Δφ_N>0.5, Δφ'_N>0.5) are described in §3.4 as 'empirically determined,' and the same DVP waveforms are then used to report the 97.5% recall and 'no false positive' results in §3.3 and §3.4. There is no train/validation split, independent test set, or out-of-sample check. The false-positive claim is especially problematic: §2.1 states that only waveform segments around cataloged events are analyzed, so the study never scans continuous noise-only intervals; 'no false positives within the analyzed time windows' does not establish a false-positive rate for continuous detection.","section":"§3.3 and §3.4"},{"comment":"The detection rate is reported as 268 out of 275 cataloged events, but Figure 3a shows 2403 DVP waveforms. It is not stated how the 275 events map to these waveforms (multi-station recordings per event), nor what criterion counts an event as detected (any one station? a minimum number of stations? a single window?). Without this aggregation rule, the 97.5% recall figure is ambiguous and cannot be independently verified.","section":"§3.3"},{"comment":"Eq. (7) defines Δφ'_N(t) = 1 − (φ_max − φ(t))/(100 − φ(t)). Algebraically this is (100 − φ_max)/(100 − φ(t)), which measures how close φ(t) is to its maximum in units of the distance to 100%; it is not a temporal gradient. The text repeatedly states that Δφ'_N 'emphasizes the gradient of φ' and 'highlight[s] the initial transition from noise-dominated to signal-dominated behavior,' but the formula contains no time derivative or finite-difference operator. This mismatch between the stated purpose and the actual definition is not merely cosmetic, since Δφ'_N is one of the three required detection criteria.","section":"§3.4, Eq. (7)"}],"minor_comments":[{"comment":"Typo: 'parameter-lite' should be 'parameter-light' (also in the abstract the same phrase is correctly hyphenated).","section":"§4"},{"comment":"A full paragraph beginning 'While Δφ_N effectively quantifies...' is repeated nearly verbatim in §3.3 and §3.4. Please remove the duplicate.","section":"§3.3/§3.4"},{"comment":"Δφ_N(t) is defined using φ_max, the maximum of the entire trace, so at times before the event the metric uses future information. If the intended use is real-time detection, the authors should state whether φ_max is replaced by a causal running maximum in practice; as written, the definition is not causal.","section":"§2.2, Eq. (4)"},{"comment":"The color scale is labeled φmax(%) but the axis label and caption do not explain how the color value relates to the plotted points; please clarify.","section":"Figure 3c"},{"comment":"The statement that 'φ values remain consistently low during the pre-event noise window' is only supported by example figures; no aggregate statistics for pre-event φ are provided. A histogram or median curve over all waveforms would strengthen the claim.","section":"§3.1"}],"recommendation":"reject","confidential_remarks":"The reader's report and my own reading converge: Eq. (3) is internally inconsistent with Eq. (2), and the performance evaluation is circular. The paper would need a fundamentally rewritten methodology and an independent validation protocol before it could be reconsidered; that goes beyond a routine revision. I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'm inclined to agree with the reader: this manuscript has a load-bearing flaw in its central formula. The Goodness-of-Fit metric φ (Eq. 3) is defined as (1−χ)×100%, where χ is the chi-square statistic from Eq. 2. If χ is the usual chi-square sum, it is nonnegative and typically around 8 for 9 digits, so (1−8)×100% is negative. Yet the paper reports φ values near 90–100%. This means the equation as written cannot generate the traces shown in Figures 2, 4, 5, or 6. Either the symbol χ denotes something else (a p-value? reduced chi-square?) or the figures came from an unstated normalization. Since every downstream quantity — Δφ_N, Δφ′_N, the φ>90% threshold, and the 97.5% recall — is computed from φ, this is a fundamental reproducibility problem, not a minor typo.\n\nWhat the paper does well: the idea itself is worth a look. The authors apply Benford's Law to two contrasting datasets in India, present a careful window-length sensitivity analysis, and compare with STA/LTA. The STA/LTA illustration showing how parameter choices affect triggering is a useful pedagogical point. The empirical observation that the BL conformity rises near P onsets, visible in several figures, is suggestive and may be real. The manuscript is clearly structured and the writing is direct.\n\nThe soft spots go beyond the formula issue. The thresholds in §3.4 are described as 'empirically determined' and then applied to the same DVP waveforms to claim 97.5% recall and no false positives. That is circular. False positives are only assessed within segments already known to contain cataloged events, not on continuous untriggered data, so the 'no false positives' claim is not established. Using TauP/AK135 travel times as the reference for onset accuracy is weaker than manual picks, though the paper acknowledges this.\n\nIn my view, the paper deserves to be sent back to the authors rather than passed to reviewers in this form. The central equation needs to be fixed and the evaluation redesigned. I'd desk reject this version and invite a resubmission once the methodology is clarified.","headline":"The Benford-detection idea is appealing, but the central Goodness-of-Fit definition in Eq. 3 is mathematically inconsistent, so the reported 97.5% recall is not reproducible.","tokens_in":11206,"tokens_out":3248,"would_cite":false,"duration_ms":31849,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simple digit-counting rule can detect local earthquakes and pinpoint P-wave onsets without training.","keywords":["Benford's Law","earthquake detection","P-wave onset","goodness of fit","seismic waveforms","leading digits","chi-square","STA/LTA comparison"],"falsifier":"Compute φ on 200 seconds of pre-event noise from the same stations using the paper's exact Eq. 3: if the 95th percentile φ_th approaches the 90% detection threshold, or if φ frequently exceeds 90% on noise, the detector's baseline is not separating noise from signal. Alternatively, compare BL-derived onsets against analyst-picked P arrivals for a set of high-SNR local events; if the median |Δt| exceeds one window length, the onset-alignment claim fails.","tokens_in":10160,"feed_emoji":"🌍","tokens_out":4158,"duration_ms":37061,"temperature":0.7,"pith_summary":"This paper argues that the leading digits of seismic waveform amplitudes follow Benford's Law during the onset of earthquake energy but not during pre-event noise, and that this contrast can serve as a detection and onset-timing tool. Across a year-long intraplate array in the Deccan Volcanic Province and a Himalayan aftershock network, a sliding-window goodness-of-fit measure rises sharply at P-wave arrivals, detecting 268 of 275 cataloged events (≈97.5%) with no false positives. Crucially, the method requires no training, templates, or amplitude thresholds—only a window length—making it a candidate pre-filter for noisy or data-sparse settings such as planetary seismology. The paper positions the result as a statistical-transition detector rather than a sample-level phase picker.","feed_headline":"Benford's Law detects 97.5% of local earthquakes","feed_subtitle":"A sliding-window digit test on waveforms spots P-wave onsets using only a window length, with zero false positives in the study.","key_machinery":"The carrying object is the sliding-window first-digit goodness-of-fit: each 1-second-stepped window tallies the leading digits of absolute amplitudes, compares them to Benford's distribution via a chi-square statistic, and converts the misfit into a Goodness-of-Fit measure φ(t). Two derived metrics—the noise-normalized deviation Δφ_N and the gradient-emphasizing Δφ′_N—convert raw φ into a detection trigger. Together they transform a digit-frequency histogram into an onset detector whose only tunable parameter is window length; the work this does is to separate 'noise-like' digit statistics from 'Benford-like' signal statistics.","core_discovery":"The central claim is that local earthquake waveforms consistently conform to Benford's Law—the logarithmic distribution of first significant digits—during the onset of seismic energy, while pre-event noise does not. Using a sliding window to track the temporal evolution of a chi-square-based goodness-of-fit, the authors define normalized deviation metrics that compare signal windows against station-specific noise baselines. Applied to two contrasting tectonic datasets, the metrics show maximum BL conformity within about one window length of theoretical P-wave arrivals, and a combined detection criterion recovers 97.5% of cataloged events with no false positives. The authors conclude that BL","pith_inferences":["If BL conformity is genuinely a property of transient broadband seismic energy, the same sliding-window test might be adapted to detect tremor, volcanic signals, or debris flows—any source that changes the amplitude distribution's statistical character—without modifying the core metric.","The claim that noise does not conform to BL could be made into a sharper test: measure φ on noise-only records from different stations; a stable φ_th near 100% would undermine the separation.","The onset precision limited by window length suggests a two-pass design: BL to find candidate windows, then a short-window re-analysis or an ML picker inside those windows—implicit in the paper but not developed.","The 97.5% recall is measured against cataloged events; the method's practical value for unknown events depends on false positives over long continuous time, which the paper only evaluates in event-triggered windows."],"forward_implications":["A threshold-free, training-free detector that can be run on continuous records at trivial computational cost.","BL-based onset timing aligns with theoretical P arrivals within one window length, so it can seed phase association or ML pickers.","Because only window length matters, the method transfers across tectonic environments and noise conditions without retuning.","In planetary or remote deployments with no templates or labeled data, the method offers a first-pass event detection layer.","Compared with STA/LTA, it avoids false triggers that arise from amplitude-threshold tuning."],"fun_headline_variants":["Benford's Law catches quakes at 97.5% with zero false positives","No training needed: Benford's Law detects local earthquakes","Sliding-window digit test spots P-wave onsets","Quake detection via Benford's Law: simple, no tuning"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The results assume that the Goodness-of-Fit formula φ=(1−χ)×100% really converts the chi-square misfit into a percentage near 100% for signal windows, and that the theoretical TauP/AK135 P-wave arrival times used as reference onsets are accurate to within a sliding-window length; if either fails, the 97.5% recall and onset-alignment claims are not established.","fun_headline_variants_meta":{"raw":{"variants":["Benford's Law catches quakes at 97.5% with zero false positives","No training needed: Benford's Law detects local earthquakes","Sliding-window digit test spots P-wave onsets","Quake detection via Benford's Law: simple, no tuning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1550,"prompt_tokens":701,"completion_tokens":849,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":774}},"tokens_in":445,"tokens_out":849,"duration_ms":6774,"temperature":1.0,"reasoning_tokens":774,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:46:22.248268+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute φ on 200 seconds of pre-event noise from the same stations using the paper's exact Eq. 3: if the 95th percentile φ_th approaches the 90% detection threshold, or if φ frequently exceeds 90% on noise, the detector's baseline is not separating noise from signal. Alternatively, compare BL-derived onsets against analyst-picked P arrivals for a set of high-SNR local events; if the median |Δt| exceeds one window length, the onset-alignment claim fails.","supporting_citations":[],"review_version":1}