{"id":"608a0634-9543-4395-8bb3-6811fa9008a3","arxiv_id":"2606.16511","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A pre-registered protocol rejects tail-shape claims in standard LLM toxicity evaluation as false positives under two different scorer families.","lead":"The paper proposes a pre-registered protocol with gates for admissibility, goodness-of-fit, threshold stability, and effect size to test tail-index claims in LLM evaluation. A smart generalist might read it to learn why statistical tail claims in AI testing can produce false positives even when they look convincing.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Protocol's admissibility/GOF/threshold-stability gates may reject genuine tail shape in the small-tail regimes typical of LLM toxicity data.","rationale":"The reader's weakest assumption is precisely the load-bearing point: whether the gates are calibrated to avoid both false positives and false negatives. The proposed synthetic-data check directly tests that calibration without requiring the full manuscript details.","tokens_in":1694,"tokens_out":309,"duration_ms":13395,"concrete_test":"Generate 500 synthetic toxicity-score vectors from a Pareto(α=1.8) distribution with tail mass and sample size matched to the paper's real data; apply the exact pre-registered gates (admissibility, GOF, threshold stability, effect size) and report the fraction that pass all four. If the pass rate is below 20 %, the protocol is too conservative for the claimed use case.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the pre-registered protocol correctly diagnoses false positives and thereby rejects tail-shape claims. For that to hold, the four gates must have high power against spurious tail-index estimates while retaining power against real heavy tails. In EVT, with tail sample sizes of a few dozen (standard in toxicity scoring), even exact Pareto draws routinely fail strict Kolmogorov-Smirnov or Anderson-Darling GOF at conventional levels and exhibit threshold instability; the protocol therefore risks systematic false negatives on valid tail shape rather than only catching the three false-positive modes described.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes a pre-registered diagnostic protocol for tail-index (tail-shape) claims in LLM evaluation, consisting of four gates: admissibility, goodness-of-fit, threshold-stability, and effect-size. The protocol is applied as a demonstration to a standard toxicity-evaluation setup using two structurally different scorer families; the authors report that it identifies three distinct false-positive modes that a naive analysis would accept and rejects the headline tail-shape claim on both scorers, concluding that tail-shape estimation is fragile in these setups.","tokens_in":1831,"tokens_out":512,"duration_ms":19788,"significance":"If the protocol can be shown to have high power against the three claimed false-positive modes while retaining reasonable power against genuine heavy tails in the small-tail-sample regimes typical of LLM toxicity data, it would be a useful methodological contribution for the emerging literature on tail-aware LLM metrics. The paper correctly identifies that naive tail-index reporting is prone to overinterpretation, but the strength of this conclusion depends on validation of the gates themselves.","major_comments":[{"comment":"Protocol definition section: the four gates are asserted to correctly diagnose false positives without introducing selection bias or missing genuine tail-shape signal, yet the manuscript reports no power simulations or calibration on synthetic data drawn from exact Pareto (or other EVT) distributions with tail sizes matching the LLM toxicity setup (typically a few dozen exceedances). This is load-bearing for the central claim because standard EVT results show that Kolmogorov-Smirnov or Anderson-Darling GOF tests at conventional levels routinely reject even i.i.d. heavy-tailed samples of this size, creating a risk that the protocol produces systematic false negatives on valid tail shape rather than only catching the three false-positive modes.","section":"Protocol definition and empirical demonstration"}],"minor_comments":[{"comment":"Abstract and results section: the three distinct false-positive modes are referenced but not given explicit operational definitions, thresholds, or fit statistics (e.g., specific GOF p-values or stability ranges) in the text provided, making it difficult to reproduce or extend the demonstration.","section":"Abstract"},{"comment":"The manuscript would benefit from citing standard EVT references on small-sample behavior of tail-index estimators and GOF tests (e.g., work on the Pickands estimator or finite-sample properties of Hill plots) to situate the protocol's design choices.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful and constructive review. The central contribution is the pre-registered protocol itself; the LLM toxicity application is presented explicitly as a demonstration. We respond to the single major comment below.","responses":[{"response":"We agree that the manuscript contains no Monte Carlo power analysis on synthetic Pareto (or other EVT) data calibrated to the small exceedance counts typical of LLM toxicity tails. The protocol was pre-registered to combine four standard EVT diagnostics into a conservative filter whose purpose is to reject tail-shape claims that fail any gate; the real-data demonstration illustrates three concrete false-positive modes that survive naive analysis. Because the GOF gate can indeed produce elevated rejection rates for modest sample sizes even under exact heavy-tailed laws, the absence of synthetic calibration leaves open the possibility that the protocol is overly conservative. We will therefore add a new subsection containing targeted power simulations: synthetic samples will be drawn from Pareto distributions with tail indices 0.6–1.8 and exceedance counts 25–60 (matching the observed LLM toxicity tails), and the full four-gate protocol will be applied to quantify (i) rejection rates under genuine heavy tails and (ii) specificity against the three documented false-positive mechanisms. This addition will appear in the revised manuscript.","revision_made":"yes","referee_comment":"Protocol definition section: the four gates are asserted to correctly diagnose false positives without introducing selection bias or missing genuine tail-shape signal, yet the manuscript reports no power simulations or calibration on synthetic data drawn from exact Pareto (or other EVT) distributions with tail sizes matching the LLM toxicity setup (typically a few dozen exceedances). This is load-bearing for the central claim because standard EVT results show that Kolmogorov-Smirnov or Anderson-Darling GOF tests at conventional levels routinely reject even i.i.d. heavy-tailed samples of this size, creating a risk that the protocol produces systematic false negatives on valid tail shape rather than only catching the three false-positive modes."}],"tokens_in":1344,"tokens_out":422,"duration_ms":27962,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is a pre-registered protocol with four gates—admissibility, goodness-of-fit, threshold-stability, and effect-size—for any positive tail-shape claim in LLM work. The authors apply it to a standard toxicity setup with two different scorers, say it catches three false-positive modes that a naive analysis would have let through, and reject the headline tail-shape result on both.\n\nThat protocol is the actual contribution, and framing it as pre-registered multi-gate checks is a reasonable response to the known fragility of tail estimates when tail samples are small. It does the field a service by making explicit the checks that should precede any claim about tail index adding information beyond mean and tail mass.\n\nThe soft spots are clear from the abstract. No fit statistics, no thresholds, no definitions of the three modes, and no data are given, so the claim that the protocol correctly rejects the finding cannot be checked. The stress-test worry about the gates producing false negatives on real heavy tails in small-sample regimes like toxicity scoring lands; with only a few dozen tail observations, standard GOF tests routinely reject even exact Pareto draws, and nothing in the text shows the protocol has been calibrated to avoid that.\n\nThis is for researchers who use or build tail-aware metrics in LLM evaluation, especially on rare events. A reader who wants a concrete checklist for extreme-value claims would get something usable from the protocol idea itself.\n\nIt deserves a serious referee because the problem it targets is real and the proposed structure could be a starting point, provided the full paper supplies the missing empirical details and power checks.","headline":"The paper offers a pre-registered four-gate protocol to vet tail-index claims in LLM evaluation and applies it to reject one in toxicity scoring, but supplies no numbers or definitions to assess whether the gates work as claimed.","tokens_in":2296,"tokens_out":414,"would_cite":false,"duration_ms":22412,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A pre-registered protocol of four checks rejects tail-index estimates as adding no value beyond mean and tail magnitude in LLM toxicity evaluation.","keywords":["LLM evaluation","tail index","extreme value theory","false positives","toxicity evaluation","protocol","tail shape"],"falsifier":"An LLM evaluation setup in which the tail-index estimate satisfies all four protocol requirements on the same toxicity data and scorer families would falsify the rejection; repeated failure across additional setups would strengthen it.","tokens_in":2588,"feed_emoji":"⚠️","tokens_out":641,"duration_ms":14511,"temperature":0.7,"pith_summary":"The paper examines whether the extreme-value-theory tail-index parameter supplies extra information in LLM evaluation beyond the mean and a conventional tail-magnitude measure. It supplies a pre-registered protocol that any positive tail-shape claim must satisfy on four criteria: admissibility, goodness-of-fit, threshold-stability, and effect size. When the protocol is run on a standard toxicity-evaluation task with two structurally different scorer families, it identifies three distinct patterns of false positives that a naive analysis would have accepted and rejects the tail-shape claim for both families. The protocol itself is offered as the reusable contribution; the empirical demonstration shows the kinds of invalid claims it is designed to block.","feed_headline":"Protocol rejects tail-index claims in LLM toxicity tests","feed_subtitle":"Pre-registered checks for admissibility, fit, stability and effect size catch three modes of false positives on two scorers","key_machinery":"The pre-registered protocol that enforces admissibility, goodness-of-fit, threshold-stability, and effect-size requirements before any tail-shape claim is accepted.","core_discovery":"The canonical extreme-value-theory tail-index parameter does not add discriminative information beyond the mean and a standard tail-magnitude statistic in the examined LLM toxicity-evaluation setups, because it fails to meet the pre-registered admissibility, goodness-of-fit, threshold-stability, and effect-size requirements on both of the scorer families tested.","pith_inferences":["The same gates could be applied to other tail-aware metrics such as conditional value-at-risk to test their robustness.","Fragility may appear in evaluation domains outside toxicity once the protocol is used.","Meeting the stability and effect-size gates may require substantially larger evaluation samples than current practice supplies.","Different scorer architectures may systematically produce different apparent tail behaviors even on identical underlying data."],"forward_implications":["Any tail-index claim in comparable LLM evaluation must pass the four gates or be treated as unsupported.","Naive tail analyses in these setups produce three identifiable modes of false positives.","The tail-index estimate is rejected under both scorer families examined.","Future tail-index work in similar evaluation pipelines should begin with the protocol rather than reporting raw estimates."],"fun_headline_variants":["Tail-shape estimation fragile in LLM toxicity tests","Protocol catches false positives in LLM tail estimates","Tail-index fails checks in LLM toxicity evaluation","Diagnostic protocol rejects LLM tail-index claims"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The four pre-registered requirements are enough to detect false positives without themselves creating selection bias or overlooking a genuine tail-shape signal.","fun_headline_variants_meta":{"raw":{"variants":["Tail-shape estimation fragile in LLM toxicity tests","Protocol catches false positives in LLM tail estimates","Tail-index fails checks in LLM toxicity evaluation","Diagnostic protocol rejects LLM tail-index claims"]},"model":"grok-4.3","cost_usd":0.004135,"raw_usage":{"total_tokens":2076,"prompt_tokens":629,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":41349500,"prompt_tokens_details":{"text_tokens":629,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1395,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":629,"tokens_out":52,"duration_ms":9736,"temperature":1.0,"reasoning_tokens":1395,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-02T22:06:12.757128+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An LLM evaluation setup in which the tail-index estimate satisfies all four protocol requirements on the same toxicity data and scorer families would falsify the rejection; repeated failure across additional setups would strengthen it.","supporting_citations":[],"review_version":1}