{"id":"5f937263-3182-403e-997b-7c6432865019","arxiv_id":"2512.03116","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A gambling-style test over the largest observed rainfalls selects the extreme-value threshold that won the EVA2025 challenge — except on one target, where the authors overrode the game and paid for it.","lead":"This paper recounts the winning strategy of the EVA2025 extreme-precipitation challenge: fit simple exponential tails to daily rainfall extremes after removing seasonal effects, and use a betting-game statistical test to choose the threshold instead of expert judgment. The generalist hook is a new way to decide where 'extreme' begins when the events you are predicting never occur in the training data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbounded score increments invalidate the wealth process as a martingale test (§3.2); K=3 transfer remains unverified","rationale":"The reader's weakest assumption correctly identifies the K=3 transfer from observed to unfeasible order statistics as the central unsupported move. I agree that this is load-bearing. My stress-test sharpens the concern: before the transfer is even meaningful, the wealth process must be a valid nonnegative martingale/e-value; the unbounded, unscaled differences in §3.2 can break this, and §3.3's power guarantees explicitly require bounded errors that are never checked. The paper's own admissions—'arbitrarily chose K=3', scores 'very sensitive to K', single simulated path, and 'lacking rigorous guarantees' in §5—support this reading. Since the reader already reached CONDITIONAL, my analysis does not move the verdict; it adds a precise, checkable condition that should be part of the required revisions. I selected 'partial' agreement because my primary attack is the unbounded-increment validity of the martingale process, whereas the reader framed the main issue as the K=3 transfer assumption, although the two are connected.","tokens_in":8408,"tokens_out":13074,"duration_ms":144696,"concrete_test":"On the challenge data (or a faithful synthetic replicate), for every p in {0.9,0.99,0.995,0.999,0.9992,0.9995,0.9997,0.9999} and K∈{3,5}, draw at least 10,000 independent model samples and record the error sequence D_k=~Y_{(K−k)}−Y_{(K−k)}. Compute (i) max|D_k|, (ii) the minimum of W^C_k, and (iii) the fraction of samples with log(W) undefined or W≤0. If any max|D_k|>2, or W≤0 occurs with non-negligible frequency, the wealth process is not a valid nonnegative martingale test as defined in §3.2; if max|D_k|>1, the §3.3 power guarantee does not apply. This directly settles whether the proposed score is a legitimate martingale test on the actual data scale.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that terminal wealth W^C_{K−1} is an agnostic score of extrapolation. This requires W to be a nonnegative capital process under H0. In §3.2, W^C_k is a product of factors 1+(γ_{1,k}−0.5)(~Y_{(K−i)}−Y_{(K−i)}). If the error D is outside [−2,2], the factor with γ=0 or γ=1 becomes negative; the EWA weights can leave [0,1], W or L can become nonpositive, log(W) is undefined, and Ville's inequality is inapplicable. §3.3's imported power guarantee explicitly requires |~Y−Y|≤1 and E[(~Y−Y)^2|F]≤m/4, yet no such bound or normalization is reported for precipitation order statistics on the Leadbetter scale. Even with a valid e-process, the paper asserts without derivation (§3.1) that coherence over the first K−1 observed rounds assesses extrapolation at unfeasible steps k≥K. The implementation makes this especially fragile: K=3 is 'arbitrarily' chosen (§4.2), scores are admitted to be 'very sensitive to K' (§4.1), and only one simulated path is used (§4.1). Thus both the mathematical 'martingale testing' property and the extrapolation-transfer premise are unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the winning strategy for the EVA2025 Data Challenge, where the task was to estimate frequencies of extreme precipitation events from four climate-model runs. The authors reduce each target to a univariate variable, fit a Peaks Over Threshold model with exponential exceedances after seasonal adjustment, and use a martingale testing-by-betting procedure to select the POT quantile level p. The claimed novelty is that the terminal wealth of an adversarial gambler betting on the sign of the prediction error between model-simulated and observed top order statistics is an agnostic score of extrapolation power, and that minimizing this wealth selects the threshold. The paper reports the resulting estimates, discusses the final choices, and includes a post-competition assessment showing that the martingale-selected thresholds would have improved Target 1.","tokens_in":8720,"tokens_out":2475,"duration_ms":29945,"significance":"If the central claim is valid, the paper offers a principled, data-driven way to choose the POT threshold in extreme value analysis, replacing expert judgment with an agnostic sequential-testing score. The EVA2025 win is concrete external evidence that the overall pipeline performs well in a realistic extrapolation task, and the post-hoc Target 1 analysis is a useful case study. However, the theoretical core is thin: the martingale property is not established for the actual wealth process used, the imported power guarantees rely on unverified boundedness assumptions, and the transfer from observed top order statistics to unobservable extrapolation steps is asserted rather than proven. The paper itself acknowledges several of these limitations, which is honest but does not remove their load-bearing status for the novelty claim.","major_comments":[{"comment":"The wealth process W_k^C is defined as a product of factors 1 + (γ_{1,k} − 0.5)(\\tilde{Y}_{(K−i)} − Y_{(K−i)}). This is a valid nonnegative martingale only if the factors remain nonnegative. For γ = 0 or γ = 1, a prediction error outside [−2, 2] makes the factor negative; the EWA weights can then leave [0,1], and log(W_k^C) can become undefined. No bound or normalization is reported for precipitation order statistics on the Leadbetter scale, and the data include values above 5.7. Since the terminal wealth W_{K−1}^C is the score used to select p, this invalidates the martingale-testing interpretation of the selection criterion.","section":"§3.2"},{"comment":"The power guarantees are imported from reference [10] under assumptions |\\tilde{Y} − Y| ≤ 1 and E[(\\tilde{Y} − Y)^2 | F_k] ≤ m/4. The paper states 'These results are valid in our non-standard game' without proving that these conditions hold for the order-statistics game with exponential-tail models and observed precipitation extremes. Since [10] is a preprint co-authored by one of the present authors, this is not independent support. The O(log(1/α)/m) rate and Equation (3) therefore do not currently apply to the procedure as implemented.","section":"§3.3"},{"comment":"The load-bearing transfer assumption — that coherence of the top K−1 order statistics assesses extrapolation at unfeasible steps k ≥ K — is asserted without derivation. The implementation makes the assumption fragile: K = 3 is chosen 'arbitrarily' (§4.2), the scores are 'very sensitive to the choice of K' (§4.1), and the game is run on a single simulated sample of top order statistics. No experiment or argument shows that minimizing wealth at round K−1 selects thresholds that extrapolate better beyond the observed range. This is the central novelty, and it remains unsupported.","section":"§3.1"},{"comment":"The paper explicitly validates the model on the training set: 'we validate the model on the training set due to the lack of top order statistics at the Targets level' (§4.1). An in-sample self-coherence check with one realization and an arbitrary K does not make the threshold selection agnostic. The conclusion concedes that the method is 'lacking rigorous guarantees' and that 'the literature lacks clear interpretations of the terminal wealth value, particularly for tests that should be rejected at unfeasible steps k larger than the deterministic stopping time K − 1.' These admissions are consistent with the referee's reading that the statistical claim is not yet established.","section":"§4.1 and §5"}],"minor_comments":[{"comment":"The text refers to 'using (7)' but no Equation (7) appears in the manuscript; the numbering appears to be missing or broken.","section":"§4.2"},{"comment":"Several Springer references list '???' as the place of publication; this needs completion.","section":"References"},{"comment":"Figure 7 shows scores for K=3 and K=5, but the text states these are 'very sensitive' without quantifying the sensitivity or providing any uncertainty assessment. Since only one simulation is used, the figure should at least be described as a single realization, not a stable score surface.","section":"§4.1"},{"comment":"The display '1\\tilde Y^{(3)} = ...' appears to contain a typo or formatting artifact; the meaning of the leading '1' is unclear.","section":"§2.4"},{"comment":"The paper alternates between 'EV A' and 'EVA' for the conference/challenge; please standardize the spelling.","section":"§1 and §5"}],"recommendation":"major_revision","confidential_remarks":"The paper is best viewed as a competition report with a promising but underdeveloped methodological idea. The empirical win is real evidence, but the central novelty — agnostic threshold selection via martingale wealth — is not rigorously established. The authors themselves concede the main gaps. A revision that either proves the martingale property under explicit boundedness conditions, or reframes the paper as a case study with clearly stated heuristic status, would be more defensible. I would not reject outright because the challenge result and post-hoc Target 1 analysis are of genuine interest to the EVA community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is an honest, useful competition paper with one genuinely new idea, but the central methodological claim—that the wealth process is a martingale test—does not hold as stated.\n\nWhat's new: the game on increasing order statistics, where a gambler bets on the sign of the difference between model-sampled and observed top order statistics, is a real departure from sequential testing by betting. The authors correctly flag this as their originality. The idea is attractive because threshold selection in POT is otherwise heuristic, and tying it to e-values could open a new line.\n\nWhat they do well: they won the EVA2025 challenge, they transparently report where the method failed (Target 1 override), they admit arbitrary K=3, single simulated path, in-sample validation, and they show that the score's chosen candidates would have improved Target 1. The write-up is refreshingly direct about limitations.\n\nSoft spots: the most serious is the construction of W. The factors are 1 + (γ−0.5)(Ytilde−Y). With γ ∈ {0,1}, if the error is outside [−2,2] the factor goes negative; W is not a nonnegative capital process, log W is undefined, and Ville's inequality doesn't apply. The imported power guarantees from [10] require bounded errors |·|≤1 and conditional variance ≤ m/4, conditions neither checked nor plausibly true for extreme precipitation on the Leadbetter scale. So the mathematical 'martingale testing' claim is unsupported. The second soft spot is the transfer: coherence over K−1 observed top order statistics (with K=3) is simply asserted to assess extrapolation beyond the sample. That may be a working heuristic, but it is not an argument. The evaluation uses one simulated path and scores are sensitive to K; the paper should report robustness across K and multiple samples.\n\nThese do not make the paper worthless. As a competition write-up, it is a legitimate contribution; as a methodology proposal, it needs repairs. The authors themselves acknowledge 'major limitations' and 'lacking rigorous guarantees.'\n\nWho for: people working on e-values/testing by betting and extreme value threshold selection. It would be a good reading-group paper to dissect. It deserves a serious referee—the idea is important enough—but the referee should demand the martingale fix and code/data.\n\nRecommendation: send it out for review, expecting major revision.","headline":"A genuinely new idea—betting-game wealth on top order statistics to choose a POT threshold— wrapped in a winning competition write-up, but the martingale test as stated is not valid and the extrapolation-transfer claim is asserted, not shown.","tokens_in":9247,"tokens_out":3190,"would_cite":true,"duration_ms":36937,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","60G42","62L10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a gambler betting on the sign of discrepancies between model-simulated and observed top order statistics yields an agnostic score of extrapolation power, and that choosing the Peaks-Over-Threshold quantile that minimi","keywords":["Peaks Over Threshold","extreme value extrapolation","martingale testing","testing by betting","threshold selection","EVA2025 Data Challenge","top order statistics","precipitation extremes"],"falsifier":"Take a large climate-model dataset, hold out 46 of 50 runs, fit the POT model on 4 runs, and compute the terminal wealth at K=3 for the candidate p-grid; if the p with minimal wealth does not give frequency estimates closer to the 50-run truth than, say, the median grid p, the score isn't selecting extrapolation power. A cheaper test: on Target 1, where the true frequency is now known to be 0.24, verify that the K=3 wealth-minimizing level (p≈0.995 or 0.9) beats the Q-Q-chosen p=0.999.","tokens_in":8224,"feed_emoji":"🎲","tokens_out":4224,"duration_ms":43058,"temperature":0.7,"texified_at":"2026-08-05T20:40:36.489297+00:00","pith_summary":"The paper tries to show that threshold selection in Peaks Over Threshold extreme-value modeling—normally a matter of expert judgment—can be automated by treating extrapolation as a sequential betting game. A gambler bets on the sign of the error between the model's largest simulated order statistics and the observed ones; under a correct model, the gambler's capital is a martingale, and a wealth that grows means the model predicts the tail wrongly. The authors select the quantile level p that minimizes the final wealth, claiming this agnostic score chose the high thresholds that won the EVA2025 data challenge. If they are right, the key obstacle to routine extreme-value extrapolation—where to put the threshold—has a principled, assumption-light answer.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":3170,"prompt_tokens":748,"completion_tokens":2422,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":748,"completion_tokens_details":{"reasoning_tokens":1771}},"feed_headline":"A betting game picks the best extreme-value threshold","feed_subtitle":"Winning EVA2025 by letting a gambler's wealth, not expert judgment, choose the Peaks-Over-Threshold quantile.","key_machinery":"The capital process $$W_k = \\prod_{i=0}^k \\left(1 + (\\gamma_{1,k} - 0.5)(\\tilde{Y}_{(K-i)} - Y_{(K-i)})\\right)$$, where γ is the exponential-weighting betting fraction; it is a martingale under the null that the model's and the data's top order statistics have the same conditional expectations, so wealth grows only when the model systematically over- or under-predicts the largest observations. Testing by betting converts extrapolation checking into a game on the K largest order statistics; the method is called 'agnostic' because it does not specify an alternative hypothesis.","core_discovery":"The central claim is that the terminal wealth of an exponentially weighted betting strategy on the signs of discrepancies between model-sampled and observed top order statistics is a valid agnostic measure of an extrapolation model's quality, and that minimizing this wealth over the grid of candidate quantile levels selects a threshold whose extrapolations fare well. In the EVA2025 challenge, this procedure selected p=0.9995 for targets 2 and 3, which the authors state contributed to winning; for target 1, the score pointed to lower levels, but the authors overrode it based on Q-Q plots, a choice they later report was wrong.","pith_inferences":["The betting score could be repurposed as a general validation layer for AI or machine-learning extreme-value generators, letting a nonparametric test decide whether a fitted generator extrapolates.","The choice of K=3 is admitted arbitrary and the scores are sensitive to K; a stability-based heuristic (e.g., choose p that is optimal across several K) is a natural testable extension.","The terminal wealth has no calibrated interpretation; converting it into an e-value or a valid p-value would give error control and could let the method replace subjective threshold choice in operational settings.","Since the paper's power guarantees require bounded errors and conditional variance bounds never verified for precipitation data, a diagnostic check of these bounds on exceedances would tell when the guarantees actually apply."],"forward_implications":["If true, threshold selection in Peaks Over Threshold can be based on an automated score instead of Q-Q plots or expert heuristics.","The score picks very high quantiles (e.g., 0.9995) where only a few hundred exceedances remain, enabling extrapolation from sparse data.","The approach extends to any generative model of extremes, not just exponential POT, provided top order statistics can be simulated.","The authors' Target 1 error shows the score can be more reliable than graphical diagnostics, because the Gumbel bias distorts Q-Q comparisons."],"fun_headline_variants":["Betting game chooses extreme-value threshold, wins EVA2025","Martingale testing picks best threshold for extremes","A gambler's wealth, not experts, sets the high quantile","Winning strategy: let a betting game set your threshold","Extrapolation quality judged by a betting game's payoff"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The method assumes that a three-round betting game on a single simulated sample of the largest order statistics reveals whether the model will extrapolate correctly beyond the observed range, and the power guarantees assume bounded prediction errors that the authors never verify for extreme precipitation.","fun_headline_variants_meta":{"raw":{"variants":["Betting game chooses extreme-value threshold, wins EVA2025","Martingale testing picks best threshold for extremes","A gambler's wealth, not experts, sets the high quantile","Winning strategy: let a betting game set your threshold","Extrapolation quality judged by a betting game's payoff"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1380,"prompt_tokens":670,"completion_tokens":710,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":627}},"tokens_in":414,"tokens_out":710,"duration_ms":8886,"temperature":1.0,"reasoning_tokens":627,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:57:42.536564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a large climate-model dataset, hold out 46 of 50 runs, fit the POT model on 4 runs, and compute the terminal wealth at K=3 for the candidate p-grid; if the p with minimal wealth does not give frequency estimates closer to the 50-run truth than, say, the median grid p, the score isn't selecting extrapolation power. A cheaper test: on Target 1, where the true frequency is now known to be 0.24, verify that the K=3 wealth-minimizing level (p≈0.995 or 0.9) beats the Q-Q-chosen p=0.999.","supporting_citations":[],"review_version":1}