{"id":"60ea2cc3-8a88-483f-bbf7-da15e2180aa3","arxiv_id":"2607.07146","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A three-number reporting protocol—balanced-test, operational-prior, and post-deployment precision—exposes the prior-shift gap in an operational satellite classifier, which reaches 0.927 precision at the operational prior.","lead":"This paper shows that a balanced test set badly overstates the precision a satellite-image wave detector will have in real operation, and proposes reporting three numbers instead: performance on a balanced set, on a test set at the true event rate, and on real post-deployment adjudications. Applied to the Sentinel-1 Internal Waves Service, the promoted classifier reaches 0.927 precision at the operational prior, and the authors show that the fixed operating point decays out-o","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Operational-prior cell is built from a confidence-ordered review queue, so the 0.927 lockbox number is not yet tied to the live stream; a random-sample check is required.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the operational prior and verified pools are drawn from a confidence-ordered review queue, so the quoted precisions may measure the queue rather than the stream. I agree with this and would sharpen it: the lockbox's 0.927 figure and the out-of-time 0.810 figure are both computed on review-selected data, so neither validates the method's central promise that figure (2) tracks figure (3). The paper is transparent about this limitation in Sect. 6, which is a point in its favor, but transparency does not remove the load-bearing nature of the assumption. A random-sample adjudication study is the concrete check that would settle whether the selection bias is material. The rest of the argument—the monotone-transform point, the leverage-controlled development cycle, the sealed single-read protocol—appears internally consistent and does not need to be overturned. The reader's CONDITIONAL verdict already appropriately reflects that the central empirical claim is conditional on the unverified representativeness of the review-selected pool, so I recommend no change to the verdict.","tokens_in":15708,"tokens_out":8035,"duration_ms":86729,"concrete_test":"Adjudicate a uniform random sample of ~5,000–10,000 Sentinel-1 Wave-mode vignettes from a post-freeze window (e.g., after 2026-07-01), with experts labeling every vignette regardless of model confidence. Compute gem's precision at the fixed threshold θ=0.935 and at the recall-0.80 operating point, plus the observed positive rate. Compare these to the lockbox values (0.927, recall 0.827) and the out-of-time values (0.810, FPR 1.0%). If the random-sample precision and prior match within a pre-registered tolerance, the selection-bias concern is resolved; if not, the operational-prior cell must be re-estimated on the random sample and the 0.927 headline figure revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central corrective claim is that the operational-prior figure predicts fielded precision. That claim rests on the frozen test set at π=0.05 being a faithful sample of the live stream. But the 126,742 'confirmed negatives' and the 6,868 positives come from a validator review queue organized by the model's own confidence (Sect. 2.1, 5.1), not from a uniform sample of the ~17M-vignette archive. The operational prior π≈0.05 is itself measured from the same adjudicated detections (Sect. 2.3). The paper explicitly states this selection runs optimistic: validators preferentially review high-confidence flags, so lower-confidence flags—which fail more often—are under-reviewed, and the quoted precisions lean high against the full stream (Sect. 6). If a random stream sample contains many easy negatives the queue never surfaces, the lockbox negatives are conditionally harder, and precision at the fixed threshold 0.935 could be quite different from 0.927; if the true stream prior is below 0.05, the same threshold yields lower precision still. The out-of-time read (0.810) is also on the same review-selected pool, so it does not resolve the bias. Thus the load-bearing link 'operational-prior cell → real operational cell' is not yet established; the paper itself leaves cell 3 open, but the problem is not only that the loop is unclosed—the pre-deployment estimate is built on a selection mechanism that can break the link.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes that operational Earth-observation classifiers be reported by three numbers rather than one: balanced-test performance, performance on a frozen test set at the operational prior, and real post-deployment performance. It argues that at a fixed recall floor, prior correction and calibration are monotone score transforms and therefore cannot change precision, so the balanced-to-operational gap is an evaluation artifact, not a training defect. The method is demonstrated on the Internal Waves Service: the incumbent reports 0.794 balanced precision vs 0.192 real operational precision; through a leakage-controlled, pre-registered development cycle (negative variety, capacity, GeM pooling), the promoted model 'gem' reports 0.927 precision at the operational prior on a sealed lockbox, with an out-of-time check showing discrimination transfer but fixed-threshold decay. The paper is explicit that the real-operational cell remains pre-registered future work.","tokens_in":15960,"tokens_out":7689,"duration_ms":72447,"significance":"The contribution is significant if it holds: three-number reporting is simple, portable, and directly addresses a common mismatch between reported and fielded performance in operational remote sensing. The paper's strengths are concrete and should be credited: the monotone-transform argument is correct; the lockbox is genuinely sealed and read once at a pre-registered threshold; the split is pinned and leakage controlled at footprint level; the development levers are isolated with pre-registered promotion margins; and the negative results are reported. These reproducibility practices raise the bar for the field. However, the central corrective claim — that the operational-prior figure predicts the real operational figure — is not yet demonstrated, and the current pre-deployment estimate is built on a review-selected pool, so the headline numbers are conditional on the queue rather than on the stream.","major_comments":[{"comment":"The load-bearing link of the method is that the operational-prior cell predicts the real operational cell. That link is not yet established because the prior (pi approx 0.05) and the positive/negative pools are generated by the model's own confidence-ordered review queue, not by a uniform sample of the live stream. The lockbox's 29,203 negatives are 'confirmed negatives' from adjudicated detections, so the 0.927 precision at pi=0.05 measures the queue, not the stream. The paper states the selection runs optimistic (Sect. 6), and the out-of-time check (Sect. 5.3) uses the same review-selected pool, so it does not resolve the bias. I would require a random-sample stream check, or a clearly labelled queue-conditional interpretation of 0.927, before accepting the central corrective claim.","section":"Sect. 2.3, 5.1, 5.2, 6"},{"comment":"The headline cautionary contrast is incumbent balanced-test precision 0.794 vs real operational 0.192. However, the 0.794 cell is computed as a parity projection over the full unanimous verified pool, with the paper acknowledging the incumbent's training membership is unrecoverable. The balanced cell may therefore overlap training data, inflating the gap. Please provide a leakage-controlled balanced evaluation of the incumbent on a held-out split, or re-label this cell as a retrospective parity projection and soften the abstract's 'scores 0.794... scores 0.192' claim.","section":"Sect. 5.1, Table 1"},{"comment":"The paper's own framing is honest that cell 3 is pre-registered future work. But the abstract and conclusions present 0.927 as the number validators should expect. Since the out-of-time check is a single ten-day window on the same review-selected pool, with only 20 positives in the footprint-disjoint new-site subset, it cannot substitute for cell 3. The claim that discrimination transfers is supported by AUC, but the fixed-operating-point decay is based on one small window. I recommend framing the paper as a protocol proposal with an open validation cell, and either adding a second out-of-time window or restricting the abstract's predictive claim.","section":"Sect. 5.3 and 6"}],"minor_comments":[{"comment":"The phrase 'reported exploratory and pre-registered' is confusing: if the scoring script and test were committed before the numbers were seen, the result is confirmatory with respect to that protocol, not exploratory. Please clarify the intended meaning.","section":"Sect. 5.3"},{"comment":"The caption says the 126,742 confirmed negatives are 'a count rather than a geography and are not mapped.' A small inset showing the lockbox negatives' spatial distribution would help the reader assess the hotspot-clustering concern, though this is not essential.","section":"Fig. 2 caption"},{"comment":"Define 'gem' and 'incumbent' in the table caption, and state explicitly that the balanced-test cells are not all held-out evaluations; the incumbent's cell is a retrospective parity projection.","section":"Table 1"},{"comment":"The sentence 'A separate calibration arm confirmed the inertness by construction' would benefit from one sentence on how the arm was constructed, so the reader can distinguish the empirical arm from the theoretical argument.","section":"Sect. 4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually transparent about its limitations, and I see no integrity concerns. The main risk is that the headline 0.927 number will be quoted without the queue-conditional caveat. The editor may wish to require a random-sample or early-deployment validation before accepting the strong predictive claim; without that, the paper should be framed as a protocol proposal whose central validation cell remains open."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this for the three-number reporting axis and the unusual care of the development cycle. The core claim is correct: for a classifier deployed at a prior far from its training balance, a balanced-test precision figure is systematically optimistic, and at a fixed recall floor prior correction and calibration cannot move precision because they are monotone transforms. The paper demonstrates that 0.794 balanced precision becomes 0.192 in operation, and that the promoted model reads 0.927 at the operational prior versus 0.996 balanced. The lever-by-lever development with pre-registered promotion margins, the sealed single-read lockbox, and the open code/data are all genuinely good practice; the honest negatives (ratio shift doesn't help, extra capacity stops paying, attention pooling doesn't beat GeM) are as valuable as the gains.\n\nThe main soft spot is exactly what the stress-test says, and the paper acknowledges it. The verified pool — the 126,742 negatives and 6,868 positives — comes from a review queue organized by the model's own confidence, so high-confidence flags are over-represented. The operational prior of 0.05 is measured from that same adjudicated set. The frozen operational-prior test set therefore is not a uniform sample of the live stream, and the quoted precisions — 0.927 on the lockbox, 0.810 out-of-time — are best-case estimates against the full stream. The out-of-time check, though a nice robustness result for discrimination transfer, is also on the review-selected pool, so it does not resolve the selection bias. The link between the operational-prior cell and the real post-deployment cell is pre-registered future work, not a closed demonstration. This does not break the methodological argument — the three-number instrument is sound — but it does mean the headline empirical contrast is not yet tied to the true stream.\n\nThe citation pattern is reasonable: the prior-shift evaluation ideas trace to Saito & Rehmsmeier, Dockès, Maxwell, and the spatial-splitting methods to Kattenborn and Roberts; the self-citation to the predecessor paper is appropriate since this is a direct continuation.\n\nI'd send this to peer review. The instrument is worth a serious referee, and the careful gatekeeping deserves credit. The main review question should be whether the authors can add a random-sample validation from the live stream, or at least characterize the selection bias explicitly enough that readers can judge how much the 0.927 should be discounted.","headline":"The three-number reporting axis is a genuinely useful instrument and the development-cycle discipline is unusually careful, but the headline operational-prior figure rests on a review-selected pool, so the link to the live stream is a pre-registered commitment rather than a closed result.","tokens_in":16590,"tokens_out":2632,"would_cite":true,"duration_ms":23796,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scoring an operational classifier on a balanced test set hides the real fielded skill: the same model drops from 0.794 precision to 0.192 in operation, and the paper proposes three-number reporting to make that gap visible.","keywords":["prior-matched evaluation","operational classifier","precision-recall","class imbalance","rare-event detection","internal solitary waves","Sentinel-1 Wave-mode","sealed lockbox evaluation"],"falsifier":"After deployment, adjudicate a uniform random sample of the live stream, including low-confidence flags and unflagged scenes, over a pre-registered window and compare the measured real-operational precision at recall at least 0.80 with the certified 0.927; if the 95% confidence interval excludes 0.927, the operational-prior cell does not predict fielded precision.","tokens_in":15497,"feed_emoji":"🌊","tokens_out":6614,"duration_ms":61711,"temperature":0.7,"pith_summary":"The paper's claim is that an operational rare-event classifier should never be certified by a balanced-test score alone, because the reporting prior, not the model, can be the main determinant of measured precision. Demonstrated on a Sentinel-1 internal-wave detection service whose true positive rate is about 0.05, the same model at the same threshold scores 0.794 balanced-test precision but only 0.192 in real operation; the paper shows this is a systematic artefact of evaluating at a 50/50 prior. At a fixed recall floor, prior correction and calibration are monotone transforms of the score and cannot change precision, so the mismatch is an evaluation problem wearing a training costume. The remedy is a three-figure report—balanced-test, operational-prior, and real post-deployment—and the developed model certifies 0.927 precision at the operational prior in a sealed, single-read lockbox. The real post-deployment figure is pre-registered and still outstanding, so the corrective half of the claim is a falsifiable commitment rather than a closed result.","feed_headline":"A classifier scoring 0.794 on balanced tests hit 0.192 in operation","feed_subtitle":"The gap is a reporting-prior artefact; the paper proposes a three-number method that predicts fielded precision before deployment.","key_machinery":"The central instrument is the three-number reporting axis: (1) balanced-test performance at equal class counts, (2) operational-prior performance on a frozen, spatially de-correlated test set drawn at the true positive rate pi = 0.05, and (3) real post-deployment performance from prospective adjudication. The argument-carrying identity is that, at a recall-pinned operating point, prior correction and probability calibration are monotone transformations of the score—so they can relocate the decision threshold but cannot move the precision/recall curve, hence cannot move precision at all. Carrying the demonstration is a sealed lockbox read exactly once at a pre-registered threshold, a footprin","core_discovery":"On its own terms, the paper establishes that evaluating a classifier on equally balanced classes when the deployed stream is skewed guarantees an overstated precision, and that the gap is not removable by prior-matched retraining or calibration. The central demonstration is a pair of numbers: the deployed network reads 0.794 precision on a balanced test, and 0.192 precision on its own adjudicated detections at the same 0.5 threshold; neither number is wrong, the prior is different. After a precision-first development cycle, the promoted model reports 0.996 balanced precision and 0.927 operational-prior precision at recall 0.827, and an out-of-time check shows discrimination transfers (AUC 0.","pith_inferences":["My inference: because the adjudication queue is ordered by the model's own confidence, the quoted operational and lockbox precisions are likely optimistic for the full stream; a uniform random-sample audit of unflagged scenes is the cheapest way to quantify this and should precede reliance on the 0.927 figure.","My inference: the three-number axis could be turned into a portable rule for any detector whose base rate is far from 0.5—first quote precision at the deployment prior, treat the balanced figure as descriptive only, and compute the expected gap from the PR curve without retraining.","My inference: the out-of-time decay suggests rolling threshold re-estimation may be a general operational requirement for classifiers under a drifting stream, and the pre-registered out-of-time check is a reusable template for measuring it.","My inference: the finding that negative variety, rather than prior-matched training, carried the data lever suggests other rare-event projects should invest in hard-negative collection before rebalancing; that ordering is directly testable in other domains."],"forward_implications":["Rare-event operational services should report balanced-test, operational-prior, and real post-deployment performance together; the contrast, not any single cell, is the honest measure.","Pending the post-deployment read, validators should expect roughly 0.927 precision from the promoted model at the chosen operating point if the verified pools represent the live stream.","A balanced F1 of 0.904 crossing an operational F1 of 0.874 shows that model comparisons must hold the reporting prior fixed or the ranking is confounded.","The fixed threshold decays out-of-time while ranking holds, so deployed operating points must be revalidated against a current catalogue slice on schedule instead of being inherited from the development set.","The binding constraint has moved from architecture to data: more validated negative variety, not more width, is what lifts precision next."],"fun_headline_variants":["Balanced test says 0.794, operation says 0.192: prior mismatch","Three-number method predicts fielded precision before deployment","Prior-matched reporting lifts operational precision from 0.192 to 0.927","Operational precision gap exposed: balanced 0.794, fielded 0.192"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim collapses if the verified positive/negative pools and the measured one-in-twenty operational rate do not represent the live stream, because the review queue is organized by the model's own confidence and therefore over-represents the high-confidence flags on which precision is naturally highest—so the quoted operational and lockbox precisions lean optimistic against the full stream.","fun_headline_variants_meta":{"raw":{"variants":["Balanced test says 0.794, operation says 0.192: prior mismatch","Three-number method predicts fielded precision before deployment","Prior-matched reporting lifts operational precision from 0.192 to 0.927","Operational precision gap exposed: balanced 0.794, fielded 0.192"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000943,"raw_usage":{"total_tokens":3929,"prompt_tokens":873,"completion_tokens":3056,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":2981}},"tokens_in":617,"tokens_out":3056,"duration_ms":19619,"temperature":1.0,"reasoning_tokens":2981,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:06:57.808230+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"After deployment, adjudicate a uniform random sample of the live stream, including low-confidence flags and unflagged scenes, over a pre-registered window and compare the measured real-operational precision at recall at least 0.80 with the certified 0.927; if the 95% confidence interval excludes 0.927, the operational-prior cell does not predict fielded precision.","supporting_citations":[],"review_version":2}