{"id":"38009cd7-2a3f-465a-a723-32ec0884adc1","arxiv_id":"2607.28577","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Shadow Before Swap promotes warm-refit challengers only after a paired one-week NLL edge, beating calendar, blind, and continuous-maintenance baselines while cutting deployments 78%.","lead":"A deployment rule called Shadow Before Swap only installs a retrained crypto forecaster after it beats the live model on the next week of delayed labels. It slightly improves probability scores while cutting model swaps by about four-fifths in long Binance replays.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged scope limits.","rationale":"The strongest claim is supported by the reported contrasts, hierarchical aggregation, episode stratification, and sensitivity tables as written. The softest point is external validity / economic translation of NLL-plus-turnover gains—the same point the reader already treats as the condition for CONDITIONAL rather than unconditional ACCEPT. That does not invalidate the internal policy comparison on the delayed-label LOB stream. Stress-testing does not surface a distinct correctness failure (timing leak, shared mutable state, or post-hoc threshold fishing after lock) that would move the verdict to REJECT or force a harder downgrade. Keeping CONDITIONAL with the reader's artifact, transport, and release-cost caveats is appropriate; no adjustment.","tokens_in":14389,"tokens_out":481,"duration_ms":9874,"concrete_test":"Re-run the full 48-week recursive SBS vs maintenance/blind/calendar pipelines with the locked τ=10^{-4} rule but replace equal asset/contract weights by a single pre-registered capital- or volume-weighted scheme frozen before outcomes; if any canonical relative NLL contrast in Table 1 changes sign or its four-week block interval covers zero, the deployment-value claim weakens under economically weighted estimands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a carefully scoped empirical policy comparison under historical replay, not a universal optimality result. The paper's own design already isolates authorization (SBS vs blind), timing (blind vs calendar), and value of accepted refits (SBS vs maintenance), with positive four-week block intervals, seed replication, margin/duration grids, a 20-asset panel, and a Temporal-CNN safety containment case. The reader's weakest assumption correctly notes that development-frozen τ, head-only maintenance, equal weighting, and NLL (without frozen economic weights or release costs) limit transport to live P&L governance; those are stated limitations (§5.4), not internal contradictions that overturn the reported recursive NLL and turnover contrasts on the stated estimand. No stronger load-bearing flaw (e.g., label leakage, non-causal pairing, or comparator collapse) is evident in the text.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes Shadow Before Swap (SBS), a causal deployment policy for production forecasters: at each scheduled boundary a challenger is deep-copied from the maintained incumbent, warm-refit off path, advanced with the incumbent on the same next week of delayed labels, and promoted only if paired mean NLL improves by at least τ=10^{-4}. Four recursive replays (maintenance, calendar, schedule-matched blind promotion, SBS) isolate timing, authorization, and the value of accepted refits. On two nonoverlapping Binance episodes (48 UTC weeks, 3 seeds, 8 underlyings, USD-M and COIN-M), SBS reduces hierarchically aggregated NLL by 0.1472% vs calendar, 0.0755% vs blind, and 0.0428% vs maintenance, with positive episode-stratified four-week block intervals, while accepting 114/528 challengers (78.4% fewer deployed changes). Directional consistency is reported across seeds, margins, trial lengths, a 20-asset panel, a supervised objective, and a Temporal-CNN failure-containment stress test.","tokens_in":14572,"tokens_out":1223,"duration_ms":36155,"significance":"If the reported recursive contrasts hold, the paper cleanly separates retraining from release authorization—an operational decision that production-ML and model-risk guidance often leave underspecified for delayed-label, nonstationary streams. Strengths include the schedule-matched blind contrast, full state isolation, development freeze of the gate before the primary episodes, multi-seed and multi-episode replication, margin/duration grids, block-bootstrap inference with episode stratification, and an explicit safety containment case under a misspecified Temporal-CNN. The contribution is a deployment policy rather than a new architecture, which is appropriately scoped and useful to applied forecasting systems. The work does not claim universal optimality or direct trading P&L; its value is a falsifiable, replayable authorization rule with measured turnover reduction.","major_comments":[{"comment":"§3.5 and Eq. (5): the primary estimand equal-weights underlyings, contract types, and weeks. That is a coherent release-policy estimand, but the practical claim in the abstract and §5.3 (a deployment policy that improves serving quality) is left somewhat open without at least one complete recursive sensitivity under frozen volume- or risk-based weights. A single pre-specified capital-weighted replay would show whether the sign and the 78.4% turnover reduction survive the weighting that many production desks would actually use; absence of that check is the main load-bearing gap between the reported NLL contrasts and operational adoption.","section":"§3.5, Eq. (5), §5.3"},{"comment":"§4.1 and §5.2: absolute gains vs maintenance are ~4.55×10^{-4} NLL per forecast. The conversion to ~455 log-loss units per million predictions and the matched-coverage Brier/diagnostic are helpful, but the manuscript still leans on relative percentages that are easy to over-read. Please state absolute NLL (or average per-week NLL) for all four policies in Table 1, and keep the service-scale interpretation explicitly tied to proper scoring rather than implying decision value without a frozen downstream rule—consistent with the limitation already noted in §5.4.","section":"§4.1, Table 1, §5.2"}],"minor_comments":[{"comment":"Figure 1 pseudocode step (2) uses W', b', μ', σ' without defining the incumbent (W,b,μ,σ) in the main text notation of §2.1; a one-line alignment with S_t=(ϕ_t,W_t,N_t,Q_t,O_t) would help.","section":"Figure 1, §2.1"},{"comment":"§3.1: “development-selected and frozen” is carefully worded; still, briefly state whether any hyperparameter of the learner (Adam lr, patience, 28-day window) was touched after the lock, or only the gate rule.","section":"§3.1"},{"comment":"Table 2: P/R counts are useful; adding the corresponding promotion rate next to the 48-week 114/528 figure in the main text would make the turnover claim easier to audit.","section":"Table 2, §4.1"},{"comment":"§6 / Table 4: ChaCha and the Digalakis et al. challenger-switching framework are well placed; a sentence on whether progressive-validation bounds could replace the fixed τ deadband would round out the related-work contrast.","section":"§6"},{"comment":"Minor copy edits: “b𝑣_b” / “b𝐴_b” rendering in Eq. (2)–(3) is hard to read in the text build; “Forward–Forward” hyphenation is inconsistent in one place.","section":"§2.1"}],"recommendation":"minor_revision","confidential_remarks":"Fit for a methods/applications venue in computational finance or production ML is good. I do not see a load-bearing causal or comparator flaw; the equal-weight and absolute-scale points are fixable in revision without new theory. The second episode’s “retrospective robustness” framing is acceptable given the stated freeze, but editors may want the authors’ lock artifacts available for audit if the journal supports that."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bit is the framing. Dutta treats replacement as authorization against a continuously maintained incumbent, not a stale checkpoint, and measures the full recursive serving path after states diverge. SBS clones off-path, warm-refits, runs a one-week paired NLL trial on mature delayed labels, and promotes only past a fixed deadband. That is the right operational question for production forecasters under label delay.\n\nWhat is new is not champion–challenger or canaries—those are cited—but holding generation and cadence fixed while isolating three things with matched replays: timing (calendar vs blind), authorization (blind vs SBS), and value of accepted refits (SBS vs maintenance). The 48-week Binance result is coherent: ~0.15%/0.08%/0.04% NLL cuts vs those three, positive four-week block intervals, 114/528 promotions, directionally stable across seeds, margins, trial lengths, a 20-asset panel, and a supervised twin. The Temporal-CNN stress case is the clearest safety win—gate rejects almost everything and stays near maintenance while blind/calendar blow up. Design hygiene is above average for cs.CE: causal isolation, mature-label scoring, episode freeze of the rule, hierarchical equal-weight aggregation, block bootstrap.\n\nSoft spots are real but mostly the ones the paper already flags. τ and the one-week budget were development-chosen (sensitivities help). Evidence is Binance perps plus one head-only maintenance operator; NLL and turnover are not frozen-weight P&L or release-cost accounting. No public code/data in the text, so runs are not independently checkable from the PDF alone. None of that collapses the stated estimand.\n\nThis is for people who ship and govern forecasting systems—quant ML ops, model risk, production ML—not for core learning theory. Math is elementary and appropriate; citations sit in the right neighborhood (prequential, delayed feedback, canary, ChaCha, proper scores, block bootstrap). I would send it to referees. Engage if you care about deployment gates under delayed labels; skip if you only want new predictors.","headline":"Solid ops paper: delayed-label paired gate on recursive serving trajectories, not another LOB architecture, with clean baselines and honest scope.","tokens_in":15201,"tokens_out":521,"would_cite":true,"duration_ms":9670,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Warm-refit challengers off-path, promote them only after a fixed paired NLL edge over the live incumbent, and you get better crypto forecasts with far fewer model swaps.","keywords":["model replacement","nonstationarity","shadow evaluation","delayed labels","limit order books","probabilistic forecasting","online learning","deployment policy"],"falsifier":"Replay the same recursive policies on a held-out nonoverlapping market episode (or live frozen-weight deployment) and check whether SBS’s NLL edge over blind promotion and continuous maintenance stays positive with the locked one-week trial and τ = 10^{-4}; a zero or negative maintenance contrast with many accepted bad promotions would falsify the claim.","tokens_in":15224,"feed_emoji":"📉","tokens_out":1041,"duration_ms":20698,"temperature":0.7,"pith_summary":"Production forecasters keep learning between rebuilds, so a newly trained model can beat an old checkpoint and still lose to the state that is actually serving. This paper argues that replacement is therefore a sequential authorization problem, not just a training schedule. It proposes Shadow Before Swap: clone the incumbent, warm-refit the clone on mature history, run challenger and incumbent side by side on the next week of delayed labels, and promote only if the challenger’s mean negative log-likelihood is better by a fixed deadband. In 48 weeks of Binance perpetual-futures replay across seeds, underlyings, and contract types, that gate beats calendar replacement, blind promotion on the same schedule, and pure continuous maintenance, while accepting only about one in five challengers. A sympathetic reader cares because the policy sits between an existing retrain pipeline and the registry: you can train often, yet install only when forward evidence justifies changing the parent of every later update.","feed_headline":"Train often, swap only after a one-week shadow win","feed_subtitle":"A forward NLL gate beats calendar, blind, and live maintenance while cutting model changes 78%","key_machinery":"Shadow Before Swap (SBS): a causal shadow trial that deep-copies the full incumbent state, warm-fits the clone off the serving path, advances both branches on the same next week of delayed labels, and authorizes promotion only when the paired mean NLL advantage clears a fixed deadband (τ = 10^{-4}). The continuously maintained incumbent is the deployment null; a schedule-matched blind policy isolates the value of that authorization decision.","core_discovery":"Shadow Before Swap improves the recursive serving trajectory of probabilistic LOB forecasts relative to calendar replacement, schedule-matched automatic promotion, and continuous maintenance, while cutting deployed model changes by roughly four-fifths. On two nonoverlapping Binance episodes totaling 48 UTC weeks, three seeds, eight underlyings, and two perpetual contract types, SBS reduces hierarchically aggregated NLL by 0.1472%, 0.0755%, and 0.0428% against those three baselines, with positive episode-stratified four-week block intervals, by promoting 114 of 528 challengers.","pith_inferences":["Any domain with delayed labels and a continuously adapting serving state—not only crypto LOBs—could reuse the same clone-shadow-compare release pattern between training and registry.","Organizations that price rollback and validation work highly may want a wider deadband: the paper’s margin grid already shows fewer promotions with stable sign, which hints at a tunable ops cost knob.","Equal asset and contract weighting makes this a governance estimand; capital-weighted live value would need weights frozen before outcomes and full recursive replay, as the author flags but does not run.","The Temporal-CNN stress case suggests forward authorization is especially useful as a safety layer when validation repeatedly prefers unsafe replacements."],"forward_implications":["Retraining cadence and release authorization can be separated: pipelines may propose refits on schedule while a forward gate controls which state enters service.","Schedule-matched waiting alone is not enough; the paper’s blind comparator shows most of the remaining gain comes from refusing weak challengers.","Accepted refits can beat a strong online-updated incumbent, so the safest policy is not simply never to full-refit.","Deployed-state turnover can fall by about 78% without sacrificing probabilistic forecast quality under the reported setup.","When candidate generators are badly misspecified, the same gate can keep serving near the incumbent by rejecting almost all proposals."],"fun_headline_variants":["Shadow Before Swap: gate challengers on one-week forward NLL","Train often, promote only after a paired shadow NLL win","Forward-gated replacement cuts model swaps 78% vs calendar","SBS beats calendar, blind promo, and live maintenance on NLL","Promote 114 of 528 challengers after one-week shadow edge"],"cache_read_input_tokens":128,"weakest_assumption_plain":"A one-week paired NLL advantage under this fixed deadband and head-only delayed maintenance, measured in historical equal-weighted replay, is enough to decide which refits should become the parent of future live updates.","fun_headline_variants_meta":{"raw":{"variants":["Shadow Before Swap: gate challengers on one-week forward NLL","Train often, promote only after a paired shadow NLL win","Forward-gated replacement cuts model swaps 78% vs calendar","SBS beats calendar, blind promo, and live maintenance on NLL","Promote 114 of 528 challengers after one-week shadow edge"]},"model":"grok-4.5","effort":"low","cost_usd":0.002041,"raw_usage":{"total_tokens":984,"prompt_tokens":857,"num_sources_used":0,"completion_tokens":75,"cost_in_usd_ticks":20408000,"prompt_tokens_details":{"text_tokens":857,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":52,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":857,"tokens_out":75,"duration_ms":2708,"temperature":1.0,"reasoning_tokens":52,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T03:16:30.094724+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Replay the same recursive policies on a held-out nonoverlapping market episode (or live frozen-weight deployment) and check whether SBS’s NLL edge over blind promotion and continuous maintenance stays positive with the locked one-week trial and τ = 10^{-4}; a zero or negative maintenance contrast with many accepted bad promotions would falsify the claim.","supporting_citations":[],"review_version":1}