{"id":"01ae7f3a-bf44-4ba4-9e49-0752b5d59505","arxiv_id":"2607.08273","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"On a 10-bearing PHME subset, a residual-calibrated fusion model hits ~0.15 normalized MAE and 0.90 average 90% coverage under leave-regime-out splits, while conditional diagnostics expose 0.666 coverage and raw-channel-loss collapse.","lead":"A carefully bounded study shows that bearing remaining-life forecasts can look well-calibrated on average yet fail in specific load-speed cells and under raw vibration loss. It offers a regime-disjoint evaluation protocol that makes those reliability failures visible rather than promising a deployable digital twin.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"The primary leave-regime-out numbers rest on full-set tertile cutpoints that already include future test windows, so the reported reliability protocol is not a pure train-only partition.","rationale":"The paper’s strongest claim is carefully bounded: a reliability-evaluation protocol on a documented 10-bearing subset, not full-PHME SOTA or deployment guarantees. Evidence largely matches that claim—competitive calibrated neural performance, no uniform dominance, explicit conditional undercoverage (0.666 in low-load/high-speed) and raw-channel-loss collapse. The softest load-bearing condition is exactly the one the reader flags: Section 3.3’s full-set empirical tertiles define the held-out units. That is not ordinary train/test leakage into model inputs (regime codes are withheld; models see only continuous load/speed), but it does mean the primary split structure is not a pure train-only operating-condition partition. Because the contribution is the protocol itself, sensitivity of the headline MAE/coverage pair and of the signature undercoverage cell to train-only re-binning is the single check that would settle whether the concern lands. If numbers are stable, the CONDITIONAL verdict stands for the stated reasons (subset scope, post-hoc conditional calibration, retrospective absolute-step scaling). If they move materially, the “strict leave-regime-out” framing needs tightening before the protocol can be treated as a field reference. No stronger objection (e.g., fabricated results or internal contradiction) is supported by the manuscript. Verdict remains CONDITIONAL; agreement with the reader is full on the weakest assumption.","tokens_in":28076,"tokens_out":792,"duration_ms":7324,"concrete_test":"Recompute the nine regime labels using only the training-regime windows’ load/speed tertiles for each of the nine strict rotations (never using the held-out test regime’s windows to set cutpoints). Re-run the primary calibrated predictive-representation pipeline under otherwise identical four-way splits; if normalized MAE moves by more than ~0.01 or mean empirical 90% coverage leaves the 0.85–0.95 band (or the low-load/high-speed cell ceases to be the undercoverage cell), the current leave-regime-out reliability numbers are partition-sensitive.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that, under a strict leave-one-operating-regime-out protocol with separate train/validation/calibration/test roles, residual-calibrated fusion yields normalized MAE 0.1477 and empirical 90% coverage 0.900 on the 10-bearing PHME subset, with conditional undercoverage and raw-channel loss reported as failure modes. That claim is load-bearing on the definition of the held-out units themselves. Section 3.3 states that load and speed are binned by the one-third and two-third quantiles of the frozen analysis set, and that these cutpoints are an evaluation-design choice rather than a deployable train-only classifier. Because those quantiles are computed on the full 14,297-window set, every held-out test cell participates in defining the regime boundaries that later isolate it. The paper already notes this and withholds regime codes from model inputs, so the issue is not hidden leakage into features; it is whether the reported leave-regime-out reliability numbers are an artifact of a non-deployable, full-set partition of the same frozen subset used to fix the analysis set before comparison. If re-binning from train-only quantiles materially reassigns windows or changes the hard low-load/high-speed cell, the headline accuracy–coverage pair and the conditional-failure narrative rest on a softer foundation than the “strict protocol” framing suggests.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript studies remaining-useful-life (RUL) point accuracy and interval reliability for bearings under time-varying load and speed on a documented 10-bearing PHME subset. Derived load–speed regimes define leave-one-regime-out evaluation units, while models receive only measured load and speed as context. A fused predictive-representation model (raw vibration, engineered descriptors, operating context) is trained under a strict train/validation/calibration/test separation and converted to intervals by empirical residual calibration. On this protocol the model reports normalized MAE 0.1477, empirical 90% coverage 0.900, and retrospective absolute-step MAE 285.26, close to a 400-tree random forest (0.1538 / 0.871 / 294.57). Conditional diagnostics show non-uniform reliability (notably 0.666 coverage in low-load/high-speed), a post-hoc pooled regime-conditioned residual diagnostic is offered only as motivation for future conditional calibration, and raw-channel loss is identified as the largest tested reliability failure mode. The stated contribution is a bounded reliability-evaluation protocol with explicit failure modes, not full-archive PHME SOTA or deployment guarantees.","tokens_in":28465,"tokens_out":1429,"duration_ms":26253,"significance":"If the reported protocol and diagnostics hold under the stated evidence boundary, the paper is a useful reliability-facing contribution rather than another architecture claim. Its main value is methodological honesty: regime-disjoint four-way splitting, residual calibration separated from model selection, strong tabular comparators (including a 400-tree forest), conditional coverage with concentration diagnostics, ablations, prefix/noise/raw-loss stress, and explicit non-claims about physics twins and maintenance policy. Public code and configuration provenance further strengthen reproducibility within the processed subset. The work will matter most to PHM readers who care whether intervals remain trustworthy under operating-condition shift; the significance is therefore as a careful evaluation template and failure-mode inventory on a transparent 10-bearing PHME slice, not as a general RUL breakthrough.","major_comments":[{"comment":"Section 3.3 defines the nine held-out regimes by empirical one-third/two-third load and speed quantiles of the frozen full 14,297-window analysis set, so every later test cell participates in the cutpoints that isolate it. The paper correctly labels this an evaluation-design choice rather than a deployable train-only classifier, but the primary claim is still framed as a strict leave-operating-regime-out reliability protocol (Abstract; §3.4; Table 8). Because those cutpoints are load-bearing for the headline MAE/coverage pair and for the hard low-load/high-speed cell, a sensitivity analysis that re-derives tertiles from training regimes only (or from non-test regimes only) and re-runs the primary endpoint is needed. If re-binning materially reassigns windows or changes the 0.666 LL/HS coverage, the “strict protocol” numbers rest on a softer partition than currently implied.","section":null},{"comment":"Table 10 and Figure 4 show that the critical low-load/high-speed undercoverage cell (coverage 0.666) has dominant-bearing share 0.947 from B17; Table 11’s exclusion diagnostic then rests on only 119 remaining windows. Conditional undercoverage is central to the paper’s reliability narrative, yet this cell is effectively a near-single-trajectory estimate. The manuscript should either (i) reweight or restate that cell as bearing-confounded rather than as a pure regime failure mode, or (ii) provide additional regime-definition / leave-bearing-within-regime checks so that the 0.666 figure is not over-read as multi-bearing regime unreliability.","section":null},{"comment":"§4.2 and Table 6 make clear that absolute-step MAE multiplies normalized predictions by each bearing’s known maximum step count and is therefore a retrospective offline metric. Tables 8–9 and the Abstract still lead with absolute-step MAE alongside normalized MAE and coverage. For a reliability-evaluation paper whose decision relevance is repeatedly emphasized, the primary quantitative claims should lead with normalized MAE/coverage (and interval score), with absolute-step figures demoted or consistently labeled “retrospective” in every main table caption so readers cannot mistake them for deployment-time lifetime units.","section":null}],"minor_comments":[{"comment":"Figure 1 is useful but dense; the matched-sensitivity branch is easy to miss relative to the primary four-way split. A one-line caption callout that only the strict four-way split is the primary endpoint would help.","section":null},{"comment":"Table 13 residual-only coverages (e.g., 0.868 at nominal 0.90 for the predictive representation) sit below the primary ensemble-dispersion coverage of 0.900 in Table 8; a short sentence in §6.2 clarifying why residual-only and primary intervals differ would prevent confusion.","section":null},{"comment":"§5.1 notes that bitwise-identical CUDA replay is not asserted. For a reproducibility-oriented reliability paper, stating which metrics were verified from saved prediction files versus re-trained runs would be helpful.","section":null},{"comment":"Leave-bearing-out coverage of 0.821 (Table 18) falls outside the stated 0.85–0.95 band; this is already treated as supporting evidence, but a brief cross-reference in the Abstract or Conclusion would better match the paper’s otherwise careful non-overclaim style.","section":null},{"comment":"Minor wording: “Cross-Transformer fusioning” in the Hou et al. citation title (§2.6 / References) appears to be a transcription artifact; verify against the source title.","section":null},{"comment":"Eq. (6) first-passage surrogate is clear, but the role of ε = 10^{-4} and the independent-window encoding limitation could be restated once near the monotonicity diagnostic so readers do not expect trajectory-level monotone H_i.","section":null}],"recommendation":"major_revision","confidential_remarks":"This is a carefully bounded evaluation paper rather than an architecture novelty claim; that is a feature for a reliability-oriented venue, but some readers may still ask whether the contribution is large enough without the train-only regime-cutpoint sensitivity. I would not reject on novelty alone if the authors deliver that check and tighten the absolute-step / single-bearing-cell framing. Scope fit is good for reliability/PHM evaluation methodology; less so for a pure deep-learning architecture track."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: this is not a new RUL architecture paper. It is a carefully bounded evaluation protocol on a fixed 10-bearing PHME subset that forces leave-operating-regime-out testing, separates residual calibration from model selection, and then reports conditional undercoverage and sensor-stress failures instead of hiding them. That packaging is the real contribution.\n\nWhat they do well is match claim to evidence. The strict four-way split, load/speed-only context, strong tabular baselines (including a 400-tree forest that stays competitive on MAE), paired regime statistics, ablations, prefix/noise/raw-loss probes, and the explicit 0.666 low-load/high-speed coverage cell are all cleanly documented. They refuse to claim uniform dominance, full-PHME SOTA, physics-twin status, or deployment guarantees. The public code release helps. For readers who care about whether intervals stay trustworthy when duty cycles change, this is the right kind of paper.\n\nSoft spots exist and are proportionate. The regime bins are empirical tertiles on the frozen full analysis set, so the held-out cells help define their own boundaries; the authors flag this as an evaluation-design choice rather than a train-only classifier, but it still softens the “strict protocol” framing. Absolute-step MAE is retrospective (bearing max-step scaling). The subset is only ten of seventeen public runs. The pooled regime-conditioned residual fix is post-hoc. None of these break the central empirical story; they just keep the paper where the authors already place it—bounded protocol, not field standard.\n\nMath and residual calibration are standard split-conformal style, citations are appropriate and not padded, and the work is coherent on its own terms. This is for PHM/reliability people who evaluate interval methods under operating-condition shift, and for anyone tired of mixed-split leaderboards. It deserves a serious referee. I would engage with it, cite the protocol when I need a clean leave-regime-out baseline, and send it to peer review rather than desk-reject.","headline":"Honest, bounded reliability-evaluation protocol for PHME bearing RUL under regime shift; useful packaging of leave-regime-out calibration and failure-mode reporting, not a new architecture or full-archive result.","tokens_in":29146,"tokens_out":532,"would_cite":true,"duration_ms":12006,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Average RUL coverage can look fine under changing load and speed while specific regimes and raw-channel loss still fail.","keywords":["reliability","prognostics","remaining useful life","calibration","conformal prediction","operating-regime shift","bearing"],"falsifier":"Recompute the nine load-speed tertile cuts using only training windows for each rotation (or process the remaining seven public bearings under the same protocol) and check whether primary coverage, the low-load/high-speed undercoverage cell, and the raw-channel-loss collapse still hold at comparable magnitudes.","tokens_in":28906,"feed_emoji":"⚙️","tokens_out":1019,"duration_ms":13724,"temperature":0.7,"pith_summary":"Remaining useful life estimates only help maintenance if both the point forecasts and their intervals stay trustworthy when load and speed change. Mixed train-test splits often hide that failure. This paper fixes a documented 10-bearing subset with time-varying conditions, holds out entire derived load-speed regimes, and asks whether a residual-calibrated model that fuses raw vibration, engineered descriptors, and measured load and speed still covers at the nominal rate. On the strict four-way split the model reaches normalized MAE 0.1477 and empirical 90% coverage 0.900, close to a strong random forest, but conditional checks show coverage falling to 0.666 in a low-load/high-speed cell and raw-channel loss as the largest tested reliability failure. The point is not a new twin or a full-archive leaderboard win; it is a bounded protocol that reports those failure modes explicitly instead of treating average accuracy as a deployment guarantee.","feed_headline":"Average RUL coverage hides regime and sensor failures","feed_subtitle":"On a 10-bearing load-speed subset, target-band intervals still undercover one cell and collapse when the raw channel is lost.","key_machinery":"Leave-operating-regime-out residual calibration: derived load-speed regimes define held-out evaluation units; models see only measured load and speed; intervals are formed from empirical absolute residuals on a separate calibration regime (with optional ensemble dispersion), then judged by conditional coverage and stress probes rather than average MAE alone.","core_discovery":"Under strict leave-one-operating-regime-out evaluation with separate train, validation, calibration, and test roles on the processed 10-bearing PHME subset, a residual-calibrated predictive-representation model attains normalized MAE 0.1477 with empirical 90% coverage 0.900, while a 400-tree random forest is close on point error but undercovers. The same protocol exposes non-uniform reliability—coverage 0.666 in low-load/high-speed—and identifies raw-channel loss as the largest tested reliability failure mode. The contribution is therefore that bounded reliability-evaluation protocol with failure modes reported as evidence, not deployment guarantees or full-archive superiority.","pith_inferences":["The same regime-disjoint residual-calibration checklist could be applied to other time-varying rotating-machinery archives before architecture claims are ranked.","If single-bearing concentration drives hard cells, joint leave-bearing-and-regime-out splits will likely shrink apparent coverage further and should be the next stress test.","Sensor-fault designs that degrade engineered features together with the raw stream would probably expose a larger reliability cliff than the retained-feature raw-channel ablation alone.","Pre-specified Mondrian-style residual radii by load-speed cell may be the shortest path from the reported post-hoc diagnostic to usable conditional intervals."],"forward_implications":["Mixed or random splits are insufficient for bearing RUL claims under time-varying duty; regime-held-out coverage must be reported.","Nominal 90% intervals that average near target can still leave maintenance-critical cells undercovered and must be broken out by regime and bearing.","Raw vibration loss is a first-class reliability failure mode even when engineered features and load/speed context remain available.","Post-hoc regime-conditioned residual scaling can repair undercoverage cells only as motivation for pre-specified conditional calibration, not as a silent fix.","Maintenance-trigger value of intervals should be scored at the bearing-regime unit, not only as window-level width or score."],"fun_headline_variants":["Average RUL coverage hides regime undercoverage and sensor loss","Regime-conditioned residuals expose 0.666 cell coverage gaps","Raw-channel drop is largest tested RUL reliability failure mode","Strict regime-out protocol shows non-uniform interval trust","Point accuracy holds but conditional coverage fails under shift"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The load-speed regime bins that define every held-out test cell are cut using quantiles of the frozen full 10-bearing analysis set, so the split structure is an evaluation design choice rather than a train-only, deployable operating-condition rule.","fun_headline_variants_meta":{"raw":{"variants":["Average RUL coverage hides regime undercoverage and sensor loss","Regime-conditioned residuals expose 0.666 cell coverage gaps","Raw-channel drop is largest tested RUL reliability failure mode","Strict regime-out protocol shows non-uniform interval trust","Point accuracy holds but conditional coverage fails under shift"]},"model":"grok-4.5","effort":"low","cost_usd":0.006068,"raw_usage":{"total_tokens":1608,"prompt_tokens":895,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":60680000,"prompt_tokens_details":{"text_tokens":895,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":649,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":895,"tokens_out":64,"duration_ms":5933,"temperature":1.0,"reasoning_tokens":649,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T10:14:53.529166+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Recompute the nine load-speed tertile cuts using only training windows for each rotation (or process the remaining seven public bearings under the same protocol) and check whether primary coverage, the low-load/high-speed undercoverage cell, and the raw-channel-loss collapse still hold at comparable magnitudes.","supporting_citations":[],"review_version":1}