Pith. sign in

REVIEW 3 major objections 6 minor 52 references

Average RUL coverage can look fine under changing load and speed while specific regimes and raw-channel loss still fail.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 10:14 UTC pith:IVF7CNWV

load-bearing objection Honest, bounded reliability-evaluation protocol for PHME bearing RUL under regime shift; useful packaging of leave-regime-out calibration and failure-mode reporting, not a new architecture or full-archive result. the 3 major comments →

arxiv 2607.08273 v1 pith:IVF7CNWV submitted 2026-07-09 cs.CE

Empirical Calibration and Conditional-Reliability Diagnostics for Bearing RUL Prediction under Operating-Regime Shift

classification cs.CE
keywords reliabilityprognosticsremaining useful lifecalibrationconformal predictionoperating-regime shiftbearing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Remaining useful life estimates only help maintenance if both the point forecasts and their intervals stay trustworthy when load and speed change. Mixed train-test splits often hide that failure. This paper fixes a documented 10-bearing subset with time-varying conditions, holds out entire derived load-speed regimes, and asks whether a residual-calibrated model that fuses raw vibration, engineered descriptors, and measured load and speed still covers at the nominal rate. On the strict four-way split the model reaches normalized MAE 0.1477 and empirical 90% coverage 0.900, close to a strong random forest, but conditional checks show coverage falling to 0.666 in a low-load/high-speed cell and raw-channel loss as the largest tested reliability failure. The point is not a new twin or a full-archive leaderboard win; it is a bounded protocol that reports those failure modes explicitly instead of treating average accuracy as a deployment guarantee.

Core claim

Under strict leave-one-operating-regime-out evaluation with separate train, validation, calibration, and test roles on the processed 10-bearing PHME subset, a residual-calibrated predictive-representation model attains normalized MAE 0.1477 with empirical 90% coverage 0.900, while a 400-tree random forest is close on point error but undercovers. The same protocol exposes non-uniform reliability—coverage 0.666 in low-load/high-speed—and identifies raw-channel loss as the largest tested reliability failure mode. The contribution is therefore that bounded reliability-evaluation protocol with failure modes reported as evidence, not deployment guarantees or full-archive superiority.

What carries the argument

Leave-operating-regime-out residual calibration: derived load-speed regimes define held-out evaluation units; models see only measured load and speed; intervals are formed from empirical absolute residuals on a separate calibration regime (with optional ensemble dispersion), then judged by conditional coverage and stress probes rather than average MAE alone.

Load-bearing premise

The load-speed regime bins that define every held-out test cell are cut using quantiles of the frozen full 10-bearing analysis set, so the split structure is an evaluation design choice rather than a train-only, deployable operating-condition rule.

What would settle it

Recompute the nine load-speed tertile cuts using only training windows for each rotation (or process the remaining seven public bearings under the same protocol) and check whether primary coverage, the low-load/high-speed undercoverage cell, and the raw-channel-loss collapse still hold at comparable magnitudes.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Mixed or random splits are insufficient for bearing RUL claims under time-varying duty; regime-held-out coverage must be reported.
  • Nominal 90% intervals that average near target can still leave maintenance-critical cells undercovered and must be broken out by regime and bearing.
  • Raw vibration loss is a first-class reliability failure mode even when engineered features and load/speed context remain available.
  • Post-hoc regime-conditioned residual scaling can repair undercoverage cells only as motivation for pre-specified conditional calibration, not as a silent fix.
  • Maintenance-trigger value of intervals should be scored at the bearing-regime unit, not only as window-level width or score.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same regime-disjoint residual-calibration checklist could be applied to other time-varying rotating-machinery archives before architecture claims are ranked.
  • If single-bearing concentration drives hard cells, joint leave-bearing-and-regime-out splits will likely shrink apparent coverage further and should be the next stress test.
  • Sensor-fault designs that degrade engineered features together with the raw stream would probably expose a larger reliability cliff than the retained-feature raw-channel ablation alone.
  • Pre-specified Mondrian-style residual radii by load-speed cell may be the shortest path from the reported post-hoc diagnostic to usable conditional intervals.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript studies remaining-useful-life (RUL) point accuracy and interval reliability for bearings under time-varying load and speed on a documented 10-bearing PHME subset. Derived load–speed regimes define leave-one-regime-out evaluation units, while models receive only measured load and speed as context. A fused predictive-representation model (raw vibration, engineered descriptors, operating context) is trained under a strict train/validation/calibration/test separation and converted to intervals by empirical residual calibration. On this protocol the model reports normalized MAE 0.1477, empirical 90% coverage 0.900, and retrospective absolute-step MAE 285.26, close to a 400-tree random forest (0.1538 / 0.871 / 294.57). Conditional diagnostics show non-uniform reliability (notably 0.666 coverage in low-load/high-speed), a post-hoc pooled regime-conditioned residual diagnostic is offered only as motivation for future conditional calibration, and raw-channel loss is identified as the largest tested reliability failure mode. The stated contribution is a bounded reliability-evaluation protocol with explicit failure modes, not full-archive PHME SOTA or deployment guarantees.

Significance. If the reported protocol and diagnostics hold under the stated evidence boundary, the paper is a useful reliability-facing contribution rather than another architecture claim. Its main value is methodological honesty: regime-disjoint four-way splitting, residual calibration separated from model selection, strong tabular comparators (including a 400-tree forest), conditional coverage with concentration diagnostics, ablations, prefix/noise/raw-loss stress, and explicit non-claims about physics twins and maintenance policy. Public code and configuration provenance further strengthen reproducibility within the processed subset. The work will matter most to PHM readers who care whether intervals remain trustworthy under operating-condition shift; the significance is therefore as a careful evaluation template and failure-mode inventory on a transparent 10-bearing PHME slice, not as a general RUL breakthrough.

major comments (3)
  1. Section 3.3 defines the nine held-out regimes by empirical one-third/two-third load and speed quantiles of the frozen full 14,297-window analysis set, so every later test cell participates in the cutpoints that isolate it. The paper correctly labels this an evaluation-design choice rather than a deployable train-only classifier, but the primary claim is still framed as a strict leave-operating-regime-out reliability protocol (Abstract; §3.4; Table 8). Because those cutpoints are load-bearing for the headline MAE/coverage pair and for the hard low-load/high-speed cell, a sensitivity analysis that re-derives tertiles from training regimes only (or from non-test regimes only) and re-runs the primary endpoint is needed. If re-binning materially reassigns windows or changes the 0.666 LL/HS coverage, the “strict protocol” numbers rest on a softer partition than currently implied.
  2. Table 10 and Figure 4 show that the critical low-load/high-speed undercoverage cell (coverage 0.666) has dominant-bearing share 0.947 from B17; Table 11’s exclusion diagnostic then rests on only 119 remaining windows. Conditional undercoverage is central to the paper’s reliability narrative, yet this cell is effectively a near-single-trajectory estimate. The manuscript should either (i) reweight or restate that cell as bearing-confounded rather than as a pure regime failure mode, or (ii) provide additional regime-definition / leave-bearing-within-regime checks so that the 0.666 figure is not over-read as multi-bearing regime unreliability.
  3. §4.2 and Table 6 make clear that absolute-step MAE multiplies normalized predictions by each bearing’s known maximum step count and is therefore a retrospective offline metric. Tables 8–9 and the Abstract still lead with absolute-step MAE alongside normalized MAE and coverage. For a reliability-evaluation paper whose decision relevance is repeatedly emphasized, the primary quantitative claims should lead with normalized MAE/coverage (and interval score), with absolute-step figures demoted or consistently labeled “retrospective” in every main table caption so readers cannot mistake them for deployment-time lifetime units.
minor comments (6)
  1. Figure 1 is useful but dense; the matched-sensitivity branch is easy to miss relative to the primary four-way split. A one-line caption callout that only the strict four-way split is the primary endpoint would help.
  2. Table 13 residual-only coverages (e.g., 0.868 at nominal 0.90 for the predictive representation) sit below the primary ensemble-dispersion coverage of 0.900 in Table 8; a short sentence in §6.2 clarifying why residual-only and primary intervals differ would prevent confusion.
  3. §5.1 notes that bitwise-identical CUDA replay is not asserted. For a reproducibility-oriented reliability paper, stating which metrics were verified from saved prediction files versus re-trained runs would be helpful.
  4. Leave-bearing-out coverage of 0.821 (Table 18) falls outside the stated 0.85–0.95 band; this is already treated as supporting evidence, but a brief cross-reference in the Abstract or Conclusion would better match the paper’s otherwise careful non-overclaim style.
  5. Minor wording: “Cross-Transformer fusioning” in the Hou et al. citation title (§2.6 / References) appears to be a transcription artifact; verify against the source title.
  6. Eq. (6) first-passage surrogate is clear, but the role of ε = 10^{-4} and the independent-window encoding limitation could be restated once near the monotonicity diagnostic so readers do not expect trajectory-level monotone H_i.

Circularity Check

1 steps flagged

Mild evaluation-design dependence on full-set tertiles for defining held-out regimes; no prediction or coverage result reduces to its inputs by construction.

specific steps
  1. other [Section 3.3 (Operating-Regime Definition)]
    "The preprocessing code divides load and speed into empirical low, middle, and high bins using the one-third and two-third quantiles of the analysis set. ... Because these cutpoints are derived from the frozen analysis set, the regime definition should be read as an evaluation-design choice rather than a deployable train-only operating-condition classifier."

    The primary leave-one-operating-regime-out units are defined by quantiles computed on the full 14 297-window analysis set, so every test cell participates in setting the boundaries that later isolate it. This is not feature leakage (models receive only continuous load/speed) and does not force coverage or MAE by construction, but it means the reported “strict” regime-disjoint reliability numbers rest on a non-train-only partition of the same frozen subset fixed before comparison.

full rationale

This is an empirical reliability-evaluation paper on a fixed 10-bearing PHME subset, not a first-principles derivation. Point predictions and residual quantiles are obtained under explicit train/validation/calibration/test separation; empirical coverage is measured on held-out regimes rather than forced to the nominal level by the fit. Absolute-step MAE is transparently retrospective (normalized predictions scaled by known trajectory length). The sole mild circularity burden is that the nine load–speed cells used as primary evaluation units are formed from one-third/two-third quantiles of the entire frozen analysis set (including future test windows). The paper itself labels this an evaluation-design choice rather than a deployable train-only classifier and withholds regime codes from model inputs, so the issue does not collapse the reported MAE/coverage numbers into a tautology. No self-citation chain, uniqueness import, or ansatz smuggling carries the central claim. Score 2 reflects that single acknowledged design dependence without elevating it to a constructional reduction of the results.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The central claim is empirical and protocol-bound: it rests on standard supervised RUL labeling, pragmatic regime binning, residual-quantile interval construction related to split conformal practice, and many ordinary training hyperparameters. No new physical particle or force is introduced; the 'predictive representation' and first-passage surrogate are modeling devices. The main non-standard loads are evaluation-design choices (full-set tertiles, 10-bearing freeze, retrospective lifetime scaling) and the exchangeability-like hope that calibration residuals transfer under regime shift—which the paper itself stress-tests and partially falsifies conditionally.

free parameters (5)
  • load/speed empirical tertile cutpoints (1/3 and 2/3 quantiles)
    Define the nine held-out regimes that are the primary evaluation units; computed on the frozen analysis set rather than train-only.
  • neural hidden width / ensemble size / learning rate / weight decay / batch / patience
    Primary model uses width 96, 3-member ensemble, AdamW 1e-3, weight decay 1e-4, batch 64, patience 18, up to 120 epochs—standard fitted training choices that affect reported MAE/coverage.
  • nominal miscoverage α = 0.1 residual quantile
    Sets the empirical interval half-width via the calibration residual quantile; target coverage band 85–95% is also a chosen reporting rule.
  • maintenance trigger threshold τ = 0.20 and cost ratios ρ
    Late-group definition and stylized risk L_ρ depend on hand-chosen τ and ρ grid; diagnostic only but free design parameters.
  • window length 512 samples, single channel, per-window standardization
    Fixed preprocessing design that shapes all features and labels; not optimized or ablated for amplitude cues.
axioms (5)
  • domain assumption RUL labels can be constructed retrospectively from known run-to-failure endpoints within each bearing trajectory for offline evaluation.
    Stated in Sections 3.1 and 4.1–4.2; standard for offline RUL studies but not true online.
  • domain assumption Empirical residual quantiles from a held-out calibration regime yield useful interval radii under leave-regime-out shift (split-conformal spirit without claiming finite-sample guarantees under non-exchangeability).
    Section 4.5; paper measures empirical coverage and reports conditional failures rather than proving coverage.
  • ad hoc to paper Derived regime labels may define splits/reporting while models receive only continuous load and speed.
    Core protocol choice in Sections 3.3–3.4 and 5.1 to avoid giving models the evaluation keys.
  • domain assumption Weak monotonicity/smoothness regularizers are optional inductive biases, not validated fatigue physics.
    Sections 4.4 and 7.2; authors correctly refuse mechanism claims.
  • standard math Standard supervised learning and bootstrap/sign-flip/Wilcoxon paired tests over nine regimes support comparative inference.
    Section 5.4; small n=9 regimes limits power, which the paper notes.
invented entities (2)
  • Calibrated predictive-representation RUL model (fused encoders + first-passage health/rate surrogate + residual calibration) no independent evidence
    purpose: Provide a competitive neural baseline that produces point RUL and empirically calibrated intervals under the protocol.
    Architecture is a study-specific assembly of known parts (1D CNN, MLPs, residual quantiles, optional weak regularizers), not an independently evidenced physical damage coordinate.
  • Pooled regime-conditioned residual diagnostic (cross-rotation) no independent evidence
    purpose: Probe whether group-specific residual scale explains conditional undercoverage and motivate future pre-specified conditional calibration.
    Explicitly post-hoc and not a primary Mondrian conformal guarantee (Section 4.5, Table 14).

pith-pipeline@v1.1.0-grok45 · 32008 in / 4124 out tokens · 44070 ms · 2026-07-10T10:14:53.529166+00:00 · methodology

0 comments
read the original abstract

Remaining useful life (RUL) estimates support reliability and maintenance decisions only if both point accuracy and prediction intervals remain trustworthy when operating conditions change. Convenient mixed splits can hide that failure. This paper studies the question on a documented 10-bearing PHME subset with time-varying load and speed. Derived load-speed regimes define the held-out evaluation units, while models receive only measured load and speed as context. A calibrated predictive-representation model fuses raw vibration windows, engineered descriptors, and operating context, then forms intervals by empirical residual calibration. Under strict train/validation/calibration/test separation, the model reaches normalized MAE 0.1477, empirical 90% coverage 0.900, and retrospective absolute-step MAE 285.26; a 400-tree random forest reaches 0.1538, 0.871, and 294.57. The results do not show uniform dominance: conditional diagnostics expose non-uniform reliability, including 0.666 coverage in a low-load/high-speed cell, and a post-hoc pooled regime-conditioned residual diagnostic raises that cell to 0.941 only as motivation for future pre-specified conditional calibration. Stress tests further identify raw-channel loss as the largest tested reliability failure mode. The contribution is therefore a bounded reliability-evaluation protocol for the processed 10-bearing subset, with conditional undercoverage and raw-channel loss reported explicitly as failure modes rather than deployment guarantees.

Figures

Figures reproduced from arXiv: 2607.08273 by Jun Wang, Shaoliang Yang, Yunsheng Wang.

Figure 1
Figure 1. Figure 1: Four-part protocol schematic for the bounded PHME subset, derived load-speed regimes, [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Small-multiple load-speed phase planes for the ten evaluated PHME bearing runs. [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Three-by-three derived regime heatmap for the 10-bearing analysis set. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Bearing-by-regime window-count heatmap in the 10-bearing PHME analysis set. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Calibrated predictive-representation architecture with explicit post-hoc residual interval [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Per-regime conditional coverage heatmap for the separate-calibrated predictive [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Strict-endpoint MAE–coverage tradeoff and per-regime relative MAE gains against the [PITH_FULL_IMAGE:figures/full_fig_p023_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Per-bearing 10-bearing leave-bearing-out diagnostic performance. [PITH_FULL_IMAGE:figures/full_fig_p025_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Matched-protocol diagnostic prefix-observation and robustness MAE profile for the cali [PITH_FULL_IMAGE:figures/full_fig_p027_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Matched-protocol diagnostic empirical-interval coverage and interval score across prefix [PITH_FULL_IMAGE:figures/full_fig_p027_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Bearing-regime maintenance-risk curve under miss-to-early-inspection cost ratios. [PITH_FULL_IMAGE:figures/full_fig_p029_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 52 canonical work pages · 1 internal anchor

  1. [1]

    Run-to-failure data set of ball bearings subjected to time-varying load and speed conditions

    Osarenren Kennedy Aimiyekagbon. Run-to-failure data set of ball bearings subjected to time-varying load and speed conditions. Zenodo dataset, 2024. URLhttps://zenodo.org/ records/10868257

  2. [2]

    Angelopoulos and Stephen Bates

    Anastasios N. Angelopoulos and Stephen Bates. A gentle introduction to conformal prediction and distribution-free uncertainty quantification, 2021. URLhttps://arxiv.org/abs/2107. 07511

  3. [3]

    An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

    Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. An empirical evaluation of generic convolu- tional and recurrent networks for sequence modeling.arXiv preprint arXiv:1803.01271, 2018. doi: 10.48550/arXiv.1803.01271. URLhttps://arxiv.org/abs/1803.01271

  4. [5]

    Random forests.Machine Learning, 45(1):5–32, 2001

    Leo Breiman. Random forests.Machine Learning, 45(1):5–32, 2001. doi: 10.1023/A: 1010933404324

  5. [6]

    Digital twin-driven graph domain adaptation neural network for remaining useful life prediction of rolling bearing.Re- liability Engineering & System Safety, 245:109991, 2024

    Lingli Cui, Yongchang Xiao, Dongdong Liu, and Honggui Han. Digital twin-driven graph domain adaptation neural network for remaining useful life prediction of rolling bearing.Re- liability Engineering & System Safety, 245:109991, 2024. doi: 10.1016/j.ress.2024.109991

  6. [7]

    Remain- ing useful lifetime prediction via deep domain adaptation.Reliability Engineering & System Safety, 195:106682, 2020

    Paulo Roberto de Oliveira da Costa, Alp Akçay, Yingqian Zhang, and Uzay Kaymak. Remain- ing useful lifetime prediction via deep domain adaptation.Reliability Engineering & System Safety, 195:106682, 2020. doi: 10.1016/j.ress.2019.106682

  7. [8]

    Remain- ing useful life prediction based on physics-informed data augmentation.Reliability Engineering & System Safety, 252:110451, 2024

    Martin Hervé de Beaulieu, Mayank Shekhar Jha, Hugues Garnier, and Farid Cerbah. Remain- ing useful life prediction based on physics-informed data augmentation.Reliability Engineering & System Safety, 252:110451, 2024. doi: 10.1016/j.ress.2024.110451

  8. [9]

    Remaining useful life estimation using deep metric transfer learning for kernel regression.Reliability Engineering & System Safety, 212:107583, 2021

    Yifei Ding, Minping Jia, Qiuhua Miao, and Peng Huang. Remaining useful life estimation using deep metric transfer learning for kernel regression.Reliability Engineering & System Safety, 212:107583, 2021. doi: 10.1016/j.ress.2021.107583

  9. [10]

    Olga Fink, Qin Wang, Markus Svensén, Pierre Dersin, Wan-Jui Lee, and Melanie Ducoffe. Potential, challenges and future directions for deep learning in prognostics and health man- agement applications.Engineering Applications of Artificial Intelligence, 92:103678, 2020. doi: 10.1016/j.engappai.2020.103678

  10. [11]

    The Annals of Statistics , volume=

    Jerome H. Friedman. Greedy function approximation: A gradient boosting machine.The Annals of Statistics, 29(5):1189–1232, 2001. doi: 10.1214/aos/1013203451. 32

  11. [12]

    Digital twin: Enabling technologies, challenges and open research.IEEE Access, 8:108952–108971, 2020

    Aidan Fuller, Zhong Fan, Charles Day, and Chris Barlow. Digital twin: Enabling technologies, challenges and open research.IEEE Access, 8:108952–108971, 2020. doi: 10.1109/ACCESS. 2020.2998358

  12. [13]

    Fengjin Gong, Ping Ma, Hongli Zhang, Cong Wang, Xinkai Li, and Yinfei Wu. Rolling bearings remaining useful life estimation using digital twin and physics-informed methods with uncer- tainty quantification.Engineering Applications of Artificial Intelligence, 154:111070, 2025. doi: 10.1016/j.engappai.2025.111070

  13. [14]

    A recurrent neural network based health indicator for remaining useful life prediction of bearings.Neurocomputing, 240:98–109,

    Liang Guo, Naipeng Li, Feng Jia, Yaguo Lei, and Jing Lin. A recurrent neural network based health indicator for remaining useful life prediction of bearings.Neurocomputing, 240:98–109,

  14. [15]

    doi: 10.1016/j.neucom.2017.02.045

  15. [16]

    Yuxuan He, Huai Su, Enrico Zio, Shiliang Peng, Lin Fan, Zhaoming Yang, Zhe Yang, and Jinjun Zhang. A systematic method of remaining useful life estimation based on physics- informed graph neural networks with multisensor data.Reliability Engineering & System Safety, 237:109333, 2023. doi: 10.1016/j.ress.2023.109333

  16. [17]

    Aiwina Heng, Sheng Zhang, Andy C. C. Tan, and Joseph Mathew. Rotating machinery prognostics: State of the art, challenges and opportunities.Mechanical Systems and Signal Processing, 23(3):724–739, 2009. doi: 10.1016/j.ymssp.2008.06.009

  17. [18]

    Dongxiao Hou, JiaHui Chen, Rongcai Cheng, Xue Hu, and Peiming Shi. A bearing remaining lifepredictionmethodundervariableoperatingconditionsbasedoncross-transformerfusioning segmented data cleaning.Reliability Engineering & System Safety, 245:110021, 2024. doi: 10.1016/j.ress.2024.110021

  18. [19]

    Prognostics and health man- agement: A review from the perspectives of design, development and decision.Reliability Engineering & System Safety, 217:108063, 2022

    Yang Hu, Xuewen Miao, Yong Si, Ershun Pan, and Enrico Zio. Prognostics and health man- agement: A review from the perspectives of design, development and decision.Reliability Engineering & System Safety, 217:108063, 2022. doi: 10.1016/j.ress.2021.108063

  19. [20]

    Andrew K. S. Jardine, Daming Lin, and Dragan Banjevic. A review on machinery diagnostics and prognostics implementing condition-based maintenance.Mechanical Systems and Signal Processing, 20(7):1483–1510, 2006. doi: 10.1016/j.ymssp.2005.09.012

  20. [21]

    Conformal prediction intervals for remaining useful lifetime estimation.International Journal of Prognostics and Health Management, 14(2), 2023

    Alireza Javanmardi and Eyke Hüllermeier. Conformal prediction intervals for remaining useful lifetime estimation.International Journal of Prognostics and Health Management, 14(2), 2023. doi: 10.36001/ijphm.2023.v14i2.3417

  21. [22]

    Remaining useful lifetime estimation of bearings operating under time-varying conditions.PHM Society European Conference, 8(1):9, 2024

    Alireza Javanmardi, Osarenren Kennedy Aimiyekagbon, Amelie Bender, James Kuria Ki- motho, Walter Sextro, and Eyke Hüllermeier. Remaining useful lifetime estimation of bearings operating under time-varying conditions.PHM Society European Conference, 8(1):9, 2024. doi: 10.36001/phme.2024.v8i1.4101

  22. [23]

    Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang

    George Em Karniadakis, Ioannis G. Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang. Physics-informed machine learning.Nature Reviews Physics, 3(6):422–440, 2021. doi: 10.1038/s42254-021-00314-5

  23. [24]

    Journal of the American Statistical Association , author =

    Jing Lei, Max G’Sell, Alessandro Rinaldo, Ryan J. Tibshirani, and Larry Wasserman. Distribution-free predictive inference for regression.Journal of the American Statistical Asso- ciation, 113(523):1094–1111, 2018. doi: 10.1080/01621459.2017.1307116. 33

  24. [25]

    Machinery health prognostics: A systematic review from data acquisition to RUL prediction.Mechanical Systems and Signal Processing, 104:799–834, 2018

    Yaguo Lei, Naipeng Li, Liang Guo, Ningbo Li, Tao Yan, and Jun Lin. Machinery health prognostics: A systematic review from data acquisition to RUL prediction.Mechanical Systems and Signal Processing, 104:799–834, 2018. doi: 10.1016/j.ymssp.2017.11.016

  25. [26]

    A review on physics-informed data- driven remaining useful life prediction: Challenges and opportunities.Mechanical Systems and Signal Processing, 209:111120, 2024

    Huiqin Li, Zhengxin Zhang, Tianmei Li, and Xiaosheng Si. A review on physics-informed data- driven remaining useful life prediction: Challenges and opportunities.Mechanical Systems and Signal Processing, 209:111120, 2024. doi: 10.1016/j.ymssp.2024.111120

  26. [27]

    Naipeng Li, Nagi Gebraeel, Yaguo Lei, Linkan Bian, and Xiaosheng Si. Remaining useful life prediction of machinery under time-varying operating conditions based on a two-factor state- space model.Reliability Engineering & System Safety, 186:88–100, 2019. doi: 10.1016/j.ress. 2019.02.017

  27. [28]

    Hierarchical attention graph convolutional network to fuse multi-sensor signals for remaining useful life prediction

    Tianfu Li, Zhibin Zhao, Chuang Sun, Ruqiang Yan, and Xuefeng Chen. Hierarchical attention graph convolutional network to fuse multi-sensor signals for remaining useful life prediction. Reliability Engineering & System Safety, 215:107878, 2021. doi: 10.1016/j.ress.2021.107878

  28. [29]

    The emerging graph neural networks for intelligent fault diagnostics and prognostics: A guideline and a benchmark study.Mechanical Systems and Signal Processing, 168:108653, 2022

    Tianfu Li, Zheng Zhou, Sinan Li, Chuang Sun, Ruqiang Yan, and Xuefeng Chen. The emerging graph neural networks for intelligent fault diagnostics and prognostics: A guideline and a benchmark study.Mechanical Systems and Signal Processing, 168:108653, 2022. doi: 10.1016/ j.ymssp.2021.108653

  29. [30]

    A reliable bearing remaining useful life prediction method based on multi-hierarchy dynamic evaluation and uncertainty amelioration

    Wenjie Li, Dongdong Liu, Xin Wang, and Lingli Cui. A reliable bearing remaining useful life prediction method based on multi-hierarchy dynamic evaluation and uncertainty amelioration. Reliability Engineering & System Safety, 263:111270, 2025. doi: 10.1016/j.ress.2025.111270

  30. [31]

    Remaining useful life estimation in prognostics using deep convolution neural networks.Reliability Engineering & System Safety, 172:1–11, 2018

    Xiang Li, Qian Ding, and Jian-Qiao Sun. Remaining useful life estimation in prognostics using deep convolution neural networks.Reliability Engineering & System Safety, 172:1–11, 2018. doi: 10.1016/j.ress.2017.11.021

  31. [32]

    Deep learning-based remaining useful life estimation of bearings using multi-scale feature extraction.Reliability Engineering & System Safety, 182: 208–218, 2019

    Xiang Li, Wei Zhang, and Qian Ding. Deep learning-based remaining useful life estimation of bearings using multi-scale feature extraction.Reliability Engineering & System Safety, 182: 208–218, 2019. doi: 10.1016/j.ress.2018.11.011

  32. [33]

    Remaining useful life with self- attention assisted physics-informed neural network.Advanced Engineering Informatics, 58: 102195, 2023

    Xinyuan Liao, Shaowei Chen, Pengfei Wen, and Shuai Zhao. Remaining useful life with self- attention assisted physics-informed neural network.Advanced Engineering Informatics, 58: 102195, 2023. doi: 10.1016/j.aei.2023.102195

  33. [34]

    Digital twin-driven remaining useful life prediction for rolling element bearing.Machines, 11(7):678, 2023

    Quanbo Lu and Mei Li. Digital twin-driven remaining useful life prediction for rolling element bearing.Machines, 11(7):678, 2023. doi: 10.3390/machines11070678

  34. [35]

    IMS Bearings dataset

    NASA Prognostics Center of Excellence. IMS Bearings dataset. NASA Open Data Portal,

  35. [36]

    Accessed 30 April 2026

    URLhttps://data.nasa.gov/dataset/ims-bearings. Accessed 30 April 2026

  36. [37]

    PRONOSTIA: An experimental plat- form for bearings accelerated degradation tests

    Patrick Nectoux, Rafael Gouriveau, Kamal Medjaher, Emmanuel Ramasso, Brigitte Chebel- Morello, Noureddine Zerhouni, and Christophe Varnier. PRONOSTIA: An experimental plat- form for bearings accelerated degradation tests. InIEEE International Conference on Prog- nostics and Health Management, pages 1–8, 2012

  37. [38]

    Scikit-learn: Machine learning in python.Journal of Machine Learning Research, 12:2825–2830, 2011

    Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and 34 Édouard Duchesnay. Scikit-learn: Machine learning in python.Journal of Machine Learning Resea...

  38. [39]

    Maziar Raissi, Paris Perdikaris, and George Em Karniadakis. Physics-informed neural net- works: A deep learning framework for solving forward and inverse problems involving nonlin- ear partial differential equations.Journal of Computational Physics, 378:686–707, 2019. doi: 10.1016/j.jcp.2018.10.045

  39. [40]

    Prediction of bearing remaining useful life with deep convolution neural network.IEEE Access, 6:13041–13049, 2018

    Lei Ren, Yaqiang Sun, Hao Wang, and Lin Zhang. Prediction of bearing remaining useful life with deep convolution neural network.IEEE Access, 6:13041–13049, 2018. doi: 10.1109/ ACCESS.2018.2804930

  40. [41]

    Conformalized quantile regression

    Yaniv Romano, Evan Patterson, and Emmanuel Candès. Conformalized quantile regression. InAdvances in Neural Information Processing Systems, volume 32, pages 3538–3548, 2019

  41. [42]

    Significance, interpretation, and quantification of uncertainty in prog- nostics and remaining useful life prediction.Mechanical Systems and Signal Processing, 52–53: 228–247, 2015

    Shankar Sankararaman. Significance, interpretation, and quantification of uncertainty in prog- nostics and remaining useful life prediction.Mechanical Systems and Signal Processing, 52–53: 228–247, 2015. doi: 10.1016/j.ymssp.2014.05.029

  42. [43]

    Remaining useful life es- timation: A review on the statistical data driven approaches.European Journal of Operational Research, 213(1):1–14, 2011

    Xiao-Sheng Si, Wenbin Wang, Chang-Hua Hu, and Dong-Hua Zhou. Remaining useful life es- timation: A review on the statistical data driven approaches.European Journal of Operational Research, 213(1):1–14, 2011. doi: 10.1016/j.ejor.2010.11.018

  43. [44]

    Sikorska, Melinda Hodkiewicz, and Lin Ma

    Joanna Z. Sikorska, Melinda Hodkiewicz, and Lin Ma. Prognostic modelling options for re- maining useful life estimation by industry.Mechanical Systems and Signal Processing, 25(5): 1803–1836, 2011. doi: 10.1016/j.ymssp.2010.11.018

  44. [45]

    Fei Tao, He Zhang, Ang Liu, and Andrew Y. C. Nee. Digital twin in industry: State-of-the-art. IEEE Transactions on Industrial Informatics, 15(4):2405–2415, 2019. doi: 10.1109/TII.2018. 2873186

  45. [46]

    Deep separable convolutional network for remaining useful life prediction of machinery.Mechanical Systems and Signal Processing, 134: 106330, 2019

    Biao Wang, Yaguo Lei, Naipeng Li, and Tao Yan. Deep separable convolutional network for remaining useful life prediction of machinery.Mechanical Systems and Signal Processing, 134: 106330, 2019. doi: 10.1016/j.ymssp.2019.106330

  46. [47]

    A hybrid prognostics approach for esti- mating remaining useful life of rolling element bearings.IEEE Transactions on Reliability, 69 (1):401–412, 2020

    Biao Wang, Yaguo Lei, Naipeng Li, and Ningbo Li. A hybrid prognostics approach for esti- mating remaining useful life of rolling element bearings.IEEE Transactions on Reliability, 69 (1):401–412, 2020. doi: 10.1109/TR.2018.2882682

  47. [48]

    Xin Wang, Yongbo Li, Khandaker Noman, and Asoke K. Nandi. Multi-task learning mixture density network for interval estimation of the remaining useful life of rolling element bearings. Reliability Engineering & System Safety, 251:110348, 2024. doi: 10.1016/j.ress.2024.110348

  48. [49]

    Remaining useful life prediction with uncertainty quantification based on multi-distribution fusion structure

    Yuling Zhan, Ziqian Kong, Ziqi Wang, Xiaohang Jin, and Zhengguo Xu. Remaining useful life prediction with uncertainty quantification based on multi-distribution fusion structure. Reliability Engineering & System Safety, 251:110383, 2024. doi: 10.1016/j.ress.2024.110383

  49. [50]

    Jianjing Zhang, Peng Wang, Ruqiang Yan, and Robert X. Gao. Long short-term memory for machine remaining life prediction.Journal of Manufacturing Systems, 48:78–86, 2018. doi: 10.1016/j.jmsy.2018.05.011. 35

  50. [51]

    Wei Zhang, Xiang Li, Hui Ma, Zhong Luo, and Xu Li. Transfer learning using deep represen- tation regularization in remaining useful life prediction across operating conditions.Reliability Engineering & System Safety, 211:107556, 2021. doi: 10.1016/j.ress.2021.107556

  51. [52]

    Rui Zhao, Ruqiang Yan, Zhenghua Chen, Kezhi Mao, Peng Wang, and Robert X. Gao. Deep learning and its applications to machine health monitoring.Mechanical Systems and Signal Processing, 115:213–237, 2019. doi: 10.1016/j.ymssp.2018.05.050

  52. [53]

    Prognostics and health management (PHM): Where are we and where do we (need to) go in theory and practice.Reliability Engineering & System Safety, 218:108119, 2022

    Enrico Zio. Prognostics and health management (PHM): Where are we and where do we (need to) go in theory and practice.Reliability Engineering & System Safety, 218:108119, 2022. doi: 10.1016/j.ress.2021.108119. 36