REVIEW 2 major objections 6 minor 25 references
Clock-linear RUL labels disagree with vibration; a three-stage degradation-state target fits the measured reference better and can be learned on held-out bearings.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 17:24 UTC pith:PBKIL375
load-bearing objection Honest target-vs-predictor separation with a real but descriptive IMS shape gain; useful hygiene, not a new RUL law. the 2 major comments →
When Linear RUL Labels Disagree with Vibration Degradation: A Stage-Aware Target and Dual-Scale Predictor Evaluated on XJTU-SY and IMS
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For the evaluated vibration-derived degradation references, a continuous three-stage linear–quadratic–exponential target describes remaining-life state better than the strongest comparable global linear target, and that stage-aware target is learnable from causal feature sequences on bearings excluded from every fitted development step.
What carries the argument
The stage-aware degradation-state target: an oriented PCA health indicator is clustered into chronological early–middle–late stages, then fitted with a continuous linear–quadratic–exponential curve to the complementary normalized indicator, and compared against a best anchored linear fit to the same reference.
Load-bearing premise
The whole case rests on treating a smoothed, oriented first principal component of screened vibration features, forced into three contiguous stages, as a faithful enough picture of real degradation that fitting curves to it tests whether labels are adequate.
What would settle it
On additional complete failed-bearing runs from a different rig, check whether the same stage-aware curve still cuts RMSE and MAE versus the best anchored linear fit to that run’s vibration health reference, and whether blocked-bootstrap intervals for the squared-error gain stay above zero.
If this is right
- RUL benchmarks should report target-shape checks against a vibration-derived reference, not only predictor error under a fixed clock-linear label.
- Strong sequence models cannot fix a mismatched label; target design must be audited separately from architecture.
- A better condition-consistent target still leaves cross-run and cross-fault shift unresolved, so transfer methods remain necessary.
- Deployed systems should treat the output as a normalized degradation-state score that later needs machine-specific time calibration and uncertainty checks.
Where Pith is reading between the lines
- If target audits become standard, many published RUL gains may shrink once models are scored against condition-consistent labels rather than 1−u alone.
- Online stage detection without full-run hindsight is the practical bottleneck between this retrospective supervision recipe and live maintenance use.
- Comparing LQE against monotone splines or change-point families on the same HI reference would show whether three fixed pieces are special or just one workable nonlinear family.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript argues that clock-linear RUL labels can disagree with vibration-observed degradation and therefore separates target construction from prediction. A development-only pipeline builds an oriented PCA health indicator, repairs k-means stages into chronological early–middle–late segments, and fits a continuous linear–quadratic–exponential degradation-state target Rb(u) to the complementary normalized HI reference zb. Causal CNN–LSTM and Transformer branches learn this target from feature sequences, with validation-fitted two-expert OWA fusion. On a bearing-wise XJTU-SY hold-out (Bearings 1–4 development, Bearing 5 test per condition, excluded from all fitted steps), fused RMSE/MAE/R² are 0.0608/0.0392/0.9617, dominated by the Transformer. Independently, on three documented IMS failed-bearing runs, the stage-aware curve reduces RMSE by 3.6–18.2% (mean 10.2%) and MAE by 3.1–31.1% (mean 15.0%) versus the best anchored linear fit to the same zb, with conservative ΔBIC 128.8–368.1 favoring LQE, while moving-block bootstrap intervals cross zero. The authors explicitly limit claims to descriptive target adequacy and learnability, not a universal nonlinear TTF law or robust cross-domain prediction.
Significance. If the bounded claims hold, the paper’s main value is methodological rather than architectural: it treats the RUL label as a testable measurement choice, compares stage-aware and fitted-linear representations against a shared vibration-derived reference on a second rig, and keeps target validity, held-out learnability, and transfer as separate research questions with matching claim boundaries (Table 3). Strengths include leakage-audited development-only preprocessing, Bearing-5 exclusion from normalization/screening/training/early-stopping/OWA, a stricter linear baseline than 1−u (Eqs. 14–15), conservative five-parameter BIC, dependence-aware bootstrap with open non-significance, and transparent negative exploratory IMS leave-one-run-out results. That honesty and protocol design are useful for Measurement/prognostics practice even though n=3 IMS runs and non-significant blocked bootstrap keep RQ1 descriptive. The dual-scale OWA predictor is secondary; the contribution is a target-validity framework, not SOTA architecture dominance.
major comments (2)
- [§2.4–2.6, Table 6, Fig. 8] §2.4–2.6 and Table 6: RQ1’s reference zb is built from the same pipeline’s oriented PCA-HI, smoothing, three-stage k-means with chronological repair, and end-of-run min–max normalization (Eqs. 7–10). Relative gains of LQE vs Rlin(a*) on that shared zb are well-specified and not tautological, but they still test functional form against a surrogate the authors define. Stagewise results already show non-uniform superiority (e.g., IMS2-B1 middle-stage RMSE worse by 14.5%). A load-bearing sensitivity check is needed: alternative HIs (e.g., RMS/kurtosis-only, different q, unsmoothed, or monotone composite), two- vs four-stage repairs, and at least one other parametric family (logistic/Gompertz/monotone spline) on the same three official runs, reporting whether mean RMSE/MAE gains and sign of ΔBIC persist. Without this, “stage-dependent targets better describe the evaluated vibration-derived de
- [§5 Limitations; Data Availability] §5 (Limitations) and Data Availability: the arXiv package omits final selected feature list, window/step, smoothing span, dropout/weight-decay, sequence length, seeds, and numerical OWA weights. For a measurement-oriented target-validity paper whose XJTU numbers and IMS fits are central evidence, these are not optional extras—they are part of the fitted model (§2.3). Either release a persistent-identifier artifact bundle with frozen preprocessing objects, stage boundaries, and exact configs, or demote quantitative XJTU/IMS point estimates to illustrative status until reproduction objects are available. This is fixable but currently load-bearing for verifiability.
minor comments (6)
- [§3.3, Table 9–10] Table 9 vs text: OWA MAE equals Transformer MAE at four decimals (0.0392) while MSE/R² improve only slightly; state explicitly that fusion gain is negligible on MAE and avoid language that could be read as strong ensemble benefit.
- [§2.7] §2.7: random 80/20 window-level validation within development bearings is correctly labeled optimistic; consider adding one blocked/group validation metric in the supplement so early-stopping behavior is auditable.
- [Figs. 10–12] Figs. 10–12: residual envelopes are descriptive only—add a one-line caption reminder that they are not calibrated prediction intervals (already noted in §5).
- [§2.5–2.6, Eq. (11)] Eq. (11): clarify whether u1, u2 are held fixed from clustering during least-squares for α,β,γ or jointly optimized; the conservative BIC counts five parameters, but the fitting paragraph should match that accounting exactly.
- [Table 5, Fig. 9] Table 5 / exploratory extended run: keep the strong 56.4% RMSE figure strictly out of abstract/confirmatory aggregates (already done); ensure Fig. 9 caption remains the sole home for that result.
- [§2.7–2.8] Minor prose: standardize “OWA” vs “OW A” spacing; fix “V alidation” line-break artifacts in §2.7–2.8.
Circularity Check
RQ1 ‘target adequacy’ is largely goodness-of-fit of a stage-aligned LQE ansatz to a self-built HI reference (zb), not an independent degradation measurement—real GOF content remains, but the central shape claim is partly by construction.
specific steps
-
self definitional
[§2.4–2.6, Eqs. (7)–(11), (14)–(16); RQ1 / Table 6 / Fig. 8]
"hb(t)=PCA1({fj(t)}q_j=1) ... ehb(u)=(hb(u)−min hb)/(max hb−min hb+ε), and define the complementary degradation reference zb(u)=1−ehb(u) ... Rb(u)=1−αu [S1]; 1−αu−β(u−u1)2 [S2]; Rb(u2)exp[−γ(u−u2)] [S3]. The transition locations u1 and u2 are obtained from the repaired stage sequence. Parameters α,β,γ≥0 are fitted by least squares to zb(u) ... The main target-shape comparison is therefore between Rlin(u;a⋆) and the three-stage curve in Eq. (11) ... Thus, stage-dependent targets better describe the evaluated vibration-derived degradation states"
The ‘vibration-derived degradation reference’ zb is defined from the paper’s own oriented PCA-HI (and end-of-run min–max). Stage boundaries used inside Rb are obtained by clustering/repairing that same HI. Rb is then LS-fitted to zb. Claiming the stage-aware target better describes that reference is therefore largely claiming that a flexible, stage-tied fit to the pipeline’s surrogate is closer to the surrogate than a global line—i.e., the quantity being validated is constructed from the same objects that define the proposed target, not an external physical RUL or independent wear measurement. Relative GOF can still favor either form, so this is partial self-definitional circularity, not a full identity.
-
fitted input called prediction
[§2.5 Eq. (11); §2.6 Eqs. (14)–(15); Abstract / §3.2 IMS target-shape results]
"Against the best anchored linear fit to the same vibration-derived reference, the stage-aware curve reduces RMSE by 3.6–18.2% and MAE by 3.1–31.1%; mean reductions are 10.2% and 15.0%. Conservative BIC differences of 128.8–368.1 favour the stage-aware representation"
Both Rlin(a⋆) and Rb(α,β,γ,u1,u2) are fitted to the identical within-run zb; u1,u2 are taken from the HI stage procedure and BIC still assigns them as effective parameters. The reported ‘reductions’ and ΔBIC are in-sample descriptive superiority of a higher-capacity, stage-aligned parameterization over a one-parameter anchored line on a reference the authors built—not out-of-sample prediction of an independent label. The manuscript mostly labels this correctly as descriptive target-shape assessment, but the abstract’s ‘better describe … degradation states’ framing still presents fitted adequacy on a self-defined surrogate as the headline scientific result.
full rationale
The paper’s strongest claim is descriptive: on three IMS runs, a continuous linear–quadratic–exponential curve fits the same vibration-HI-derived reference zb better than the best anchored linear fit (RMSE/MAE gains, ΔBIC). That comparison is not a pure tautology—both competitors are scored against identical zb, so a more flexible form can lose if zb is nearly linear, and the authors report blocked-bootstrap intervals that cross zero and limit the claim to descriptive target shape. Circularity is only partial: zb is defined as 1 minus min–max-normalized oriented PCA of the paper’s own screened features; chronological three-stage labels are obtained by k-means on that same HI; Rb is then least-squares-fitted to zb with transitions u1,u2 taken from those stages. Thus ‘stage-aware better describes the vibration-derived degradation state’ reduces largely to ‘a three-piece, stage-tied fit to our HI surrogate beats a one-parameter line on that same surrogate.’ RQ2 (XJTU held-out learnability of the constructed target) and OWA fusion are ordinary supervised evaluation against a retrospective label and are not circular in the strong sense. No load-bearing self-citation, uniqueness import, or renamed theorem chain supports the central claim. Score 4 reflects one structural fitted-input/self-definitional step around RQ1 while leaving independent GOF content and honest caveats intact.
Axiom & Free-Parameter Ledger
free parameters (7)
- LQE coefficients α, β, γ (≥0) =
Per-run LS fits (not globally tabulated)
- Stage boundaries u1, u2 =
e.g. IMS1-B3: 0.074, 0.845; IMS1-B4: 0.077, 0.748; IMS2-B1: 0.584, 0.842
- Anchored linear slope a* =
Per-run (not listed numerically beyond resulting RMSE/MAE)
- OWA weights w1*, w2* =
Not numerically reported in manuscript
- Feature-screening thresholds τvar, τρ, Crob threshold; q top features =
IMS: 10 features, redundancy 0.98; others not fully specified
- HI smoothing fraction / min stage fraction =
0.035 smoothing; 0.06 min stage fraction (IMS)
- Window length W, step S, sequence length L, training hyperparameters =
Partial: Adam lr 0.001, batch 32, patience 10; W/S/L/dropout incomplete
axioms (5)
- domain assumption An oriented PCA1 health indicator from screened multi-domain vibration features is a valid monotone proxy for degradation state within a run.
- ad hoc to paper Bearing degradation is adequately summarized by exactly three chronological stages after k-means and contiguous repair.
- ad hoc to paper A continuous linear–quadratic–exponential piecewise map is an appropriate phenomenological form for the normalized remaining-life state.
- domain assumption Retrospective full-run labels may supervise causal predictors if no future windows enter model inputs and test bearings are excluded from fitted development steps.
- standard math Standard signal-processing and learning tools (Welch PSD, wavelets, PCA, k-means, CNN–LSTM, Transformer, Adam, OWA) behave as usual.
invented entities (2)
-
Continuous linear–quadratic–exponential degradation-state target Rb(u)
no independent evidence
-
Development-only stage-aware target + dual-scale OWA predictor pipeline
no independent evidence
read the original abstract
Remaining useful life (RUL) studies commonly treat the label as fixed, although clock-linear labels may decline while measured vibration remains nearly stable and then changes rapidly near failure. We separate target design from prediction. A development-only pipeline constructs an oriented vibration health indicator, identifies chronological early, middle, and late stages, and fits a continuous linear-quadratic-exponential degradation-state target. A compact CNN-LSTM and Transformer learn the target from causal feature sequences, and validation-fitted Ordered Weighted Averaging combines their outputs. In a bearing-wise XJTU-SY hold-out, all bearings ending in 5 are excluded from fitted preprocessing, training, early stopping, and fusion. The fused predictor obtains an RMSE of 0.0608, an MAE of 0.0392, and an R-squared value of 0.9617, with the Transformer providing most of the accuracy. Target shape is assessed independently on three documented IMS failed-bearing trajectories. Against the best anchored linear fit to the same vibration-derived reference, the stage-aware curve reduces RMSE by 3.6-18.2% and MAE by 3.1-31.1%; the mean reductions are 10.2% and 15.0%, respectively. Conservative BIC differences of 128.8-368.1 favor the stage-aware representation, whereas moving-block bootstrap intervals cross zero. Thus, stage-dependent targets better describe the evaluated vibration-derived degradation states, but the evidence remains descriptive because only three official IMS runs are available. The study establishes a measurement-oriented target-validity framework, not a universal nonlinear law for physical time-to-failure or robust cross-domain prediction.
Figures
Reference graph
Works this paper leans on
-
[1]
Conformal prediction: A gentle introduction,
A. N. Angelopoulos and S. Bates, “Conformal prediction: A gentle introduction,”Foundations and Trends in Machine Learning, vol. 16, no. 4, pp. 494–591, 2023, doi: 10.1561/2200000101
-
[2]
A. Ayman, A. Onsy, O. Attallah, H. Brooks, and I. Morsi, “Feature learning for bearing prognos- tics: A comprehensive review of machine/deep learning methods, challenges, and opportunities,” Measurement, vol. 245, art. 116589, 2025, doi: 10.1016/j.measurement.2024.116589
arXiv 2025
-
[3]
A. Bott, B. Liu, L. Nuding, J. Wachsmuth, A. Puchta, and J. Fleischer, “Uncertainty-aware prognostics of ball bearings using physics-based simulation and conditional normalizing flows,” IEEE Access, vol. 14, pp. 20100–20110, 2026, doi: 10.1109/ACCESS.2026.3661174
arXiv 2026
-
[4]
N. Gebraeel, Y. Lei, N. Li, X. Si, and E. Zio, “Prognostics and remaining useful life prediction of machinery: Advances, opportunities and challenges,”Journal of Dynamics, Monitoring and Diagnostics, vol. 2, no. 1, pp. 1–12, 2023, doi: 10.37965/jdmd.2023.148
-
[5]
J. Guo, Z. Wang, H. Li, Y. Yang, C.-G. Huang, M. Yazdi, and H. S. Kang, “A hybrid prognosis scheme for rolling bearings based on a novel health indicator and nonlinear Wiener process,”Reli- ability Engineering & System Safety, vol. 245, art. 110014, 2024, doi: 10.1016/j.ress.2024.110014. 26
arXiv 2024
-
[6]
Bearing Data Set,
J. Lee, H. Qiu, G. Yu, J. Lin, and Rexnord Technical Services, “Bearing Data Set,” IMS, University of Cincinnati, NASA Prognostics Data Repository, NASA Ames Research Center, Moffett Field, CA, USA, 2007. [Online]. Available: NASA PCoE Data Repository. Accessed: Jul. 21, 2026
2007
-
[7]
XJTU-SY rolling element bearing accelerated life test datasets: A tutorial,
Y. Lei, T. Han, B. Wang, N. P. Li, T. Yan, and J. Yang, “XJTU-SY rolling element bearing accelerated life test datasets: A tutorial,”Journal of Mechanical Engineering, vol. 55, no. 16, pp. 1–6, 2019, doi: 10.3901/JME.2019.16.001
-
[8]
H. Li, Z. Zhang, T. Li, and X. Si, “A review on physics-informed data-driven remaining useful life prediction: Challenges and opportunities,”Mechanical Systems and Signal Processing, vol. 209, art. 111120, 2024, doi: 10.1016/j.ymssp.2024.111120
arXiv 2024
-
[9]
X. Li, W. Teng, Y. Zhang, D. Peng, and Y. Liu, “Dynamic normalized health indicator construction and Bayesian recurrent state estimation for remaining useful life prediction of high-speed bearings in wind turbine drivetrain,”Measurement, vol. 246, art. 116725, 2025, doi: 10.1016/j.measurement.2025.116725
arXiv 2025
-
[10]
C. M. Lillelund, F. Pannullo, M. O. Jakobsen, M. Morante, and C. F. Pedersen, “RULSurv: A probabilistic survival-based method for early censoring-aware prediction of remaining useful life in ball bearings,” arXiv:2405.01614v3 [cs.LG], 2025, doi: 10.48550/arXiv.2405.01614
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2405.01614 2025
-
[11]
B. Mousaei Shir-Mohammad, B. Moshiri, and A. Yaghmaei, “Topology-aware hybrid Wi- Fi/BLE fingerprinting via evidence-theoretic fusion and persistent homology,” arXiv preprint arXiv:2510.16557, 2025, doi: 10.48550/arXiv.2510.16557
-
[12]
Q. Ni, J. C. Ji, and K. Feng, “Data-driven prognostic scheme for bearings based on a novel health indicator and gated recurrent unit network,”IEEE Transactions on Industrial Informatics, vol. 19, no. 2, pp. 1301–1311, 2023, doi: 10.1109/TII.2022.3169465
arXiv 2023
-
[13]
H. Peng, B. Jiang, Z. Mao, and S. Liu, “Local enhancing Transformer with temporal con- volutional attention mechanism for bearings remaining useful life prediction,”IEEE Trans- actions on Instrumentation and Measurement, vol. 72, art. 3522312, pp. 1–12, 2023, doi: 10.1109/TIM.2023.3291787
arXiv 2023
-
[14]
H. Qiu, Y. Niu, J. Shang, L. Gao, and D. Xu, “A piecewise method for bearing remaining useful life estimation using temporal convolutional networks,”Journal of Manufacturing Systems, vol. 68, pp. 227–241, 2023, doi: 10.1016/j.jmsy.2023.04.002
-
[15]
H. Qiu, J. Lee, J. Lin, and G. Yu, “Wavelet filter-based weak signature detection method and its application on rolling element bearing prognostics,”Journal of Sound and Vibration, vol. 289, no. 4–5, pp. 1066–1090, 2006, doi: 10.1016/j.jsv.2005.03.007
-
[16]
L. Shuang, X. Shen, J. Zhou, H. Miao, Y. Qiao, and G. Lei, “Bearings remaining useful life prediction across equipment-operating conditions based on multisource-multitarget domain adaptation,”Measurement, vol. 236, art. 115026, 2024, doi: 10.1016/j.measurement.2024.115026
arXiv 2024
-
[17]
N. Sun, J. Tang, X. Ye, C. Zhang, S. Zhu, S. Wang, and Y. Sun, “Remaining useful life prognostics of bearings based on convolution attention networks and enhanced transformer,” Heliyon, vol. 10, no. 19, art. e38317, 2024, doi: 10.1016/j.heliyon.2024.e38317. 27
-
[18]
Y. Tang, R. Liu, C. Li, and N. Lei, “Remaining useful life prediction of rolling bearings based on time convolutional network and Transformer in parallel,”Measurement Science and Technology, vol. 35, no. 12, art. 126102, 2024, doi: 10.1088/1361-6501/ad73ee
-
[19]
Knowledge informed machine learning using a Weibull- based loss function,
T. von Hahn and C. K. Mechefske, “Knowledge informed machine learning using a Weibull- based loss function,”Journal of Prognostics and Health Management, vol. 2, no. 1, pp. 9–44, 2022, doi: 10.22215/jphm.v2i1.3162
-
[20]
A hybrid prognostics approach for estimating remaining useful life of rolling element bearings,
B. Wang, Y. Lei, N. P. Li, and N. B. Li, “A hybrid prognostics approach for estimating remaining useful life of rolling element bearings,”IEEE Transactions on Reliability, vol. 69, no. 1, pp. 401–412, 2020, doi: 10.1109/TR.2018.2882682
arXiv 2020
-
[21]
Z. Wang, Y. Ta, W. Cai, and Y. Li, “Research on a remaining useful life prediction method for degradation angle identification two-stage degradation process,”Mechanical Systems and Signal Processing, vol. 184, art. 109747, 2023, doi: 10.1016/j.ymssp.2022.109747
arXiv 2023
-
[22]
Z. Xu, C. W. K. Chow, M. M. Rahman, R. Rameezdeen, and Y. W. Law, “Remaining useful life prediction across conditions based on a health indicator-weighted subdomain alignment network,”Sensors, vol. 25, no. 15, art. 4536, 2025, doi: 10.3390/s25154536
-
[23]
On ordered weighted averaging aggregation operators in multicriteria decision- making,
R. R. Yager, “On ordered weighted averaging aggregation operators in multicriteria decision- making,”IEEE Transactions on Systems, Man, and Cybernetics, vol. 18, no. 1, pp. 183–190, 1988, doi: 10.1109/21.87068
doi:10.1109/21.87068 1988
-
[24]
Y. Zhang, X. Zhao, Z. Peng, R. Xu, and Y. Hui, “Cross-domain remaining useful life prediction for rolling bearings based on wavelet decomposition and dynamic calibrated domain adaptive networks,”Measurement, vol. 251, art. 117278, 2025, doi: 10.1016/j.measurement.2025.117278
arXiv 2025
-
[25]
H. Zhou, X. Huang, G. Wen, Z. Lei, S. Dong, P. Zhang, and X. Chen, “Construction of health indicators for condition monitoring of rotating machinery: A review of the research,”Expert Systems with Applications, vol. 203, art. 117297, 2022, doi: 10.1016/j.eswa.2022.117297. 28
arXiv 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.