{"id":"ad6dc9be-5206-4cf1-9642-e716eb7000aa","arxiv_id":"2411.15185","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A hybrid LSTM and Gaussian process model produces RUL point and interval forecasts on C-MAPSS with narrower intervals than baselines, but with inconsistent coverage across sub-datasets.","lead":"The paper describes a hybrid machine-learning model that forecasts the remaining useful life of aircraft engines and gives a confidence interval for every forecast. It matters because engineering maintenance decisions need both a point estimate and a measure of uncertainty, and the authors report large reductions in interval width on the standard C-MAPSS benchmark.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Under the paper's own CWC definition, the FD004 entry (NAW=28, CWC=468.43) is impossible: CWC ≤ e·NAW ≈ 76.1, so the central interval-improvement claim is unsupported.","rationale":"The reader's verdict is REJECT, and I agree that the central interval-prediction claim is unsupported. The reader identified the FD004 CWC of 468.43 as evidence that the latent GP variance is miscalibrated and that the claim of consistent improvement fails. My stress-test found a more specific and decisive problem: under the paper's own Eq. (13), the reported pair (NAW=28, CWC=468.43) is mathematically impossible because the exponential factor is bounded above by e. This is an internal arithmetic contradiction, not a matter of comparing against external baselines or consensus. It means the evaluation metrics as reported cannot be trusted, and the headline statement that the method 'consistently reduces NAW values' and 'also significantly improve[s]' CWC is not supported by the paper's own data. No further assumptions about calibration or latent-space smoothness are needed to see that the central quantitative evidence fails. Since this only reinforces the reader's REJECT verdict rather than changing it, I set verdict_should_be to UNCHANGED and mark agreement as partial: the reader pointed at the FD004 CWC as a symptom of miscalibration, while I identify it as an impossibility under the stated metric definition.","tokens_in":14233,"tokens_out":6340,"duration_ms":69808,"concrete_test":"Recompute the FD004 CWC from the raw test-set prediction intervals using the formula as printed in Eq. (13). If any computed value exceeds e×NAW (specifically, if FD004 CWC is 468.43 with NAW=28), then the printed formula and the table cannot both be correct; obtain the corrected metric or the actual coverage probability to see whether the claimed 'consistent outperformance' survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that HRP 'consistently' reduces NAW and improves CWC across all four C-MAPSS subsets. Table 2 reports FD004 with NAW=28.00 and CWC=468.43. But Eq. (13) defines CWC = NAW × exp(1 − coverage/α). Since coverage is a probability in [0,1] and α>0, the exponential factor is at most e ≈ 2.718, so CWC ≤ e·NAW for any possible coverage value. For FD004 this gives CWC ≤ e·28 ≈ 76.1, yet the table reports 468.43, which is more than six times the maximum possible under the stated formula. This is not merely a sign of miscalibrated intervals; it means the paper's quantitative evidence for interval quality is internally inconsistent. Either Eq. (13) is not the metric that produced Table 2, or the FD004 CWC entry is erroneous. Because the most general claim ('Out of the four sub-datasets, our method consistently outperforms existing approaches') rests directly on these numbers, the reported evaluation cannot establish the claimed interval performance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes HRP, a hybrid method that feeds LSTM hidden states into a Gaussian process regressor to produce point and interval predictions of remaining useful life (RUL) on the C-MAPSS benchmark. The paper claims that HRP consistently narrows normalized average width (NAW) and improves the coverage width criterion (CWC) relative to several baselines across all four sub-datasets, while providing feature importance analysis. The evaluation reports RMSE, NAW, and CWC for FD001–FD004.","tokens_in":14520,"tokens_out":5321,"duration_ms":46684,"significance":"If the central claims were supported, the paper would offer a practical hybrid of deep temporal feature extraction and Bayesian nonparametric uncertainty quantification, with an interpretability layer, for aeroengine prognostics. The idea of using LSTM hidden states as GP inputs is a reasonable direction. However, the paper's own numerical evidence contains an internal contradiction that invalidates the reported interval comparisons, and the lack of repeated-run statistics further weakens the empirical claims.","major_comments":[{"comment":"The FD004 row of Table 2 reports NAW = 28.00 and CWC = 468.43. Under the stated definition CWC = NAW × exp(1 − coverage/α), and since coverage is a probability in [0,1] and α > 0, the exponential factor is at most e, so CWC ≤ e·NAW ≈ 76.1. The reported value 468.43 exceeds this bound by a factor of six. Either Eq. (13) is not the metric used to produce Table 2, or the FD004 CWC entry is erroneous. Because the claim that HRP 'consistently outperforms existing approaches' on all four sub-datasets rests directly on these numbers, the reported evaluation does not currently establish the paper's central interval-performance claim.","section":"Table 2 and Eq. (13)"},{"comment":"Even setting aside the internal inconsistency, the FD004 CWC of 468.43 is far worse than every baseline, including AGCNN's 95.56. The text acknowledges improvement only 'on the first three sub-datasets' for CWC but then concludes 'Out of the four sub-datasets, our method consistently outperforms existing approaches.' This is a direct contradiction of the table. The authors must either correct the values or revise the claim.","section":"Comparison with state-of-the-art methods (Table 2)"},{"comment":"The sliding-window sizes are reported together with 'Min life cycle in test datasets' and 'Number of testing time windows' for each sub-dataset. If these test-set statistics were used to choose the window lengths, the evaluation is compromised by information leakage. The manuscript does not explain how the window sizes were selected or justify their dependence on test-set properties. The authors should describe the selection procedure and confirm that the test data were not used for hyperparameter choice.","section":"Table 6 and 'Evaluation setting'"},{"comment":"The posterior mean in Eq. (7) contains µ(h) and µ(h*), but the prediction equation (9) is written as K(h*)⊺(K + δ²I)^{-1}y, which corresponds to a zero-mean GP. The manuscript states 'assuming a GP(0, K) prior' immediately after presenting the general mean formulation, but then the narrative continues to refer to µ. This inconsistency needs to be resolved, and the GP hyperparameters η, τ, and δ should be reported with their fitting procedure for reproducibility.","section":"Eqs. (7)–(9)"},{"comment":"The paper reports only NAW and CWC, not the coverage probability itself. A narrow interval that misses the true RUL most of the time yields a low NAW but is not a valid 95% interval. The claim that 'the predicted intervals consistently cover the real RUL' is supported only by selected example engines (Fig. 4), not by quantitative coverage on the full test sets. The authors should report coverage probabilities alongside NAW and CWC for all four datasets.","section":"Prognostic results analysis"}],"minor_comments":[{"comment":"The coverage indicator is written as Cj = 1 if RULj ∈ [yU_j, yL_j]; the interval endpoints should be reversed to [yL_j, yU_j].","section":"Eq. (13)"},{"comment":"The importance analysis section describes permutation-based feature importance, but the Conclusion states that 'by adaptively employing additional random forest regression' the influence of sensors is assessed. These two accounts should be reconciled.","section":"Importance analysis and Conclusion"},{"comment":"There is a typo 'standard deviation deviation' in the normalization paragraph.","section":"Data preprocessing"},{"comment":"The Huber loss definition uses 'delta' in the second case; this should be the same symbol δ used elsewhere.","section":"Eq. (2)"},{"comment":"References [16] and [17] are the same paper by Hochreiter and Schmidhuber; one should be removed or replaced.","section":"References"},{"comment":"The caption refers to red and green segments indicating over- and under-prediction, but these colors are not shown or explained in the figure itself.","section":"Figure 4"}],"recommendation":"reject","confidential_remarks":"The impossible CWC value in Table 2 is a serious red flag; even if it is a typo, the authors need to provide the underlying data and corrected evaluation. The manuscript also lacks reproducibility details (no code, no seed/run statistics) and contains inconsistencies that make it unsuitable for publication in its current form. Given the journal's standards, I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: the central interval-prediction claim collapses on arithmetic. Table 2 reports FD004 with NAW=28.00 and CWC=468.43, but Eq. (13) caps CWC at e·NAW ≈ 76.1 for any coverage probability. That is not a small inconsistency; it means the reported evidence for the paper's main claim is internally invalid. The text's assertion that HRP “consistently outperforms existing approaches” on interval quality is simply false for FD004, and the claimed average improvements rest on the same broken numbers.\n\nTo give credit where it is due: the pipeline is clearly described, the idea of feeding LSTM hidden states into a GP is coherent (though not novel — that is deep kernel learning, which the paper does not cite or compare against), and the FD002 RMSE of 12.33 is the one number that looks competitive with the cited baselines. The feature importance analysis via permutation is a reasonable addition and is properly explained. But these positives do not compensate for a load-bearing evaluation error.\n\nThe soft spots are severe. Beyond the impossible CWC, there are no error bars or multiple seeds, so we cannot tell whether the FD002 result is real or noise. The sliding-window sizes in Table 6 are presented without a criterion for selection; they may have been tuned on the test set. The “modified GPR” is just standard GPR with LSTM-extracted features; the modification is not mathematically new. And the paper's claim of “consistent improvement” is contradicted by its own Table 2, not just by a rival method.\n\nIn short, this is a paper with a standard architecture, one interesting point-forecast result, and an interval evaluation that is demonstrably wrong. The internal contradiction with Eq. (13) is enough to reject without needing to go further. If the authors fix the CWC entry and provide a proper multi-seed evaluation with uncertainty, the core idea might be worth a second look. As it stands, I would not send it to peer review — the arithmetic failure is a desk-reject signal, and the novelty is too thin to justify referee time.","headline":"The paper's central interval-prediction claims collapse on arithmetic: the FD004 CWC value in Table 2 is impossible under the paper's own formula, so the core evaluation is unsupported.","tokens_in":15087,"tokens_out":2722,"would_cite":false,"duration_ms":31079,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hybrid LSTM-GPR model narrows aeroengine RUL intervals by more than 50% while keeping point forecasts competitive.","keywords":["remaining useful life prediction","Gaussian process regression","prediction intervals","LSTM temporal feature extraction","aeroengine prognostics","C-MAPSS dataset","uncertainty quantification","feature importance"],"falsifier":"Compute the empirical coverage of the HRP 95% intervals on the FD004 test engines, separately for each of the six operating conditions, and compare it with the nominal 95%; if coverage falls well below 95% in any condition while NAW stays at 28%, the narrow intervals are a symptom of miscalibration rather than calibrated uncertainty.","tokens_in":14000,"feed_emoji":"🛩️","tokens_out":8870,"duration_ms":87340,"temperature":0.7,"pith_summary":"This paper argues that replacing raw sensor windows with LSTM-learned temporal features as inputs to a Gaussian process regressor (GPR) gives aeroengine remaining-useful-life (RUL) predictions that stay competitive with deep learning while adding 95% confidence intervals that are far narrower than those of existing interval methods. On the four C-MAPSS sub-datasets the paper reports normalized averaged width (NAW) reductions of at least 50.70% on average and improved coverage-width criterion (CWC) values on the first three sub-datasets. The model also ranks sensor contributions by permuting each feature, giving engineers a partial interpretability layer for maintenance planning. The central claim is that a GPR operating on a compressed temporal latent space can quantify degradation uncertainty in a structured way.","feed_headline":"RUL intervals cut by half with hybrid Gaussian process","feed_subtitle":"LSTM temporal features plus Gaussian process regression tighten 95% RUL intervals while keeping point forecasts competitive.","key_machinery":"The load-bearing object is the pair formed by LSTM gating and the squared-exponential Gaussian process posterior. The LSTM equations define forget, input, and output gates that compress each sliding window into a hidden state $h$; the GP then places a multivariate normal prior on RUL values with kernel $k(h,h')$ and noise variance $\\delta^2$, and the posterior mean and variance generate the point forecast and the interval endpoints. A permutation-based feature-importance component measures the drop in prediction accuracy when each sensor is shuffled, producing the $\\lambda$ ranking. The mechanism matters because it lets a Bayesian non-parametric regressor work in a compact temporal latent space rather than on high-dimensional raw sensor streams, which is what the paper credits for both narrow intervals and partial interpretability.","core_discovery":"On its own terms, the paper's central claim is that the hybrid HRP model learns a probabilistic mapping from multivariate sensor histories to RUL such that point predictions match or beat deep learners and the 95% prediction interval is tighter than existing interval-prediction baselines. The LSTM compresses a sliding window of 14 selected sensors into hidden states $h$; the GP prior is $y \\sim N(\\mu, K + \\delta^2 I)$ with squared-exponential kernel $k(h,h') = \\tau^2 \\exp(-\\|h-h'\\|^2/(2\\eta^2))$, and the predictive interval for a new test point is the posterior mean plus or minus $1.96$ times the square root of the posterior variance. The paper reports RMSE values of 13.09, 12.33, 13.49, and 19.65 on FD001-FD004, with the FD002 value the best among all compared models, and NAW values of 21%, 29%, 23%, and 28%, which it interprets as a consistent interval-narrowing advantage. CWC improvements are claimed on FD001, FD002, and FD003.","pith_inferences":["A direct consequence we draw from Table 2 is that on FD004, where six operating conditions and two fault modes combine, the narrow intervals are not backed by coverage: the CWC of 468.43, versus 95.56 for AGCNN, indicates miscalibration far more than good uncertainty quantification.","The paper never reports empirical coverage probabilities separately from CWC; reporting coverage per sub-dataset and per operating condition would settle whether the NAW reductions come from calibrated narrowness or overconfidence.","Since the LSTM is trained on a Huber pointwise loss, no objective ties the GP variance to observed residuals; adding a calibration or maximum-likelihood term on the GP's predictive distribution is a natural, testable extension that could repair the FD004 case.","A practical extension consistent with the paper's architecture is post-hoc recalibration of the GP variance, for instance by scaling the posterior standard deviation, before constructing intervals; this would preserve the point forecasts while restoring coverage."],"forward_implications":["On FD001-FD003, the reported NAW values of 21-29% would give maintenance planners considerably tighter 95% failure windows than the compared LSTM-BS, IESGP, and AGCNN baselines.","The FD002 RMSE of 12.33, the best among compared models, suggests that a GP head on LSTM features can improve point accuracy under multiple operating conditions while also providing uncertainty.","Permutation-based feature rankings give a concrete sensor priority list, notably sensors 6, 10, and 13, that can be used for condition monitoring and maintenance focus.","Because the GP posterior updates with new data points, the method implies a natural online loop: as more run-to-failure history accumulates, the interval for a given engine should narrow."],"supporting_citations":[{"why":"supplies the C-MAPSS turbofan simulator dataset and its four sub-datasets used in all experiments.","marker":"[14]"},{"why":"provides the LSTM gating equations that the temporal feature extractor is based on.","marker":"[16]"},{"why":"motivates the sliding-window segmentation used to form training and test inputs.","marker":"[19]"},{"why":"is the LSTM-BS baseline whose interval width and CWC are compared in Table 2.","marker":"[24]"},{"why":"is the IESGP baseline whose interval width and CWC are compared in Table 2.","marker":"[25]"},{"why":"is the AGCNN baseline whose interval width and CWC are compared in Table 2.","marker":"[27]"},{"why":"supplies the Gaussian process regression formulation and posterior equations the modified GPR adapts.","marker":"[40]"},{"why":"provides the piecewise linear degradation model and exponential smoothing used in preprocessing.","marker":"[41]"}],"fun_headline_variants":["Hybrid LSTM-GP model tightens RUL prediction intervals","Reliable aeroengine RUL intervals via hybrid Gaussian process","Interpretable RUL interval prediction with LSTM-GP fusion","GP-LSTM hybrid narrows uncertainty in RUL forecasts","Partially interpretable RUL intervals from temporal GP"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the hidden states of an LSTM trained only on a Huber pointwise RUL loss form a space in which the Gaussian process posterior variance is a well-calibrated 95% interval across all operating conditions and fault modes.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid LSTM-GP model tightens RUL prediction intervals","Reliable aeroengine RUL intervals via hybrid Gaussian process","Interpretable RUL interval prediction with LSTM-GP fusion","GP-LSTM hybrid narrows uncertainty in RUL forecasts","Partially interpretable RUL intervals from temporal GP"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000324,"raw_usage":{"total_tokens":1818,"prompt_tokens":946,"completion_tokens":872,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":788}},"tokens_in":562,"tokens_out":872,"duration_ms":9664,"temperature":1.0,"reasoning_tokens":788,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:49:53.825182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the empirical coverage of the HRP 95% intervals on the FD004 test engines, separately for each of the six operating conditions, and compare it with the nominal 95%; if coverage falls well below 95% in any condition while NAW stays at 28%, the narrow intervals are a symptom of miscalibration rather than calibrated uncertainty.","supporting_citations":[{"cited_title":"User’s guide for the commercial modular aero-propulsion system simulation (c-mapss)","cited_arxiv_id":null,"evidence_quote":"supplies the C-MAPSS turbofan simulator dataset and its four sub-datasets used in all experiments."},{"cited_title":"Time series data prediction using sliding window based rbf neural network","cited_arxiv_id":null,"evidence_quote":"motivates the sliding-window segmentation used to form training and test inputs."},{"cited_title":"Uncertainty prediction of remaining useful life using long short-term memory network based on bootstrap method","cited_arxiv_id":null,"evidence_quote":"is the LSTM-BS baseline whose interval width and CWC are compared in Table 2."},{"cited_title":"Multiple sensors based prognostics with prediction interval optimization via echo state gaussian process","cited_arxiv_id":null,"evidence_quote":"is the IESGP baseline whose interval width and CWC are compared in Table 2."},{"cited_title":"Uncertainty quantification and interval prediction of equipment remaining useful life based on semi-supervised learning","cited_arxiv_id":null,"evidence_quote":"is the AGCNN baseline whose interval width and CWC are compared in Table 2."},{"cited_title":"A tutorial on gaussian process regression: Modelling, exploring, and exploiting functions","cited_arxiv_id":null,"evidence_quote":"supplies the Gaussian process regression formulation and posterior equations the modified GPR adapts."},{"cited_title":"A dual attention lstm lightweight model based on exponential smoothing for remaining useful life prediction","cited_arxiv_id":null,"evidence_quote":"provides the piecewise linear degradation model and exponential smoothing used in preprocessing."}],"review_version":1}