{"id":"b7383d86-ddfe-479a-90ad-c35fd2579dcd","arxiv_id":"2504.12156","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Near-optimal survival models can give conflicting failure-risk estimates for the same equipment, and the proposed ambiguity, discrepancy, and obscurity metrics quantify this on CMAPSS engine data.","lead":"This paper applies the idea of predictive multiplicity, where equally accurate models make different predictions, to survival models used in maintenance. It defines three measures of that disagreement and shows, on aircraft engine data, that several good models can give very different failure-risk estimates for the same machine.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The multiplicity metrics are evaluated at each unit's observed event/censoring time, so the headline conflict may reflect realized failure times rather than disagreement about fixed-horizon risks.","rationale":"I read the abstract's strongest claim as existential: multiple accurate survival models may yield conflicting risk estimates, and this matters for maintenance. The Rashomon-set representativeness concern raised by the reader is real but not decisive for an existential claim, because RSFs are survival models, so one family with multiple accurate members can demonstrate 'may'. The evaluation-time issue is more load-bearing because it questions whether the demonstrated conflict is the kind that matters for maintenance decisions. Equations (8)-(10) explicitly use t_i, and Section 4.2 fixes the censoring time, making t_i for censored units a constant and for failures their realized event time. At the realized failure time, the CDF is a probability integral transform uniform under a calibrated model, so any two calibrated models can differ substantially in f(x_i,t_i) merely because t_i is stochastic. This is not a purely theoretical point: it changes the quantity being measured. The concrete test is inexpensive and directly settles it. I therefore keep the reader's CONDITIONAL verdict: the formal apparatus may be sound, but the empirical demonstration needs this analysis before the maintenance-relevant conclusion can be accepted.","tokens_in":13247,"tokens_out":8060,"duration_ms":91014,"concrete_test":"Recompute A, D, and O on the same Rashomon sets for the four CMAPSS subsets using f(x_i,h) for two fixed horizons, e.g., h=100 and h=200 cycles, keeping the same epsilons, deltas, and test split. If the metric values collapse or change sign of the trend relative to Table 3 and Figure 2, the central claim must be restricted to outcome-dependent evaluation; if values remain comparable, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that accurate survival models disagree about failure risk for the same equipment. The formal metrics (Eqs. 8-10) compare f(x_i,t_i) where t_i is the observation's actual event or censoring time, not a fixed prediction horizon. With censoring time fixed at 250 (Section 4.2), t_i=250 for every censored unit and t_i equals the realized failure time for uncensored units. For an uncensored unit, f(x_i,t_i)=Pr(T<=t_i|x_i) is the CDF evaluated at the realized failure time; under a well-calibrated model, these values are marginally uniform and will vary even between equally accurate models because the evaluation point is the random outcome, not a decision target. A maintenance planner instead needs Pr(T<=h|x) for a chosen horizon h (e.g., failure within the next 100 cycles). The reported ambiguity/discrepancy/obscurity therefore establish conflict about an outcome-dependent quantity, not about failure risk within a fixed maintenance window. The 'degradation progression' half of the conclusion is even less supported: no metric evaluates the full survival curve, only point estimates at t_i. If fixed-horizon recomputation eliminates the conflict, the central empirical claim is an artifact of evaluation times; if it persists, the claim is robust.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the notion of predictive multiplicity to survival analysis, defining three metrics—ambiguity, discrepancy, and obscurity—over a Rashomon set of near-optimal survival models. The methodology is applied to the CMAPSS predictive-maintenance benchmark using Random Survival Forests with a large hyperparameter grid, and the metrics are computed at various Rashomon tolerances epsilon and conflict thresholds delta. The authors report that ambiguity and discrepancy increase with epsilon and reach values near or equal to 1 across multiple datasets, and they conclude that multiple accurate survival models can yield conflicting failure-risk estimates and degradation predictions for the same equipment.","tokens_in":13506,"tokens_out":4360,"duration_ms":40348,"significance":"If the empirical results were robust, this would be a valuable contribution: it formalizes predictive multiplicity for survival data, addresses a real need in maintenance decision-making, and provides an open code repository. The formal definitions in Section 3 are mathematically coherent, and the exploration of a 22,500-configuration model grid is a strength. However, the central empirical claims are undercut by several load-bearing issues: the monotonic increase of ambiguity and discrepancy with epsilon is a tautological consequence of the max-based definitions, the metrics are evaluated at each unit's observed event/censoring time rather than at a fixed decision horizon, the abstract's headline 40-45% figure is inconsistent with Table 3 values that reach 1.0, and no uncertainty quantification (error bars, repeated seeds) is provided for the reported numbers. These issues prevent the paper from currently supporting its general conclusions, though they are addressable within the manuscript's scope.","major_comments":[{"comment":"The paper presents the increase of ambiguity and discrepancy with epsilon as an empirical finding, but this monotonicity is a direct mathematical consequence of the definitions: both metrics take a maximum (or a ratio derived from a maximum) over the Rashomon set H_epsilon, and enlarging the set cannot decrease the maximum. The text in Section 5 ('Overall, model uncertainty is limited when epsilon is small but increases markedly as epsilon grows') is therefore circular. Please reframe the monotonicity as an axiomatic property of the definitions, or recompute the metrics on fixed model sets of controlled size to separate the effect of set expansion from the multiplicity phenomenon.","section":"Section 3.3, Eqs. (8)-(9); Section 5"},{"comment":"The abstract states that ambiguity 'reaching up to 40-45% of observations', but Table 3 reports ambiguity values of 1.0 at several (epsilon, delta) combinations across all four datasets (e.g., epsilon=0.05, delta=0.01 and epsilon=0.10, delta=0.01). The 40-45% figure does not match any value in the table as far as can be determined from the text. Please either correct the abstract to reflect the actual range of the metrics or identify explicitly the specific configuration to which the 40-45% figure refers.","section":"Abstract; Table 3"},{"comment":"The metrics evaluate f(x_i, t_i) at each observation's realized event or censoring time. With censoring time fixed at 250 cycles, t_i=250 for all censored units and t_i equals the realized failure time for uncensored units. Under a well-calibrated model, the values f(x_i, t_i) for uncensored units are marginally uniform, so the reported disagreement may reflect randomness in the realized failure times rather than genuine conflict about failure risk within a fixed maintenance horizon (e.g., Pr(T <= h | x) for a chosen h). The conclusion that models 'yield conflicting estimations of failure risk and degradation progression' is not supported because no metric evaluates the full survival curve or a prespecified decision horizon. Please recompute the three metrics at one or more fixed horizons (e.g., h=50, 100, 150 cycles) and, ideally, also report an integral metric over the survival curve.","section":"Section 3.3, Eqs. (8)-(10); Section 4.2"},{"comment":"The censoring time is fixed at t=250 'to eliminate the censoring sensitivity', but no sensitivity analysis is provided. This is particularly concerning because the authors themselves cite Yardimci and Cavus (2025) reporting that censoring time significantly affects prediction uncertainty in this setting. The limitation statement in Section 6 acknowledges dependence on the performance metric, and the experiments use only the Brier score. Please add sensitivity analyses that vary the censoring time and use an alternative performance metric (e.g., concordance index) to assess whether the reported multiplicity values and trends are robust to these choices.","section":"Section 4.2; Section 6"},{"comment":"No error bars, confidence intervals, or repeated runs are reported. Random Survival Forests are randomized (bootstrap sampling, random split points), and the paper trains 22,500 configurations without specifying seeds. The reported values, especially the exact numbers in Table 3 (e.g., 0.8875, 0.9028), may be unstable across runs. Please repeat the full pipeline over multiple random seeds and report means and standard deviations (or confidence intervals) for the multiplicity metrics.","section":"Section 5, Table 3"}],"minor_comments":[{"comment":"The sentence 'a commonly accepted framework for uncertainty quantification in remains elusive' is missing a word; it should read 'in predictive maintenance' or similar.","section":"Section 2.3"},{"comment":"In the text following Eq. (9), 'as a result of replacing f0 with another model' refers to the reference model; please use fR consistently instead of f0.","section":"Section 3.3.2"},{"comment":"The sentence 'Because it is recognized as the benchmark dataset' is a sentence fragment; it should be integrated into the previous sentence or completed.","section":"Section 4.1"},{"comment":"The phrase 'δ sets the threshold for how much a model's prediction must differ from the reference to be considered conflicting' appears mid-caption without clear punctuation; please integrate it grammatically.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's scope fits the journal, but the circularity of the headline monotonicity result and the mismatch between the abstract and Table 3 are serious concerns that the authors must address. The fixed-time evaluation issue is also fundamental to the paper's practical relevance for maintenance scheduling. I would not recommend rejection, as the proposed metrics and the general framing could be useful after a substantial revision that adds fixed-horizon evaluations, sensitivity analyses, and proper uncertainty quantification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the honest take: the formal contribution is real and worth a referee's time. The paper takes ambiguity, discrepancy, and obscurity from classification and defines them for survival model outputs. The definitions are mathematically coherent, and measuring disagreement over a Rashomon set of survival models is a sensible extension, especially for maintenance applications. The authors also ship a reproducibility repo.\n\nThe problems are concentrated in the experiments and the interpretation. First, the metrics in Eqs. (8)-(10) evaluate |f(x_i,t_i) - f_R(x_i,t_i)| at the observed event or censoring time t_i, not at a fixed horizon. For an uncensored unit, t_i is a random draw from the true failure time distribution; under a well-calibrated model, F(t_i|x_i) is marginally uniform, so equally good models can disagree substantially at t_i even if they agree on, say, Pr(T<=200|x). The claim about \"conflicting failure risk\" for the same equipment is therefore weaker than advertised. A maintenance planner wants a fixed-horizon risk, and the paper doesn't show the disagreement survives that re-evaluation. The \"degradation progression\" half of the conclusion is even less supported -- no metric looks at the whole survival curve.\n\nSecond, the abstract's headline number (40-45%) does not match Table 3, where ambiguity and discrepancy reach 1.0 for several datasets and thresholds. That's a contradiction a careful reader will hit immediately.\n\nThird, the monotonic increase of ambiguity and discrepancy with epsilon is largely definitional, since Eqs. (8)-(9) take a maximum over a growing set. The paper presents it as an empirical finding; the authors should acknowledge this or restructure the claim.\n\nThere are also no error bars or repeated runs, and only RSF models are used, so generalizability beyond this model class is untested. Those are standard limitations the authors do acknowledge.\n\nOverall: the definitions and the framework are a legitimate contribution, and with a fixed-horizon evaluation and a corrected abstract this could be a solid application paper. As it stands, the central empirical claim is not yet established. I would send it out -- a serious referee could turn this into something usable -- but I would not take the current results at face value.","headline":"The formal adaptation of predictive multiplicity metrics to survival models is real and clean, but the experiments evaluate at each unit's realized event time rather than a fixed horizon, which undermines the headline claims.","tokens_in":14011,"tokens_out":3423,"would_cite":false,"duration_ms":33301,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62N02","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Equally accurate survival models can produce conflicting failure-risk predictions for the same equipment, and the paper introduces three metrics that quantify how often this happens.","keywords":["predictive multiplicity","survival analysis","Rashomon effect","model uncertainty","predictive maintenance","random survival forests","CMAPSS","time-to-failure"],"falsifier":"Recompute ambiguity, discrepancy, and obscurity on the same CMAPSS subsets using Cox proportional-hazards and deep survival networks for the Rashomon set instead of Random Survival Forests, keeping the same $\\epsilon$, $\\delta$, and Brier-score rule; if the metrics do not rise with $\\epsilon$, the claimed multiplicity is an artifact of the forest family rather than a general property of accurate survival models.","tokens_in":13025,"feed_emoji":"⚙️","tokens_out":7644,"duration_ms":79831,"temperature":0.7,"pith_summary":"Predictive maintenance relies on survival models to estimate when equipment will fail, but the paper argues that choosing a single best-performing survival model hides a real risk: many models with nearly identical accuracy can disagree sharply on which units are in danger. It transfers the predictive-multiplicity framework from classification to survival analysis and defines three measures—ambiguity, discrepancy, and obscurity—that count how often near-optimal survival models conflict. On the four CMAPSS aircraft-engine datasets, the paper finds that once the performance tolerance $\\epsilon$ widens, the ambiguous observations grow until nearly all observations have at least one plausible model assigning a different risk, with discrepancy slightly lower and obscurity concentrated in tight model sets. A sympathetic reading: this gives maintenance engineers a concrete way to measure whether their risk estimates are trustworthy, rather than assuming a single accuracy score settles the question.","feed_headline":"Accurate survival models can disagree on equipment failure risk","feed_subtitle":"Near-optimal aircraft-engine risk models often conflict; ambiguity, discrepancy, and obscurity measure the danger.","key_machinery":"The mechanism that carries the argument is the Rashomon set, the set of models whose performance score is within a tolerance $\\epsilon$ of the best available model, together with the conflict threshold $\\delta$ that decides when two risk estimates should be called conflicting. The survival-risk output being compared is the conditional cumulative distribution function $f(x_i,t_i)=\\Pr(T\\le t_i\\mid x_i)$, evaluated at each observation's event or censoring time. The three metrics are then summary statistics over the Rashomon set: ambiguity asks whether any plausible model changes an observation's risk, discrepancy asks how many observations the most divergent single model would flip, and obscurity averages disagreement over all plausible models. What makes the argument load-bearing is the choice of this model set and the scoring metric: differences among models are only meaningful if the models are genuinely near-optimal, so the entire measurement is conditioned on how the Rashomon set is constructed.","core_discovery":"The central discovery is that predictive multiplicity—previously defined for binary, probabilistic, and multi-target classification—also occurs in survival models, and that it is quantifiable with three adapted metrics. Given a reference survival model $f_R$, a performance metric $\\Phi$, and a Rashomon parameter $\\epsilon$, the Rashomon set $H_\\epsilon(f_R)$ collects all models whose performance is within $\\epsilon$ of the reference. For a conflict threshold $\\delta$, the paper defines ambiguity as the fraction of observations whose risk estimate $f(x_i,t_i)$ differs from $f_R(x_i,t_i)$ by at least $\\delta$ under some model in the set, discrepancy as the largest single-model conflict fraction, and obscurity as the average conflict fraction across the set. Applied to Random Survival Forests on CMAPSS, these metrics grow with $\\epsilon$ and shrink with $\\delta$, with ambiguity and discrepancy often reaching values near one; the paper takes this as direct evidence that accurate survival models can yield conflicting failure-risk and degradation estimates for the same equipment.","pith_inferences":["Beyond the paper, $\\delta$ can be calibrated to maintenance costs: if a missed failure is more expensive than a false alarm, the relevant conflict threshold is the risk difference that changes the optimal action, not a statistical convention.","The same three metrics apply verbatim to other survival outputs—predicted time-to-failure or remaining-useful-life quantiles—if $f(x_i,t_i)$ is replaced by the corresponding functional; the paper tests only cumulative risk.","Observation-level ambiguity could act as an acquisition function for active learning or sensor placement: units whose risk estimates are least stable under the Rashomon set are the ones worth labeling or inspecting first.","A testable implication is that ambiguity at tiny $\\epsilon$ measures model underspecification; datasets with high tight-set ambiguity should benefit more from additional features than from further hyperparameter search."],"forward_implications":["A maintenance team that picks a single survival model by Brier score cannot infer that the chosen risk rankings are reliable; ambiguity and discrepancy give the possible spread of failure-risk estimates.","Reporting ambiguity alongside a point risk estimate turns model uncertainty into a decision quantity: high ambiguity at a chosen $\\delta$ is a signal to inspect, add sensors, or defer maintenance actions.","Evaluating a survival model family by its multiplicity profile at fixed $\\epsilon$ is a complement to accuracy ranking; a family with lower ambiguity at the same accuracy is more decision-reliable.","The tight-set behavior of obscurity means the most dangerous disagreements can hide inside a small cluster of nearly tied models, so Rashomon-set size alone is not a safe proxy for uncertainty."],"supporting_citations":[{"why":"Supplies the original predictive-multiplicity framework and the ambiguity, discrepancy, and obscurity definitions that the paper adapts to survival risk estimates.","marker":"Marx et al. (2020)"},{"why":"Introduces the Rashomon effect, the existence of many near-optimal models, from which the paper's Rashomon set definition is taken.","marker":"Breiman (2001)"},{"why":"Provides Random Survival Forests, the only model family used to build the Rashomon sets in the experiments.","marker":"Ishwaran et al. (2008)"},{"why":"Releases the CMAPSS benchmark datasets used for all four experimental subsets.","marker":"Saxena et al. (2008)"},{"why":"Extends multiplicity to probabilistic classification and serves as the template for adapting the metrics to survival probabilities.","marker":"Watson-Daniels et al. (2023a)"},{"why":"Prior Rashomon-curve uncertainty analysis on the same CMAPSS survival setting; this paper generalizes it into formal multiplicity metrics.","marker":"Yardimci and Cavus (2025)"},{"why":"Motivates fixing the censoring time at 250 cycles to remove censoring sensitivity from the Rashomon set construction.","marker":"Baskay et al. (2025)"},{"why":"Shows similar multiplicity metrics in imbalanced classification and is cited as inspiration for the adapted measures.","marker":"Cavus and Biecek (2024)"}],"fun_headline_variants":[],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that the Random Survival Forest variants scored by the integrated Brier score under a censoring time fixed at 250 cycles adequately represent all plausible near-optimal survival models for predictive maintenance.","fun_headline_variants_meta":{"error":"Client error '402 Payment Required' for url 'https://api.deepseek.com/chat/completions'\nFor more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/402"},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:36:15.341686+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute ambiguity, discrepancy, and obscurity on the same CMAPSS subsets using Cox proportional-hazards and deep survival networks for the Rashomon set instead of Random Survival Forests, keeping the same $\\epsilon$, $\\delta$, and Brier-score rule; if the metrics do not rise with $\\epsilon$, the claimed multiplicity is an artifact of the forest family rather than a general property of accurate survival models.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original predictive-multiplicity framework and the ambiguity, discrepancy, and obscurity definitions that the paper adapts to survival risk estimates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Rashomon effect, the existence of many near-optimal models, from which the paper's Rashomon set definition is taken."},{"cited_title":"(2008, October)","cited_arxiv_id":null,"evidence_quote":"Releases the CMAPSS benchmark datasets used for all four experimental subsets."},{"cited_title":"Rashomon perspective for measuring uncertainty in the survival predictive maintenance models","cited_arxiv_id":"2502.15772","evidence_quote":"Prior Rashomon-curve uncertainty analysis on the same CMAPSS survival setting; this paper generalizes it into formal multiplicity metrics."}],"review_version":1}