{"id":"63ceaca3-fc6a-4c39-9744-fda074d8a2d9","arxiv_id":"1908.04839","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper proposes treating classifier scores as a time-like axis and fitting survival-analysis hazard models to produce global and score-dependent local explanations for time series classification.","lead":"This paper reinterprets a binary classifier's scores as a survival-analysis timeline and uses Cox proportional hazards regression to generate global and score-dependent explanations for time series data. The idea is demonstrated on SMART hard-drive data, but the key Markov claim is assumed rather than proven and no comparison to existing explanation methods is provided.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Markov property is asserted, not proven, and the LSTM's score depends on a history window, so the assumption that lower scores are uninformative is exactly what needs testing.","rationale":"The survival-analysis pipeline stands or falls on the score-indexed process being Markov: without it, the product-limit estimator's conditional probabilities and the CPH/GAM hazard regressions are not the quantities the paper claims to estimate. The reader identifies the same premise in Sec. 4, so I agree. My proposed check would settle it empirically on the actual system. If the test rejects Markovity, the 'proof' claim is false and the reported explanations are not trustworthy. If the test fails to reject, the assumption is empirically supported, though the paper still needs a fidelity benchmark; but the current manuscript provides no such support. The sensitivity-analysis mismatch (SMART 184/7 vs SMART 197/242) is a further sign of unfaithfulness but is secondary to the model-specification issue. Overall, the reader's REJECT is appropriate; the central novelty is an unsupported, potentially false premise.","tokens_in":8237,"tokens_out":11529,"duration_ms":127044,"concrete_test":"On the Backblaze test set with the trained LSTM, test the conditional-independence premise directly: for every device-day in the at-risk set, fit a Cox model for failure with current-day score and current SMART covariates, then a second model that additionally includes lagged scores and lagged SMART values over the preceding 5 days as time-dependent covariates. Compare the two by likelihood-ratio test (or AIC/partial likelihood); if the history-augmented model significantly improves fit (e.g., p<0.05), the Markov assumption is rejected and the explanation is not valid for this model. Report the test statistic and p-value.","verdict_should_be":"REJECT","load_bearing_attack":"The central novelty claim (Sec. 1) is that 'the model score and response has been proven to be a Markov process state space model.' Section 4 does not prove this; it defines X(s) and then says 'The assumption holds as long as the value of X at lower scores is uninformative when predicting outcomes of X at higher scores...' That conditional-independence condition is a substantive premise. It is not automatic for the system studied: the LSTM produces its score from a 5-day lookback window (Sec. 6.2), so at any score index the hidden state and earlier SMART statistics can carry information about later failure that is not contained in the current score. No test or argument is offered to show the premise holds on the Backblaze data. If it fails, the product-limit estimator and the CPH/GAM fits are estimating a misspecified intensity: Table 1's hazard ratios and Figures 3-4's score-dependent coefficients lose the probabilistic meaning required for them to count as faithful explanations of the black box. The text itself flags the premise as an assumption, which directly undercuts the 'proven' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a post-hoc explanation method for binary classifiers on time-series data. It treats the classifier's score as an index analogous to time, models the transition of an observation from a non-responder to a responder as a Markov process, and then uses survival-analysis machinery — the product-limit estimator, Cox proportional hazards regression, and an additive Cox/GAM extension — to associate feature covariates with the hazard of a positive response. The method is demonstrated on an LSTM trained on Backblaze SMART hard-drive data, with SMART 242 and the SMART 197 indicator reported as the significant explanatory variables. The paper claims, in Section 1, that this is the first proof that the model score and response form a Markov process state space model, and that the resulting explanations incorporate time-dependent data.","tokens_in":8461,"tokens_out":7808,"duration_ms":80813,"significance":"The idea of using the score axis as a time axis is potentially interesting and connects model explanation to a mature statistical toolbox. The paper correctly reproduces standard product-limit and martingale equations and the CPH setup, and it explicitly attempts to provide global and local explanations rather than only feature-attribution heatmaps. However, the central theoretical claim is not established: the Markov property is asserted as an assumption, not proven, and the paper's own wording in Sections 3 and 4 admits this. The interpretation of the continuous-covariate hazard ratio is incorrect, the local score-dependent GAM curves are presented without uncertainty estimates, and no fidelity check ties the fitted survival models to the LSTM's actual score-generating process. The experimental evaluation therefore does not support the stated novelty, and the proposed explanations cannot be considered faithful to the black-box model without substantial additional validation. No code or data are provided, which further limits reproducibility.","major_comments":[{"comment":"The claim that 'the model score and response has been proven to be a Markov process state space model' is not supported by the manuscript. In Section 4 the authors define X(s) and then state 'The assumption holds as long as the value of X at lower scores is uninformative when predicting outcomes of X at higher scores, or lower scores and higher scores are independent given the scores.' This is a substantive conditional-independence assumption, not a derived theorem. Section 3 even says the method is 'derived using the assumption of an underlying Markov process,' which contradicts the 'proven' claim in Section 1. The distinction is load-bearing because the LSTM in Section 6.2 uses a 5-day lookback and a hidden state; nothing in the paper rules out that lower-score information (e.g., earlier SMART statistics or the hidden state) is informative about failure at higher scores beyond the current score. A test of the Markov property — for example, comparing product-limit transition estimates conditioned on lower-score histories — is required before the hazard estimates in Table 1 and the curves in Figures 3-4 can be interpreted as explanations of the black box. As written, the central derivation treats a premise as a theorem.","section":"Section 1 and Section 4"},{"comment":"The interpretation of the continuous covariate SMART 242 is incorrect. The text says a hard drive is '3.8486 times the normalized SMART 242 statistic times more likely than the baseline hazard rate.' For a continuous covariate Z, exp(beta) is the hazard ratio for a one-unit increase in Z. Since SMART 242 is normalized to [0,1], a one-unit increase spans the entire observed range of that covariate, so the statement is not a general claim about the presence of the feature. The sentence should be rewritten as 'a one-unit increase in normalized SMART 242 multiplies the hazard by 3.85,' or the hazard ratio for a practically meaningful increment should be reported. This is not a purely cosmetic issue because the explanatory conclusion in Section 6.4 depends on the magnitude and meaning of the reported hazard ratio.","section":"Section 4.4, Section 6.4, Table 1"},{"comment":"The score-dependent coefficient curves are the only evidence for the local explanations, but they are plotted without confidence intervals or standard-error bands. The estimation procedure for the additive Cox model is described only at a high level: the text does not specify the smoothing parameter selection, basis dimension, penalization, or how significance of the smooth terms was assessed. Without uncertainty quantification, the claim in Section 6.4 that 'SMART 197 leads to a greater probability of failure only for higher score regions while in lower score regions indicates a lower probability of failure than the baseline hazard rate' is unsupported. The apparent reversal in Figure 4 could be noise, and the reader has no way to assess this from the manuscript.","section":"Figures 3 and 4, Section 5.3"},{"comment":"The 'MSE Ratio' column in Table 1 is never defined or discussed in the text. If it is intended as a fidelity measure of the explanation, it needs a clear definition and an indication of what level of fidelity is acceptable; if it is not, it should be removed. More generally, the paper does not check whether the fitted proportional-hazards or additive models reproduce the LSTM's score-to-response mapping. Without such a check, the reported coefficients describe associations between covariates and model scores in the training data, and calling them 'explanations' of the black box requires an independent fidelity test that is missing. This is the circularity concern: the survival models are fitted to score-response pairs generated by the LSTM and then used to explain the LSTM, with no validation that the surrogate captures the model's actual decision logic.","section":"Section 6.4 and Table 1"}],"minor_comments":[{"comment":"The text says 'Relu activation functions'; this should be 'ReLU activation functions.'","section":"Section 6.2"},{"comment":"The caption and text contain a stray 'i' in 'SMART 197 i in Figure 4'; the notation should be cleaned up.","section":"Section 6.4 and Figure 4"},{"comment":"The company name is misspelled as 'Blackblaze'; it should be 'Backblaze.'","section":"Section 6.1"},{"comment":"The product-limit estimator is plotted without confidence bands, even though Section 4.3 provides a variance formula; adding confidence bands would help readers assess the uncertainty in the inclusion curve.","section":"Figure 1"},{"comment":"The abstract uses 'Generalized Additive Model' but the body mostly says 'Additive Cox Proportional Hazards Model'; the terminology should be consistent throughout.","section":"Abstract and Section 5.3"},{"comment":"Algorithm 1 is under-specified: it returns 'maxL(β)' but the text describes returning cumulative regression functions B_i(s); the algorithm and the surrounding text should be aligned.","section":"Algorithm 1"}],"recommendation":"reject","confidential_remarks":"The paper extends a workshop presentation and makes an overstrong novelty claim in Section 1. The central Markov property is an assumption rather than a proof, and the experiments do not test it; for the specific LSTM architecture with a 5-day lookback, the assumption is unlikely to hold automatically. The continuous-covariate hazard-ratio misstatement and the absence of uncertainty on the GAM curves are additional substantive problems. In my view the manuscript would require a fundamental reworking of its theoretical foundation and a new validation strategy before it could be considered for publication, so I recommend rejection rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper's core idea is worth your time: it maps a binary classifier's output score to a survival-analysis time axis, then uses product-limit estimators, Cox proportional hazards, and a GAM extension to produce global and score-local feature explanations for time-series black boxes. To my knowledge, that specific mapping is novel in the explainability literature, and the authors deserve credit for borrowing a well-developed statistical toolkit rather than inventing shaky heuristics.\n\nWhat it does well: the mathematics is standard survival analysis, applied carefully; the use of censored observations in the lookback window is a sensible touch; and the paper is unusually candid in Section 4, where the Markov property is stated as an assumption with a clear condition: lower scores are uninformative given the current score. The Backblaze case study is concrete and reproducible in principle.\n\nThe soft spots are real. The Introduction says the score-response process has been \"proven to be a Markov process,\" but Section 4 does not prove it—it asserts an assumption that is likely false for this setup. The LSTM uses a 5-day lookback; its hidden state and earlier SMART values can carry failure-relevant information beyond the current score, so the conditional independence premise needs testing, not assertion. If the premise fails, the hazard ratios and GAM curves lose their probabilistic meaning as explanations of the black box. The paper also never checks fidelity: the CPH coefficients are fitted to the score-response data and then taken to explain the model, but no baseline or held-out validation connects those coefficients to the model's actual decision boundary. The sensitivity analysis pointing to SMART 184 and SMART 7 as most impactful, while the CPH highlights SMART 197 and 242, is not directly contradictory—different questions—but it deserves discussion rather than silence. And the GAM plots have no confidence bands.\n\nNone of these are fatal to the underlying idea; they are fixable with an honest rewrite. The \"proven\" claim should become \"assumed,\" the Markov premise should be empirically tested or explicitly scoped, and the explanation should be benchmarked against something. As it stands, I would not accept it, but it deserves a serious referee rather than a desk reject—there is a usable contribution here after revision.","headline":"Score-as-time survival analysis is a genuinely different idea for explaining time series classifiers, but the paper's proof claim is really an assumption and the validation is missing.","tokens_in":8968,"tokens_out":2575,"would_cite":false,"duration_ms":25915,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a binary classifier's score and the true response form a Markov process, so survival analysis can explain time series black boxes.","keywords":["model explanation","survival analysis","Cox proportional hazards","time series classification","long short-term memory","Markov process","product-limit estimator","score-dependent explanations"],"falsifier":"Take a trained binary classifier and a scored dataset. Split the observations by a feature measured at low scores (e.g., the value of SMART 242 in the first look-back day). Estimate the transition probability $P(\\text{response at score } s_2 \\mid \\text{score } s_1)$ separately within each low-score group, with $s_1 < s_2$. If the estimated transition probabilities differ materially across groups after conditioning on the score at $s_1$, the Markov property fails, and the product-limit estimator and Cox hazard model are not valid descriptions of the score process.","tokens_in":8045,"feed_emoji":"⏳","tokens_out":8134,"duration_ms":75668,"temperature":0.7,"pith_summary":"This paper tries to establish that the score a black-box binary classifier assigns to an observation, together with the observation's true class, forms a Markov process indexed by score. If that holds, the score-to-response journey can be analyzed with survival analysis: a product-limit estimator gives the probability that a true responder has not yet appeared by a given score, and a Cox proportional hazards regression attaches feature weights as multiplicative risk factors. The authors use this to produce both global explanations (constant feature effects) and local, score-dependent explanations (effects that change as the score rises) for time series classifiers, including long short-term memory networks. The payoff is an explanation that respects the ordered, temporal structure of the input features instead of treating each time point as an isolated attribute.","feed_headline":"Survival analysis turns a classifier's score into a time axis","feed_subtitle":"Hazard models fit over model scores reveal which features push a hard drive toward failure, and when.","key_machinery":"The load-bearing mechanism is the score-as-time survival model: the model's score $S$ is declared analogous to time, and the inclusion function $I(s)=P(S>s)$ is rebuilt as a product-integral $I(s)=\\prod_{u\\le s}(1-dA(u))$, estimated by the product-limit estimator. On top of that sit the Cox proportional hazards model, which expresses the hazard of a positive response as a baseline hazard times multiplicative feature effects $e^{\\beta Z}$, and a generalized additive model that lets the $\\beta$'s themselves depend on score. Together they convert a black box's ranking of observations into a probabilistic account of which covariates push an observation toward the modeled class, and at what score regions that push is strongest.","core_discovery":"On its own terms, the paper's central discovery is a translation: the ordered output scores of a binary classifier are reinterpreted as a time index, and the event 'the observation is a true positive' is treated as a transition in a two-state Markov process with an absorbing responder state. Under the Markov assumption, the cumulative probability of a response by score $s$ is $I(s)=\\prod_{u\\le s}(1-dA(u))=\\exp(-\\int_0^s \\alpha(u)\\,du)$, and this survival curve is estimated nonparametrically with the product-limit estimator, commonly called the Kaplan-Meier estimator. A Cox proportional hazards model then explains the hazard $\\alpha(s|Z)=\\alpha_0(s)\\exp(\\beta Z)$ in terms of the original input features, and a generalized additive extension allows the coefficients $\\beta(s)$ to vary with score, producing local explanations. In the hard-drive experiment, the global explanation finds that the indicator for SMART 197 and the normalized SMART 242 count raise failure risk by factors of 1.6484 and 3.8486, while the score-dependent explanation shows SMART 197's risk contribution is only positive in high-score regions and negative at low scores.","pith_inferences":["If the score-as-time Markov assumption holds only approximately, the hazard ratios are still descriptive summaries of the score-response association, but their causal reading weakens; a natural next step is a diagnostic test of the Markov assumption on held-out scores.","The same setup could be applied to multiclass classifiers by treating each class as a competing risk, or to regression models by defining 'response' as a large residual.","Score-dependent explanations suggest a practical auditing tool: comparing the hazard curves of protected subgroups could reveal score regions where a model's reliance on a feature differs across groups.","Because the method needs more true positives than covariates, regularized or penalized survival models are a natural extension for rare-event datasets where the current estimator would be unstable."],"forward_implications":["Any binary classifier that emits a usable score, not just neural networks, can in principle be explained by fitting this survival model over its scores.","Time-dependent covariates enter the explanation naturally: values in a look-back window are treated as censored observations, so the temporal structure of time series is part of the explanation rather than discarded.","Score-dependent coefficients turn global feature importance into a local statement, e.g., a feature can be protective at low scores and harmful at high scores.","The product-limit estimator over scores gives a survival-theoretic definition of the recall curve: the cumulative probability that a true positive has appeared by a score segment.","The method's output is directly usable by domain experts as hazard ratios: a feature with coefficient $\\beta$ multiplies the baseline failure probability by $e^\\beta$."],"supporting_citations":[{"why":"Supplies the survival-analysis definitions, counting process, martingale, and product-integral framework that the method adapts.","marker":"[Aalen et al., 2008]"},{"why":"Supplies the Markov property and memorylessness definition used to justify treating score as a time index.","marker":"[Paul and Baschnagel, 2013]"},{"why":"Supplies the Chapman-Kolmogorov equations used to derive the transition probability matrix and the product-integral solution.","marker":"[Ross, 2014]"},{"why":"Defines the LSTM cell architecture used as the black-box classifier in the hard-drive experiment.","marker":"[Hochreiter and Schmidhuber, 1997]"},{"why":"Cited in the paper as the source of the product-limit estimator used to estimate the cumulative probability of response over score.","marker":"[Kaplan and Meier, 1958]"}],"fun_headline_variants":["Turn model scores into a survival curve for sharper explanations","Survival analysis explains black-box time series predictions","Explain any binary classifier with a hazard model over scores","From score to survival: local explanations using Cox regression","Replace score thresholds with survival curves for feature attribution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the score-to-response process is memoryless: once an observation's current score is known, its lower-score feature values add no information about whether it will respond at higher scores.","fun_headline_variants_meta":{"raw":{"variants":["Turn model scores into a survival curve for sharper explanations","Survival analysis explains black-box time series predictions","Explain any binary classifier with a hazard model over scores","From score to survival: local explanations using Cox regression","Replace score thresholds with survival curves for feature attribution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000554,"raw_usage":{"total_tokens":2655,"prompt_tokens":980,"completion_tokens":1675,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":1601}},"tokens_in":596,"tokens_out":1675,"duration_ms":11193,"temperature":1.0,"reasoning_tokens":1601,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:32:11.332071+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a trained binary classifier and a scored dataset. Split the observations by a feature measured at low scores (e.g., the value of SMART 242 in the first look-back day). Estimate the transition probability $P(\\text{response at score } s_2 \\mid \\text{score } s_1)$ separately within each low-score group, with $s_1 < s_2$. If the estimated transition probabilities differ materially across groups after conditioning on the score at $s_1$, the Markov property fails, and the product-limit estimator and Cox hazard model are not valid descriptions of the score process.","supporting_citations":[{"cited_title":"Survival and event history analysis: a process point of view","cited_arxiv_id":null,"evidence_quote":"Supplies the survival-analysis definitions, counting process, martingale, and product-integral framework that the method adapts."},{"cited_title":"Stochastic processes , volume","cited_arxiv_id":null,"evidence_quote":"Supplies the Markov property and memorylessness definition used to justify treating score as a time index."},{"cited_title":"Introduction to probability models","cited_arxiv_id":null,"evidence_quote":"Supplies the Chapman-Kolmogorov equations used to derive the transition probability matrix and the product-integral solution."}],"review_version":1}