{"id":"6ff6f3f0-2201-4bae-b7db-e2134c37452e","arxiv_id":"2607.29092","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"MATERO-RCA jointly optimizes root set, root-effect modes, and repaired counterfactual trajectories, with a MILP-accelerated best-bound search that reports strong benchmark RCA results.","lead":"MATERO-RCA picks the set of variables that caused an industrial alarm, decides how each root acted (only on its own signal, or by physically affecting downstream variables), and rewrites the event timeline until alarms and causal relations look normal. It reports the best root-set recovery on five benchmarks, but no code or data are released and several settings were chosen using test outcomes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline results assume the supplied causal graph is complete for root-alarm reachability; this premise is unverified on causRCA, so the 97.8% Set F1 is conditional on the candidate ceiling C=∪an(A_j).","rationale":"The reader's weakest-assumption analysis identifies the correct load-bearing point. The paper is internally scrupulous: Theorem 1 is explicitly about the fixed inner solver, and S4.4 states the certificate does not cover causal completeness outside X. But the central empirical claim—state-of-the-art RCA performance—can only be interpreted as conditional on the candidate set containing true roots. On generated datasets the condition is true by construction; on causRCA it is neither tested nor reported. The QT graph-sensitivity experiment does not address this because it preserves root-to-alarm reachability, and the failure analysis attributes residual errors to exactly this assumption. This is a boundary condition rather than an internal inconsistency, so it does not justify rejecting the paper; but it does justify the reader's CONDITIONAL verdict. The proposed check directly settles whether the premise holds on the one real benchmark and quantifies how much of the reported performance depends on the supplied graph.","tokens_in":30774,"tokens_out":11363,"duration_ms":133530,"concrete_test":"Using the released causRCA events, compute C(event)=∪_{A_j∈A^⊤} an(A_j) from the provided graph and verify every annotated root belongs to C; report the per-event pass rate. Then re-run MATERO-RCA on causRCA under graph edits that remove one true root-to-alarm edge at a time for a random subset of events, keeping checkpoints, K, λ, γ, ε, and all other settings fixed. If any annotated root is outside C, or if deleting a single reachability edge causes MATERO-RCA to drop that root or add a non-root, the headline result is conditional on graph completeness rather than on the trajectory-level objective. Release of code/data is required to run this check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Problem Setting restricts every root set to C = ∪_{A_j∈A^⊤} an(A_j): 'Since non-ancestors cannot causally reach a top-level active alarm, candidates are restricted to C=...' Theorems 1 and the MILP search therefore certify optimality only within this finite candidate space, and S4.4 explicitly disclaims 'causal completeness outside X.' If an event has a true observed root that is not an ancestor of any top-level active alarm because G omits or mis-specifies an edge, that root is not in C and MATERO-RCA cannot return it, regardless of CompatNet/RepairNet quality. The Conclusion's statement that remaining failures arise from 'causal-graph misspecification' confirms the assumption is load-bearing.\n\nThe empirical support for the premise is incomplete. Synthetic TA/TS/LM/QT graphs are generator- or physics-defined, so all roots are in C by construction. For causRCA, the only real dataset, the paper does not report whether every annotated root lies in C for its events, nor whether the root annotations were produced independently of the expert graph. S7.5's graph perturbation study is run only on QT and explicitly keeps 'declared root-to-alarm reachability' fixed, so it never tests the candidate ceiling. Thus the main-table causRCA ExactSet results do not establish behavior under the realistic condition where the supplied graph is incomplete.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"MATERO-RCA proposes an RCA method that, given a causal graph G and normal training windows, jointly optimizes a root set, per-root effect modes (observation-only vs physical propagation), and counterfactual trajectories. CompatNet supplies calibrated energies for alarm-context plausibility and local parent–child compatibility; RepairNet proposes diverse counterfactuals refined by gradient descent under the same objective. The outer combinatorial search is driven by a residual-cover lower bound encoded as an exact MILP, with anytime optimality guarantees for a fixed inner solver. Experiments on one real industrial dataset (causRCA), three synthetic generators, and the Quadruple-Tank simulation report mean Set F1 97.8% and Exact Set 94.3%, together with ablations, sensitivity studies, and failure analyses. The appendix contains proofs of Propositions 1–2, Theorem 1, and Corollary 1, with the guarantees explicitly scoped to the finite admissible root–mode space and the fixed inner solver.","tokens_in":31217,"tokens_out":7937,"duration_ms":88201,"significance":"The proposed framework is a coherent and nontrivial integration of learned trajectory-level compatibility, mode-aware counterfactual generation, and certified combinatorial search. If the empirical claims hold, this is a meaningful advance over score- and rollout-based temporal RCA baselines. The paper is commendably careful in hedging its theoretical claims: the proofs are in the appendix, the certificates are stated for a fixed inner solver, and the authors include failure analysis and extensive sensitivity studies rather than only favorable tables. The principal weakness is that the central empirical claim is conditional on the supplied graph being complete for root–alarm reachability, a condition that is not verified on the real dataset; in addition, no code/data or uncertainty quantification accompany the headline numbers.","major_comments":[{"comment":"The candidate restriction C = ∪_{A_j∈A^⊤} an(A_j) is load-bearing. Any true root that is not an ancestor of a top-level active alarm under the supplied graph is unrecoverable, regardless of CompatNet/RepairNet quality. The paper itself disclaims causal completeness outside X in S4.4, and the Conclusion attributes remaining failures to causal-graph misspecification. For the synthetic groups the roots are in C by construction; for the only real group, causRCA, the paper does not report whether every annotated root lies in C, nor whether the annotations were produced independently of the expert graph. The S7.5 perturbation study explicitly keeps 'declared root-to-alarm reachability' fixed, so it cannot probe the candidate ceiling. The headline claims should therefore be qualified as conditional on graph completeness, and the authors should add a per-event check that annotated roots are in C","section":"Problem Setting and Notation; S4.4; S7.5"},{"comment":"The central empirical claim rests on single point estimates. Table 1 (and Tables 2–3) report no error bars, confidence intervals, or per-seed variation; the appendix's sensitivity experiments fix model seed 1 (e.g., S7.5, S7.4). With test sets of 25–50 events, a difference of one or two events changes the percentages by 2–4 points, so the statement that MATERO-RCA 'outperforms all baselines' needs uncertainty quantification. No code or data are provided, making the results in Tables 1–2 impossible to reproduce independently. I do not regard this as a correctness error, but it is necessary support for the claimed state-of-the-art performance.","section":"Experiments: Table 1; Implementation"}],"minor_comments":[{"comment":"The inequality 0 ≤ bU_t − bU⋆ ≤ [bU_t − B_t]_+ is stated for every iteration t, but at t = 0, bU_0 = +∞ and the expression is not meaningful. The proof correctly notes that the finite anytime gap starts at t = 1; the theorem statement should be rephrased accordingly.","section":"Theorem 1, Eq. (11)"},{"comment":"The main text does not explain that the causRCA column is a macro average over four overlapping views (Probe, Coolant, Hydraulics, Full) of the same 100 physical events. This should be stated next to Table 1 to avoid the impression that the column aggregates independent test events.","section":"Table 1 and Table S4"},{"comment":"Some entries use inconsistent abbreviations ('Stable.', 'MATERO.', 'Smooth.'). Use the full method names or a uniform abbreviation scheme.","section":"Table 2"},{"comment":"A data/code availability statement is missing. Given the diversity of datasets and the dependence on many hyperparameters (ε, λ, γ, ρ_C, ρ_A, bin counts, etc.), releasing code and trained configurations would substantially strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you should know: this is a real method paper, not an incremental tweak. MATERO-RCA jointly optimizes the root set, root-effect modes, and counterfactual trajectories under a graph-factored energy, and it adds a residual-cover MILP that certifies fixed-oracle optimality over the finite admissible space. That is genuinely new relative to the trajectory-level RCA baselines I know, which mostly rank candidates by forward simulation. The mode distinction (observation-only vs physical-propagation) is also a useful modeling choice. The appendix is honest: proofs, ablations, sensitivity studies, and a failure analysis are all there, and the certificate scope is clearly stated.\n\nThe weak spots are mostly around verification. No code or data are released, the main tables have no error bars, and a few design knobs—ε, bin count, ρ_A, even the baseline α—are justified by their effect on the test set. That makes the headline 97.8% Set F1 plausible but not independently checked. The stress-test concern about the candidate ceiling C = ∪ an(A_j) is valid but not fatal, because the paper explicitly says the certificate does not cover causal completeness outside X and attributes remaining failures to graph misspecification. What is missing is an empirical check on the real dataset: do all annotated roots actually fall inside C? They do by construction on the synthetic sets, and S7.5 only perturbs QT while keeping declared root-to-alarm reachability fixed, so it never tests the ceiling. That is a real gap.\n\nThe central argument, though, holds up: within the stated assumptions, the fixed-oracle certificate is sound, and the empirical story is coherent. The paper is for researchers working on industrial temporal RCA and on counterfactual/energy-based explanation methods. It deserves a serious referee. My recommendation: send it to review, but the authors should be required to release artifacts and to run model selection on a proper validation split, not on test-set outcomes.","headline":"A genuinely new joint root-set/mode/counterfactual optimization with a certified search, but the empirical headline is conditional on a graph-completeness assumption and missing artifacts.","tokens_in":31639,"tokens_out":2307,"would_cite":true,"duration_ms":23558,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MATERO-RCA jointly optimizes root sets, root-effect modes, and counterfactual trajectories to resolve industrial alarms, achieving 97.8% mean Set F1 and 94.3% Exact Set across five dataset groups.","keywords":["root cause analysis","industrial time series","energy-based model","counterfactual trajectories","best-bound search","temporal compatibility","mixed-integer linear program"],"falsifier":"Construct an event where the true root is not an ancestor of any top-level active alarm in the given graph (e.g., a latent cause) and show that MATERO-RCA reports a root set that omits the true root while achieving a low objective value; this would directly violate the paper's central claim that the root set is recoverable from the graph-factored objective.","tokens_in":30677,"feed_emoji":"⚙️","tokens_out":1919,"duration_ms":22112,"temperature":0.7,"pith_summary":"The paper proposes a root cause analysis method for industrial time series where a signal can look normal in isolation but violate its operating context. Instead of ranking anomalies or simulating forward effects, MATERO-RCA jointly selects a set of root variables, assigns each root either an observation-only or physical-propagation mode, and optimizes counterfactual trajectories that would resolve the active alarms. The graph-wide objective combines alarm resolution with temporal compatibility across local causal relations, and an exact MILP search with a residual-cover lower bound certifies that no better root set was missed under the fixed inner solver. On five simulated and real industrial dataset groups, the method reports the highest set-level accuracy across all baselines.","feed_headline":"Root-cause search recovers exact roots in 94% of industrial events","feed_subtitle":"MATERO-RCA jointly optimizes root set, effect modes, and counterfactual trajectories against a graph-wide energy objective.","key_machinery":"The central machinery is an energy-based objective J(x) = Φ_A(x) + λ Φ_C(x) with two learned components: CompatNet provides calibrated compatibility energies for local causal relations (Φ_C) and alarm-parent contexts (Φ_A), while RepairNet generates diverse counterfactual trajectory proposals that are refined by straight-through gradient optimization. The outer search uses a residual-cover lower bound, exactly encoded as a binary MILP, to prioritize unseen root sets and certify fixed-oracle outer optimality.","core_discovery":"The paper introduces MATERO-RCA, which treats root cause identification as minimizing an energy function over counterfactual trajectories rather than ranking anomalies or performing candidate-wise forward simulation. A Temporal Compatibility Network (CompatNet) learns calibrated energies for local causal relations and alarm-parent contexts, and a Counterfactual Repair Network (RepairNet) initializes mode-aware counterfactual trajectories that are refined by gradient descent under the same objective. The outer search minimizes a graph-wide objective with a cardinality penalty, and a residual-cover lower bound enables certified best-bound search over the finite root–mode space. The paper repor","pith_inferences":["The load-bearing assumption is that the supplied causal graph correctly captures all root-alarm reachability; the method cannot propose a root that is not an ancestor of a top-level active alarm in the graph.","The approach could be extended to incorporate graph uncertainty by jointly optimizing over a distribution of causal graphs or by adding graph-structure refinement.","The certified search's guarantees are relative to the learned compatibility model, not the true physical system; failures attributed to out-of-distribution propagation suggest a need for open-set response models."],"forward_implications":["RCA becomes a trajectory-level optimization problem rather than anomaly-first ranking, directly handling multi-root, multi-alarm events.","The certified best-bound search guarantees that no better root set exists within the admissible root–mode space, given the learned energy model and fixed inner solver.","The method handles observation-only root effects, which previous methods excluded, by explicitly modeling root-effect modes.","The ablations show that alarm resolution and temporal compatibility provide complementary evidence; removing either term degrades set recovery."],"fun_headline_variants":["Certified best-bound search pinpoints exact industrial root causes","Counterfactual trajectories enable mode-aware root-set optimization","Energy-based model uncovers multi-root causes in time series","MATERO-RCA: exact root recovery in 94% of industrial events","Trajectory-level energy optimization for certified RCA"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The causal graph provided must be correct for root-to-alarm reachability, because the search restricts candidates to ancestors of top-level active alarms and has no way to propose a true root that lies outside that ancestor set.","fun_headline_variants_meta":{"raw":{"variants":["Certified best-bound search pinpoints exact industrial root causes","Counterfactual trajectories enable mode-aware root-set optimization","Energy-based model uncovers multi-root causes in time series","MATERO-RCA: exact root recovery in 94% of industrial events","Trajectory-level energy optimization for certified RCA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00031,"raw_usage":{"total_tokens":1594,"prompt_tokens":724,"completion_tokens":870,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":788}},"tokens_in":468,"tokens_out":870,"duration_ms":9835,"temperature":1.0,"reasoning_tokens":788,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T13:47:40.190987+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct an event where the true root is not an ancestor of any top-level active alarm in the given graph (e.g., a latent cause) and show that MATERO-RCA reports a root set that omits the true root while achieving a low objective value; this would directly violate the paper's central claim that the root set is recoverable from the graph-factored objective.","supporting_citations":[],"review_version":1}