{"id":"484c6f6a-4b5c-4a50-929a-50d5470b3696","arxiv_id":"2603.29183","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"IMPACT applies influence modeling to generate semantically divergent yet realistic unseen anomalies and to repurpose influential samples for decontamination in open-set time series anomaly detection.","lead":"The paper introduces IMPACT, a framework that learns influence scores from training data to create realistic pseudo-anomalies for unseen classes and to clean contaminated labels in open-set time series anomaly detection. Smart readers might examine it for practical improvements in detecting novel anomalies in sequential data such as sensor streams or transaction logs.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Influence estimation accuracy for non-convex time series models is unverified, risking unreliable anomaly generation and decontamination.","rationale":"The reader's weakest assumption correctly flags the generation step; the deeper prerequisite is whether the influence scores themselves are trustworthy for the non-convex models used. Verifying that correlation would directly test the load-bearing link.","tokens_in":1707,"tokens_out":288,"duration_ms":26537,"concrete_test":"On a held-out subset of 200 training windows, compute exact leave-one-out retraining loss change for each sample; compare rank correlation (Spearman) with the paper's influence scores. If correlation < 0.4, the influence modeling step does not support the downstream claims.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The framework first learns an influence function to score training samples, then uses high-influence scores both to synthesize unseen anomalies and to select samples for decontamination. Time series anomaly detectors are typically non-convex (RNNs, transformers, or autoencoders). Standard influence-function approximations rely on Hessian-vector products that become unstable or biased under non-convexity and sequential dependencies. If the estimated scores do not faithfully reflect true leave-one-out effects, both the generated anomalies and the decontamination step rest on noisy or misleading rankings, so reported gains under contamination could be artifacts of the particular training dynamics rather than the claimed mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces IMPACT, a framework for open-set time series anomaly detection (OSAD). It first trains a base anomaly detector and learns an influence function to score the impact of individual training samples on the model. High-influence scores are then used in two ways: (i) to synthesize semantically divergent yet realistic unseen anomalies that preserve sequential structure, and (ii) to identify and repurpose contaminated samples as supervised anomalies for decontamination. Experiments across multiple datasets, OSAD settings, and contamination rates report that IMPACT outperforms existing state-of-the-art methods.","tokens_in":1846,"tokens_out":512,"duration_ms":24811,"significance":"If the influence-based mechanisms are shown to be reliable, the work addresses a clear gap: standard augmentation techniques fail to preserve temporal dependencies in time series, and contamination handling is rarely addressed in OSAD. The explicit code release supports reproducibility and allows direct inspection of baseline implementations and metric definitions.","major_comments":[{"comment":"§3.2 (Influence Function Estimation): The manuscript applies the standard Hessian-vector-product approximation of influence functions to non-convex time-series models (RNNs, transformers, autoencoders). No leave-one-out retraining experiments or correlation analysis between estimated and true influence scores are reported. Because both the anomaly-synthesis step and the decontamination step rest directly on the ranking induced by these scores, this verification is load-bearing for the central performance claims.","section":"§3.2"},{"comment":"§5 (Experiments): The superiority claims under varying contamination rates rely on specific choices for baseline implementations, metric definitions, and contamination simulation. The manuscript does not include statistical significance tests (e.g., paired t-tests or Wilcoxon tests across runs) or ablation on the influence-function hyperparameters, leaving open the possibility that reported gains are sensitive to these implementation details.","section":"§5"}],"minor_comments":[{"comment":"Notation for the influence function (Eq. 3) and the anomaly-generation procedure could be clarified with an explicit algorithmic box showing the exact steps from influence scores to synthesized sequences.","section":"§3"},{"comment":"Figure 3 (qualitative anomaly examples) would benefit from side-by-side comparison with the simple augmentation baselines mentioned in the introduction to illustrate the claimed preservation of sequential structure.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed feedback. The comments highlight important aspects of influence function validation and experimental rigor that we address below. We have revised the manuscript to incorporate additional analyses where feasible while maintaining the core contributions.","responses":[{"response":"We acknowledge that explicit verification of the influence approximations via leave-one-out retraining would strengthen confidence in the rankings. However, full leave-one-out retraining is computationally prohibitive for the RNN and transformer models used across our datasets. The Hessian-vector product approximation follows standard practice in the influence function literature for non-convex settings. To address the concern, we have added a limited correlation study on a smaller synthetic time series dataset where leave-one-out is tractable, along with a discussion of approximation limitations and their potential impact on the anomaly synthesis and decontamination steps. The consistent outperformance across diverse OSAD settings provides indirect empirical support for the utility of the estimated influences.","revision_made":"partial","referee_comment":"[§3.2] §3.2 (Influence Function Estimation): The manuscript applies the standard Hessian-vector-product approximation of influence functions to non-convex time-series models (RNNs, transformers, autoencoders). No leave-one-out retraining experiments or correlation analysis between estimated and true influence scores are reported. Because both the anomaly-synthesis step and the decontamination step rest directly on the ranking induced by these scores, this verification is load-bearing for the central performance claims."},{"response":"We agree that additional statistical validation and hyperparameter sensitivity analysis would improve the robustness of the experimental claims. We have revised the experiments section to include paired t-tests across multiple random seeds for the key performance comparisons under different contamination rates. We have also added an ablation study examining the influence function hyperparameters (damping term and conjugate gradient iterations). These additions confirm that the reported gains remain consistent. We have further clarified the baseline implementations and metric definitions by referencing the exact prior works and code used, which are already released in the repository.","revision_made":"yes","referee_comment":"[§5] §5 (Experiments): The superiority claims under varying contamination rates rely on specific choices for baseline implementations, metric definitions, and contamination simulation. The manuscript does not include statistical significance tests (e.g., paired t-tests or Wilcoxon tests across runs) or ablation on the influence-function hyperparameters, leaving open the possibility that reported gains are sensitive to these implementation details."}],"tokens_in":1417,"tokens_out":517,"duration_ms":25441,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is that IMPACT applies influence modeling in two ways for open-set time series anomaly detection: to create pseudo-anomalies that stay realistic and sequential, and to pick high-influence samples for decontamination when labels are noisy. That dual application looks like the actual new piece relative to prior augmentation work on images or simpler time series methods.","headline":"IMPACT uses influence scores to generate pseudo-anomalies and clean contaminated data for open-set time series anomaly detection, but the reliability of those scores on non-convex models remains a question.","tokens_in":2326,"tokens_out":153,"would_cite":false,"duration_ms":34533,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"learn an influence function that can accurately estimate the impact of individual training samples on the modeling... leverage these influence scores to generate semantically divergent yet realistic unseen anomalies"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"multi-channel deviation loss... entropy minimization of the latent distribution H(S) ∝ log σ²"}],"headline":"Influence-function anomaly scoring and risk-reduction generation in time-series OSAD share no structural overlap with RS distinction-forcing or J-cost machinery","alignment":"orthogonal","rationale":"The paper's core apparatus (TIS influence scoring via Hessian-vector products on multi-channel deviation loss, RADG label-flipping and feature perturbation to minimize test risk, Theorems 1-4 on entropy minimization and distribution shift) operates entirely within standard empirical-risk and influence-function analysis for ML. No reference appears to reciprocal cost J(x), golden-ratio fixed points, 8-tick periodicity, or the single-distinction forcing chain. The loss L(zi,θ) and influence IL(zi,zt) are conventional convex approximations, not instances of the RS cost functional or its uniqueness theorems.","tokens_in":64136,"confidence":"high","tokens_out":335,"duration_ms":11965,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Influence scores from a trained model generate realistic unseen anomalies in time series data while cleaning contaminated training sets.","keywords":["open-set anomaly detection","time series anomaly detection","influence modeling","anomaly decontamination","unseen anomalies","sequential data","pseudo anomaly generation"],"falsifier":"A controlled test in which the anomalies synthesized by IMPACT are fed to a detector and produce no accuracy gain over simple augmentation baselines on standard time-series benchmarks with injected unseen anomalies.","tokens_in":2630,"feed_emoji":"📈","tokens_out":707,"duration_ms":35317,"temperature":0.7,"pith_summary":"Open-set anomaly detection tries to identify both anomalies seen during training and entirely new ones using only limited labels for the seen types. In time series this is hard because simple augmentation methods destroy sequential structure and produce trivial or unrealistic patterns, and the problem gets worse when the training data already contains unlabeled anomalies. IMPACT learns an influence function that measures how much each training sample affects the model, then applies the resulting scores to create new anomalies that stay semantically different yet keep the original time-series shape and to treat the most influential samples as extra supervised signals for removing contamination. A reader would care because many real monitoring tasks involve sequential measurements where both known fault types and completely novel ones must be caught despite imperfect labels.","feed_headline":"Influence scores create realistic unseen time series anomalies","feed_subtitle":"The method also repurposes high-impact samples to clean contaminated training data and raises accuracy over prior open-set detectors.","key_machinery":"An influence function that scores the impact of each training sample on the learned model, used both to synthesize new anomalies and to select samples for decontamination.","core_discovery":"The paper claims that learning an influence function to estimate the effect of individual training samples on the model allows two things at once: the scores can drive the synthesis of semantically divergent yet realistic unseen anomalies that respect sequential structure, and the same high-influence samples can be repurposed as supervised anomalies to decontaminate the training data, yielding higher accuracy than prior open-set methods across varying contamination rates and settings.","pith_inferences":["The same influence-based generation step could be tested on other ordered data such as sensor streams or financial ticks where sequential realism matters.","If the influence function generalizes, it might reduce reliance on domain-specific augmentation rules in broader open-set learning settings.","A follow-up experiment could measure how much of the reported gain comes from the decontamination step versus the anomaly synthesis step by ablating each component separately.","The approach suggests a route to make anomaly detectors more robust without requiring perfectly clean training sets, a common practical constraint."],"forward_implications":["Detection accuracy rises for both seen and unseen anomaly classes even when training data contains unlabeled anomalies.","Sequential structure is preserved in the generated pseudo-anomalies, avoiding the trivial patterns produced by prior augmentation techniques.","High-influence training samples can be directly reused as additional labeled anomalies, reducing the need for extra manual labeling.","Performance remains stable across different levels of contamination and different open-set configurations."],"fun_headline_variants":["Influence functions synthesize realistic time series anomalies","High-influence samples decontaminate training data","Modeling influence improves open-set anomaly detection","Influence scores generate realistic unseen anomalies","Impact estimation aids open-set time series anomaly tasks"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That scores derived from how much each training point changes the model can be turned into new anomaly examples that are both realistic for time series and genuinely different from what was already seen.","fun_headline_variants_meta":{"raw":{"variants":["Influence functions synthesize realistic time series anomalies","High-influence samples decontaminate training data","Modeling influence improves open-set anomaly detection","Influence scores generate realistic unseen anomalies","Impact estimation aids open-set time series anomaly tasks"]},"model":"grok-4.3","cost_usd":0.005094,"raw_usage":{"total_tokens":2412,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":50940500,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1654,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":63,"duration_ms":18583,"temperature":1.0,"reasoning_tokens":1654,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T10:50:58.069394+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which the anomalies synthesized by IMPACT are fed to a detector and produce no accuracy gain over simple augmentation baselines on standard time-series benchmarks with injected unseen anomalies.","supporting_citations":[],"review_version":1}