REVIEW 4 major objections 6 minor 103 references
Temporal cross-validation impacts multivariate time series subsequence anomaly detection evaluation
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read How you slice time in cross-validation changes anomaly detection benchmark results: sliding-window validation yields higher median AUC-PR and lower variance than walk-forward on multivariate CAN fault data, especially for deep classifiers.
desk verdict Useful but internally inconsistent: the reported sliding-window advantage may be an artifact of uneven fold exclusion, since the paper's own Definitions 2 and 3 define identical test windows for WF and SW. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the temporal partition itself. Walk-forward defines fold $k$'s training subsequence as $S^k_{\text{train}}=\{x_1,\dots,x_{\omega+(k-1)\delta}\}$ and test subsequence as $S^k_{\text{test}}=\{x_{\omega+(k-1)\delta+1},\dots,x_{\omega+k\delta}\}$, so the training window keeps growing and never revisits old data. Sliding window keeps a fixed-length training window $S^k_{\text{train}}=\{x_{1+(k-1)\delta},\dots,x_{\omega+(k-1)\delta}\}$ with the same test-set definition, so consecutive training windows overlap in time. That overlap is the operative mechanism: it gives the classifier repeated, locally contiguous views of intermittent fault signatures and keeps recent context, which the paper credits for higher AUC-PR and lower variance, while the expanding WF window loses recency and creates test-time discontinuity.
What would settle it
Re-run the same eight classifiers and $K$ values on the two CAN datasets while keeping single-class folds (e.g., scoring them with a metric defined for degenerate test sets) or after matching WF and SW folds to have identical positive-class-ratio distributions; if the sliding-window median AUC-PR advantage disappears or reverses under either correction, the paper's central claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that temporal validation design materially changes benchmark conclusions for subsequence anomaly detection in multivariate time series. On two CAN fault datasets (power injector and spark plug), all eight classifiers—SVM, random forest, XGBoost, ResNet, TCN, LSTM+FCN, InceptionTime, and ROCKET—achieve higher median AUC-PR under SW than under WF except random forest, whose small drop from 0.80 to 0.77 is not statistically significant. The differences are significant at $\alpha=0.05$ for the other seven classifiers, with the largest shift from 0.50 to 0.76 for SVM. The paper also finds that the SW advantage is clearest at $K=3$ through $K=6$, weakens at $K=7$, and re-emerges at $K=8$ and $K=9$. It concludes that overlap-based sliding-window evaluation, rather than growing-window walk-forward evaluation, should be preferred for benchmarking fault-like subsequence anomalies in streaming multivariate data.
Load-bearing premise
The comparison assumes that dropping test folds that contain only one class produces an unbiased comparison between WF and SW; in the CAN data those folds are more common under SW at higher fold counts, and the two strategies produce different class-imbalance profiles, so the reported SW advantage could be an artifact of which folds were excluded rather than a property of the validation method.
Editorial extensions
If this is right
- Benchmark studies of subsequence anomaly detectors should report the temporal cross-validation scheme and fold count, because the choice alone moves median AUC-PR substantially (e.g., SVM from 0.50 to 0.76).
- At lower fold counts ($K=3$ through $K=6$), sliding-window validation gives a statistically significant advantage at $\alpha=0.05$; at $K=7$ the advantage loses significance but returns at $K=8$ and $K=9$, so conclusions drawn from a single fold count should not be generalized.
- Deep classifiers (ResNet, TCN, LSTM+FCN, InceptionTime, ROCKET) all improve significantly under SW, while random forest stays flat, meaning the validation scheme can alter which architecture appears best in a benchmark.
- The overlapping windows of SW preserve fault signatures more effectively at low fold counts, so evaluations meant to mimic streaming fault detection should prefer fixed-window overlap at modest $K$.
- Because the two strategies produce different class-imbalance profiles across folds, any comparison of validation schemes should state how degenerate single-class folds were handled.
Reading between the lines
- The class-imbalance asymmetry between WF and SW means pooled AUC-PR comparisons may overstate the SW advantage; a fairer evaluation would weight folds by positive-class ratio or use a metric robust to degenerate test folds.
- If the SW advantage is driven by overlap, then window length $\omega$ and offset $\delta$ should modulate the effect; sweeping those parameters, which the paper deliberately fixed, is a direct testable extension.
- Because the datasets are two small CAN recordings, the conclusion that SW is generally better for MTS subsequence anomaly detection is a hypothesis about other domains; applying the same protocol to larger public MTS fault datasets would show whether it transfers.
- Walk-forward's failure despite more training data suggests that recency of the training distribution, not volume, matters for intermittent faults; adding an explicit forgetting mechanism to WF could close that gap.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper conducts an empirical comparison of two temporal cross-validation strategies, walk-forward (WF) and sliding window (SW), for subsequence anomaly detection in multivariate time series, using two CAN-bus fault datasets. Eight classifiers (SVM, RF, XGBoost, ResNet, TCN, LSTM+FCN, InceptionTime, ROCKET) are evaluated with AUC-PR under K-folds 3 through 9. The central claim is that SW yields consistently higher median AUC-PR and lower fold-to-fold variance than WF, especially for deep architectures, and that the choice of TSCV strategy materially changes benchmark conclusions. The authors release code and data and report a series of Mann-Whitney tests in support of the claim, while also acknowledging limitations around class imbalance and fold comparability.
Significance. If the main claim is correct, the paper addresses a real gap: benchmark evaluations of multivariate time series anomaly detectors often ignore how the temporal validation scheme interacts with classifier performance and class imbalance. The study is one of the few to systematically vary TSCV strategy, fold count, and classifier family on a controlled fault-injection dataset, and it ships public code and datasets, which is a concrete strength for reproducibility. However, the central comparison currently rests on a contradiction in the formal definitions and on statistical analyses that pool non-independent fold-level scores; these issues must be resolved before the significance of the SW advantage can be assessed.
major comments (4)
- [Definitions 2–3; §4.3; §6.4] The manuscript contradicts its own formal setup. Definition 2 (WF) and Definition 3 (SW) specify identical test subsequences, S^k_test = {x_{ω+(k−1)δ+1}, ..., x_{ω+kδ}}, for fixed k, ω, δ. Consequently, the positive-class ratios of the test folds, and hence the set of single-class test folds excluded from AUC-PR evaluation, cannot differ between WF and SW under these definitions. Yet §4.3 states that single-class test folds are 'prevalent under the sliding window configuration, especially at higher fold counts,' and §6.4 says class imbalance is different for WF and SW, limiting comparability. Either the implementation deviates from the definitions (e.g., different ω or different window starts for SW), in which case the pooled comparison in §5.1 is over different fold populations and the reported SW advantage may be an artifact of which folds were dropped, or the implementation matches the definitions, in which case §4.3 and §6.4 are inaccurate. The authors must verify the fold construction in the released code and re-run the comparison on a matched, non-degenerate set of folds, reporting how many folds were excluded per strategy and K.
- [§5.1; §5.3; Table 2] The statistical evidence for the central claim is not robust. The Mann-Whitney U tests in Section 5.1 and Table 2 treat fold-level AUC-PR scores as independent observations, but folds from the same time series are strongly dependent, and SW windows are explicitly overlapping. This dependence can inflate apparent differences and produce p-values reported as 0.00. The manuscript should use a cluster-robust permutation test or a mixed-effect model with fold/sequence as a random effect, report confidence intervals for median differences, and address the multiple comparisons across K values and classifiers (e.g., with a false-discovery-rate correction). Without this, the 'consistently' in the abstract is not supported by the reported statistics.
- [§5; §4.1 (Definitions 2–3)] The comparison between WF and SW is confounded by the training window size. Section 5 says ω is 'fluctuating and dependant on TSCV strategy and K-fold number,' but the paper never reports the actual ω values used for each strategy and K. In Definition 2 the training window grows as ω+(k−1)δ, while in Definition 3 the training window has fixed length ω but starts at 1+(k−1)δ; if the ω assigned to SW differs from the initial ω assigned to WF, then the comparison in §5.1 is partly a comparison of training set sizes and recency, not of validation strategy per se. The authors should report the full schedule of ω and δ for both strategies and, if possible, include an analysis that matches training set sizes or conditions on them.
- [§6.4] The limitation section concedes that class imbalance 'limits comparability of the two TSCV methods as class imbalance is different for WF and SW.' This concession directly qualifies the paper's main empirical conclusion. The manuscript should either provide evidence that the differing imbalance profiles do not drive the SW advantage (e.g., by stratifying on test positive-class ratio or restricting to folds with comparable ratios) or soften the abstract's claim accordingly. As written, the conclusion in Section 7 that 'SW consistently outperformed WF' goes beyond what the acknowledged limitations and the reported analysis can support.
minor comments (6)
- [Abstract] The final sentence is grammatically incomplete: 'TSCV design in benchmarking anomaly detection models on streaming time series and provide guidance' should be rephrased, e.g., 'This study demonstrates that TSCV design matters in benchmarking anomaly detection models on streaming time series and provides guidance for selecting evaluation strategies.'
- [§5] The notation 'K∈{3,9}' is ambiguous; the experiments and Table 3 clearly include K = 3,4,5,6,7,8,9. Please write K ∈ {3,4,...,9}.
- [Table 2; §5.3] The RF p-value is reported as 0.35 in the text but as 0.18 in Table 2; these should be reconciled, and exact p-values should be reported rather than rounded to 0.00.
- [§5.5] There are several typos, including 'medidan' for 'median' and 'subsequnce' for 'subsequence', which should be corrected throughout.
- [§4.3] The statement 'Extreme cases such as test folds with only a single class were prevalent under the sliding window configuration' conflicts with Definitions 2–3, as noted in the major comments; once the fold-construction issue is resolved, this sentence should be revised to match the actual implementation.
- [§2.2] The subsection begins 'This section we contextualize our work', which is ungrammatical; it should be 'This section contextualizes our work'.
Circularity Check
No circularity: the comparison is empirical, the metric is external to the split construction, and the acknowledged fold-exclusion issue is a validity concern rather than a derivation from the paper's own inputs.
full rationale
This paper reports an empirical benchmark rather than a derivation: predefined WF and SW temporal cross-validation splits are combined with fixed/default classifier configurations, and performance is measured with AUC-PR, a metric that is external to how the splits were constructed. No parameter is tuned to maximize the WF-vs-SW difference, no fitted quantity is later relabeled as a prediction, and the central claim (SW yields higher median AUC-PR on these datasets) is an observed statistic rather than a formal consequence of the definitions. The main threat to the comparison is the exclusion of single-class test folds: Section 6.4 explicitly concedes that 'class imbalance is different for WF and SW,' which could bias the pooled comparison. That is a confounding/validity concern, not a circularity, because the fold-exclusion rule is applied to both strategies before any model is evaluated and is not derived from the outcome. I also note that Definitions 2 and 3 define identical test windows for fixed k, omega, and delta; if the released implementation actually uses the same omega for both strategies, then Section 4.3's claim that single-class test folds are more prevalent under SW would be internally inconsistent, and the code would need to be checked. But even in that scenario the reported performance differences would be an implementation or experimental-design artifact, not an equivalence between the result and its inputs. The paper's comparisons are self-contained against the data and the chosen split definitions, so no circularity score above 0 is warranted.
Assumptions & free parameters
free parameters (3)
- offset delta =
150
- training window length omega =
not reported
- classifier hyperparameters =
defaults (e.g., 100 estimators, learning rate 0.01)
assumptions (4)
- domain assumption Relay-toggle timestamps mark the true start and end of each injected fault.
- domain assumption CAN-D reverse engineering and 100 Hz resampling produce a faithful MTS representation.
- ad hoc to paper Excluding single-class test folds yields a valid comparison between WF and SW.
- standard math A multivariate time series is non-stationary if at least one component is non-stationary.
Cite this review
Pith. "Pith review of Temporal cross-validation impacts multivariate time series subsequence anomaly detection evaluation." pith.science (2026). https://pith.science/paper/7NUXSXDN
@misc{pith2026250612183,
author = {Pith},
title = {Pith review of: Temporal cross-validation impacts multivariate time series subsequence anomaly detection evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/7NUXSXDN}},
note = {Machine review of arXiv:2506.12183}
}
read the original abstract
Evaluating anomaly detection in multivariate time series (MTS) requires careful consideration of temporal dependencies, particularly when detecting subsequence anomalies common in fault detection scenarios. While time series cross-validation (TSCV) techniques aim to preserve temporal ordering during model evaluation, their impact on classifier performance remains underexplored. This study systematically investigates the effect of TSCV strategy on the precision-recall characteristics of classifiers trained to detect fault-like anomalies in MTS datasets. We compare walk-forward (WF) and sliding window (SW) methods across a range of validation partition configurations and classifier types, including shallow learners and deep learning (DL) classifiers. Results show that SW consistently yields higher median AUC-PR scores and reduced fold-to-fold performance variance, particularly for deep architectures sensitive to localized temporal continuity. Furthermore, we find that classifier generalization is sensitive to the number and structure of temporal partitions, with overlapping windows preserving fault signatures more effectively at lower fold counts. A classifier-level stratified analysis reveals that certain algorithms, such as random forests (RF), maintain stable performance across validation schemes, whereas others exhibit marked sensitivity. This study demonstrates that TSCV design in benchmarking anomaly detection models on streaming time series and provide guidance for selecting evaluation strategies in temporally structured learning environments.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Anam Abid, Muhammad Tahir Khan, and Javaid Iqbal. 2021. A review on fault detection and diagnosis techniques: basics and beyond. Artificial Intelligence Review54, 5 (2021), 3639–3664
2021
-
[2]
Julien Audibert, Pietro Michiardi, Fr ´ed´eric Guyard, S ´ebastien Marti, and Maria A Zuluaga. 2020. Usad: Unsupervised anomaly detection on multivariate time series. InProceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining. Association for Computing Machin- ery, Virtual Event CA USA, 3395–3404
2020
-
[3]
Christoph Bergmeir, Rob J Hyndman, and Bonsoo Koo. 2018. A note on the validity of cross-validation for evaluating autore- gressive time series prediction.Computational Statistics & Data Analysis 120 (2018), 70–83
2018
-
[4]
Albert Bifet, Gianmarco de Francisci Morales, Jesse Read, Ge- off Holmes, and Bernhard Pfahringer. 2015. E fficient online evaluation of big data stream classifiers. In Proceedings of the 21th ACM SIGKDD international conference on knowledge dis- covery and data mining. Association for Computing Machinery, Sydney NSW Australia, 59–68
2015
-
[5]
2023.Machine learning for data streams: with prac- tical examples in MOA
Albert Bifet, Ricard Gavalda, Geo ffrey Holmes, and Bernhard Pfahringer. 2023.Machine learning for data streams: with prac- tical examples in MOA. MIT press, Cambridge, Massachusetts. 85–127 pages
2023
-
[6]
Anil Kumar Biswal, Debabrata Singh, Binod Kumar Pattanayak, Debabrata Samanta, Shehzad Ashraf Chaudhry, and Azeem Ir- shad. 2021. Adaptive Fault-Tolerant System and Optimal Power Allocation for Smart Vehicles in Smart Cities Using Controller Area Network. Security and Communication Networks 2021, 1 (2021), 2147958
2021
-
[7]
Ane Bl ´azquez-Garc´ıa, Angel Conde, Usue Mori, and Jose A Lozano. 2021. A review on outlier /anomaly detection in time series data. ACM computing surveys (CSUR) 54, 3 (2021), 1– 33
2021
-
[8]
Avrim Blum, Adam Kalai, and John Langford. 1999. Beat- ing the hold-out: Bounds for k-fold and progressive cross- validation. In Proceedings of the twelfth annual conference on Computational learning theory. Association for Computing Ma- chinery, Santa Cruz California USA, 203–208
1999
Show all 103 references
-
[9]
Jan Brabec and Lukas Machlica. 2018. Bad practices in evalua- tion methodology relevant to class-imbalanced problems. arXiv preprint arXiv:1812.01388 (2018)
2018 arXiv
-
[10]
Leo Breiman. 2001. Random forests. Machine learning 45 (2001), 5–32
2001
-
[11]
Alessio Buscemi, Ion Turcanu, German Castignani, Andriy Panchenko, Thomas Engel, and Kang G Shin. 2023. A survey on controller area network reverse engineering. IEEE Commu- nications Surveys & Tutorials 25, 3 (2023), 1445–1481
2023
-
[12]
Vitor Cerqueira, Luis Torgo, and Igor Mozetiˇc. 2020. Evaluating time series forecasting models: An empirical study on perfor- mance estimation methods. Machine Learning 109, 11 (2020), 1997–2028
2020
-
[13]
Lingli Chen, Xin Gao, Jing Liu, Yunkai Zhang, Xinping Diao, Taizhi Wang, Jiawen Lu, and Zhihang Meng. 2025. A multivari- ate time series anomaly detection method with Multi-Grain Dy- namic Receptive Field. Knowledge-Based Systems 309 (2025), 112768
2025
-
[14]
Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data min- ing. Association for Computing Machinery, San Francisco Cal- ifornia USA, 785–794
2016
-
[15]
Haibin Cheng, Pang-Ning Tan, Christopher Potter, and Steven Klooster. 2009. Detection and characterization of anomalies in multivariate time series. In Proceedings of the 2009 SIAM inter- national conference on data mining . SIAM, Society for Indus- trial and Applied Mathemati...
2009
-
[16]
Corinna Cortes and Vladimir Vapnik. 1995. Support-vector net- works. Machine learning 20 (1995), 273–297. 19
1995
-
[17]
Mayur Datar, Aristides Gionis, Piotr Indyk, and Rajeev Mot- wani. 2002. Maintaining stream statistics over sliding windows. SIAM journal on computing 31, 6 (2002), 1794–1813
2002
-
[18]
Jesse Davis and Mark Goadrich. 2006. The relationship be- tween Precision-Recall and ROC curves. In Proceedings of the 23rd international conference on Machine learning. Association for Computing Machinery, Pittsburgh Pennsylvania USA, 233– 240
2006
-
[19]
Angus Dempster, Franc ¸ois Petitjean, and Geo ffrey I. Webb
-
[20]
Ailin Deng and Bryan Hooi. 2021. Graph neural network-based anomaly detection in multivariate time series. In Proceedings of the AAAI conference on artificial intelligence , V ol. 35. AAAI Press, Virtual, 4027–4035
2021
-
[21]
Grace Deng, Cuize Han, Tommaso Dreossi, Clarence Lee, and David S Matteson. 2022. Ib-gan: A unified approach for multi- variate time series classification under class imbalance. In Pro- ceedings of the 2022 SIAM International Conference on Data Mining (SDM). SIAM, Society for ...
2022
-
[22]
David A Dickey and Wayne A Fuller. 1979. Distribution of the estimators for autoregressive time series with a unit root. Journal of the American statistical association 74, 366a (1979), 427–431
1979
-
[23]
Chang Dong, Jianfeng Tao, Qun Chao, Honggan Yu, and Chengliang Liu. 2023. Subsequence time series clustering- based unsupervised approach for anomaly detection of axial pis- ton pumps. IEEE Transactions on Instrumentation and Mea- surement 72 (2023), 1–12
2023
-
[24]
Quang-Huy Duong, Heri Ramampiaro, and Kjetil Nørvåg
-
[25]
Peter Ebbes, Dominik Papies, and Harald J Van Heerde. 2011. The sense and non-sense of holdout sample validation in the presence of endogeneity.Marketing Science 30, 6 (2011), 1115– 1122
2011
-
[26]
Donovan Fuqua and Steven Hespeler. 2022. Commodity de- mand forecasting using modulated rank reduction for humani- tarian logistics planning. Expert Systems with Applications 206 (2022), 117753
2022
-
[28]
Joao Gama, Raquel Sebastiao, and Pedro Pereira Rodrigues
-
[29]
Jo ˜ao Gama, Indr ˙e ˇZliobait˙e, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia. 2014. A survey on concept drift adaptation. ACM computing surveys (CSUR) 46, 4 (2014), 1– 37
2014
-
[30]
Alexander J Gates and Yong-Yeol Ahn. 2019. CluSim: a python package for calculating clustering similarity. Journal of Open Source Software 4, 35 (2019), 1264
2019
-
[31]
Heitor M Gomes, Albert Bifet, Jesse Read, Jean Paul Barddal, Fabr´ıcio Enembreck, Bernhard Pfharinger, Geo ff Holmes, and Talel Abdessalem. 2017. Adaptive random forests for evolving data stream classification. Machine Learning 106 (2017), 1469– 1495
2017
-
[32]
James D Hamilton. 2020. Time series analysis. Princeton uni- versity press, Princeton, New Jeresy
2020
-
[33]
Trevor Hastie, Robert Tibshirani, Jerome Friedman, Trevor Hastie, Robert Tibshirani, and Jerome Friedman. 2009. Ran- dom forests. The elements of statistical learning: Data mining, inference, and prediction (2009), 587–604
2009
-
[34]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
-
[35]
Steven Hespeler, Ehsan Dehghan-Niri, Michael Juhasz, Kevin Luo, and Harold S Halliday. 2022. Deep learning for in-situ layer quality monitoring during laser-based directed energy de- position (LB-DED) additive manufacturing process. Applied Sciences 12, 18 (2022), 8974
2022
-
[36]
Steven Hespeler and Donovan Fuqua. 2021. Online state of charge prediction in next generation vehicle batteries using deep recurrent neural networks and continuous model size control. Journal of Energy and Power Technology3, 1 (2021), 1–24
2021
-
[37]
Steven C Hespeler, Hamidreza Nemati, Nihar Masurkar, Fer- nando Alvidrez, Hamidreza Marvi, and Ehsan Dehghan-Niri
-
[38]
Sepp Hochreiter and J ¨urgen Schmidhuber. 1997. Long short- term memory. Neural computation 9, 8 (1997), 1735–1780
1997
-
[39]
Rob J Hyndman and George Athanasopoulos. 2018. Forecast- ing: principles and practice. OTexts, OTexts
2018
-
[40]
Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhas- sane Idoumghar, and Pierre-Alain Muller. 2019. Deep learning for time series classification: a review. Data mining and knowl- edge discovery 33, 4 (2019), 917–963
2019
-
[41]
Hassan Ismail Fawaz, Benjamin Lucas, Germain Forestier, Charlotte Pelletier, Daniel F Schmidt, Jonathan Weber, Geof- frey I Webb, Lhassane Idoumghar, Pierre-Alain Muller, and Franc ¸ois Petitjean. 2020. Inceptiontime: Finding alexnet for time series classification. Data Mining...
2020
-
[42]
Yonghyeok Ji and Hyeongcheol Lee. 2022. Event-based anomaly detection using a one-class SVM for a hybrid elec- tric vehicle. IEEE Transactions on Vehicular Technology 71, 6 (2022), 6032–6043
2022
-
[43]
Gaoxia Jiang and Wenjian Wang. 2017. Markov cross-validation for time series model evaluations. Information Sciences 375 (2017), 219–233
2017
-
[44]
Fazle Karim, Somshubra Majumdar, Houshang Darabi, and Shun Chen. 2017. LSTM fully convolutional networks for time series classification. IEEE access 6 (2017), 1662–1669
2017
-
[45]
Fazle Karim, Somshubra Majumdar, Houshang Darabi, and Samuel Harford. 2019. Multivariate LSTM-FCNs for time se- ries classification. Neural networks 116 (2019), 237–245
2019
-
[46]
Ron Kohavi et al. 1995. A study of cross-validation and boot- strap for accuracy estimation and model selection. In Ijcai, V ol. 14. Montreal, Canada, IJCAI, Montreal, Quebec, Canada, 1137–1145
1995
-
[47]
Denis Kwiatkowski, Peter CB Phillips, Peter Schmidt, and Yongcheol Shin. 1992. Testing the null hypothesis of station- arity against the alternative of a unit root: How sure are we that economic time series have a unit root? Journal of econometrics 54, 1-3 (1992), 159–178
1992
-
[48]
Jeongsu Lee, Young Chul Lee, and Jeong Tae Kim. 2020. Fault detection based on one-class deep learning for manufacturing applications limited to an imbalanced database.Journal of Man- ufacturing Systems 57 (2020), 357–366
2020
-
[49]
Jinbo Li, Witold Pedrycz, and Iqbal Jamal. 2017. Multivariate 20 time series anomaly detection: A framework of Hidden Markov Models. Applied Soft Computing 60 (2017), 229–240
2017
-
[50]
Haoran Liang, Lei Song, Jianxing Wang, Lili Guo, Xuzhi Li, and Ji Liang. 2021. Robust unsupervised anomaly detection via multi-time scale DCGANs with forgetting mechanism for in- dustrial multivariate time series. Neurocomputing 423 (2021), 444–462
2021
-
[51]
Quangao Liu, Ruiqi Li, Maowei Jiang, Wei Yang, Chen Liang, Longlong Pang, and Zhuozhang Zou. 2025. Mstvi: Multi-scale Time-Variable Interaction for multivariate time series forecast- ing. Knowledge-Based Systems (2025), 113551
2025
-
[52]
Pierre-Xavier Loe ffel. 2017. Adaptive machine learning algo- rithms for data streams subject to concept drifts. Ph. D. Disser- tation. Universit´e Pierre et Marie Curie-Paris VI
2017
-
[53]
Helmut L ¨utkepohl. 2005. New introduction to multiple time se- ries analysis. Springer Science & Business Media, Berlin, Ger- many
2005
-
[54]
Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics 18, 1 (1947), 50–60
1947
-
[55]
Pablo Moriano, Steven C Hespeler, Mingyan Li, and Robert A Bridges. 2024. Benchmarking Unsupervised Online IDS for Masquerade Attacks in CAN. arXiv preprint arXiv:2406.13778 (2024)
2024
-
[56]
Pablo Moriano Salazar, Robert Bridges, and Michael Iannacone
-
[57]
Kenniy Olorunnimbe and Herna Viktor. 2023. Deep learning in the stock market—a systematic survey of practice, backtesting, and applications. Artificial Intelligence Review 56, 3 (2023), 2057–2109
2023
-
[58]
Vasilis Papastefanopoulos, Pantelis Linardatos, Theodor Pana- giotakopoulos, and Sotiris Kotsiantis. 2023. Multivariate time- series forecasting: A review of deep learning methods in internet of things applications to smart cities. Smart Cities 6, 5 (2023), 2519–2552
2023
-
[59]
Kostas Patroumpas and Timos Sellis. 2006. Window specifica- tion over data streams. In International Conference on Extend- ing Database Technology . Springer, Springer, Berlin, Heidel- berg, 445–464
2006
-
[60]
Dragutin Petkovic, Russ Altman, Mike Wong, and Arthur Vigil
-
[61]
Patr ´ıcia Ramos and Jos ´e Manuel Oliveira. 2016. A procedure for identification of appropriate state space and ARIMA models based on time-series cross-validation. Algorithms 9, 4 (2016), 76
2016
-
[62]
David R Roberts, V olker Bahn, Simone Ciuti, Mark S Boyce, Jane Elith, Gurutzeta Guillera-Arroita, Severin Hauenstein, Jos´e J Lahoz-Monfort, Boris Schr ¨oder, Wilfried Thuiller, et al
-
[63]
Alejandro Pasos Ruiz, Michael Flynn, and Anthony Bagnall
-
[64]
Alejandro Pasos Ruiz, Michael Flynn, James Large, Matthew Middlehurst, and Anthony Bagnall. 2021. The great multivariate time series classification bake o ff: a review and experimental evaluation of recent algorithmic advances. Data Mining and Knowledge Discovery 35, 2 (2021), 401–449
2021
-
[65]
Mulyana Saripuddin, Azizah Suliman, Sera Syarmila Sameon, and Bo Norregaard Jorgensen. 2021. Random undersampling on imbalance time series data for anomaly detection. In Pro- ceedings of the 2021 4th International Conference on Machine Learning and Machine Intelligence. Associ...
2021
-
[66]
Robert P Sheridan. 2013. Time-split cross-validation as a method for estimating the goodness of prospective prediction. Journal of chemical information and modeling 53, 4 (2013), 783–790
2013
-
[67]
Richard Simon. 2007. Resampling strategies for model assess- ment and selection. InFundamentals of data mining in genomics and proteomics. Springer, New York, NY , 173–186
2007
-
[68]
In Pacific Symposium on Biocomputing
Improving the explainability of Random Forest classifier– user centered approach. In Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing , V ol. 23. NIH Public Ac- cess, World Scientific Publishing Company, Singapore, 204
-
[69]
Ya Su, Youjian Zhao, Chenhao Niu, Rong Liu, Wei Sun, and Dan Pei. 2019. Robust anomaly detection for multivariate time series through stochastic recurrent neural network. In Proceed- ings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. Ass...
2019
-
[70]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Van- houcke, and Andrew Rabinovich. 2015. Going deeper with con- volutions. In Proceedings of the IEEE conference on computer vision and pattern recognition. IEEE C...
2015
-
[71]
Pengfei Tang, Peijun Du, Junshi Xia, Peng Zhang, and Wei Zhang. 2021. Channel attention-based temporal convolutional network for satellite image time series classification.IEEE Geo- science and Remote Sensing Letters 19 (2021), 1–5
2021
-
[72]
Andreas Theissler. 2017. Detecting known and unknown faults in automotive systems using ensemble-based anomaly detec- tion. Knowledge-Based Systems 123 (2017), 163–173
2017
-
[73]
arXiv preprint arXiv:2007.13156 (2020)
Benchmarking multivariate time series classification al- gorithms. arXiv preprint arXiv:2007.13156 (2020)
2020 arXiv
-
[74]
Miki E Verma, Robert A Bridges, Jordan J Sosnowski, Samuel C Hollifield, and Michael D Iannacone. 2021. CAN-D: A modular four-step pipeline for comprehensively decoding controller area network data. IEEE Transactions on Vehicular Technology 70, 10 (2021), 9685–9700
2021
-
[75]
Juliane Verwiebe, Philipp M Grulich, Jonas Traub, and V olker Markl. 2023. Survey of window types for aggregation in stream processing systems. The VLDB Journal 32, 5 (2023), 985– 1011
2023
-
[76]
Longkai Wang, Yi Yang, and Yong Lei. 2023. Physical Layer Identification of Intermittent Open and Short Connection Faults in CAN Network. In 2023 42nd Chinese Control Conference (CCC). IEEE, Institute of Electrical and Electronics Engineers (IEEE), Tianjin, China, 4998–5003
2023
-
[77]
Longkai Wang, Yi Yang, Li Li, Shilun Du, and Yong Lei. 2022. Comparison study of error patterns of intermittent open and short circuit connection faults in CAN network. In 2022 34th Chinese Control and Decision Conference (CCDC) . IEEE, In- stitute of Electrical and Electronic...
2022
-
[78]
CAN Specification. 1991. Bosch. Robert Bosch GmbH, Post- fach 50 (1991), 15
1991
-
[79]
Zhiguang Wang, Weizhong Yan, and Tim Oates. 2017. Time series classification from scratch with deep neural networks: A strong baseline. In 2017 International joint conference on neu- 21 ral networks (IJCNN). IEEE, IEEE Xplore, Anchorage, Alaska, 1578–1585
2017
-
[80]
Tianming Xie, Qifa Xu, and Cuixia Jiang. 2023. Anomaly detec- tion for multivariate times series through the multi-scale convo- lutional recurrent variational autoencoder. Expert Systems with Applications 231 (2023), 120725
2023
-
[81]
Weixuan Xiong, Peng Wang, Xiaochen Sun, and Jun Wang
-
[82]
Jiehui Xu, Haixu Wu, Jianmin Wang, and Mingsheng Long
-
[83]
Miki E Verma, Robert A Bridges, Michael D Iannacone, Samuel C Hollifield, Pablo Moriano, Steven C Hespeler, Bill Kay, and Frank L Combs. 2024. A comprehensive guide to CAN IDS data and introduction of the ROAD dataset.PLoS one 19, 1 (2024), e0296879
2024
-
[84]
Ling-rui Yu, Qiu-hong Lu, and Yang Xue. 2024. DTAAD: Dual TCN-attention networks for anomaly detection in multi- variate time series data. Knowledge-Based Systems 295 (2024), 111849
2024
-
[85]
Leiming Zhang, Fan Yang, and Yong Lei. 2019. Tree-based intermittent connection fault diagnosis for controller area net- work. IEEE Transactions on Vehicular Technology68, 9 (2019), 9151–9161
2019
-
[86]
Hang Zhao, Yujing Wang, Juanyong Duan, Congrui Huang, Defu Cao, Yunhai Tong, Bixiong Xu, Jing Bai, Jie Tong, and Qi Zhang. 2020. Multivariate time-series anomaly detection via graph attention network. In 2020 IEEE international conference on data mining (ICDM). IEEE, Institute...
2020
-
[87]
Mengmeng Zhao, Haipeng Peng, Lixiang Li, and Yeqing Ren
-
[88]
Longkai Wang, Leiming Zhang, and Yong Lei. 2023. Diagnosis of intermittent connection faults for CAN networks with com- plex topology. IEEE Access 11 (2023), 52199–52213
2023
-
[89]
Hao Zhou, Ke Yu, Xuan Zhang, Guanlin Wu, and Anis Yazidi
-
[90]
Xianzhe Zhou and Arturo Del Valle. 2020. Range based con- fusion matrix for imbalanced time series classification. In 2020 6th Conference on Data Science and Machine Learning Appli- cations (CDMA). IEEE, Institute of Electrical and Electronics Engineers (IEEE), Riyadh, Saudi A...
2020
-
[92]
Knowledge-Based Sys- tems 296 (2024), 111928
SiET: Spatial information enhanced transformer for mul- tivariate time series anomaly detection. Knowledge-Based Sys- tems 296 (2024), 111928
2024
-
[95]
Qing-Song Xu and Yi-Zeng Liang. 2001. Monte Carlo cross validation. Chemometrics and Intelligent Laboratory Systems 56, 1 (2001), 1–11
2001
-
[100]
Sensors 24, 5 (2024), 1522
Graph Attention Network and Informer for Multivariate Time Series Anomaly Detection. Sensors 24, 5 (2024), 1522
2024
-
[101]
Pu Zhao, Chuan Luo, Bo Qiao, Lu Wang, Saravan Rajmohan, Qingwei Lin, and Dongmei Zhang. 2022. T-SMOTE: Temporal- oriented Synthetic Minority Oversampling Technique for Imbal- anced Time Series Classification.. In IJCAI. International Joint Conferences on Artificial Intelligenc...
2022
-
[103]
Information Sciences 610 (2022), 266–280
Contrastive autoencoder for anomaly detection in multi- variate time series. Information Sciences 610 (2022), 266–280
2022
-
[2009]
InPro- ceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining
Issues in evaluation of stream learning algorithms. InPro- ceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining. Association for Comput- ing Machinery, Paris France, 329–338
-
[2013]
On evaluating stream learning algorithms.Machine learn- ing 90 (2013), 317–346
2013
-
[2016]
In Pro- ceedings of the IEEE conference on computer vision and pat- tern recognition
Deep residual learning for image recognition. In Pro- ceedings of the IEEE conference on computer vision and pat- tern recognition. IEEE Computer Society Conference Publish- ing Services, Las Vegas, Nevada, 770–778
-
[2017]
Ecography 40, 8 (2017), 913–929
Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 40, 8 (2017), 913–929
2017
-
[2018]
Applied Intelligence 48 (2018), 4805–4823
Applying temporal dependence to detect changes in streaming data. Applied Intelligence 48 (2018), 4805–4823
2018
-
[2020]
Data Min- ing and Knowledge Discovery 34, 5 (Sept
ROCKET: exceptionally fast and accurate time series classification using random convolutional kernels. Data Min- ing and Knowledge Discovery 34, 5 (Sept. 2020), 1454–1495. https://doi.org/10.1007/s10618-020-00701-z
2020 doi
-
[2021]
arXiv preprint arXiv:2110.02642 (2021)
Anomaly transformer: Time series anomaly detection with association discrepancy. arXiv preprint arXiv:2110.02642 (2021)
2021 arXiv
-
[2022]
In Workshop on Automotive and Autonomous Vehicle Security
Detecting CAN Masquerade Attacks with Signal Clus- tering Similarity. In Workshop on Automotive and Autonomous Vehicle Security. Internet Society, San Diego, CA, 1–8
-
[2024]
Journal of Nondestructive Evaluation, Diagnos- tics and Prognostics of Engineering Systems 7, 1 (2024), 1–17
Deep Learning–Based Time-Series Classification for Robotic Inspection of Pipe Condition Using Non-Contact Ultra- sonic Testing. Journal of Nondestructive Evaluation, Diagnos- tics and Prognostics of Engineering Systems 7, 1 (2024), 1–17
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.