Pith. sign in

REVIEW 2 major objections 2 minor 27 references

Online Graph Topology Learning from Matrix-valued Time Series

T0 review · 2 major / 2 minor · reviewed 2026-05-24 · grok-4.3

Pith's one-line read Matrix-variate auto-regressive models with trend parameters enable online learning of sensor dependency graphs from streaming data.

desk verdict The paper gives a usable online matrix-variate VAR extension with built-in trend handling for streaming graph inference, but the justification for the online estimator's validity is the part that needs the most scrutiny. read the letter →

arxiv 2107.08020 v4 submitted 2021-07-16 stat.ML cs.LG

classification stat.MLcs.LG
keywords matrix-valuedtimeseriesgraphtopologylearningvectorauto-regressivemodelsonlineLassohomotopyalgorithmsGrangercausalitytrendestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper develops methods to learn the dependency graph among sensors that each record multiple features over time. It extends standard vector auto-regressive models to handle matrix-valued observations directly. Online update procedures are introduced for both low and high dimensional cases, including a Lasso approach with homotopy algorithms for rapid coefficient updates as data streams in. The models incorporate trend parameters to handle periodic patterns without needing separate detrending steps. Numerical experiments on synthetic and real data confirm that the procedures recover the graphs effectively.

What carries the argument

Matrix-variate vector auto-regressive model augmented with trend parameters, estimated via online Lasso-type estimators and homotopy algorithms.

What would settle it

If the graphs recovered by the online procedure on accumulating data diverge substantially from graphs obtained by standard batch processing on the same full dataset, the online updates would be shown invalid.

Watch

Extended reading notes

Core claim

The central claim is that matrix-variate auto-regressive models, extended with trend parameters particularly for periodic trends, can be fitted online using specialized procedures to learn the underlying graph structure representing dependencies between sensors from streaming data, with adaptations for high-dimensional settings via Lasso-type penalties and homotopy continuation.

Load-bearing premise

The data follows a matrix-variate auto-regressive process with additive trend components that can be learned jointly from streaming samples.

Editorial extensions

If this is right

  • Coefficient estimates update rapidly with each new sample without full recomputation.
  • Graph structure and periodic trends are learned simultaneously in an online setting.
  • High-dimensional cases use adaptive regularization without requiring full detrending.
  • The methods apply directly to streaming sensor data where batch detrending is impractical.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Real-time monitoring of changing dependencies in sensor networks becomes feasible without periodic batch restarts.
  • The online framework could be tested on non-periodic trends to check robustness beyond the periodic case emphasized.
  • Similar streaming updates might apply to other structured data types like tensors if the matrix extension holds.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper extends vector auto-regressive (VAR) models to matrix-variate auto-regressive models for inferring Granger-causal graphs from matrix-valued time series observed at sensor networks. It proposes online update procedures for both low- and high-dimensional regimes, including a novel Lasso-type estimator and homotopy algorithms in the high-dimensional case, together with an adaptive regularization scheme. To enable online operation without batch detrending, the models are augmented with periodic trend parameters that are estimated jointly with the graph coefficients; effectiveness is illustrated on synthetic and real data.

Significance. If the online estimators with joint trend augmentation are shown to be consistent with their batch counterparts, the work would provide a practical route to real-time topology learning from streaming multi-feature sensor data, a setting where standard detrending is infeasible. The provision of both low- and high-dimensional online algorithms plus numerical validation on real data constitutes a concrete contribution to online graphical modeling.

major comments (2)
  1. [Section describing the augmented model and online adaptation (around the statement that 'the online algorithms are-adapt] The central claim that the trend-augmented matrix-variate AR model permits valid online Lasso/homotopy updates without batch re-processing rests on an unproven assertion that the joint estimator converges to the same limit as the batch estimator when trends deviate from exact periodicity or when new samples introduce unmodeled non-stationarity. No convergence analysis or fixed-point argument is supplied for this joint estimation procedure.
  2. [High-dimensional online procedure and homotopy algorithm description] The high-dimensional Lasso-type estimator and its homotopy algorithm are introduced without an accompanying error bound or oracle inequality that accounts for the additional trend parameters; it is therefore unclear whether the regularization and homotopy steps remain statistically valid once the trend coefficients are estimated simultaneously.
minor comments (2)
  1. [Model definition] Notation for the matrix-variate AR coefficients and the trend parameters should be introduced with explicit dimensions and indexing to avoid ambiguity when the online recursions are written.
  2. [Numerical experiments] The numerical experiments section would benefit from an explicit comparison table of online versus batch estimates under controlled trend misspecification to quantify the practical impact of the augmentation.

Simulated Author's Rebuttal

2 responses · 2 unresolved

We thank the referee for the constructive comments and for recognizing the practical contributions of the online algorithms and trend augmentation. We address each major comment below.

read point-by-point responses
  1. Referee: The central claim that the trend-augmented matrix-variate AR model permits valid online Lasso/homotopy updates without batch re-processing rests on an unproven assertion that the joint estimator converges to the same limit as the batch estimator when trends deviate from exact periodicity or when new samples introduce unmodeled non-stationarity. No convergence analysis or fixed-point argument is supplied for this joint estimation procedure.

    Authors: We agree that the manuscript does not provide a formal convergence analysis or fixed-point argument for the joint estimator under deviations from exact periodicity or unmodeled non-stationarity. The online procedures are derived under the modeled periodic trend assumption to enable streaming operation without batch detrending, and their practical performance is validated empirically on synthetic data generated with periodic trends as well as real sensor data. In revision we will add an explicit statement in the model section clarifying the modeling assumptions and noting that theoretical convergence guarantees for the joint procedure are beyond the current scope. revision: partial

  2. Referee: The high-dimensional Lasso-type estimator and its homotopy algorithm are introduced without an accompanying error bound or oracle inequality that accounts for the additional trend parameters; it is therefore unclear whether the regularization and homotopy steps remain statistically valid once the trend coefficients are estimated simultaneously.

    Authors: We acknowledge that no oracle inequalities or error bounds are derived that explicitly account for the simultaneous estimation of the additional trend parameters. The adaptive regularization procedure and the adapted homotopy algorithm are presented for the augmented model, with statistical behavior assessed via simulations rather than new theoretical bounds. In the revision we will insert a brief discussion in the high-dimensional section noting this reliance on empirical validation and the effect of the trend parameters on the effective regularization. revision: partial

standing simulated objections not resolved
  • No convergence analysis or fixed-point argument is supplied for the joint trend-augmented estimator.
  • No error bounds or oracle inequalities are provided that account for the additional trend parameters in the high-dimensional estimator.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity detected

full rationale

The paper extends standard VAR models to a matrix-variate formulation, augments them with periodic trend parameters to enable online estimation, and introduces Lasso-type penalties plus homotopy algorithms for streaming updates. No equations or fitting procedures are exhibited that reduce a claimed prediction or uniqueness result to a fitted input by construction, nor do any load-bearing steps rely on self-citations whose content is itself unverified. The algorithmic adaptations are presented as independent contributions whose validity rests on external statistical assumptions rather than definitional equivalence, so the derivation chain remains self-contained.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Only the abstract is available; no explicit free parameters, axioms, or invented entities can be extracted or verified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Online Graph Topology Learning from Matrix-valued Time Series." pith.science (2026). https://pith.science/paper/2107.08020

@misc{pith2026210708020,
  author       = {Pith},
  title        = {Pith review of: Online Graph Topology Learning from Matrix-valued Time Series},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2107.08020}},
  note         = {Machine review of arXiv:2107.08020}
}
read the original abstract

The focus is on the statistical analysis of matrix-valued time series, where data is collected over a network of sensors, typically at spatial locations, over time. Each sensor records a vector of features at each time point, creating a vectorial time series for each sensor. The goal is to identify the dependency structure among these sensors and represent it with a graph. When only one feature per sensor is observed, vector auto-regressive (VAR) models are commonly used to infer Granger causality, resulting in a causal graph. The first contribution extends VAR models to matrix-variate models for the purpose of graph learning. Additionally, two online procedures are proposed for both low and high dimensions, enabling rapid updates of coefficient estimates as new samples arrive. In the high-dimensional setting, a novel Lasso-type approach is introduced, and homotopy algorithms are developed for online learning. An adaptive tuning procedure for the regularization parameter is also provided. Given that the application of auto-regressive models to data typically requires detrending, which is not feasible in an online context, the proposed AR models are augmented by incorporating trend as an additional parameter, with a particular focus on periodic trends. The online algorithms are adapted to these augmented data models, allowing for simultaneous learning of the graph and trend from streaming samples. Numerical experiments using both synthetic and real data demonstrate the effectiveness of the proposed methods.

Figures

Figures reproduced from arXiv: 2107.08020 by the authors.

Figure 1
Figure 1. Monthly climatological records of weather stations in California. On the left is the network of weather stations in California. On the right are demonstrated the vectorial observations on a certain station i, where a vector xit P IR4 of min/max/avg temperature and precipitation is recorded per time t at i, leading to 4 scalar time series. We are interested in learning a graph of weather dependency for the network. i… view at source ↗
Figure 2
Figure 2. Comparison of the Cartesian and the tensor products of graphs. The node set of both product graphs is the Cartesian product of the components’ node sets, yet follows the different adjacencies. The example is based on Sandryhaila & Moura [24, [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Matrices pUkqk as entry locators, which characterise the structure of KG. orthogonal projection onto KG and provide an explicit formula to calculate it using pUk{}Uk}Fqk in Proposition 3.2. Proposition 3.2. For a matrix A P IRNF ˆNF , its orthogonal projection onto KG is defined by ProjG pBq “ arg min MPKG }B ´ M} 2 F . (3.4) Then given the orthonormal basis Uk{}Uk}F, k P K, the projections can be calculated explici… view at source ↗
Figures from the paper (17 more)
Figure 4
Figure 4. Figure 4: Top is the stationary time series from Model [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]
Figure 5
Figure 5. Figure 5: Initial spatial graph estimates which start the online procedures. True AN (left), AyN,91 of the low-dimensional procedure (middle), and ANp20, 0.05q of the high-dimensional procedure (right) are represented by heatmaps. Simulation settings: N “ 10, F “ 4, number of mo…
Figure 6
Figure 6. Figure 6: Initial feature graph estimates which start the online procedures. True AF (left), AyF,91 of the low-dimensional procedure (middle), and AFp20, 0.05q of the high-dimensional procedure (right). Simulation settings: N “ 10, F “ 4, number of model parameters = 571. Figure…
Figure 7
Figure 7. Figure 7: Spatial graph estimated at the arrival of the 182-th sample. True AN (left), AyN,182 of the low-dimensional procedure (middle), and ANp182, 0.0286q of the high-dimensional procedure (right) are represented by heatmaps. Simulation settings: N “ 10, F “ 4, number of mode…
Figure 8
Figure 8. Figure 8: Spatial graph estimated at the arrival of the 591-th sample. True AN (left), AyN,591 of the low-dimensional procedure (middle), and ANp591, 0.0130q of the high-dimensional procedure (right) are represented by heatmaps. Simulation settings: N “ 10, F “ 4, number of mode…
Figure 9
Figure 9. Figure 9: Trend of the first node, first feature, estimated at different times. Estimation at t “ 182 (left), t “ 273 (middle), t “ 591 (right). Simulation settings: N “ 10, F “ 4, M “ 12, number of model parameters = 571. the regularization parameter of the Lasso estimator towa…
Figure 10
Figure 10. Figure 10: Root mean square deviation. The red curves are the mean RMSD of the high￾dimensional procedure, taken over 10 simulations each. The blue curve is the mean RMSD of the low-dimensional procedure, taken over the same 30 simulations. The shaded areas represent the corresp…
Figure 11
Figure 11. Figure 11: Average one step prediction error. The red curves are the mean prediction error of the high-dimensional procedure, taken over 10 simulations each. The blue curve is the mean prediction error of the low-dimensional procedure, taken over the same 30 simulations. The sha…
Figure 12
Figure 12. Figure 12: Regularization parameter evolution. The red curves are the mean regularization parameter values, taken over 10 simulations each. The shaded areas represent the corresponding one standard deviations. Other simulation settings: N “ 20, F “ 5, M “ 12, number of model par…
Figure 13
Figure 13. Figure 13: Running time of each online update. The red curves are the mean running time of the high-dimensional procedure, taken over 10 simulations each. The blue curve is the mean running time of the low-dimensional procedure, taken over the same 30 simulations. The shaded are…
Figure 14
Figure 14. Figure 14: Updated spatial graph by the low-dimensional procedure at different times. t “ 507 (left), t “ 1015 (middle), and t “ 1522 (right). Experiment settings: N “ 27, F “ 4, M “ 12, number of model parameters = 1761, significance level of χ 2 test “ 0.1, η “ 10´5 , t0 “ 20,…
Figure 15
Figure 15. Figure 15: Updated spatial graph by the high-dimensional procedure at different times. t “ 507 (left), t “ 1015 (middle), and t “ 1522 (right). Experiment settings: N “ 27, F “ 4, M “ 12, number of model parameters = 1761, significance level of χ 2 test “ 0.1, η “ 10´5 , t0 “ 20…
Figure 16
Figure 16. Figure 16: Regularization parameter evolution. Experiment settings: N “ 27, F “ 4, M “ 12, number of model parameters = 1761, significance level of χ 2 test “ 0.1, η “ 10´5 , t0 “ 20, λ0 “ 0.03. Next we show the last updated feature graphs in [PITH_FULL_IMAGE:figures/full_fig_p…
Figure 17
Figure 17. Figure 17: Average one step prediction error of raw time series. the low-dimensional procedure (top), and the high-dimensional procedure (bottom). tmin tmax tavg prcp tmin tmax tavg prcp 0.08 0.00 0.08 0.16 0.24 3.0 1.5 0.0 1.5 [PITH_FULL_IMAGE:figures/full_fig_p037_17.png]
Figure 18
Figure 18. Figure 18: Updated feature graph at t “ 1522. the low-dimensional procedure (left), and the high-dimensional procedure (right). Experiment settings: N “ 27, F “ 4, M “ 12, number of model parameters = 1761, significance level of χ 2 test “ 0.1, η “ 10´5 , t0 “ 20, λ0 “ 0.03. 6 C…
Figure 19
Figure 19. Figure 19: Estimated trends along years. On the left, middle, right are the estimated trends at different years of Station USH00040693 for minimal temperature, average temperature, and precipitation respectively. Experiment settings: N “ 27, F “ 4, M “ 12. 040693 041715 041912 0…
Figure 20
Figure 20. Figure 20: Overlap spatial graph. On the left is the adjacency matrix of an unweighed undirected graph which is the overlap of the two last updated spatial graphs in [PITH_FULL_IMAGE:figures/full_fig_p038_20.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

27 extracted references · 27 canonical work pages

  1. [1]

    Bach, F. R. and Jordan, M. I. Learning graphical models for stationary time series.IEEE transactions on signal processing, 52(8):2189–2199, 2004

  2. [2]

    and Teboulle, M

    Beck, A. and Teboulle, M. A fast iterative shrinkage-thresholding algorithm for linear inverse problems. SIAM journal on imaging sciences, 2(1):183–202, 2009

  3. [3]

    D., and Nowak, R

    Bolstad, A., Van Veen, B. D., and Nowak, R. Causal network inference via group sparse regularization. IEEE transactions on signal processing, 59(6):2628–2641, 2011

  4. [4]

    V., Chai, K., and Williams, C

    Bonilla, E. V., Chai, K., and Williams, C. Multi-task gaussian process prediction.Advances in neural information processing systems, 20, 2007

  5. [5]

    Autoregressive models for matrix-valued time series

    Chen, R., Xiao, H., and Yang, D. Autoregressive models for matrix-valued time series. Journal of Econometrics, 222(1):539–560, 2021

  6. [6]

    and Chen, X

    Chen, S. and Chen, X. Weak connectedness of tensor product of digraphs.Discrete Applied Mathematics, 185:52–58, 2015

  7. [7]

    Learning graphs from data: A signal representation perspective

    Dong, X., Thanou, D., Rabbat, M., and Frossard, P. Learning graphs from data: A signal representation perspective. IEEE Signal Processing Magazine, 36(3):44–63, 2019

  8. [8]

    Least angle regression.Annals of statistics, 32(2):407–499, 2004

    Efron, B., Hastie, T., Johnstone, I., Tibshirani, R., et al. Least angle regression.Annals of statistics, 32(2):407–499, 2004

Show all 27 references
  1. [9]

    Sparse inverse covariance estimation with the graphical lasso

    Friedman, J., Hastie, T., and Tibshirani, R. Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3):432–441, 2008

  2. [10]

    Regularization paths for generalized linear models via coordinate descent.Journal of statistical software, 33(1):1, 2010

    Friedman, J., Hastie, T., and Tibshirani, R. Regularization paths for generalized linear models via coordinate descent.Journal of statistical software, 33(1):1, 2010. 39

  3. [11]

    and Ghaoui, L

    Garrigues, P. and Ghaoui, L. An homotopy algorithm for the lasso with online observations. Advances in neural information processing systems, 21:489–496, 2008

  4. [12]

    Tensor graphical lasso (teralasso).Journal of the Royal Statistical Society: Series B (Statistical Methodology), 81(5):901–931, 2019

    Greenewald, K., Zhou, S., and Hero III, A. Tensor graphical lasso (teralasso).Journal of the Royal Statistical Society: Series B (Statistical Methodology), 81(5):901–931, 2019

  5. [13]

    H., Imrich, W., Klavžar, S., Imrich, W., and Klavžar, S.Handbook of product graphs, volume 2

    Hammack, R. H., Imrich, W., Klavžar, S., Imrich, W., and Klavžar, S.Handbook of product graphs, volume 2. CRC press Boca Raton, 2011

  6. [14]

    H., and Friedman, J

    Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H.The elements of statistical learning: data mining, inference, and prediction, volume 2. Springer, 2009

  7. [15]

    and Peterin, I

    Imrich, W. and Peterin, I. Cartesian products of directed graphs with loops.Discrete Mathematics, 341(5):1336–1343, 2018

  8. [16]

    D., and Zhou, S

    Kalaitzis, A., Lafferty, J., Lawrence, N. D., and Zhou, S. The bigraphical lasso. In International Conference on Machine Learning, pp. 1229–1237. PMLR, 2013

  9. [17]

    Springer Science & Business Media, 2005

    Lütkepohl, H.New introduction to multiple time series analysis. Springer Science & Business Media, 2005

  10. [18]

    M., Cetin, M., and Willsky, A

    Malioutov, D. M., Cetin, M., and Willsky, A. S. Homotopy continuation for sparse signal representation. InProceedings.(ICASSP’05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005., volume 5, pp. v–733. IEEE, 2005

  11. [19]

    and Moura, J

    Mei, J. and Moura, J. M. Signal processing on graphs: Causal modeling of unstructured data. IEEE Transactions on Signal Processing, 65(8):2077–2092, 2016

  12. [20]

    High-dimensional graphs and variable selection with the lasso

    Meinshausen, N., Bühlmann, P., et al. High-dimensional graphs and variable selection with the lasso. Annals of statistics, 34(3):1436–1462, 2006

  13. [21]

    P., Anagnostopoulos, C., and Montana, G

    Monti, R. P., Anagnostopoulos, C., and Montana, G. Adaptive regularization for lasso models in the context of nonstationary data streams.Statistical Analysis and Data Mining: The ASA Data Science Journal, 11(5):237–247, 2018

  14. [22]

    R., Presnell, B., and Turlach, B

    Osborne, M. R., Presnell, B., and Turlach, B. A. A new approach to variable selection in least squares problems.IMA journal of numerical analysis, 20(3):389–403, 2000

  15. [23]

    and Boyd, S

    Parikh, N. and Boyd, S. Proximal algorithms.Foundations and Trends in optimization, 1(3): 127–239, 2014

  16. [24]

    and Moura, J

    Sandryhaila, A. and Moura, J. M. Big data analysis with signal processing on graphs: Representation and processing of massive data sets with irregular structure.IEEE Signal Processing Magazine, 31(5):80–90, 2014. 40

  17. [25]

    and Vandenberghe, L

    Songsiri, J. and Vandenberghe, L. Topology selection in graphical models of autoregressive processes. The Journal of Machine Learning Research, 11:2671–2705, 2010

  18. [26]

    The sylvester graphical lasso (syglasso)

    Wang, Y., Jang, B., and Hero, A. The sylvester graphical lasso (syglasso). InInternational Conference on Artificial Intelligence and Statistics, pp. 1943–1953. PMLR, 2020

  19. [27]

    ř k,k1PKNvecpUkqJ pΣols,tvecpUk1q ` svecpEkqsvecpEk1qJ˘ is the consistent estimator of ΣN. Then by continuous mapping theorem: t ˆαJ t

    Zaman, B., Ramos, L. M. L., Romero, D., and Beferull-Lozano, B. Online topology identification from vector autoregressive time series.IEEE Transactions on Signal Processing, 69:210–225, 2020. A Proof of Results in Section 3.2 and the CLT forpAt Proof of Theorem 3.3.By Cramér-Wo...

Pith tools

Reviewed May 24, 2026 · model on record in the stance chip above.