REVIEW 3 major objections 4 minor 28 references
Learning Control for LQR with Unknown Packet Loss Rate Using Finite Channel Samples
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A finite-sample error threshold decides whether the learned LQR controller stabilizes a lossy channel.
desk verdict The scalar analysis and optimality-gap identity are solid, but the n-dimensional stability threshold and sample-complexity bounds rest on an unjustified matrix inequality plus a missing nonsingularity assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the modified Riccati equation for LQR over a lossy channel, $P=Q+A^\top P A-(1-q)A^\top P B(R+B^\top P B)^{-1}B^\top P A$, whose positive-definite solution $P$ defines the optimal gain $K=-(R+B^\top P B)^{-1}B^\top P A$. The certainty-equivalence controller uses the same two equations with the estimated rate $\hat{q}$ in place of $q$, producing $\hat{P}$ and $\hat{K}$. Stability is certified by the Lyapunov function $\hat{V}(x_t)=\mathbb{E}[x_t^\top \hat{P}x_t]$; the paper shows that it strictly decreases exactly when $Q+(1-q)\hat{K}^\top R\hat{K}-(q-\hat{q})A^\top \hat{P}B(R+B^\top \hat{P}B)^{-1}B^\top \hat{P}A>0$ holds (Theorem 2 for $n$ dimensions, Theorem 1 in the scalar case). The threshold and sample-complexity bounds are obtained by controlling this condition with the monotonicity of the Riccati map ($\hat{P}\le P$ when $\hat{q}<q$) and with Hoeffding's inequality connecting the sample count to the estimation error. Theorem 7 replaces the unknown $q$ by a small semidefinite program whose feasible point certifies stabilization.
What would settle it
Take a stabilizable pair with singular $A$, e.g. $A=0$, $B=I$ in dimension two with $Q,R>0$: the expressions in Theorems 3 and 5 contain $(A^\top P^2 A)^{-1}$ and $(A^\top P A)^{-1}$, which do not exist, so the paper's general lower bounds cannot even be evaluated. A concrete numerical search on such a system, comparing the true stabilizing region with the claimed lower bound, would settle whether the general theorem's domain must be restricted to nonsingular $A$.
Extended reading notes
Core claim
The paper claims that for a discrete-time linear time-invariant system controlled over a Bernoulli packet-loss channel, the certainty-equivalence LQR gain $\hat{K}=-(R+B^\top \hat{P}B)^{-1}B^\top \hat{P}A$ is mean-square stabilizing at least when the estimation error $q-\hat{q}$ is below a stability threshold; the threshold is necessary and sufficient in the scalar case, and the paper gives sufficient explicit lower bounds for general systems (Theorem 3) and sharper ones for scalar systems and for systems with invertible input matrix (Theorems 4 and 5). It further claims that the sample complexity for stabilization with probability at least $1-\beta$ is at most $c_1^2 \log(2/\beta)/(2\lambda_{\min}\{Q^{1/2}(A^\top P^2 A)^{-1}Q^{1/2}\}^2)$. The closing performance claim is the exact optimality-gap identity $J(x_0,\hat{u})-J^*(x_0)=\operatorname{tr}\{(q-\hat{q})X_{\hat{K}}+(\hat{P}-P)X_0\}$, which makes the gap vanish as $\hat{q}\to q$.
Load-bearing premise
The general n-dimensional theorems require the state matrix $A$ to be invertible because their formulas contain $(A^\top P^2 A)^{-1}$ and $(A^\top P A)^{-1}$; the paper never states this, and stabilizability with an invertible $B$ does not force $A$ to be invertible.
Editorial extensions
If this is right
- If the estimate overestimates the loss rate ($\hat{q}\ge q$), the certainty-equivalence controller always mean-square stabilizes the system, though at a larger cost.
- The allowable estimation error shrinks as $q$ approaches the critical loss rate $q_c$; consequently the sample complexity grows without bound near $q_c$.
- For sufficiently small true loss rates, any estimate in $[0,q_c)$ stabilizes the system, so no channel samples are needed at all.
- The optimality gap between the learned controller and the true optimal controller is bounded linearly by the estimation error or by the matrix difference $\hat{P}-P$, and the gap tends to zero as the estimate approaches the true rate.
- Theorem 7 gives a check that depends only on the sample estimate and the confidence parameter, not on the unknown $q$, so an operator can certify stabilization from the data actually available.
Reading between the lines
- Editorial extension: the same Lyapunov decrement condition suggests an online variant in which $\hat{q}$ is updated sample by sample; the paper only treats a fixed batch estimate, but the stability condition is stated for whatever $\hat{q}$ is used.
- Editorial extension: the scalar necessary-and-sufficient condition is a natural benchmark; dimension-by-dimension or structure-exploiting bounds could shrink the gap between the sufficient lower bounds and the true stability threshold in $n>1$.
- Editorial extension: the optimality-gap identity separates the cost penalty into a term proportional to $q-\hat{q}$ and a term proportional to $\hat{P}-P$, which suggests allocating samples where that penalty is largest, i.e. near $q_c$.
- Editorial extension: since the general formulas assume $A$ invertible, the natural boundary case $A=0, B=I$ is not covered; checking whether the threshold admits a finite limit as $A$ becomes singular would test whether the theorem's domain is genuinely all stabilizable pairs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies infinite-horizon LQR for discrete-time linear systems over a Bernoulli packet-loss channel with unknown loss probability q. The loss rate is estimated from Nq i.i.d. channel samples, and a certainty-equivalence controller is synthesized from the estimate q̂ through the modified Riccati equation. The authors give a Lyapunov sufficient condition for mean-square stabilization (Theorem 2), a necessary and sufficient scalar condition (Theorem 1), explicit lower bounds on the admissible estimation error q−q̂ for general, scalar, and invertible-B systems (Theorems 3–5), corresponding sample-complexity upper bounds (Theorem 6 and Eqs. (24)–(25)), a q-independent verification condition (Theorem 7), and an exact optimality-gap identity (Theorem 8). Numerical simulations illustrate the conservatism of the bounds.
Significance. If the technical gaps identified below are repaired, the paper would be a useful contribution: it gives explicit finite-sample stabilization guarantees for certainty-equivalence LQR over unknown Bernoulli loss channels, going beyond the existing confidence-interval and worst-case analyses. The scalar characterization and the Lyapunov condition in Theorem 2 are sound, and Theorem 8's exact optimality-gap identity is clean and checkable. The paper contains no fitted constants and is transparent about the conservative, sufficient nature of the general conditions. However, two load-bearing steps in the general and invertible-B results are not justified as written, so those advertised thresholds and sample-complexity bounds are not yet established.
major comments (3)
- [Theorem 3, Eq. (16)] The general n-dimensional results require A to be nonsingular, but this is never stated. Equation (16) contains (A^T P^2 A)^{-1}, and Theorem 5's Eq. (20), Theorem 6's Eq. (23), and Eq. (25) contain (A^T P A)^{-1}; these are undefined when A is singular. Stabilizability plus invertibility of B does not imply invertibility of A: A=0, B=I is stabilizable and satisfies the standing assumptions, yet the displayed formulas are undefined. The statements should either assume A invertible explicitly or provide a meaningful reformulation for singular A.
- [Theorem 3, proof of Eq. (18)] The proof uses the step A^T P̂^2 A ≤ A^T P^2 A as a consequence of P̂ ≤ P. For positive semidefinite matrices, X ≤ Y does not imply X^2 ≤ Y^2 unless additional structure such as commutativity holds; for example X=diag(1,0.1) and Y=[[5,2],[2,1.1]] satisfy X ≤ Y, but Y^2−X^2 has a negative eigenvalue. The paper gives no Riccati-specific argument that P and P̂ are exceptional. Since this is exactly the comparison that converts Q−(q−q̂)c1 A^T P^2 A > 0 into condition (13), the stability threshold (16), Corollary 1, and the sample-complexity bound (23) are unsupported as written.
- [Theorem 5, Eq. (21)] The same square-monotonicity defect appears in the proof of the invertible-B result: the lower bound Q+(1−q)c2 A^T P0^2 A is reached from A^T P̂^2 A ≥ A^T P0^2 A, inferred from P0 ≤ P̂. Again, PSD order does not imply order of squares for noncommuting matrices, and no commutativity or other Riccati-specific property is established. Additionally, the denominator in the first inequality of (21) is written λmax(R+B P B^T), while c2 is defined with λmax(R+B^T P B); these are not equal in general (e.g., B=[[2,1],[0,1]], R=diag(1,100), P=I). The proof should use a consistent, valid denominator, and the square-monotonicity step needs a rigorous justification.
minor comments (4)
- [Notation] The notation section says "spectral radium"; this should read "spectral radius".
- [Assumption 1, Eq. (5)] The displayed bound on qc is hard to read: "1Q_i |λu_i(A)|^2" should be typeset as 1/∏_i |λu_i(A)|^2. Please clarify the product notation.
- [Proof of Theorem 3] The proof states "we know P0 ≤ P̂ ≤ P from Lemma 1", but Lemma 1 only states P̂ ≤ P under q̂ < q. The inequality P0 ≤ P̂ follows by applying the same monotonicity argument with loss rate 0 and q̂, but this should be stated explicitly.
- [Remark 4] The scalar comparison in Remark 4 is difficult to parse because of missing parentheses and line breaks; please rewrite it so the displayed inequality can be verified line by line.
Circularity Check
No significant circularity: the stability threshold, sample-complexity bound, and optimality-gap identity are derived from an explicit Lyapunov argument, Hoeffding's inequality, and algebraic manipulation rather than from the results they claim to establish.
full rationale
The paper's derivation chain is self-contained. The packet-loss rate q is estimated by the sample average q̂, and the certainty-equivalence controller is defined directly from the estimated Riccati equation (8); no parameter is fitted to the stability or cost conclusions. The stability threshold lower bound in Theorem 3 follows from the Lyapunov decrease condition (13) combined with Lemma 1, which cites the external Riccati monotonicity result of Sinopoli et al. The tailored thresholds in Theorems 4 and 5 are obtained by algebraic manipulation of the scalar condition (9) and a matrix inequality based on the Riccati comparison. The sample-complexity results in Lemma 3 and Theorem 6 are direct compositions of Hoeffding's inequality with the previously derived threshold, not circular restatements of the desired stabilization claim. The optimality-gap identity (28) in Theorem 8 is derived by telescoping the coupled Riccati cost recursion, and its limit to zero uses the continuity lemma for P̂ in q̂, which is justified by monotone convergence and uniqueness of the Riccati solution. There is no fitted input renamed as a prediction, no target quantity defined in terms of the derived bound, and no load-bearing self-citation: the only self-referential citation is [26] for the standard sample-average estimator, and it is not used to justify any central theorem. Potential technical gaps, such as the need for A to be nonsingular in expressions like (A^T P^2 A)^{-1} and the unjustified matrix step A^T P̂^2 A ≤ A^T P^2 A from P̂ ≤ P, are correctness risks rather than circularity, because they do not make the outputs equivalent to the inputs by construction.
Assumptions & free parameters
assumptions (5)
- domain assumption (A,B) is stabilizable and q < q_c, where q_c satisfies the bounds in Eq. (5).
- domain assumption The estimate satisfies \hat q < q_c.
- ad hoc to paper A is nonsingular in Theorems 3 and 5.
- standard math Hoeffding's inequality for Bernoulli samples.
- standard math Monotonicity of Riccati-map iterates with respect to the loss rate (Lemma 1 in [9]).
Cite this review
Pith. "Pith review of Learning Control for LQR with Unknown Packet Loss Rate Using Finite Channel Samples." pith.science (2026). https://pith.science/paper/PLO4ZOWD
@misc{pith2026250102899,
author = {Pith},
title = {Pith review of: Learning Control for LQR with Unknown Packet Loss Rate Using Finite Channel Samples},
year = {2026},
howpublished = {\url{https://pith.science/paper/PLO4ZOWD}},
note = {Machine review of arXiv:2501.02899}
}
read the original abstract
This paper studies the linear quadratic regulator (LQR) problem over an unknown Bernoulli packet loss channel. The unknown loss rate is estimated using finite channel samples and a certainty-equivalence (CE) optimal controller is then designed by treating the estimate as the true rate. The stabilizing capability and sub-optimality of the CE controller critically depend on the estimation error of loss rate. For discrete-time linear systems, we provide a stability threshold for the estimation error to ensure closed-loop stability, and analytically quantify the sub-optimality in terms of the estimation error and the difference in modified Riccati equations. Next, we derive the upper bound on sample complexity for the CE controller to be stabilizing. Tailored results with less conservatism are delivered for scalar systems and n-dimensional systems with invertible input matrix. Moreover, we establish a sufficient condition, independent of the unknown loss rate, to verify whether the CE controller is stabilizing in a probabilistic sense. Finally, numerical examples are used to validate our results.
Reference graph
Works this paper leans on
-
[22]
Learning to Control over Unknown Wireless Channels,
K. Gatsis and G. J. Pappas, “Learning to Control over Unknown Wireless Channels,” IFAC-PapersOnLine, vol. 53, no. 2, pp. 2600–2605, 2020, 21st IFAC World Congress
work page 2020
-
[1]
Networked control systems: A survey of trends and techniques,
X.-M. Zhang et al., “Networked control systems: A survey of trends and techniques,” IEEE/CAA J. Automatica Sinica , vol. 7, no. 1, pp. 1-17, 2019
work page 2019
-
[2]
The many facets of information in networked estimation and control,
M. Franceschetti, M. J. Khojasteh, and M. Z. Win, “The many facets of information in networked estimation and control,” Annu. Rev. of Control Robot. and Auton. Syst. , vol. 6, no. 1, pp. 233-259, 2023
work page 2023
-
[3]
Iterative learning control for sampled-data systems: From theory to practice,
K. Abidi, and J. Xu, “Iterative learning control for sampled-data systems: From theory to practice,” IEEE Trans. on Ind. Electron. , vol. 58, no. 7, pp. 3002-3015, 2014
work page 2014
-
[4]
A new model for autonomous, networked control systems,
G. Pratl, D. Dietrich, G. P. Hancke, and W. T. Penzhorn, “A new model for autonomous, networked control systems,” IEEE Trans. on Ind. Informat. , vol. 3, no. 1, pp. 21-32, 2007
work page 2007
-
[5]
Time-varying gain controller synthesis of piecewise homogeneous semi-Markov jump linear systems,
Y . Tian, H. Yan, H. Zhang, M. Wang, and J. Yi, “Time-varying gain controller synthesis of piecewise homogeneous semi-Markov jump linear systems,” Automatica, vol. 146, 2022, Art. no. 110594
work page 2022
-
[6]
Network-induced constraints in networked control systems—A survey,
L. Zhang, H. Gao, and O. Kaynak, “Network-induced constraints in networked control systems—A survey,” IEEE Trans. on Ind. Informat. , vol. 9, no. 1, pp. 403-416, 2013
work page 2013
-
[7]
S. Weerakkody, O. Ozel, Y . Mo, and B. Sinopoli, “Resilient control in cyber-physical systems: Countering uncertainty, constraints, and adver- sarial behavior,” Found. Trends Syst. Control, vol. 7, nos. 1-2, pp. 1-252, 2019
work page 2019
Show all 28 references
-
[8]
Limitations of linear control over packet drop networks,
N. Elia and J. N. Eisenbeis, “Limitations of linear control over packet drop networks,” IEEE Trans. Autom. Control, vol. 56, no. 4, pp. 826-841, 2011
2011
-
[9]
Kalman Filtering With Intermittent Observations,
B. Sinopoli, L. Schenato, M. Franceschetti, K. Poolla, M. I. Jordan, and S. S. Sastry, “Kalman Filtering With Intermittent Observations,” IEEE Trans. Automat. Contr., vol. 49, no. 9, pp. 1453–1464, Sep. 2004
2004
-
[10]
Optimal control of LTI systems over unreliable communication links,
O. C. Imer, S. Y ¨uksel, and T. Bas ¸ar, “Optimal control of LTI systems over unreliable communication links,” Automatica, vol. 42, no. 9, pp. 1429–1439, Sep. 2006
2006
-
[11]
Foundations of Control and Estimation Over Lossy Networks,
L. Schenato, B. Sinopoli, M. Franceschetti, K. Poolla, and S. S. Sastry, “Foundations of Control and Estimation Over Lossy Networks,” Proc. IEEE, vol. 95, no. 1, pp. 163–187, Jan. 2007
2007
-
[12]
Stabilization of linear systems over networks with bounded packet loss,
J. Xiong, and J. Lam, “Stabilization of linear systems over networks with bounded packet loss,” Automatica vol. 43, no. 1,pp. 80-87, Jan. 2007
2007
-
[13]
Modelling and control of networked control sys- tems with both network-induced delay and packet-dropout,
W. Zhang, and L. Yu, “Modelling and control of networked control sys- tems with both network-induced delay and packet-dropout,” Automatica vol. 44, no.12, pp. 3206-3210, 2008
2008
-
[14]
Mean square stabilization for sampled-data T–S fuzzy systems with random packet dropout,
Z. Hu and X. Mu, “Mean square stabilization for sampled-data T–S fuzzy systems with random packet dropout,” IEEE Trans. Fuzzy Syst. , vol. 28, no. 8, pp. 1815-1824, 2020
2020
-
[15]
V . K. N. Lau and Y . K. Kwok. Channel-Adaptive Technologies and Cross-Layer Designs for Wireless Systems with Multiple Antennas - Theory and Applications, 1st edition. Wiley John Proakis Telecom Series, 2005
2005
-
[16]
Learning optimal scheduling policy for remote state estimation under uncertain channel condition,
S. Wu, X. Ren, Q. Jia, K. H. Johansson, and L. Shi, “Learning optimal scheduling policy for remote state estimation under uncertain channel condition,” IEEE Trans. Control Netw. Syst. , vol. 7, no. 2, pp. 579–591, Jun. 2020
2020
-
[17]
Matiakis, S
T. Matiakis, S. Hirche, and M. Buss. Local and remote control mea- sures for networked control systems. In Proceedings of the 47th IEEE International Conference on Decision and Control, CDC ’08
-
[18]
Learning in wireless control systems over nonstationary channels,
M. Eisen, K. Gatsis, G. J. Pappas, and A. Ribeiro, “Learning in wireless control systems over nonstationary channels,”IEEE Trans. Signal Process., vol. 67, no. 5, pp. 1123–1137, Mar. 2019
2019
-
[19]
Thompson sampling for networked control over unknown channels,
W. Liu, A. S. Leong, and D. E. Quevedo. “Thompson sampling for networked control over unknown channels,” Automatica, vol. 165, 2024, Art. no. 111684
2024
-
[20]
Sample complexity of networked control systems over unknown channels,
K. Gatsis and G. J. Pappas, “Sample complexity of networked control systems over unknown channels,” in Proc. Conf. Decis. Control, 2018, pp. 6067–6072
2018
-
[21]
Statistical learning for analysis of net- worked control systems over unknown channels,
K. Gatsis and G. J. Pappas, “Statistical learning for analysis of net- worked control systems over unknown channels,” Automatica, vol.125, Mar. 2021, Art. no. 109386
2021
-
[23]
Tsiamis, I
A. Tsiamis, I. Ziemann, N. Matni, and G. J. Pappas, ”Statistical learning theory for control: A finite-sample perspective,” IEEE Control Systems Magazine, vol. 43, no. 6, pp. 67-97, 2023
2023
-
[24]
C.,Campi and E
M. C.,Campi and E. Weyer, ”Finite sample properties of system identi- fication methods,” IEEE Trans. Autom. Control, vol. 47, no. 8, pp. 1329-
-
[25]
Data-driven distribution- ally robust LQR with multiplicative noise,
P. Coppens, M. Schuurmans, and P. Patrinos, “Data-driven distribution- ally robust LQR with multiplicative noise,” in Proc. Learn. Dyn. Control, Jul. 2020, pp.521-530
2020
-
[26]
Finite-sample-based Spec- tral Radius Estimation and Stabilizability Test for Networked Control Systems,
L. Xu, B. Guo, and G. Ferrari-Trecate, “Finite-sample-based Spec- tral Radius Estimation and Stabilizability Test for Networked Control Systems,” in Proc. 2022 European Control Conference (ECC) , London, United Kingdom: IEEE, Jul. 2022, pp. 2087–2092
2022
-
[27]
Wasserman, All of statistics: a concise course in statistical inference
L. Wasserman, All of statistics: a concise course in statistical inference. Springer Science & Business Media, 2013
2013
-
[28]
Feedback stabilization of discrete-time networked systems over fading channels,
N. Xiao, L. Xie, and L. Qiu, “Feedback stabilization of discrete-time networked systems over fading channels,” IEEE Trans. Autom. Control , vol. 57, no. 9, pp. 2176–2189, Sep. 2012
2012
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.