Pith. sign in

REVIEW 3 major objections 4 minor 28 references

Learning Control for LQR with Unknown Packet Loss Rate Using Finite Channel Samples

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A finite-sample error threshold decides whether the learned LQR controller stabilizes a lossy channel.

desk verdict The scalar analysis and optimality-gap identity are solid, but the n-dimensional stability threshold and sample-complexity bounds rest on an unjustified matrix inequality plus a missing nonsingularity assumption. read the letter →

arxiv 2501.02899 v2 pith:PLO4ZOWD submitted 2025-01-06 eess.SY cs.SY

classification eess.SYcs.SY MSC 93C5593E2093D0593B52
keywords packetlossBernoullichannellinearquadraticregulatorcertaintyequivalencesamplecomplexitymean-squarestabilityRiccatiequationnetworkedcontrolsystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Over a communication channel that drops control packets with an unknown probability $q$, the paper asks when a controller can be learned from finitely many channel samples. It establishes that the certainty-equivalence controller—the optimal LQR formula with the estimated loss rate $\hat{q}$ plugged in for $q$—keeps the closed loop mean-square stable whenever the estimation error $q-\hat{q}$ stays below a stability threshold, and it gives explicit lower bounds on that threshold in terms of the system matrices and weights. From those bounds it derives an upper bound on the number of channel samples needed for stabilization with probability at least $1-\beta$. It also proves an exact identity for the optimality gap of the learned controller, written as a linear combination of the estimation error and the difference between the two Riccati solutions. The upshot is a finite-sample, non-asymptotic guarantee for the standard estimate-then-control recipe over unreliable channels.

What carries the argument

The central object is the modified Riccati equation for LQR over a lossy channel, $P=Q+A^\top P A-(1-q)A^\top P B(R+B^\top P B)^{-1}B^\top P A$, whose positive-definite solution $P$ defines the optimal gain $K=-(R+B^\top P B)^{-1}B^\top P A$. The certainty-equivalence controller uses the same two equations with the estimated rate $\hat{q}$ in place of $q$, producing $\hat{P}$ and $\hat{K}$. Stability is certified by the Lyapunov function $\hat{V}(x_t)=\mathbb{E}[x_t^\top \hat{P}x_t]$; the paper shows that it strictly decreases exactly when $Q+(1-q)\hat{K}^\top R\hat{K}-(q-\hat{q})A^\top \hat{P}B(R+B^\top \hat{P}B)^{-1}B^\top \hat{P}A>0$ holds (Theorem 2 for $n$ dimensions, Theorem 1 in the scalar case). The threshold and sample-complexity bounds are obtained by controlling this condition with the monotonicity of the Riccati map ($\hat{P}\le P$ when $\hat{q}<q$) and with Hoeffding's inequality connecting the sample count to the estimation error. Theorem 7 replaces the unknown $q$ by a small semidefinite program whose feasible point certifies stabilization.

What would settle it

Take a stabilizable pair with singular $A$, e.g. $A=0$, $B=I$ in dimension two with $Q,R>0$: the expressions in Theorems 3 and 5 contain $(A^\top P^2 A)^{-1}$ and $(A^\top P A)^{-1}$, which do not exist, so the paper's general lower bounds cannot even be evaluated. A concrete numerical search on such a system, comparing the true stabilizing region with the claimed lower bound, would settle whether the general theorem's domain must be restricted to nonsingular $A$.

Watch

Extended reading notes

Core claim

The paper claims that for a discrete-time linear time-invariant system controlled over a Bernoulli packet-loss channel, the certainty-equivalence LQR gain $\hat{K}=-(R+B^\top \hat{P}B)^{-1}B^\top \hat{P}A$ is mean-square stabilizing at least when the estimation error $q-\hat{q}$ is below a stability threshold; the threshold is necessary and sufficient in the scalar case, and the paper gives sufficient explicit lower bounds for general systems (Theorem 3) and sharper ones for scalar systems and for systems with invertible input matrix (Theorems 4 and 5). It further claims that the sample complexity for stabilization with probability at least $1-\beta$ is at most $c_1^2 \log(2/\beta)/(2\lambda_{\min}\{Q^{1/2}(A^\top P^2 A)^{-1}Q^{1/2}\}^2)$. The closing performance claim is the exact optimality-gap identity $J(x_0,\hat{u})-J^*(x_0)=\operatorname{tr}\{(q-\hat{q})X_{\hat{K}}+(\hat{P}-P)X_0\}$, which makes the gap vanish as $\hat{q}\to q$.

Load-bearing premise

The general n-dimensional theorems require the state matrix $A$ to be invertible because their formulas contain $(A^\top P^2 A)^{-1}$ and $(A^\top P A)^{-1}$; the paper never states this, and stabilizability with an invertible $B$ does not force $A$ to be invertible.

Editorial extensions

If this is right

  • If the estimate overestimates the loss rate ($\hat{q}\ge q$), the certainty-equivalence controller always mean-square stabilizes the system, though at a larger cost.
  • The allowable estimation error shrinks as $q$ approaches the critical loss rate $q_c$; consequently the sample complexity grows without bound near $q_c$.
  • For sufficiently small true loss rates, any estimate in $[0,q_c)$ stabilizes the system, so no channel samples are needed at all.
  • The optimality gap between the learned controller and the true optimal controller is bounded linearly by the estimation error or by the matrix difference $\hat{P}-P$, and the gap tends to zero as the estimate approaches the true rate.
  • Theorem 7 gives a check that depends only on the sample estimate and the confidence parameter, not on the unknown $q$, so an operator can certify stabilization from the data actually available.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the same Lyapunov decrement condition suggests an online variant in which $\hat{q}$ is updated sample by sample; the paper only treats a fixed batch estimate, but the stability condition is stated for whatever $\hat{q}$ is used.
  • Editorial extension: the scalar necessary-and-sufficient condition is a natural benchmark; dimension-by-dimension or structure-exploiting bounds could shrink the gap between the sufficient lower bounds and the true stability threshold in $n>1$.
  • Editorial extension: the optimality-gap identity separates the cost penalty into a term proportional to $q-\hat{q}$ and a term proportional to $\hat{P}-P$, which suggests allocating samples where that penalty is largest, i.e. near $q_c$.
  • Editorial extension: since the general formulas assume $A$ invertible, the natural boundary case $A=0, B=I$ is not covered; checking whether the threshold admits a finite limit as $A$ becomes singular would test whether the theorem's domain is genuinely all stabilizable pairs.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies infinite-horizon LQR for discrete-time linear systems over a Bernoulli packet-loss channel with unknown loss probability q. The loss rate is estimated from Nq i.i.d. channel samples, and a certainty-equivalence controller is synthesized from the estimate q̂ through the modified Riccati equation. The authors give a Lyapunov sufficient condition for mean-square stabilization (Theorem 2), a necessary and sufficient scalar condition (Theorem 1), explicit lower bounds on the admissible estimation error q−q̂ for general, scalar, and invertible-B systems (Theorems 3–5), corresponding sample-complexity upper bounds (Theorem 6 and Eqs. (24)–(25)), a q-independent verification condition (Theorem 7), and an exact optimality-gap identity (Theorem 8). Numerical simulations illustrate the conservatism of the bounds.

Significance. If the technical gaps identified below are repaired, the paper would be a useful contribution: it gives explicit finite-sample stabilization guarantees for certainty-equivalence LQR over unknown Bernoulli loss channels, going beyond the existing confidence-interval and worst-case analyses. The scalar characterization and the Lyapunov condition in Theorem 2 are sound, and Theorem 8's exact optimality-gap identity is clean and checkable. The paper contains no fitted constants and is transparent about the conservative, sufficient nature of the general conditions. However, two load-bearing steps in the general and invertible-B results are not justified as written, so those advertised thresholds and sample-complexity bounds are not yet established.

major comments (3)
  1. [Theorem 3, Eq. (16)] The general n-dimensional results require A to be nonsingular, but this is never stated. Equation (16) contains (A^T P^2 A)^{-1}, and Theorem 5's Eq. (20), Theorem 6's Eq. (23), and Eq. (25) contain (A^T P A)^{-1}; these are undefined when A is singular. Stabilizability plus invertibility of B does not imply invertibility of A: A=0, B=I is stabilizable and satisfies the standing assumptions, yet the displayed formulas are undefined. The statements should either assume A invertible explicitly or provide a meaningful reformulation for singular A.
  2. [Theorem 3, proof of Eq. (18)] The proof uses the step A^T P̂^2 A ≤ A^T P^2 A as a consequence of P̂ ≤ P. For positive semidefinite matrices, X ≤ Y does not imply X^2 ≤ Y^2 unless additional structure such as commutativity holds; for example X=diag(1,0.1) and Y=[[5,2],[2,1.1]] satisfy X ≤ Y, but Y^2−X^2 has a negative eigenvalue. The paper gives no Riccati-specific argument that P and P̂ are exceptional. Since this is exactly the comparison that converts Q−(q−q̂)c1 A^T P^2 A > 0 into condition (13), the stability threshold (16), Corollary 1, and the sample-complexity bound (23) are unsupported as written.
  3. [Theorem 5, Eq. (21)] The same square-monotonicity defect appears in the proof of the invertible-B result: the lower bound Q+(1−q)c2 A^T P0^2 A is reached from A^T P̂^2 A ≥ A^T P0^2 A, inferred from P0 ≤ P̂. Again, PSD order does not imply order of squares for noncommuting matrices, and no commutativity or other Riccati-specific property is established. Additionally, the denominator in the first inequality of (21) is written λmax(R+B P B^T), while c2 is defined with λmax(R+B^T P B); these are not equal in general (e.g., B=[[2,1],[0,1]], R=diag(1,100), P=I). The proof should use a consistent, valid denominator, and the square-monotonicity step needs a rigorous justification.
minor comments (4)
  1. [Notation] The notation section says "spectral radium"; this should read "spectral radius".
  2. [Assumption 1, Eq. (5)] The displayed bound on qc is hard to read: "1Q_i |λu_i(A)|^2" should be typeset as 1/∏_i |λu_i(A)|^2. Please clarify the product notation.
  3. [Proof of Theorem 3] The proof states "we know P0 ≤ P̂ ≤ P from Lemma 1", but Lemma 1 only states P̂ ≤ P under q̂ < q. The inequality P0 ≤ P̂ follows by applying the same monotonicity argument with loss rate 0 and q̂, but this should be stated explicitly.
  4. [Remark 4] The scalar comparison in Remark 4 is difficult to parse because of missing parentheses and line breaks; please rewrite it so the displayed inequality can be verified line by line.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the stability threshold, sample-complexity bound, and optimality-gap identity are derived from an explicit Lyapunov argument, Hoeffding's inequality, and algebraic manipulation rather than from the results they claim to establish.

full rationale

The paper's derivation chain is self-contained. The packet-loss rate q is estimated by the sample average q̂, and the certainty-equivalence controller is defined directly from the estimated Riccati equation (8); no parameter is fitted to the stability or cost conclusions. The stability threshold lower bound in Theorem 3 follows from the Lyapunov decrease condition (13) combined with Lemma 1, which cites the external Riccati monotonicity result of Sinopoli et al. The tailored thresholds in Theorems 4 and 5 are obtained by algebraic manipulation of the scalar condition (9) and a matrix inequality based on the Riccati comparison. The sample-complexity results in Lemma 3 and Theorem 6 are direct compositions of Hoeffding's inequality with the previously derived threshold, not circular restatements of the desired stabilization claim. The optimality-gap identity (28) in Theorem 8 is derived by telescoping the coupled Riccati cost recursion, and its limit to zero uses the continuity lemma for P̂ in q̂, which is justified by monotone convergence and uniqueness of the Riccati solution. There is no fitted input renamed as a prediction, no target quantity defined in terms of the derived bound, and no load-bearing self-citation: the only self-referential citation is [26] for the standard sample-average estimator, and it is not used to justify any central theorem. Potential technical gaps, such as the need for A to be nonsingular in expressions like (A^T P^2 A)^{-1} and the unjustified matrix step A^T P̂^2 A ≤ A^T P^2 A from P̂ ≤ P, are correctness risks rather than circularity, because they do not make the outputs equivalent to the inputs by construction.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on existence of Riccati solutions (q < q_c, \hat q < q_c), on a monotonicity lemma from [9], and on an unstated nonsingularity of A for the general bounds. No data-fitted parameters appear; the main free objects are problem data A, B, Q, R, q and the sample estimate \hat q.

assumptions (5)
  • domain assumption (A,B) is stabilizable and q < q_c, where q_c satisfies the bounds in Eq. (5).
    Needed for the existence of the positive definite solution P of the modified Riccati equation (4); stated in Section II-A and Assumption 1.
  • domain assumption The estimate satisfies \hat q < q_c.
    The paper treats \hat q < q_c as a premise throughout; Section II-A says this is not further discussed. Needed for \hat P in Eq. (8) to exist.
  • ad hoc to paper A is nonsingular in Theorems 3 and 5.
    The bounds (16) and (20) use inverses of A^T P^2 A and A^T P A. No such assumption appears in the theorem statements, and it is not implied by stabilizability or by B invertible.
  • standard math Hoeffding's inequality for Bernoulli samples.
    Used in Lemma 2 to translate sample size into an estimation-error confidence interval; cited as Theorem 4.5 of [27].
  • standard math Monotonicity of Riccati-map iterates with respect to the loss rate (Lemma 1 in [9]).
    Used in Lemma 1 to show \hat P \le P for \hat q < q. This is an external theorem from Sinopoli et al., not proved in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning Control for LQR with Unknown Packet Loss Rate Using Finite Channel Samples." pith.science (2026). https://pith.science/paper/PLO4ZOWD

@misc{pith2026250102899,
  author       = {Pith},
  title        = {Pith review of: Learning Control for LQR with Unknown Packet Loss Rate Using Finite Channel Samples},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PLO4ZOWD}},
  note         = {Machine review of arXiv:2501.02899}
}
read the original abstract

This paper studies the linear quadratic regulator (LQR) problem over an unknown Bernoulli packet loss channel. The unknown loss rate is estimated using finite channel samples and a certainty-equivalence (CE) optimal controller is then designed by treating the estimate as the true rate. The stabilizing capability and sub-optimality of the CE controller critically depend on the estimation error of loss rate. For discrete-time linear systems, we provide a stability threshold for the estimation error to ensure closed-loop stability, and analytically quantify the sub-optimality in terms of the estimation error and the difference in modified Riccati equations. Next, we derive the upper bound on sample complexity for the CE controller to be stabilizing. Tailored results with less conservatism are delivered for scalar systems and n-dimensional systems with invertible input matrix. Moreover, we establish a sufficient condition, independent of the unknown loss rate, to verify whether the CE controller is stabilizing in a probabilistic sense. Finally, numerical examples are used to validate our results.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 28 canonical work pages

  1. [22]

    Learning to Control over Unknown Wireless Channels,

    K. Gatsis and G. J. Pappas, “Learning to Control over Unknown Wireless Channels,” IFAC-PapersOnLine, vol. 53, no. 2, pp. 2600–2605, 2020, 21st IFAC World Congress

  2. [1]

    Networked control systems: A survey of trends and techniques,

    X.-M. Zhang et al., “Networked control systems: A survey of trends and techniques,” IEEE/CAA J. Automatica Sinica , vol. 7, no. 1, pp. 1-17, 2019

  3. [2]

    The many facets of information in networked estimation and control,

    M. Franceschetti, M. J. Khojasteh, and M. Z. Win, “The many facets of information in networked estimation and control,” Annu. Rev. of Control Robot. and Auton. Syst. , vol. 6, no. 1, pp. 233-259, 2023

  4. [3]

    Iterative learning control for sampled-data systems: From theory to practice,

    K. Abidi, and J. Xu, “Iterative learning control for sampled-data systems: From theory to practice,” IEEE Trans. on Ind. Electron. , vol. 58, no. 7, pp. 3002-3015, 2014

  5. [4]

    A new model for autonomous, networked control systems,

    G. Pratl, D. Dietrich, G. P. Hancke, and W. T. Penzhorn, “A new model for autonomous, networked control systems,” IEEE Trans. on Ind. Informat. , vol. 3, no. 1, pp. 21-32, 2007

  6. [5]

    Time-varying gain controller synthesis of piecewise homogeneous semi-Markov jump linear systems,

    Y . Tian, H. Yan, H. Zhang, M. Wang, and J. Yi, “Time-varying gain controller synthesis of piecewise homogeneous semi-Markov jump linear systems,” Automatica, vol. 146, 2022, Art. no. 110594

  7. [6]

    Network-induced constraints in networked control systems—A survey,

    L. Zhang, H. Gao, and O. Kaynak, “Network-induced constraints in networked control systems—A survey,” IEEE Trans. on Ind. Informat. , vol. 9, no. 1, pp. 403-416, 2013

  8. [7]

    Resilient control in cyber-physical systems: Countering uncertainty, constraints, and adver- sarial behavior,

    S. Weerakkody, O. Ozel, Y . Mo, and B. Sinopoli, “Resilient control in cyber-physical systems: Countering uncertainty, constraints, and adver- sarial behavior,” Found. Trends Syst. Control, vol. 7, nos. 1-2, pp. 1-252, 2019

Show all 28 references
  1. [8]

    Limitations of linear control over packet drop networks,

    N. Elia and J. N. Eisenbeis, “Limitations of linear control over packet drop networks,” IEEE Trans. Autom. Control, vol. 56, no. 4, pp. 826-841, 2011

  2. [9]

    Kalman Filtering With Intermittent Observations,

    B. Sinopoli, L. Schenato, M. Franceschetti, K. Poolla, M. I. Jordan, and S. S. Sastry, “Kalman Filtering With Intermittent Observations,” IEEE Trans. Automat. Contr., vol. 49, no. 9, pp. 1453–1464, Sep. 2004

  3. [10]

    Optimal control of LTI systems over unreliable communication links,

    O. C. Imer, S. Y ¨uksel, and T. Bas ¸ar, “Optimal control of LTI systems over unreliable communication links,” Automatica, vol. 42, no. 9, pp. 1429–1439, Sep. 2006

  4. [11]

    Foundations of Control and Estimation Over Lossy Networks,

    L. Schenato, B. Sinopoli, M. Franceschetti, K. Poolla, and S. S. Sastry, “Foundations of Control and Estimation Over Lossy Networks,” Proc. IEEE, vol. 95, no. 1, pp. 163–187, Jan. 2007

  5. [12]

    Stabilization of linear systems over networks with bounded packet loss,

    J. Xiong, and J. Lam, “Stabilization of linear systems over networks with bounded packet loss,” Automatica vol. 43, no. 1,pp. 80-87, Jan. 2007

  6. [13]

    Modelling and control of networked control sys- tems with both network-induced delay and packet-dropout,

    W. Zhang, and L. Yu, “Modelling and control of networked control sys- tems with both network-induced delay and packet-dropout,” Automatica vol. 44, no.12, pp. 3206-3210, 2008

  7. [14]

    Mean square stabilization for sampled-data T–S fuzzy systems with random packet dropout,

    Z. Hu and X. Mu, “Mean square stabilization for sampled-data T–S fuzzy systems with random packet dropout,” IEEE Trans. Fuzzy Syst. , vol. 28, no. 8, pp. 1815-1824, 2020

  8. [15]

    V . K. N. Lau and Y . K. Kwok. Channel-Adaptive Technologies and Cross-Layer Designs for Wireless Systems with Multiple Antennas - Theory and Applications, 1st edition. Wiley John Proakis Telecom Series, 2005

  9. [16]

    Learning optimal scheduling policy for remote state estimation under uncertain channel condition,

    S. Wu, X. Ren, Q. Jia, K. H. Johansson, and L. Shi, “Learning optimal scheduling policy for remote state estimation under uncertain channel condition,” IEEE Trans. Control Netw. Syst. , vol. 7, no. 2, pp. 579–591, Jun. 2020

  10. [17]

    Matiakis, S

    T. Matiakis, S. Hirche, and M. Buss. Local and remote control mea- sures for networked control systems. In Proceedings of the 47th IEEE International Conference on Decision and Control, CDC ’08

  11. [18]

    Learning in wireless control systems over nonstationary channels,

    M. Eisen, K. Gatsis, G. J. Pappas, and A. Ribeiro, “Learning in wireless control systems over nonstationary channels,”IEEE Trans. Signal Process., vol. 67, no. 5, pp. 1123–1137, Mar. 2019

  12. [19]

    Thompson sampling for networked control over unknown channels,

    W. Liu, A. S. Leong, and D. E. Quevedo. “Thompson sampling for networked control over unknown channels,” Automatica, vol. 165, 2024, Art. no. 111684

  13. [20]

    Sample complexity of networked control systems over unknown channels,

    K. Gatsis and G. J. Pappas, “Sample complexity of networked control systems over unknown channels,” in Proc. Conf. Decis. Control, 2018, pp. 6067–6072

  14. [21]

    Statistical learning for analysis of net- worked control systems over unknown channels,

    K. Gatsis and G. J. Pappas, “Statistical learning for analysis of net- worked control systems over unknown channels,” Automatica, vol.125, Mar. 2021, Art. no. 109386

  15. [23]

    Tsiamis, I

    A. Tsiamis, I. Ziemann, N. Matni, and G. J. Pappas, ”Statistical learning theory for control: A finite-sample perspective,” IEEE Control Systems Magazine, vol. 43, no. 6, pp. 67-97, 2023

  16. [24]

    C.,Campi and E

    M. C.,Campi and E. Weyer, ”Finite sample properties of system identi- fication methods,” IEEE Trans. Autom. Control, vol. 47, no. 8, pp. 1329-

  17. [25]

    Data-driven distribution- ally robust LQR with multiplicative noise,

    P. Coppens, M. Schuurmans, and P. Patrinos, “Data-driven distribution- ally robust LQR with multiplicative noise,” in Proc. Learn. Dyn. Control, Jul. 2020, pp.521-530

  18. [26]

    Finite-sample-based Spec- tral Radius Estimation and Stabilizability Test for Networked Control Systems,

    L. Xu, B. Guo, and G. Ferrari-Trecate, “Finite-sample-based Spec- tral Radius Estimation and Stabilizability Test for Networked Control Systems,” in Proc. 2022 European Control Conference (ECC) , London, United Kingdom: IEEE, Jul. 2022, pp. 2087–2092

  19. [27]

    Wasserman, All of statistics: a concise course in statistical inference

    L. Wasserman, All of statistics: a concise course in statistical inference. Springer Science & Business Media, 2013

  20. [28]

    Feedback stabilization of discrete-time networked systems over fading channels,

    N. Xiao, L. Xie, and L. Qiu, “Feedback stabilization of discrete-time networked systems over fading channels,” IEEE Trans. Autom. Control , vol. 57, no. 9, pp. 2176–2189, Sep. 2012

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.