REVIEW 3 major objections 3 minor 25 references
Remote Estimation Games with Random Walk Processes: Stackelberg Equilibrium
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The Stackelberg equilibrium of a two-player remote estimation game with random-walk states, information leakage, and sampling costs is fully characterized by evaluating the leader's cost at three candidate sampling pairs, with closed-form…
desk verdict A useful, mostly correct Stackelberg equilibrium analysis for two-player remote estimation with information leakage, but the claimed full characterization is undermined by an ill-posed K1<0 branch that should be restricted or redefined. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the age-of-information Markov chain for each player's estimate. Under independent Bernoulli sampling with probabilities p1 and p2, the age Δ(t) evolves as a two-state transition: it resets to 0 if at least one player samples, and otherwise increments by 1. This chain has a geometric stationary distribution π1,k = (1 - (1-p1)(1-p2))((1-p1)(1-p2))^k, which feeds directly into the expected squared estimation error 2α2(1-p1)(1-p2)/(1-(1-p1)(1-p2)), and hence into the closed-form costs J1,J2. The Stackelberg analysis then hinges on the follower's first-order condition ∂J2/∂p2=0, which yields the piecewise best-response function BR(p1) with boundary points pL1 and pU1, and on the piecewise convex/concave structure of the leader's objective J1(p1,BR(p1)). The final mechanism is the reduction: the global minimizer of a piecewise function with at most one interior critical point lies among the endpoints and the critical point, giving the three-point candidate sets A and B.
What would settle it
Run a direct simulation of the two random walks under the candidate equilibrium (p1,p2)=(0,0) over growing time horizons and compute the empirical long-run average of J1 and J2; if the limit does not converge to the closed-form value (or diverges), the boundary equilibrium formulas fail as a description of the true game at that point. More generally, test any claimed equilibrium by computing the empirical average cost against the formulas over horizons T=$10^{3}$,$10^{4}$,$10^{5}$ and checking whether the minimizer among the three candidates matches the simulated best responses.
Extended reading notes
Core claim
The central claim is that for the Stackelberg equilibrium with stationary probabilistic sampling, the leader's optimal sampling probability p1* and the follower's best-response probability p2* = BR(p1*) have closed-form characterizations. If K2 ≤ 0, the follower never samples and the leader samples with p1* = min{$\sqrt$(max{K1,0}),1}. If K2 > 0, the follower's best response is piecewise: sample with probability 1 for small p1, follow an interior decreasing curve between boundary points pL1 = max{0,1-1/K2} and pU1 = (-K2 + $\sqrt$($K2^{2}$+4K2))/2, and stop sampling for p1 above pU1. The leader's equilibrium is the minimizer of J1(p1,BR(p1)) over this piecewise function, which the paper shows reduces to evaluating J1 at the three candidate pairs given in Theorem 7; if K1>0 and K2>1 the equilibrium is simply (0,1).
Load-bearing premise
The load-bearing premise is that the age Markov chain reaches a steady state, which requires at least one player to sample with positive probability; the equilibria at p1=p2=0 sit exactly at the boundary where no steady state exists, and for K1<0 the alleged minimizer p1=0 is only an infimum, not an attained minimum.
Editorial extensions
If this is right
- For any parameter set (α1, α2, α, c1, c2), the leader can compute the Stackelberg equilibrium by evaluating J1 at three candidate pairs, so optimal commitment is an O(1) calculation rather than a search.
- If K2≤0, the follower never samples regardless of the leader's promise, so the leader's only decision is how often to sample alone, with p1*=min{√max{K1,0},1}.
- If K1>0 and K2>1, the leader's dominant strategy is to never sample while the follower samples every step; the leader free-rides on the follower's sampling.
- The boundary points pL1 and pU1 delimit when the follower switches between sampling always, sampling probabilistically, and not sampling; these thresholds depend solely on the follower's cost-benefit ratio K2.
- The sign of K1 (the leader's net benefit per unit sampling cost) determines whether the leader's cost on the interior region is convex or concave, which is why the candidate set differs between K1≥0 and K1<0.
Reading between the lines
- The paper's formulas are derived under a positive-recurrence condition for the age chain; a strict reading of the (0,0) equilibrium is that it is a limit of equilibria as p1→0 and p2→0, not an attained stationary equilibrium, since no steady state exists when neither player ever samples.
- The same geometric-age machinery could be applied to other age-penalty functions or to asymmetric sampling costs, where the candidate-set structure may persist but the boundary formulas would change.
- Because the leader commits first, the equilibrium can force the follower to sample even when the follower would prefer not to; if commitment is not credible (e.g., in a simultaneous-move game), the outcome would differ, so the leader's commitment power is the driver of the free-riding result.
- A natural testable extension is to allow the sampling probabilities to depend on the current estimation error; if such adaptive policies outperform the constant-probability equilibrium, then the stationary-policy restriction is binding.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies a two-player remote estimation game in which each player samples the other player's random walk state, and any sampling action reveals the sampler's own state to the opponent. Restricting to stationary probabilistic sampling policies, the authors derive closed-form cost functions in terms of age of information, define a Stackelberg equilibrium with Player 1 as leader, and claim a complete characterization: Theorem 2 for K2≤0 and Theorem 7 for K2>0, with candidate sets depending on the sign of K1. Numerical simulations illustrate the equilibrium structure for representative parameter values.
Significance. The paper contributes a tractable strategic extension of a classic remote estimation problem, explicitly modeling information leakage and commitment. The derivations are self-contained, and the positive-parameter parts (K1>0, K2>0) appear mathematically correct, yielding a simple three-candidate check for the leader's Stackelberg optimum. However, the claimed full characterization is not valid in the K1<0, K2≤0 regime, and the boundary point (0,0) lies outside the steady-state domain of the AoI analysis. If these degenerate cases are handled by restricting parameters or by an explicit extended-payoff formulation, the results for the well-posed regimes would be a useful contribution to the AoI and game-theoretic sampling literature.
major comments (3)
- [Theorem 2, K1<0 case] The proposed SE (p1*,p2*)=(0,0) is not a well-defined Stackelberg equilibrium. After substituting p2*=0, the leader's cost is J1(p1,0)=c1(K1(p1^{-1}-1)+p1), which tends to -∞ as p1→0+ when K1<0; no minimizer is attained on (0,1], and J1(0,0) is undefined in (9). The paper itself states after Corollary 3 that Ji tends to -∞, yet still lists (0,0) as the SE. The theorem should restrict to K1>0 (with K1=0 handled by a limiting argument under a well-defined payoff extension) and explicitly state that for K1<0 no SE exists in the current formulation.
- [Eq. (6) and Theorem 2/Corollary 3] The steady-state AoI distribution in (6) requires (1-p1)(1-p2)<1, i.e., at least one player samples with positive probability. At the proposed equilibrium (0,0), neither player ever samples, the age Markov chain is null recurrent, and the cost formulas (9)-(10) are not the actual long-term averages; the estimation error variances diverge and the cost is undefined. Even for K1=0, where the limit of J1(p1,0) as p1→0 is zero, the point (0,0) lies outside the domain of the derived payoff functions. Thus the claim of a full characterization over all K1,K2 is not established.
- [Theorem 7 statement] Theorem 7 is stated for K2 ≥ 0, but the analysis leading to it assumes K2 > 0. For K2=0, the follower's best response is p2*=0 for all p1, as established in Theorem 2, and quantities such as pU1 in (20) become zero, so the candidate sets A and B are not meaningful in the same way. The theorem should be restricted to K2>0, with the K2=0 case explicitly subsumed under Theorem 2 (subject to the K1 issues raised above).
minor comments (3)
- [Section 5, second paragraph] The sentence 'For K1 = 2, shown in Fig. 2(d)' should read 'for K1 = -2', since the figure caption and surrounding discussion refer to K1 = -2.
- [Notation in (1) and (3)] The symbol α is used both for the step probabilities α1, α2 in (1) and for the privacy-leakage weight α in (3)-(4). Although the usage is explicit, the overloading may confuse readers and a different symbol for the privacy weight would improve clarity.
- [Corollary 8 proof] In the proof of Corollary 8, the claim that J1(p1,BR(p1)) > 0 when K1>0 in the third region of (27) is correct for p1∈(pU1,1], but the sentence would be clearer if it stated p1>0 explicitly, since the expression J1(0,0) is not defined.
Circularity Check
No significant circularity: the Stackelberg equilibrium characterization is derived from the model's cost functions by first-order conditions and best-response calculus, with no fitted inputs or self-citation used as proof.
full rationale
The paper's derivation chain is self-contained. The costs J1 and J2 in (9)-(10) are derived from the model and the AoI Markov chain in (5)-(6); the follower best response BR(p1) in (17)-(21) comes from minimizing J2; the leader's problem is then J1(p1,BR(p1)), and Theorems 2 and 7 select minima over the resulting piecewise-defined objective via calculus in Lemmas 4-6. No parameter is fitted to data, no quantity is predicted that was used as an input, and no load-bearing theorem is imported from the authors' prior work. The only self-citation is the introductory reference to Velicheti et al. (2024), which motivates the model, but the model is fully specified in this paper and the equilibrium theorems do not depend on that citation. Potential well-posedness concerns about the K1<0 boundary case, such as the unattained infimum or the null-recurrent age chain at p1=p2=0, are correctness matters rather than circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The steady-state AoI distribution in Eq. (6) exists only when (1-p1)(1-p2) < 1; the paper proposes equilibria with p1=p2=0 where it does not.
- domain assumption Each player's estimate is the last observed state of the opponent (conditional mean estimator).
- domain assumption When a player samples, it reveals its own current state to the opponent with certainty.
- ad hoc to paper Players restrict to stationary probabilistic sampling policies independent of the estimation error.
Cite this review
Pith. "Pith review of Remote Estimation Games with Random Walk Processes: Stackelberg Equilibrium." pith.science (2026). https://pith.science/paper/64W7FAZJ
@misc{pith2026241200679,
author = {Pith},
title = {Pith review of: Remote Estimation Games with Random Walk Processes: Stackelberg Equilibrium},
year = {2026},
howpublished = {\url{https://pith.science/paper/64W7FAZJ}},
note = {Machine review of arXiv:2412.00679}
}
read the original abstract
Remote estimation is a crucial element of real time monitoring of a stochastic process. While most of the existing works have concentrated on obtaining optimal sampling strategies, motivated by malicious attacks on cyber-physical systems, we model sensing under surveillance as a game between an attacker and a defender. This introduces strategic elements to conventional remote estimation problems. Additionally, inspired by increasing detection capabilities, we model an element of information leakage for each player. Parameterizing the game in terms of uncertainty on each side, information leakage, and cost of sampling, we consider the Stackelberg Equilibrium (SE) concept where one of the players acts as the leader and the other one as the follower. By focusing our attention on stationary probabilistic sampling policies, we characterize the SE of this game and provide simulations to show the efficacy of our results.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address author booktitle chapter doi edition editor eid howpublished institution journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid.sent...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Amitai, M. (1996). Repeated games with incomplete information on both sides. Hebrew University of Jerusalem
work page 1996
-
[4]
Arafa, A., Banawan, K., Seddik, K.G., and Poor, H.V. (2021). Sample, quantize, and encode: Timely estimation over noisy channels. IEEE Trans Comm, 69(10), 6485--6499
work page 2021
-
[5]
Aumann, R.J., Maschler, M., and Stearns, R.E. (1995). Repeated G with I I. MIT P
work page 1995
-
[6]
Bastopcu, M. and Ulukus, S. (2022). Using timeliness in tracking infections. Entropy, 24(6)
work page 2022
-
[7]
Bedewy, A.M., Sun, Y., Kompella, S., and Shroff, N.B. (2019). Age-optimal sampling and transmission scheduling in multi-source systems. In ACM Mobihoc, 121--130
work page 2019
-
[8]
Bertsekas, D. and Tsitsiklis, J.N. (2008). Introduction to P, volume 1. Athena Scientific
work page 2008
Show all 25 references
-
[9]
and Ephremides, A
Chen, Y. and Ephremides, A. (2021). Minimizing age of incorrect information for unreliable channel with power constraint. In IEEE GLOBECOM, 1--6
2021
-
[10]
H \"o rner, J., Rosenberg, D., Solan, E., and Vieille, N. (2010). On a m game with one-sided information. Operations R, 58(4-part-2), 1107--1115
2010
-
[11]
Hsu, Y.P., Modiano, E., and Duan, L. (2019). Scheduling algorithms for minimizing age of information in wireless broadcast networks with random arrivals. IEEE Trans Mobile Computing, 19(12), 2903--2915
2019
-
[12]
Kadota, I., Sinha, A., Uysal-Biyikoglu, E., Singh, R., and Modiano, E. (2018). Scheduling policies for minimizing age of information in broadcast wireless networks. IEEE/ACM Trans Networking, 26(6), 2637--2650
2018
-
[13]
Kam, C., Kompella, S., and Ephremides, A. (2020). Age of incorrect information for remote estimation of a binary m source. In IEEE INFOCOM 2020-IEEE Conf Comp Comm Workshops (INFOCOM WKSHPS), 1--6
2020
-
[14]
Kaul, S.K., Yates, R.D., and Gruteser, M. (2012). Real-time status: How often should one update? In IEEE Infocom
2012
-
[15]
(2020 a )
Maatouk, A., Kriouile, S., Assaad, M., and Ephremides, A. (2020 a ). The age of incorrect information: A new performance metric for status updates. IEEE/ACM Trans on Networking, 28(5), 2215--2228
2020
-
[16]
(2020 b )
Maatouk, A., Kriouile, S., Assad, M., and Ephremides, A. (2020 b ). On the optimality of the W hittle’s index policy for minimizing the age of information. IEEE Trans Wireless Comm, 20(2), 1263--1277
2020
-
[17]
Mertens, J.F. (1990). Repeated games. In Game theory and A, 77--130. Elsevier
1990
-
[18]
and Ba s ar, T
Nar, K. and Ba s ar, T. (2014). Sampling multidimensional wiener processes. In 53rd IEEE CDC, 3426--3431
2014
-
[19]
and Papavassilopoulos, G.P
Olsder, G.J. and Papavassilopoulos, G.P. (1988). About when to use the searchlight. Journal of Mathematical Analysis and Applications, 136(2), 466--478
1988
-
[20]
Sorin, S. (1983). Some results on the existence of nash equilibria for non-zero sum games with incomplete information. Internat J of Game Theory, 12, 193--205
1983
-
[21]
Sun, Y., Polyanskiy, Y., and Uysal, E. (2020). Sampling of the wiener process for remote rstimation over a channel with random delay. IEEE Trans Information Theory, 66(2), 1118--1135
2020
-
[22]
Velicheti, R.K., Dokme, A., Bastopcu, M., Chen, A., Dorothy, M., Shishika, D., and Ba s ar, T. (2024). Strategic remote estimation games: Catch me if you can! Submitted
2024
-
[23]
Yates, R.D., Sun, Y., Brown, D.R., Kaul, S.K., Modiano, E., and Ulukus, S. (2021). Age of information: An introduction and survey. IEEE J Selected Areas in Comm, 39(5), 1183--1210
2021
-
[24]
Yun, J., Joo, C., and Eryilmaz, A. (2018). Optimal real-time monitoring of an information source under communication costs. In IEEE CDC, 4767--4772
2018
-
[25]
Zhong, J., Zhang, W., Yates, R.D., Garnaev, A., and Zhang, Y. (2019). Age-aware scheduling for asynchronous arriving jobs in edge applications. In IEEE INFOCOM, 674--679
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.