{"id":"05ae1faa-0ae2-4757-8172-630948f71656","arxiv_id":"2412.00679","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Closed-form Stackelberg equilibrium sampling probabilities are derived for a two-player remote estimation game with random walk states, sampling costs, and two-sided information leakage.","lead":"Two players track each other's random movements, and each can sample the other's position, but sampling reveals their own position. This paper finds which sampling rates the players should commit to when one moves first, balancing tracking accuracy, sampling cost, and privacy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's K1<0 case is not a well-posed SE: the leader's cost is unbounded below as p1→0+, so no minimizer is attained, and the proposed (0,0) lies outside the AoI steady-state regime.","rationale":"I read the manuscript in good faith and checked the central derivation. The closed-form costs (9)-(10), the follower best response (17), the boundary points (19)-(20), the piecewise leader cost (27), and the candidate enumeration in Theorem 7 are internally consistent for K2>0. The reader's concern is real and localized: Theorem 2's K1≤0 branch is not well-posed because J1 is unbounded below as p1→0+ when K1<0, so the claimed minimizer p1=0 is only an infimum limit, and the p1=p2=0 point is outside the region where Eq. (6) is a valid stationary AoI distribution. This is a genuine correctness issue for that branch, not merely a presentation gap, and it is explicitly acknowledged in the text's remark that the cost tends to negative infinity. However, it does not invalidate the whole paper: the K2>0 analysis survives, the K1>0,K2<=0 case has p1*>0 and is well-posed, and the proposed fix is a parameter restriction or an extended-payoff convention. Therefore the reader's CONDITIONAL verdict is appropriate, and I would not move it to ACCEPT or REJECT.","tokens_in":12317,"tokens_out":10209,"duration_ms":90202,"concrete_test":"Take K1=-1, K2=-1, c1=1. For p2=0 and any p1=ε>0, the age chain is positive recurrent and Eq. (9) gives J1(ε,0)=1+ε-1/ε. Evaluate ε=0.1, 0.01, 0.001: the cost decreases as -8.9, -98.99, -998.999, so J1 is unbounded below on (0,1]. Also ∂J1/∂p1=1+1/p1^2>0, so every smaller p1 improves the cost; hence no p1* in (0,1] minimizes, and p1=0 is not attained. This single parameter example directly refutes Theorem 2's claim for K1<0 and Corollary 3.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weak spot is Theorem 2 (and Corollary 3) in the K1<0, K2<=0 regime. The proof substitutes p2*=0 into J1 and concludes p1*=0 from ∂J1/∂p1>0, but this ignores that J1(p1,0)=c1(K1(p1^{-1}-1)+p1) is unbounded below as p1→0+ when K1<0. Since every p1>0 is feasible and yields finite cost, no minimizer is attained on (0,1]; the infimum is -∞ at an excluded limit point. The paper itself notes after Corollary 3 that Ji tends to -∞, yet still lists (0,0) as the SE. Moreover, Eq. (6) is the stationary AoI distribution only if (1-p1)(1-p2)<1; at the proposed (0,0), the age Markov chain is null recurrent and Eq. (9) is not the actual long-term average. Thus the claimed 'full characterization' is not well-posed for K1<0, and for K1=0 it rests on a boundary steady state not covered by the derivation. The positive-K2 theorem and the K1>0, K2<=0 case are unaffected; the fix is to restrict parameters where the SE is attained or to explicitly treat p1→0 as an infimum and prove the appropriate extended-payoff equilibrium notion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies a two-player remote estimation game in which each player samples the other player's random walk state, and any sampling action reveals the sampler's own state to the opponent. Restricting to stationary probabilistic sampling policies, the authors derive closed-form cost functions in terms of age of information, define a Stackelberg equilibrium with Player 1 as leader, and claim a complete characterization: Theorem 2 for K2≤0 and Theorem 7 for K2>0, with candidate sets depending on the sign of K1. Numerical simulations illustrate the equilibrium structure for representative parameter values.","tokens_in":12588,"tokens_out":9486,"duration_ms":81862,"significance":"The paper contributes a tractable strategic extension of a classic remote estimation problem, explicitly modeling information leakage and commitment. The derivations are self-contained, and the positive-parameter parts (K1>0, K2>0) appear mathematically correct, yielding a simple three-candidate check for the leader's Stackelberg optimum. However, the claimed full characterization is not valid in the K1<0, K2≤0 regime, and the boundary point (0,0) lies outside the steady-state domain of the AoI analysis. If these degenerate cases are handled by restricting parameters or by an explicit extended-payoff formulation, the results for the well-posed regimes would be a useful contribution to the AoI and game-theoretic sampling literature.","major_comments":[{"comment":"The proposed SE (p1*,p2*)=(0,0) is not a well-defined Stackelberg equilibrium. After substituting p2*=0, the leader's cost is J1(p1,0)=c1(K1(p1^{-1}-1)+p1), which tends to -∞ as p1→0+ when K1<0; no minimizer is attained on (0,1], and J1(0,0) is undefined in (9). The paper itself states after Corollary 3 that Ji tends to -∞, yet still lists (0,0) as the SE. The theorem should restrict to K1>0 (with K1=0 handled by a limiting argument under a well-defined payoff extension) and explicitly state that for K1<0 no SE exists in the current formulation.","section":"Theorem 2, K1<0 case"},{"comment":"The steady-state AoI distribution in (6) requires (1-p1)(1-p2)<1, i.e., at least one player samples with positive probability. At the proposed equilibrium (0,0), neither player ever samples, the age Markov chain is null recurrent, and the cost formulas (9)-(10) are not the actual long-term averages; the estimation error variances diverge and the cost is undefined. Even for K1=0, where the limit of J1(p1,0) as p1→0 is zero, the point (0,0) lies outside the domain of the derived payoff functions. Thus the claim of a full characterization over all K1,K2 is not established.","section":"Eq. (6) and Theorem 2/Corollary 3"},{"comment":"Theorem 7 is stated for K2 ≥ 0, but the analysis leading to it assumes K2 > 0. For K2=0, the follower's best response is p2*=0 for all p1, as established in Theorem 2, and quantities such as pU1 in (20) become zero, so the candidate sets A and B are not meaningful in the same way. The theorem should be restricted to K2>0, with the K2=0 case explicitly subsumed under Theorem 2 (subject to the K1 issues raised above).","section":"Theorem 7 statement"}],"minor_comments":[{"comment":"The sentence 'For K1 = 2, shown in Fig. 2(d)' should read 'for K1 = -2', since the figure caption and surrounding discussion refer to K1 = -2.","section":"Section 5, second paragraph"},{"comment":"The symbol α is used both for the step probabilities α1, α2 in (1) and for the privacy-leakage weight α in (3)-(4). Although the usage is explicit, the overloading may confuse readers and a different symbol for the privacy weight would improve clarity.","section":"Notation in (1) and (3)"},{"comment":"In the proof of Corollary 8, the claim that J1(p1,BR(p1)) > 0 when K1>0 in the third region of (27) is correct for p1∈(pU1,1], but the sentence would be clearer if it stated p1>0 explicitly, since the expression J1(0,0) is not defined.","section":"Corollary 8 proof"}],"recommendation":"major_revision","confidential_remarks":"The core issue is localized to the K1≤0, K2≤0 regime and the boundary point (0,0). The positive-K2 analysis and the K1>0, K2≤0 part appear sound and are the main contribution. The authors should be asked to revise by narrowing the parameter claims or by explicitly introducing an extended payoff/equilibrium concept for the degenerate cases. There is no indication of fabrication or problematic citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, narrow extension of AoI remote estimation into a two-player Stackelberg game with two-sided leakage, and the equilibrium formulas for K1>0, K2>0 are correct. But the claimed full characterization is not: the K1<0, K2<=0 branch in Theorem 2 and Corollary 3 is not a well-posed equilibrium. The stress-test concern lands.\n\nWhat is actually new: modeling both players as samplers who leak their own state when sampling, over random-walk processes, and deriving closed-form best responses and a Stackelberg candidate set. The algebra is clean, the piecewise best response in (21) is correct, and Lemmas 4-6 check out for positive K. The explicit candidate set in Theorem 7 is useful, and the corollary giving (0,1) when K1>0 and K2>1 is a nice free-riding result. This is a genuinely new combination, and while it builds on the authors' own submitted work, that is not a problem here.\n\nSoft spots, in proportion: (1) Theorem 2 with K1<0 is wrong as stated. For p2=0, J1(p1,0)=c1(K1(p1^{-1}-1)+p1), which is unbounded below as p1->0+ when K1<0. The proof concludes p1*=0 from the derivative being positive, but p1=0 is not in the domain where the cost function is defined. The paper itself notes after Corollary 3 that Ji tends to negative infinity, yet still lists (0,0) as the SE. That is internally inconsistent. (2) The proposed (0,0) equilibrium lies exactly where the AoI Markov chain is null recurrent, so the steady-state distribution in (6) and the cost formulas in (9)-(10) are not valid there. For K1=0 the same boundary issue arises even though the formal expression is minimized at p1=0. The fix is straightforward: restrict the theorem to parameters where the SE is attained, or explicitly treat p1->0 as an infimum and prove the appropriate extended-payoff equilibrium notion. The positive-K results are unaffected. (3) The numerical section plots J1(p1,BR(p1)) rather than simulating the game; that is fine as illustration but not independent verification.\n\nBottom line: this deserves a serious referee. The main contribution is useful for the AoI/games subfield, and the flaw is localized and fixable. A careful revision that restricts the parameter space would make the paper solid. I would bring it to reading group and cite the positive-K results with a caveat about the boundary cases.","headline":"A useful, mostly correct Stackelberg equilibrium analysis for two-player remote estimation with information leakage, but the claimed full characterization is undermined by an ill-posed K1<0 branch that should be restricted or redefined.","tokens_in":13156,"tokens_out":2197,"would_cite":true,"duration_ms":21096,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A65","60J10","60G50"],"pacs":[],"model":"deepseek-v4-flash","headline":"The Stackelberg equilibrium of a two-player remote estimation game with random-walk states, information leakage, and sampling costs is fully characterized by evaluating the leader's cost at three candidate sampling pairs, with closed-form…","keywords":["remote estimation games","Stackelberg equilibrium","age of information","random walk process","information leakage","sampling cost","stationary probabilistic policies","asymmetric information"],"falsifier":"Run a direct simulation of the two random walks under the candidate equilibrium (p1,p2)=(0,0) over growing time horizons and compute the empirical long-run average of J1 and J2; if the limit does not converge to the closed-form value (or diverges), the boundary equilibrium formulas fail as a description of the true game at that point. More generally, test any claimed equilibrium by computing the empirical average cost against the formulas over horizons T=$10^{3}$,$10^{4}$,$10^{5}$ and checking whether the minimizer among the three candidates matches the simulated best responses.","tokens_in":12093,"feed_emoji":"📡","tokens_out":6955,"duration_ms":56316,"temperature":0.7,"pith_summary":"This paper studies a two-player remote estimation game in which each player tracks the other's random walk by sampling, but sampling reveals the sampler's own state to the opponent. The players trade off estimation accuracy, information leakage, and sampling cost. Restricting to stationary probabilistic sampling policies, the paper derives closed-form average costs and characterizes the Stackelberg equilibrium where one player commits to a sampling probability first. The result is that the leader's optimal commitment is determined by evaluating its cost on a small candidate set of at most three sampling pairs, with formulas depending only on two cost-benefit ratios. This converts a strategic sampling problem into a finite comparison, which matters for real-time monitoring under surveillance where commitment is natural.","feed_headline":"Remote estimation game's optimal sampling is a three-candidate choice","feed_subtitle":"For random-walk tracking with leakage and sampling costs, the Stackelberg equilibrium is closed form.","key_machinery":"The carrying object is the age-of-information Markov chain for each player's estimate. Under independent Bernoulli sampling with probabilities p1 and p2, the age Δ(t) evolves as a two-state transition: it resets to 0 if at least one player samples, and otherwise increments by 1. This chain has a geometric stationary distribution π1,k = (1 - (1-p1)(1-p2))((1-p1)(1-p2))^k, which feeds directly into the expected squared estimation error 2α2(1-p1)(1-p2)/(1-(1-p1)(1-p2)), and hence into the closed-form costs J1,J2. The Stackelberg analysis then hinges on the follower's first-order condition ∂J2/∂p2=0, which yields the piecewise best-response function BR(p1) with boundary points pL1 and pU1, and on the piecewise convex/concave structure of the leader's objective J1(p1,BR(p1)). The final mechanism is the reduction: the global minimizer of a piecewise function with at most one interior critical point lies among the endpoints and the critical point, giving the three-point candidate sets A and B.","core_discovery":"The central claim is that for the Stackelberg equilibrium with stationary probabilistic sampling, the leader's optimal sampling probability p1* and the follower's best-response probability p2* = BR(p1*) have closed-form characterizations. If K2 ≤ 0, the follower never samples and the leader samples with p1* = min{$\\sqrt$(max{K1,0}),1}. If K2 > 0, the follower's best response is piecewise: sample with probability 1 for small p1, follow an interior decreasing curve between boundary points pL1 = max{0,1-1/K2} and pU1 = (-K2 + $\\sqrt$($K2^{2}$+4K2))/2, and stop sampling for p1 above pU1. The leader's equilibrium is the minimizer of J1(p1,BR(p1)) over this piecewise function, which the paper shows reduces to evaluating J1 at the three candidate pairs given in Theorem 7; if K1>0 and K2>1 the equilibrium is simply (0,1).","pith_inferences":["The paper's formulas are derived under a positive-recurrence condition for the age chain; a strict reading of the (0,0) equilibrium is that it is a limit of equilibria as p1→0 and p2→0, not an attained stationary equilibrium, since no steady state exists when neither player ever samples.","The same geometric-age machinery could be applied to other age-penalty functions or to asymmetric sampling costs, where the candidate-set structure may persist but the boundary formulas would change.","Because the leader commits first, the equilibrium can force the follower to sample even when the follower would prefer not to; if commitment is not credible (e.g., in a simultaneous-move game), the outcome would differ, so the leader's commitment power is the driver of the free-riding result.","A natural testable extension is to allow the sampling probabilities to depend on the current estimation error; if such adaptive policies outperform the constant-probability equilibrium, then the stationary-policy restriction is binding."],"forward_implications":["For any parameter set (α1, α2, α, c1, c2), the leader can compute the Stackelberg equilibrium by evaluating J1 at three candidate pairs, so optimal commitment is an O(1) calculation rather than a search.","If K2≤0, the follower never samples regardless of the leader's promise, so the leader's only decision is how often to sample alone, with p1*=min{√max{K1,0},1}.","If K1>0 and K2>1, the leader's dominant strategy is to never sample while the follower samples every step; the leader free-rides on the follower's sampling.","The boundary points pL1 and pU1 delimit when the follower switches between sampling always, sampling probabilistically, and not sampling; these thresholds depend solely on the follower's cost-benefit ratio K2.","The sign of K1 (the leader's net benefit per unit sampling cost) determines whether the leader's cost on the interior region is convex or concave, which is why the candidate set differs between K1≥0 and K1<0."],"supporting_citations":[{"why":"Defines age of information, the metric the paper's estimation-error analysis is built on.","marker":"Kaul et al. (2012)"},{"why":"Supplies the Markov chain stationary-distribution result used for Eq. (6).","marker":"Bertsekas and Tsitsiklis (2008)"},{"why":"The closest prior work on remote estimation of a random walk with sampling cost; the paper extends it to a strategic game.","marker":"Yun et al. (2018)"},{"why":"Introduces the remote estimation game formulation that this paper builds on and adapts.","marker":"Velicheti et al. (2024)"}],"fun_headline_variants":["Stackelberg sampling for estimation games: only three choices","Remote estimation game leader's best sample: one of three","Closed-form Stackelberg solution for random-walk tracking","Surveillance estimation game: leader picks among three rates","Three sampling probabilities settle the estimation game"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the age Markov chain reaches a steady state, which requires at least one player to sample with positive probability; the equilibria at p1=p2=0 sit exactly at the boundary where no steady state exists, and for K1<0 the alleged minimizer p1=0 is only an infimum, not an attained minimum.","fun_headline_variants_meta":{"raw":{"variants":["Stackelberg sampling for estimation games: only three choices","Remote estimation game leader's best sample: one of three","Closed-form Stackelberg solution for random-walk tracking","Surveillance estimation game: leader picks among three rates","Three sampling probabilities settle the estimation game"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000475,"raw_usage":{"total_tokens":2329,"prompt_tokens":887,"completion_tokens":1442,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":1365}},"tokens_in":503,"tokens_out":1442,"duration_ms":14621,"temperature":1.0,"reasoning_tokens":1365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:07:45.627015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a direct simulation of the two random walks under the candidate equilibrium (p1,p2)=(0,0) over growing time horizons and compute the empirical long-run average of J1 and J2; if the limit does not converge to the closed-form value (or diverges), the boundary equilibrium formulas fail as a description of the true game at that point. More generally, test any claimed equilibrium by computing the empirical average cost against the formulas over horizons T=$10^{3}$,$10^{4}$,$10^{5}$ and checking whether the minimizer among the three candidates matches the simulated best responses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines age of information, the metric the paper's estimation-error analysis is built on."},{"cited_title":"and Tsitsiklis, J.N","cited_arxiv_id":null,"evidence_quote":"Supplies the Markov chain stationary-distribution result used for Eq. (6)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The closest prior work on remote estimation of a random walk with sampling cost; the paper extends it to a strategic game."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the remote estimation game formulation that this paper builds on and adapts."}],"review_version":1}