{"id":"8c74fb60-0d56-49b9-a658-76e589bfab00","arxiv_id":"2607.05423","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Under targetability and exploitability, a two-phase honeypot-then-trap leader strictly exceeds the classical SSE utility ceiling against a UCB follower at O(sqrt(T ln T)) signaling cost.","lead":"A leader can beat classical Stackelberg value against a UCB follower by first inflating one arm's empirical mean (honeypot) then switching to a selfish mix that keeps the follower locked by optimism. This shows static commitment equilibria fail when the follower learns from endogenous rewards.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged exploitability premise.","rationale":"The paper's strongest claim is carefully scoped: under targetability + exploitability + a honeypot length that satisfies the UCB dominance condition of Theorem 5, the induced Delta of Theorem 6 yields strict improvement over T V_SSE whenever the inequality of Theorem 7 holds, with signaling cost at most O(sqrt(T ln T)). That implication is supported by elementary but exact algebra (mean update, index comparison, induction on frozen competitors) and by a payoff decomposition that is pathwise. The only place the implication can fail is if no such x_exp exists (or if Delta is too short relative to tau), which is precisely the reader's weakest_assumption. No hidden inconsistency appears in the deterministic proofs; stochastic extensions (Theorems 8-9) correctly add concentration radii. The experimental section is diagnostic rather than confirmatory and already shows both success (certified reversion) and failure (matrix, non-reversion, change-point), matching the theory's conditionality. Therefore the reader's CONDITIONAL verdict with high confidence and low correctness risk needs no adjustment; the contribution remains a clean constructive demonstration of a formal gap between static SSE and empirical UCB incentives inside a verifiable class of games.","tokens_in":11069,"tokens_out":697,"duration_ms":5741,"concrete_test":"On the certified toy instance of Appendix D (U_L = [[0,0],[1,0]], U_F = [[1,0],[0.45,0.55]], j*=1, x_hon=(1,0), x_exp=(0,1)), recompute V_SSE by enumeration, set L*=1, enforce tau_1=1000, measure realized Delta under deterministic UCB (c=0.2), and check whether Delta*(1 - V_SSE) > tau*(V_SSE - H_min). If the inequality holds and cumulative G_dec - T V_SSE matches the reported +55178.93 under reversion, the constructive claim is verified for that instance; if not, the accounting or lock-in formula has an implementation gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Theorem 7) is an implication under explicit assumptions, not a generic guarantee. The reader's weakest_assumption correctly isolates the load-bearing premise: Exploitability (Assumption 3) requires existence of x_exp with L* = U_L(x_exp, j*) > V_SSE while j* is strictly suboptimal for the follower by gamma_F > 0. Without that profitable trap, even perfect lock-in cannot beat T V_SSE. The deterministic proofs of honeypot inflation (Lemma 4, Theorem 5), exact lock-in duration (Theorem 6), and the cumulative-payoff decomposition (Appendix A.4 / B.4) are algebraically tight once the assumptions hold; the O(sqrt(T ln T)) signaling cost follows immediately from the trivial per-round bound of 1. Experiments honestly report cases where the inequality fails (matrix games, non-reverting certified run), which is consistent with the conditional nature of the result rather than an internal contradiction. Partial observability and defended followers are acknowledged in Appendix E and do not undermine the stated theorems.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies a finite-horizon repeated Stackelberg game in which the follower uses UCB1 on empirical rewards generated by the leader’s mixtures. Classical strong Stackelberg equilibrium (SSE) assumes immediate best response and therefore supplies no prescription against a history-dependent learner. The authors construct a two-phase deceptive mechanism: a honeypot phase that inflates the empirical mean and UCB index of a designated follower action j*, followed by a trap phase that switches to an exploitative mixture x_exp under which j* is follower-suboptimal yet leader-profitable. Under explicit Targetability (Assumption 2) and Exploitability (Assumption 3), they prove exact deterministic dominance (Theorem 5), lock-in duration (Theorem 6), and a cumulative-utility inequality (Theorem 7) showing that the leader can strictly exceed T V_SSE whenever the lock-in length satisfies Delta(L* - V_SSE) > tau(V_SSE - H_min), with honeypot cost O(sqrt(T ln T)). High-probability stochastic extensions, partial-observability remarks, and diagnostic experiments (including negative cases) are supplied.","tokens_in":11426,"tokens_out":1015,"duration_ms":9690,"significance":"If the result holds, it cleanly separates static SSE mechanism design from the incentives of empirical bandit followers and exhibits a concrete, checkable class of games in which optimism itself is a controllable state variable. The contribution is theoretical rather than universal: the paper does not claim every Stackelberg game admits profitable deception, only that under verifiable separation conditions the classical ceiling can be breached. Strengths include elementary but exact constructive proofs (Appendix A), standard sub-Gaussian extensions (Appendix B), an operational certification procedure via small LPs (Appendix D), a public code release, and experiments that honestly report both positive and negative outcomes. These features make the claim falsifiable and the mechanism reproducible.","major_comments":[{"comment":"Assumption 3 (Exploitability) is load-bearing for Theorem 7: the existence of x_exp with L* = U_L(x_exp, j*) > V_SSE while j* is strictly suboptimal for the follower by gamma_F > 0 is required for any strict improvement. The paper correctly treats this as a domain condition rather than a generic guarantee, and Appendix D supplies an LP certificate. For the claim to be usable, the main text should more prominently state that the result is conditional on this certificate being nonempty and should quantify how often such triples exist in standard security-game or random-matrix ensembles (the current experiments already show both success and failure).","section":null},{"comment":"Theorem 7 and the accompanying regret claim (Appendix B.4–B.5) establish improvement only when the leader reverts to an SSE policy after escape (or when Delta = Theta(T)). The non-reverting certified experiment produces long lock-in yet negative net advantage, which is consistent with the proof but is not emphasized in the main-text statement of the theorem. Clarifying the reversion requirement (or supplying a post-escape accounting) would prevent misreading the result as an unconditional linear gain.","section":null}],"minor_comments":[{"comment":"Section 7 and Table 1: the matrix-game negative advantage and the non-reverting certified run are valuable diagnostics; a short sentence in the main text linking them to the failure of the inequality in Theorem 7 would help readers interpret the table.","section":null},{"comment":"Notation for the switch time tau = t0 + tau_1 and the lock-in Delta / q_esc is introduced cleanly but is reused with slight variations across Theorems 5–7 and the appendices; a single consistent glossary would reduce cognitive load.","section":null},{"comment":"Appendix E correctly notes that change-point and corruption-robust followers raise the honeypot cost; a one-line quantitative remark on how large the robust radius must be relative to alpha - rho would strengthen the defensive discussion.","section":null},{"comment":"Minor typographical issues: author names contain encoding artifacts (¸, ¸s); arXiv identifier formatting in the header; and a few missing spaces around math operators in the abstract and Section 4.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a solid workshop-level contribution that is already close to journal quality for a specialized venue in algorithmic game theory or multi-agent learning. The conditional nature of the result is a feature, not a bug, provided the authors keep the exploitability premise front-and-center. Fit for a general-interest journal would be tighter if the authors added a short empirical frequency analysis of the LP certificates on standard security-game libraries."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple. Static SSE does not price a UCB follower's empirical history. This paper gives exact constructive conditions under which a leader can first inflate a target arm's UCB index (honeypot), then switch to a selfish mix and keep the follower locked long enough that cumulative leader utility exceeds T V_SSE, with signaling cost O(sqrt(T ln T)).\n\nWhat is new is the endogenous mechanism: legal Stackelberg play that manipulates the follower's sufficient statistics rather than exogenous reward poisoning. The deterministic core is elementary and tight. Lemma 4 and Theorem 5 give the honeypot length for strict index dominance; Theorem 6 gives the exact lock-in duration Delta as the first escape time of the frozen-index comparison; Theorem 7 is just the payoff decomposition that is positive when Delta(L* - V_SSE) > tau(V_SSE - H_min). Stochastic extensions are standard sub-Gaussian union bounds. The experiments are correctly labeled diagnostics and honestly report both wins (security, certified reverting UCB) and losses (matrix games, non-reverting run, change-point defense). Code is linked. Citations cover bandit poisoning and Stackelberg learning without pretending the gap does not exist.\n\nThe soft spot is real but already scoped by the authors: Exploitability (Assumption 3) requires a trap mix with L* > V_SSE while the target is follower-suboptimal. Without that profitable trap, lock-in cannot beat the SSE ceiling. The main theorems also assume the leader sees (N_j, mu_j); Appendix E notes the partial-observation fix and that change-point or robust estimators raise the cost. None of this is hidden, and none of it makes the algebra circular. The free parameters (c, tau_1, j*) are the usual design knobs, not free fudge factors.\n\nThis is for people who work on Stackelberg security games, mechanism design with learning agents, or adversarial multi-agent learning. It is a workshop-length theory note with constructive proofs and honest diagnostics, not a generic claim for every game. I would send it to peer review. Engage if you care about the static-vs-empirical gap; skip if you only want results that hold without an exploitability gap.","headline":"Clean constructive theory: under explicit targetability/exploitability, a two-phase leader can lock a UCB follower and beat T V_SSE with O(sqrt(T ln T)) signaling cost.","tokens_in":11983,"tokens_out":573,"would_cite":true,"duration_ms":4687,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Optimism-driven UCB learning can be gamed so a Stackelberg leader beats the classical SSE utility ceiling.","keywords":["Stackelberg games","bandit learning","UCB","strategic deception","reward manipulation","strong Stackelberg equilibrium","honeypot","lock-in"],"falsifier":"In a certified matrix game that meets the paper's separation conditions, compute the honeypot length that forces target-index dominance, run the trap, and check whether the observed lock-in length Delta satisfies Delta(L* - V_SSE) > tau(V_SSE - H_min); if the inequality fails on the realized path, the strict-improvement claim fails.","tokens_in":11987,"feed_emoji":"🪤","tokens_out":691,"duration_ms":6478,"temperature":0.7,"pith_summary":"Classical Stackelberg analysis assumes a follower who immediately best-responds to a committed leader mix, so the leader optimizes a static best-response map. Real followers that learn with Upper Confidence Bound (UCB) algorithms do something different: they act on empirical reward histories and an optimism bonus. This paper shows that an omniscient leader can treat those histories as a controllable state. First the leader runs a honeypot that inflates the empirical mean of a chosen target action; then it switches to a selfish mix that is profitable only while the follower stays locked on that target. Under explicit separation conditions the lock-in lasts long enough that cumulative leader utility strictly exceeds the classical strong Stackelberg equilibrium ceiling, while the honeypot cost grows only like the square root of the horizon. The result is a formal gap between static equilibrium prescriptions and the incentives that arise when a boundedly rational follower learns from data the leader itself produces.","feed_headline":"UCB optimism lets a Stackelberg leader beat the SSE ceiling","feed_subtitle":"A short honeypot inflates a target index; the follower stays locked while the leader cashes in","key_machinery":"The two-phase Deceptive Leader Mechanism: a honeypot phase that forces the target UCB index above all competitors (Theorem 5), followed by a trap phase whose exact lock-in length Delta is given by the first escape time of the frozen-index comparison (Theorem 6). The net-gain inequality of Theorem 7 converts that lock-in into a strict improvement over T times the SSE value.","core_discovery":"In a finite-horizon repeated Stackelberg game, a leader who first inflates a target follower action's UCB index with a honeypot mix and then switches to an exploitative mix can lock the UCB follower onto that action for a calculable duration. When the game satisfies targetability and exploitability, the resulting cumulative leader payoff strictly exceeds the classical strong Stackelberg value, and the signaling cost is only O(sqrt(T ln T)).","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["UCB optimism is a Stackelberg vulnerability that beats the SSE ceiling","Honeypot then trap locks UCB followers past classical SSE value","Leader inflates UCB index then exploits to exceed strong Stackelberg","Optimism vulnerability lets leader surpass SSE with O(sqrt(T ln T)) cost","Deceptive Stackelberg control of UCB locks follower beyond SSE"],"cache_read_input_tokens":128,"weakest_assumption_plain":"There must exist a trap mix that is strictly better for the leader than the classical SSE value while the target action is strictly worse for the follower; without that profitable trap the whole improvement claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["UCB optimism is a Stackelberg vulnerability that beats the SSE ceiling","Honeypot then trap locks UCB followers past classical SSE value","Leader inflates UCB index then exploits to exceed strong Stackelberg","Optimism vulnerability lets leader surpass SSE with O(sqrt(T ln T)) cost","Deceptive Stackelberg control of UCB locks follower beyond SSE"]},"model":"grok-4.5","effort":"low","cost_usd":0.004534,"raw_usage":{"total_tokens":1347,"prompt_tokens":791,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":45340000,"prompt_tokens_details":{"text_tokens":791,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":476,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":791,"tokens_out":80,"duration_ms":4447,"temperature":1.0,"reasoning_tokens":476,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T10:59:27.133804+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a certified matrix game that meets the paper's separation conditions, compute the honeypot length that forces target-index dominance, run the trap, and check whether the observed lock-in length Delta satisfies Delta(L* - V_SSE) > tau(V_SSE - H_min); if the inequality fails on the realized path, the strict-improvement claim fails.","supporting_citations":[],"review_version":1}