{"id":"9bb44e69-e84e-43e5-8bf0-1a0f1716a8e7","arxiv_id":"2502.01127","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"A new potential game shows that when influencers compete to shape a receiver's aggregate opinion, any pure Nash equilibrium forces all but at most one influencer to the most extreme allowed action.","lead":"The authors introduce a game where competing influencers each try to pull a shared receiver toward their own target, and they prove that in every equilibrium all but at most one player must push their action to the boundary of the allowed set. The result offers a formal reason why strategic preference reporting in AI alignment can favor exaggeration over truthfulness.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4 omits a nonzero-weight condition; with w_i=0 the all-but-one-interior conclusion fails, so the central exaggeration theorem is false as stated, though easily repaired.","rationale":"I read the paper as resting on the formal structure theorem for pure Nash equilibria of BIG, with Theorem 4 as the headline claim. The potential construction, convexity argument, and identification of pure Nash equilibria with minima of the potential are correct when each player has a nonzero influence weight. However, Definition 1 explicitly permits zero weights and Theorem 4 does not exclude them, so the theorem is false as stated. The two-player counterexample is simple and decisive: with w1=0, player 1 is indifferent over all actions, so an entire interval of interior actions forms a pure Nash equilibrium alongside player 2's interior best response, and the receiver does not equal player 1's target. This is a correctness gap in the paper's central claim, not merely an interpretive issue about exaggeration. It is narrow and easily repaired by assuming wi≠0, so I do not recommend rejection; the existing conditional verdict remains appropriate. I disagree with the reader's choice of weakest assumption: the affine receiver is a modeling choice that is explicit in the game definition and is reasonable for the formal results, whereas the missing nonzero-weight hypothesis directly invalidates the theorem that supplies the paper's most striking conclusion.","tokens_in":11916,"tokens_out":12854,"duration_ms":151294,"concrete_test":"Instantiate the two-player example above under Definition 1 and solve for the pure Nash equilibria either by checking condition (3) directly or by minimizing the potential function (5). The set of equilibria should be all profiles (x1,0.2) with x1∈(0,1), which already violates Theorem 4. If this counterexample is accepted, add an explicit w_i≠0 (or wi>0) condition to Theorem 4 and re-verify that every later application of the theorem is confined to that regime.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Definition 1 allows each w_i to be any real number, including 0, and Theorem 4 does not exclude zero weights. The proof breaks exactly at the step where ∇xi φ = 0 is divided by 2wi to conclude ti = Σ wk xk; that division requires wi ≠ 0. The failure is not merely formal. Take d=1, X=[0,1], w0=0, w1=0, w2=1, t1=0.8, t2=0.2. Then the receiver is x̂=x2, player 1's loss is constant in x1, and player 2 uniquely minimizes its loss at x2=0.2. Hence every profile (x1,0.2) with x1∈(0,1) is a pure Nash equilibrium. Both coordinates are interior, and the receiver value 0.2 is not t1, so both conclusions (7) and (8) of Theorem 4 fail. This is a concrete counterexample under the paper's stated definitions. The intended theorem is recovered by adding an explicit w_i≠0 (in the narrative, wi>0) assumption, and the main constructive arguments for that regime remain sound. The value-alignment transfer additionally depends on the unproven affine approximation in (28), but the zero-weight gap is the more immediate formal flaw in the paper's central claim.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Battling Influencers Game (BIG), an n-player simultaneous-move game in which each influencer chooses an action in a compact convex set and suffers a squared-distance loss between its own target and the output of a common affine receiver. The main formal results are that BIG is a potential game with a convex potential, that its pure Nash equilibria coincide with the minimizers of that potential, that the equilibrium set has cardinality one or infinity, and that, when all targets are distinct, every pure NE has at most one interior action, with the unique interior influencer, if any, being the one whose target equals the receiver's output. The paper also treats a finite-action variant, a weakly dominant strategy equilibrium for an inner-product loss, and a value-alignment experiment in which a maximum-likelihood preference-aggregation receiver is approximated as affine.","tokens_in":12149,"tokens_out":15867,"duration_ms":170188,"significance":"The intended exaggeration theorem is striking and non-obvious: rational influencers in a simple affine-receiver model are driven to the boundary of the action space, and at most one can remain interior. The potential-function construction is clean, the pNE-as-convex-minimizers reduction is useful, and the proof strategy is mostly sound. The paper is also careful to separate the formal game results from the empirical value-alignment discussion. However, the central theorem is false as stated because Definition 1 permits zero weights, and the main illustrative example contains a sign error. The value-alignment transfer rests on a single numerical assertion of affine approximation. These issues are repairable, but they require substantive revision rather than copy-editing.","major_comments":[{"comment":"The statement of Theorem 4 is false as written under Definition 1, because the weights w_i are allowed to be zero. For example, take d=1, X=[0,1], w0=0, w1=0, w2=1, t1=0.8, t2=0.2. The receiver is x_hat = x2, player 1's loss is constant in x1, and player 2's unique best response is x2=0.2. Hence every profile (x1,0.2) with x1 in (0,1) is a pure Nash equilibrium; both actions are interior, and x_hat=0.2 differs from t1, contradicting both conclusions (7) and (8). The proof divides by 2 w_i when passing from gradient_x_i phi = 0 to t_i = sum_k w_k x_k. Adding the explicit assumption w_i != 0 for all i (or, for the narrative, w_i > 0) repairs the theorem, but the theorem and the abstract's exaggeration claim need this qualification.","section":"Section 4, Theorem 4, Eq. (9)-(11)"},{"comment":"The displayed pure NE set is not an equilibrium under the receiver definition used in Example 1. With X=[-a,a]^2, t1=(1,0), t2=(-1,0), and x_hat=(x1+x2)/2, fix x2=(a,z) and let player 1 deviate to x1'=(2-a,-z), which lies in X for a>=1; the receiver then equals t1, giving player 1 a strictly lower loss than the receiver value 0 attained at the displayed profile. The line of equilibria for this example should instead be x1=(a,-z), x2=(-a,z) with z in [-a,a], up to the sign convention of the targets. This does not invalidate the proof of Theorem 4, but the example is used to illustrate the exaggeration phenomenon and must be corrected.","section":"Section 4, Example 2"},{"comment":"The transfer of the formal results to value alignment rests on the assertion that the MLE receiver in (27) is 'well-approximated' by the affine receiver x_hat = (x1+x2)/2. The paper supports this with a single configuration (n=2, X=[-1,1], t1=-0.1, t2=0.3, PY uniform on [-10,10]^2) and provides no quantitative measure of the approximation error, no variation of parameters, and no argument that the approximation persists under best-response iteration. The predictions in Section 6.1 depend on the receiver being affine, so this part should be presented as a case study or strengthened substantially. The formal game-theoretic results are independent of this point.","section":"Section 6.2, Eq. (28)"}],"minor_comments":[{"comment":"The common-knowledge parameter list includes k_0, although only k_1,...,k_n are defined; remove k_0 or define it explicitly.","section":"Section 5.2, Definition 5"},{"comment":"The phrase 'empirical based-response' should read 'empirical best-response'.","section":"Section 6.2, near Figure 7"},{"comment":"The remark following the definition correctly notes the difference between weak and strict dominance, but the definition's inequality uses <= for all opponent profiles; this is stated slightly ambiguously because the displayed condition should be read as holding for every x_{-i}, not for a fixed profile. Clarifying the quantifier would remove ambiguity.","section":"Definition 4"},{"comment":"The notation 'X ⊂Rd' is missing a space before R^d; more importantly, the paper should state explicitly that the narrative interpretation of w_i as an influence weight excludes w_i=0, since the formal definition does not.","section":"Section 3, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The zero-weight gap in Theorem 4 and the sign error in Example 2 are both easily fixable, and I do not see grounds for rejection. The value-alignment section is currently a single heuristic experiment; if the applied claim is important for the journal's scope, the authors should either expand the empirical/theoretical support or explicitly downgrade the claim to a case study. The formal game-theoretic core is sound after the nonzero-weight repair."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper has a load-bearing flaw in its central theorem, but the flaw is localized and easily repaired. The rest of the formal work is mostly solid.\n\nWhat's actually new and good: the formalization of strategic data providers in an alignment setting as a potential game is clean. Proposition 2 (pure NE = convex minima) is correct, Corollary 3 follows from convexity, and the finite-action exponential-NE example is a nice addition. The paper also gives appropriate credit to Park et al. and shows how the informal example there becomes a theorem here.\n\nThe soft spot is Theorem 4. As stated, it is false. Definition 1 allows each w_i to be any real number, including 0, and the proof divides by 2w_i to conclude t_i = sum w_k x_k = t_j. That division is invalid for zero weights. Concrete counterexample: d=1, X=[0,1], w0=w1=0, w2=1, t1=0.8, t2=0.2. Then every profile (x1, 0.2) with x1 in (0,1) is a pure NE, both coordinates are interior, and the receiver value 0.2 is not t1. Both conclusions (7) and (8) fail. The intended theorem is recovered by adding the explicit condition w_i ≠ 0 (presumably w_i > 0). The rest of the constructive arguments survive, but the statement as written is wrong.\n\nThe empirical section is illustrative rather than evidential. The claim that the MLE receiver is approximately affine (equation 28) is asserted without derivation, no code or data are shipped, and there are no error bars. That is acceptable for an application sketch, but not for a prediction you hang a conclusion on.\n\nWho this is for: algorithmic game theory readers interested in potential games, and AI alignment researchers thinking about data-collection incentives. The formal core is interesting enough to deserve a serious referee even with the flaw. A careful referee should catch the zero-weight issue and ask for the one-line repair.\n\nRecommendation: yes, send to peer review. The fix is local, and the rest is worth airing.","headline":"The paper's main structural theorem is false as stated because it ignores zero influencer weights, but the fix is a one-line condition and most of the rest is sound enough to warrant peer review.","tokens_in":12690,"tokens_out":2808,"would_cite":false,"duration_ms":30759,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A10","91A06","91A80"],"pacs":[],"model":"deepseek-v4-flash","headline":"When several influencers compete for a receiver, every pure Nash equilibrium forces all but at most one of them to a maximally extreme action.","keywords":["Battling Influencers Game","potential game","pure Nash equilibrium","extreme exaggeration","convex optimization","value alignment","preference feedback","strategic data providers"],"falsifier":"Compute a pure Nash equilibrium of any BIG instance with distinct targets and check whether more than one player's action lies in the interior of the action space; Theorem 4 predicts at most one, so a counterexample with two interior actions would refute the central claim.","tokens_in":11703,"feed_emoji":"⚔️","tokens_out":5338,"duration_ms":53018,"temperature":0.7,"pith_summary":"The Battling Influencers Game models several strategic agents who each try to pull a common receiver toward their own target, where the receiver simply takes a weighted sum of the agents' chosen actions. The paper proves this game is a potential game, so its pure Nash equilibria are exactly the minima of one convex function and can be found by convex optimization. The central structural result is that when the agents' targets are distinct, every pure Nash equilibrium pushes all but at most one agent to the boundary of the action space: rational influencers must exaggerate their positions to the maximum. If one agent does stay interior, the receiver lands exactly on that agent's target. The paper then argues this formal result can explain why people providing preference data for AI value alignment have an incentive to report exaggerated values.","feed_headline":"In influencer battles, rational play means maximal exaggeration","feed_subtitle":"In every pure Nash equilibrium, only one influencer can stay interior; the rest push to the boundary.","key_machinery":"The load-bearing object is the convex potential function $\\phi(x) = \\|\\sum_{i=0}^n w_i x_i\\|_2^2 - 2\\sum_{i=1}^n w_i t_i^\\top x_i$ on the product domain $X^n$. A unilateral change by any player changes $\\phi$ by exactly the same amount as it changes that player's loss, so pure Nash equilibria coincide with minima of $\\phi$; convexity then makes the equilibrium set convex, yielding the one-or-infinite cardinality. The boundary-exaggeration theorem follows from the gradient condition $\\nabla_{x_i}\\phi = 0$ for an interior action: if two players were both interior, their targets would have to equal the same receiver aggregate, contradicting distinctness.","core_discovery":"On its own terms, the paper establishes that for any instance of the Battling Influencers Game, in which each player minimizes $\\|w_0x_0 + \\sum_{i=1}^n w_i x_i - t_i\\|_2^2$ over a compact convex action space $X$, the pure-strategy Nash equilibria are exactly the global minima of the convex potential function $\\phi(x) = \\|\\sum_{i=0}^n w_i x_i\\|_2^2 - 2\\sum_{i=1}^n w_i t_i^\\top x_i$. Because $\\phi$ is convex, the game has either exactly one pure Nash equilibrium or infinitely many. The paper's main theorem (Theorem 4) states that if all targets $t_i$ are distinct, then at any pure Nash equilibrium at most one player's action lies in the interior of $X$; every other player must choose an extreme point of $X$. Moreover, if some player $i^*$ is interior, then the receiver's aggregate equals $t_{i^*}$, so the only non-extreme player is exactly the one whose target is realized.","pith_inferences":["We infer that the exaggeration result is likely fragile under nonlinear receivers: if the receiver's aggregate is not an affine function of the influencer actions, the convex-potential argument and the all-but-at-most-one theorem do not apply, and interior equilibria may reappear.","We infer a testable extension: measuring how far a real preference-learning receiver deviates from the affine form assumed in the paper would determine how directly the formal results transfer to value alignment.","We infer that the model offers a game-theoretic microfoundation for misinformation: even rational agents with moderate targets can end up broadcasting extreme positions, so observing extreme rhetoric need not imply extreme underlying beliefs."],"forward_implications":["In any equilibrium, all but at most one influencer will take an extreme action, so exaggeration is a rational response to competition rather than an individual bias.","Equilibria can be computed by minimizing a convex function, so finding them is tractable for compact convex action spaces.","Best-response dynamics converge to an equilibrium, and each influencer can update using only the others' current actions without knowing their targets.","The same structure extends to finite action spaces, where the game remains a potential game and can even have exponentially many pure equilibria.","Applied to value alignment, the model predicts that people supplying preference data will strategically distort their reported values, so removing that incentive is a mechanism-design problem rather than a matter of assuming truthful reporting."],"supporting_citations":[{"why":"Supplies the theorem that in a convex potential game the pure Nash equilibria coincide with the minima of the potential and form a convex set.","marker":"(Neyman, 1997)"},{"why":"Motivates the game with strategic data providers who misreport opinions; the paper formalizes their numerical example and informal theorem.","marker":"(Park et al., 2024)"},{"why":"Provides the convex optimization background used to compute pure Nash equilibria by minimizing the potential.","marker":"(Boyd & Vandenberghe, 2004)"},{"why":"Supplies the standard treatment of best-response dynamics in potential games used to find equilibria.","marker":"(Roughgarden, 2010)"},{"why":"Defines the Bradley-Terry-Luce pairwise-comparison model used in the value-alignment experiment.","marker":"(Bradley & Terry, 1952)"}],"fun_headline_variants":["In any equilibrium, all but at most one influencer go extreme","Only one influencer can avoid extremism in equilibrium","Rational play: maximal exaggeration for all but at most one","Value alignment: truthfulness loses, extremism wins in equilibrium"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the receiver's aggregate is the affine weighted sum in equation (1); the value-alignment experiment additionally assumes the maximum-likelihood receiver in equation (27) is approximately affine, so if a real receiver deviates substantially from affine behavior the formal results do not carry over.","fun_headline_variants_meta":{"raw":{"variants":["In any equilibrium, all but at most one influencer go extreme","Only one influencer can avoid extremism in equilibrium","Rational play: maximal exaggeration for all but at most one","Value alignment: truthfulness loses, extremism wins in equilibrium"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001326,"raw_usage":{"total_tokens":5383,"prompt_tokens":919,"completion_tokens":4464,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":4397}},"tokens_in":535,"tokens_out":4464,"duration_ms":33757,"temperature":1.0,"reasoning_tokens":4397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T16:29:15.270251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute a pure Nash equilibrium of any BIG instance with distinct targets and check whether more than one player's action lies in the interior of the action space; Theorem 4 predicts at most one, so a counterexample with two interior actions would refute the central claim.","supporting_citations":[{"cited_title":"Correlated equilibrium and potential games","cited_arxiv_id":null,"evidence_quote":"Supplies the theorem that in a convex potential game the pure Nash equilibria coincide with the minima of the potential and form a convex set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the game with strategic data providers who misreport opinions; the paper formalizes their numerical example and informal theorem."},{"cited_title":"and Vandenberghe, L","cited_arxiv_id":null,"evidence_quote":"Provides the convex optimization background used to compute pure Nash equilibria by minimizing the potential."},{"cited_title":"Algorithmic game theory","cited_arxiv_id":null,"evidence_quote":"Supplies the standard treatment of best-response dynamics in potential games used to find equilibria."}],"review_version":1}