{"id":"d4ec0158-f167-487d-9dec-034b87c745a8","arxiv_id":"1908.09021","paper_version":8,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A smooth variant of regret matching, moving strategies fractionally toward the regret vector, converges in some tested games but cycles in others, with no proven guarantee.","lead":"This paper proposes a new 'smooth' regret matching rule for approximating Nash equilibria, in which players move their mixed strategies only partway toward the regret vector. Numerical tests show convergence in several small games, but the paper reports a non-convergent example and provides no formal convergence proof.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The universal sufficiency claim is contradicted by the paper's own 3X3-1eq3sp example, where the unique equilibrium is a repellor and the generated sequence does not converge.","rationale":"The reader's strongest claim correctly identifies the load-bearing flaw: the abstract's sufficiency statement is contradicted by the paper's own 3X3-1eq3sp experiment, where the unique equilibrium is a repellor and the sequence cycles indefinitely. My independent reading agrees. The per-player monotonicity theorems are internally sound but do not cover the simultaneous-adjustment dynamics; the paper itself concedes this in Section III and Section V. The only way to preserve the abstract claim would be to impose an unproved and unquantified infinitesimal-rate condition, but the paper does not state the claim that way and does not provide the needed convergence theorem. In addition, Algorithm 1's practice of retaining the best regret-sum point along the trajectory makes the numerical output look more convergent than the actual sequence, which weakens the empirical support further. The paper is honest about many limitations, but the mismatch between the broad sufficiency claim and the counterexample is decisive. No adversarial reinterpretation is needed: the paper's own figure and text describe non-convergence for a game with a unique equilibrium. Therefore the reader's REJECT verdict is appropriate, and this stress-test pass does not change it.","tokens_in":13269,"tokens_out":3229,"duration_ms":37861,"concrete_test":"Run Algorithm 1 on the 3X3-1eq3sp bimatrix from Section IV with decreasing adjustment rates rr=rc=10^{-k} for k=1,...,6 and sufficiently large T (e.g., 10^6). Track the overall regret sum and the metric distance d_t between consecutive iterates. If, as r tends to zero, the sequence still enters a limit cycle with liminf overall regret sum bounded away from zero, the claimed sufficiency fails even under the infinitesimal-rate interpretation. If instead the regret sum tends to zero as r decreases, the authors' informal 'approximately unilateral' hope could rescue a restricted version of the claim, but it would still need to be proved rather than assumed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that continuously and smoothly suppressing unprofitable pure strategies is sufficient for the game to evolve towards a Nash equilibrium. This claim is not merely unproven; it is contradicted by the paper's own Section IV example 3X3-1eq3sp. There, the game has one unique equilibrium, the equilibrium is a repellor, and the sequence generated by the update rule in Eq. (5) does not converge: it develops into a seemingly perfect circle away from the equilibrium. This is exactly the smooth suppression mechanism advertised in the abstract, so the counterexample targets the sufficiency claim directly. The per-player monotonicity results in Theorems 2 and 3 concern unilateral adjustment only; Section III's hope that infinitesimal adjustment rates make simultaneous adjustments approximately unilateral is never formalized or quantified. Section V confirms that regret sums fluctuate along the joint iteration and that no function is provided to determine convergence from the game data. The algorithm also selects the best point along the trajectory as its output, which can hide non-convergence, but the sequence itself still fails to approach equilibrium. Thus the evidence in the paper supports a limited heuristic for some games, not the stated universal sufficiency claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a \"geometrical regret matching\" update, Eq. (5), in which each player's mixed strategy is moved a small step toward its current regret vector, with the step controlled by a positive adjustment rate r_i. The authors prove three local results: the update increases the angle closeness to the regret vector, weakly increases the player's own payoff, and weakly decreases the player's own regret sum when the player adjusts unilaterally (Theorems 1-3). They then consider the joint iteration in which all players update simultaneously, define the induced sequence of strategy profiles, and claim in the abstract and conclusion that continuously and smoothly suppressing unprofitable pure strategies is sufficient for the game to evolve toward a Nash equilibrium. The paper supports this claim with numerical experiments on several 3x3 games, a 60x40 game, and an n-person game, and discusses approximation accuracy through regret sums, step sizes, and a candidate contraction ratio. The paper candidly acknowledges in Section V that regret sums fluctuate, that convergence cannot be predicted from a formula, and that one of its own examples, 3X3-1eq3sp, does not converge.","tokens_in":13470,"tokens_out":3058,"duration_ms":33163,"significance":"If the central claim were established, the paper would provide a simple, decentralized learning rule with monotone payoff and regret improvement for unilateral adjustments, and it would give a general argument for why equilibrium tendencies might be pervasive. The local unilateral theorems are elementary and appear correct, and the paper is transparent about its numerical evidence, including a case of non-convergence, and provides source code. However, the advertised sufficiency claim is not proven, and it is contradicted by the paper's own 3X3-1eq3sp example, where the unique equilibrium is a repellor and the trajectory forms a periodic circle. The manuscript therefore offers a useful heuristic and a collection of numerical observations, but not a general convergence theory for the proposed iteration.","major_comments":[{"comment":"The abstract and Section VII claim that continuously and smoothly suppressing unprofitable pure strategies is sufficient for the game to evolve towards a Nash equilibrium. This claim is directly contradicted by the paper's own example 3X3-1eq3sp in Section IV: the game has a unique equilibrium, that equilibrium is described as a \"repellor\" (Fig. 11), and the generated sequence does not converge but develops into a seemingly perfect circle away from the equilibrium (Fig. 8). Since this example uses exactly the same smooth update rule advertised in the abstract, the universal sufficiency statement cannot stand as written.","section":"Abstract and Section IV (3X3-1eq3sp, Figs. 8 and 11)"},{"comment":"Theorems 2 and 3 are proven only for a unilateral adjustment of one player while the strategies of the other players are held fixed. The extension to the joint iteration, where all players update simultaneously, rests on the informal hope, stated in Section III, that \"infinitesimal\" adjustment rates make simultaneous adjustments approximately unilateral. This assumption is never quantified, and no theorem is provided to show that the per-player monotonicity survives in the joint dynamics. Section V explicitly reports that regret sums do fluctuate along the joint iteration, so the local unilateral results do not imply convergence of the coupled system.","section":"Section III, Eqs. (11)-(12) and Theorems 2-3"},{"comment":"Algorithm 1 outputs the strategy profile that achieved a new minimum of the overall regret sum during the iteration, rather than the limit of the generated sequence. This selection can hide non-convergence: for a periodic trajectory such as that of 3X3-1eq3sp, some point on the cycle may have a small regret sum even though the sequence never approaches the equilibrium. Claims about convergence or approximation to a Nash equilibrium should be evaluated from the generated sequence itself, and the paper's numerical evidence should be reported accordingly.","section":"Section IV, Algorithm 1, steps 6-9"},{"comment":"The paper explicitly concedes that it cannot provide a function chi(A,B,r_r,r_c,S0) to determine whether a given initial condition leads to convergence, nor a function chi(A,B,r_r,r_c) to determine whether every initial condition converges. This admission, together with the non-convergent 3X3-1eq3sp example, means the paper's evidence for its central claim is a small set of numerical simulations rather than a mathematically supported sufficiency result. The conclusion should be weakened to a conjecture or a heuristic claim for games whose dynamics are attracted to equilibrium.","section":"Section V, discussion after Eq. (18)"}],"minor_comments":[{"comment":"The sentence stating that there is \"less proportion of pure strategy pi_ij used in s_i than in s'_i\" when the corresponding regret component is zero appears to have the comparison reversed; the update decreases the weight of strategies with zero regret component, so the proportion should be smaller in s'_i than in s_i.","section":"Section III, after Eq. (5)"},{"comment":"The sentence \"We define each player's regret matching to be a function R_i,S : I_i -> I_i\" appears to contain a typo; the symbol being defined is the matching function psi_i,S, not the regret vector R_i,S.","section":"Section III, first paragraph of Section III"},{"comment":"The discussion of scaling factors (a_r,a_c) and offsets (b_r,b_c) would be clearer if the figures showed the payoff scales and adjustment rates used, since the comparison of regret sums across different parameter values is otherwise difficult to interpret.","section":"Section V, Figs. 15 and 16"},{"comment":"The connection between the neural firing papers cited as references 11 and 12 and the game-theoretic fixed point iteration in this paper is stated only briefly; a short explanation of how the FPI 2 results support the current claims would help the reader assess the relevance.","section":"Section VI"}],"recommendation":"reject","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is a clearly written heuristic with two correct local theorems and an overclaimed abstract. The update rule (Eq. 5) is a natural smooth variant of Hart-Mas-Colell regret matching and Nash's map, and the geometric picture is genuinely nice. Theorems 2 and 3 show that a single player's payoff increases and regret sum decreases under unilateral adjustment, and the proofs are straightforward. The paper also publishes code and is unusually honest about limitations: it admits that regret sums fluctuate, that convergence cannot be predicted from game data, and that the whole thing depends on 'naked-eye observation.'\n\nThe problem is the central claim. The abstract says that smoothly suppressing unprofitable pure strategies is 'sufficient' for the game to evolve toward Nash equilibrium. The paper's own 3X3-1eq3sp example shows a game with a unique equilibrium that is a repellor; the strategy path never converges and instead forms a circle around the equilibrium. That is a direct counterexample to the sufficiency claim. The informal assumption in Section III that infinitesimal adjustment rates make simultaneous updates 'approximately unilateral' is never formalized, and Algorithm 1's practice of reporting the best point along the trajectory masks non-convergence. The local theorems are correct, but they only cover unilateral adjustment and do not imply convergence of the joint iteration.\n\nSo the paper's real contribution is a limited heuristic for some games, not a general theorem. That said, the heuristic is simple enough to be worth studying, and the repellor example is an informative cautionary case. If this crosses your editorial desk, I would send it to a referee rather than desk-reject: the right referee can judge whether a revised version that reframes the claim as a heuristic with an open convergence question is worth publishing. As written, the mismatch between the abstract and the evidence is a serious flaw.","headline":"A simple smooth regret-matching heuristic with correct local results, but its own 3X3-1eq3sp example disproves the sufficiency claim in the abstract.","tokens_in":14017,"tokens_out":2796,"would_cite":false,"duration_ms":27827,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A06","91A10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a smooth, geometry-based regret update is sufficient to steer play toward Nash equilibrium.","keywords":["geometrical regret matching","Nash equilibrium approximation","regret matching","smooth strategy updating","fixed point iteration","mixed strategies","non-cooperative games"],"falsifier":"A simple numerical test settles it: choose a 3x3 game with a unique fully mixed equilibrium and iterate Eq. (5) from a nearby starting point; if the iterates orbit the equilibrium as they do in the paper's own 3X3-1eq3sp example, the blanket sufficiency claim is false as stated.","tokens_in":12987,"feed_emoji":"🎲","tokens_out":6976,"duration_ms":68173,"temperature":0.7,"pith_summary":"The paper argues that conventional regret matching updates strategies in a jumpy way, and that this jumpiness is not necessary. It proposes a geometrical update rule in which a player's mixed strategy is nudged toward its current regret vector by a small angle; the smaller the adjustment rate, the smoother the nudge. The central claim is that continuously and smoothly suppressing pure strategies that perform below average is enough to push a game toward Nash equilibrium. The paper derives monotonic payoff and regret improvements for a single player's update, and supports the broad claim with numerical paths that converge for two-person and many-person games. Because the update needs only regret information, not knowledge of opponents' payoffs, the claim makes equilibrium-looking behavior plausible as a spontaneous outcome of play.","feed_headline":"Smoothly pruning losing moves guides games to Nash equilibrium","feed_subtitle":"A geometric update grows payoffs and shrinks regrets, making equilibrium the natural end state of play.","key_machinery":"The load-bearing object is the regret vector $R_i(s_i)$, whose components $\\phi_{ij}(s_i) = \\max\\{0, \\bar p_i(\\pi_{ij}; S) - p_i(s_i)\\}$ are the payoff gains a player would have received by switching unilaterally to each pure strategy. The update in Eq. (5) forms a normalized combination of the current mixed strategy and this regret vector, keeping the result inside the simplex; Theorem 1 states that this operation never increases the angle between the mixed strategy and the regret vector. A small adjustment rate $r_i$ makes the movement smooth, which is what lets the paper treat simultaneous adjustments as approximately unilateral and interpret the rule behaviorally as suppressing below-average pure strategies. The fixed-point framing comes from observing that an equilibrium is exactly a point at which every player's regret vector is zero, so the iteration $\\Psi^t(S_0)$ is a fixed-point search for $\\Psi$.","core_discovery":"The paper's central claim, stated in the abstract, is that continuously and smoothly suppressing \"unprofitable\" pure strategies is sufficient for a game to evolve toward Nash equilibrium. The mechanism is the geometric update of Eq. (5), $s_i' = \\frac{s_i + r_i R_i(s_i)}{1 + r_i \\lvert R_i(s_i) \\rvert}$, which moves the mixed strategy toward the regret vector rather than jumping proportionally to positive regret. For a single player this update never lowers payoff and never increases the regret sum, unless the strategy is already optimal; the paper interprets the classic proportional-to-positive-regret rule as the limit $r_i \\to \\infty$ of this smoother rule. The paper extends the update to simultaneous play and reports numerical convergence for two-person and n-person examples, with motion toward the equilibrium largely independent of the initial mixed strategies. It also reports that for at least one game with a unique interior equilibrium the sequence cycles around that equilibrium rather than converging, which it treats as a limitation of approximation accuracy rather than a breakdown of the behavioral principle.","pith_inferences":["A testable extension would be to classify games by the stability of equilibrium under Eq. (5): if the unique equilibrium is interior and the Jacobian of $\\Psi$ at the fixed point has an eigenvalue outside the unit disk, the sequence should cycle rather than converge, matching the reported repellor case.","If the sufficiency claim held, the same smooth-suppression principle should transfer to settings with many players and noisy payoffs, since the update requires no cross-player information beyond observed play.","The generalized form of Eq. (23), with component-wise functions $\\alpha_{ij}$, suggests a family of dynamics indexed by how aggressively each pure strategy is pruned; comparing members of this family could reveal a minimal smoothing level needed for convergence.","The paper leaves open the convergence criterion that it admits it lacks; supplying a function $\\chi$ from payoff matrices to attractor/repellor status would turn the empirical observation into a theorem."],"forward_implications":["A player following the unilateral update never sees its payoff fall or its regret sum rise, and reaches a fixed point exactly when its mixed strategy is optimal against the opponents' current strategy.","With small adjustment rates, simultaneous play should approximate the unilateral property, so the paper's examples exhibit paths that land on attractor equilibria and avoid repellor equilibria.","The classical proportional-to-positive-regret rule is recovered as the limit of large $r_i$, giving that rule a rationale: it is the aggressive end of smooth suppression.","The construction needs only local regret information, so if the conclusion holds, it describes a distributed process by which players can converge to equilibrium without knowing payoff matrices or the existence of an equilibrium.","Because convergence fails for repellor equilibria, the paper's own metric-based diagnostics ($\\dot d$, $\\dot q$) can serve as empirical accuracy measures for approximate equilibria."],"supporting_citations":[{"why":"introduces the regret-matching procedure whose jumpy updating this paper replaces with a smooth geometric rule","marker":"[1]"},{"why":"extends the adaptive-strategy framework the paper positions its rule within","marker":"[2]"},{"why":"supplies the payoff-gain function used to define the regret vector","marker":"[6]"},{"why":"provides the non-adaptive equilibrium algorithm that motivates the no-payoff-knowledge setting","marker":"[3]"},{"why":"gives the fixed-point theorem used to frame convergence of the iteration","marker":"[8]"},{"why":"supplies PCA for visualizing high-dimensional strategy paths in the numerical tests","marker":"[9]"}],"fun_headline_variants":["Smooth regret matching makes Nash equilibrium inevitable","Geometric regret updates guide games to equilibrium smoothly","Smoothly suppressing losing strategies finds Nash equilibrium","Regret matching without jumps converges to Nash equilibrium"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the assumption that sufficiently small adjustment rates make simultaneous strategy changes behave like unilateral ones, so that a single player's monotone improvement carries over to the joint dynamics; the paper states this as a hope and never quantifies how small the rates must be, and its own examples show joint regret sums can rise.","fun_headline_variants_meta":{"raw":{"variants":["Smooth regret matching makes Nash equilibrium inevitable","Geometric regret updates guide games to equilibrium smoothly","Smoothly suppressing losing strategies finds Nash equilibrium","Regret matching without jumps converges to Nash equilibrium"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000912,"raw_usage":{"total_tokens":3889,"prompt_tokens":887,"completion_tokens":3002,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":2944}},"tokens_in":503,"tokens_out":3002,"duration_ms":20821,"temperature":1.0,"reasoning_tokens":2944,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:45:43.233381+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A simple numerical test settles it: choose a 3x3 game with a unique fully mixed equilibrium and iterate Eq. (5) from a nearby starting point; if the iterates orbit the equilibrium as they do in the paper's own 3X3-1eq3sp example, the blanket sufficiency claim is false as stated.","supporting_citations":[{"cited_title":"Hart and A","cited_arxiv_id":null,"evidence_quote":"introduces the regret-matching procedure whose jumpy updating this paper replaces with a smooth geometric rule"},{"cited_title":"Hart and A","cited_arxiv_id":null,"evidence_quote":"extends the adaptive-strategy framework the paper positions its rule within"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the payoff-gain function used to define the regret vector"},{"cited_title":"Lemke and J","cited_arxiv_id":null,"evidence_quote":"provides the non-adaptive equilibrium algorithm that motivates the no-payoff-knowledge setting"},{"cited_title":"Istratescu","cited_arxiv_id":null,"evidence_quote":"gives the fixed-point theorem used to frame convergence of the iteration"},{"cited_title":"Jolliffe","cited_arxiv_id":null,"evidence_quote":"supplies PCA for visualizing high-dimensional strategy paths in the numerical tests"}],"review_version":1}