Pith. sign in

REVIEW 4 major objections 6 minor 8 references

The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact

T0 review · 4 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This paper claims that the apparent failure of maximum-entropy equilibrium selection in Kuhn poker is a removable artifact: the 0.0205 coordinate gap is the curvature-amplified image of a 0.00083 entropy shortfall, not a genuine selection b

desk verdict A self-aware diagnostic that explains the Kuhn gap via curvature and shortfall, but the evidence reduces to one nonzero point and the η-sweep cannot rule out a moving objective. read the letter →

arxiv 2607.17543 v1 pith:DO4I4WLG submitted 2026-07-20 cs.AI cs.GTcs.LGcs.MA

classification cs.AIcs.GTcs.LGcs.MA
keywords maximum-entropyequilibriuminformationprojectionKuhnpokerregularizedNashdynamicsentropyshortfallcurvaturepolytopezero-sumgames
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Regularized solvers on zero-sum games with a convex set of Nash equilibria have been observed to select the maximum-entropy member of that set, with one apparent counterexample: in Kuhn poker the solver lands at a bluff frequency of 0.180 instead of the maximum-entropy 0.201. This paper argues that the gap is not a bias but a "curvature shadow": the entropy landscape near its peak is flat enough that a tiny, removable shortfall in achieved entropy appears as a visible coordinate offset. The central quantitative claim is that the gap equals the square root of twice the entropy shortfall divided by the peak curvature, gap ≈ sqrt(2δ/κ). The paper verifies this across five games, shows that only the sequential game has a nonzero shortfall, and demonstrates with a causal sweep of the regularization strength that driving the shortfall toward zero drives the gap toward zero along the predicted curve. If correct, the apparent Kuhn counterexample disappears and the maximum-entropy (I-projection) account of regularized equilibrium selection survives.

What carries the argument

The central identity is a second-order Taylor decomposition of the entropy landscape along the Nash segment: gap ≈ sqrt(2δ/κ), where δ = H⋆ − H(σ(c)) is the entropy shortfall from the maximum-entropy member, κ = −d²H/dc² at the peak is the local peak curvature, and gap is the coordinate distance between the solver's landing point and the maximum-entropy coordinate. Its work is to separate the two causes of any apparent selection failure: why the solver stops short of maximum entropy (δ) and how strongly that shortfall is amplified into a visible coordinate offset (κ). The identity also ensures that a solver landing exactly on the Nash segment has no off-manifold contribution to the gap.

What would settle it

Construct a sequential game whose Nash segment has a sharply curved entropy peak (κ several times larger than Kuhn's, e.g., κ ≈ 20) and where the solver leaves a measurable entropy shortfall near δ ≈ 1e-3. The law predicts a gap of approximately sqrt(2δ/κ) ≈ 0.010 times the appropriate factor, approaching exactness as the magnet weakens; observing a gap that persists near 0.02 regardless of κ or that fails to track sqrt(2δ/κ) as δ varies would settle against the paper's central claim.

Watch

Extended reading notes

Core claim

On a one-dimensional Nash segment, restricting the mean Shannon entropy to the segment and expanding around its interior maximum gives gap ≈ sqrt(2δ/κ), where δ is the solver's entropy shortfall and κ is the curvature of the entropy landscape at the peak. In Kuhn poker the measured shortfall is δ = 8.3e-4 and the measured curvature is κ ≈ 4.0, which predicts a gap of 0.0204, matching the observed 0.0205 gap to within 2e-4. The four matrix games studied have δ ≈ 0 and therefore no gap, even though two of them have flatter peaks than Kuhn's; only the sequential game leaves a nonzero shortfall. Weakening the magnet strength η drives δ toward zero and the gap down the predicted sqrt(2δ/κ) curve,

Load-bearing premise

The load-bearing premise is that the selection target is the member maximizing unweighted mean Shannon entropy over the Nash segment; if the solver's regularization actually targets a different objective—such as reach-weighted or Tsallis entropy—then δ is not a shortfall from the true target and the 'removable artifact' conclusion collapses.

Editorial extensions

If this is right

  • The Kuhn poker gap is explained as a removable shortfall rather than a counterexample to maximum-entropy selection, so the I-projection account of solver-dependent equilibrium selection is upheld up to a flatness-limited residual.
  • Across the five tested games, the gap law holds to within 2e-4, meaning the same decomposition applies to both matrix and sequential games with one-dimensional Nash segments.
  • In matrix games the solver reaches the maximum-entropy member exactly, so flatness alone never creates a gap; a visible offset requires both a positive entropy shortfall and a sufficiently flat peak.
  • A causal sweep of the magnet strength shows the gap is removable: as δ → 0, the gap follows the predicted square-root curve toward zero, with no fixed residual, until dynamics destabilize at a stability floor.
  • The paper makes a falsifiable prediction: a sequential game with a sharply curved entropy peak should show a tighter match to the maximum-entropy member, scaling like sqrt(κ_Kuhn/κ) relative to Kuhn.
  • The paper explicitly flags an objective-dependence limitation: the maximum-entropy target depends on the entropy used, and the analysis fixes unweighted mean Shannon entropy; if the solver's true regularizer targets a different entropy, the shortfall framing would need revision.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper hypothesizes but does not prove that reach-weighted counterfactual updates cause the shortfall in sequential games; a direct test would be to construct sequential games with increasing reach skew and see whether δ grows with skew while still following the sqrt(2δ/κ) curve.
  • The moving-target caveat for Tsallis entropy suggests that any family of regularizers indexed by an entropy parameter should be compared against its own matched maximum-entropy target; comparing against the Shannon target alone would manufacture a spurious curvature dependence.
  • Because the stability floor prevents reaching δ = 0 directly, the zero-gap conclusion is a limit statement; an independent route would be to find a sequential game with a sharp peak and a nonzero shortfall, converting the curvature half of the law into a directly observed causal sweep.
  • A practical implication for reporting solver behavior is that coordinate distance alone is misleading: two games with identical entropy shortfalls can look very different if one peak is flat and the other sharp, so reporting δ alongside coordinates would make such artifacts transparent.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper examines a single apparent counterexample to maximum-entropy equilibrium selection in regularized solvers: R-NaD in Kuhn poker converges to a Nash equilibrium at bluff coordinate 0.180 while the maximum-entropy member is at 0.201, a gap of about 0.021. The authors derive a local quadratic relation gap ≈ sqrt(2δ/κ), where δ is the entropy shortfall relative to the maximum-entropy member and κ is the curvature of the entropy landscape at its peak. They verify this relation across five games, finding that four matrix games have δ≈0 and therefore no gap, while Kuhn has δ=8.3e-4 and a predicted gap of 0.0204 against a measured 0.0205. A magnet-strength sweep on Kuhn reduces δ and the gap together with a fitted log-log slope of 0.5013, which the authors interpret as strongly supporting the hypothesis that the gap is a removable shortfall rather than a fixed selection bias. They conclude that the I-projection account survives the Kuhn counterexample up to a flatness-limited residual.

Significance. If the conclusion were established, it would resolve a nagging anomaly in the empirical study of equilibrium selection in regularized dynamics and would sharpen the distinction between shortfall-induced and curvature-induced offsets. The paper is valuable in several respects: it isolates a clean geometric quantity, the curvature κ, that converts a tiny entropy shortfall into a visible coordinate offset; it provides a fully reproducible, exact-tabular engine with no sampling noise; it honestly lists limitations, including the absence of a causal curvature sweep and the objective-dependence of the target; and it warns convincingly against a naive Tsallis-entropy experiment where the target itself moves. However, the central evidentiary weight is much weaker than the prose suggests: the relation is a Taylor identity, the nonzero part of the empirical test rests on a single game, and the objective-dependence caveat may undermine the 'removable artifact' interpretation rather than merely bound it. The paper is a useful diagnostic note, but as it stands it does not establish that the Kuhn gap is a shortfall relative to R-NaD's actual objective.

major comments (4)
  1. [§3, Eq. (4); §5.2, Table 2] The relation gap ≈ sqrt(2δ/κ) is a tautological Taylor expansion, not an empirically falsifiable law. δ is defined as H*−H(σ(c)) and gap as |c−c*|, both read from the same unweighted mean-entropy function on the same Nash segment. Therefore any on-manifold point near c* satisfies the relation by construction. The agreement in Table 2 is a consistency check, not independent evidence. The four matrix-game rows have δ≈0 and hence trivially satisfy the relation; only Kuhn provides a nonzero point, and that single point lies on the curve by definition. The statement in §3 that verifying (4) 'empirically confirms that its only deviation from c* is the entropy shortfall' is circular: the coordinate c was chosen precisely as the segment coordinate, and the converged profile was verified to coincide with the family member. This does not establish which objective the solver is optimizing.
  2. [§5.3, Table 3, Fig. 3] The magnet sweep and the fitted exponent 0.5013 do not discriminate H-flat from H-bias. If c(η) remains on the smooth segment and δ(η)→0, then continuity forces gap(η)→0 and the log-log slope must approach 1/2 regardless of whether the finite-δ deviation is a genuine optimization shortfall or a systematic difference between two distinct objectives (unweighted Shannon vs. reach-weighted entropy). A moving-target model in which the effective target approaches c* linearly in η also predicts gap∝η and δ∝η², hence the same slope. Thus the claim that the test comes out 'decisively for H-flat' is unsupported. Moreover, the alternative H-bias as defined in §1 — a nonzero gap as δ→0 — is incompatible with the quadratic expansion itself, so it is not a meaningful competing hypothesis within this framework.
  3. [§2, Eq. (2); §6; §7] The objective-dependence problem is load-bearing, not a mere limitation. R-NaD's update (2) uses reach-weighted counterfactual values, and Section 6 hypothesizes that this reach-weighting 'distorts the moving-reference QRE sequence away from the unweighted maximum-(mean-)entropy member.' If that hypothesis is correct, then at finite η the measured δ is not a shortfall from R-NaD's actual target; it is the systematic gap between two different selection targets. Because the η-sweep cannot reach η=0 (the stability floor at η≈0.15 is an observed obstruction, not a limit point), the paper never demonstrates that the reach-weighted target coincides with the unweighted maximum-entropy member in the limit. Without independent evidence — for example, computing the reach-weighted max-entropy member and comparing δ to the distance to that target, or testing a correctly matched Tsallis target — the
  4. [§5.2, §7, §8] The paper's empirical support for the general claim rests on a single sequential exemplar. Only Kuhn has δ>0; all matrix games have δ≈0 and thus provide no information about the gap law for nonzero shortfalls. The paper acknowledges this ('Single sequential exemplar') but the conclusion and title generalize from this one case. The falsifiable prediction in §8 for a sharply curved sequential game is precisely the experiment needed to give the curvature half independent support, and it is not run. A matched-target Tsallis experiment would also help. As it stands, the paper is a detailed case study of one game rather than a established general decomposition.
minor comments (6)
  1. [§5.3, Table 3] The 'stable' column reports NashConv≤3e-12 for all stable rows; consider stating the convergence criterion for the limit-cycle row (η=0.10) more explicitly, since NashConv=0.11 is far from equilibrium.
  2. [§4] The curvature-fit window is a methodological choice; the sensitivity range ±5% to ±20% is reported in §5.2, but it would help to state the exact window used for Table 1 and Figure 3 (the text says |c−c*|≤0.10, but the 'range' units are per-game and not defined until Appendix A).
  3. [§5.1, Figure 1] The shaded band 'within 99.7% of H*' is described verbally; a precise definition (i.e., the set of c satisfying H(σ(c)) ≥ 0.997 H*) would improve reproducibility.
  4. [§5.3] The phrase 'the fit's intercept independently recovers κ=3.88' could mislead: the intercept and slope are not independent in a two-parameter log-log fit; the recovered κ is a derived quantity with its own uncertainty, and the 3% agreement is consistent with the local curvature estimate but not an independent confirmation.
  5. [§7] The initialization-dependence paragraph reports coordinates 0.057–0.258 over ten random initializations; consider adding a small table or figure, since this is the only stochastic element and is relevant to the scope of Conjecture 1.
  6. [Throughout] Minor language: 'R-N aD' formatting is inconsistent (spacing after hyphen); 'numpy' should be 'NumPy'; 'Kuhn poker' is sometimes written as 'Kuhn' and sometimes 'kuhn' — unify game names.

Circularity Check

3 steps flagged · score 5.0 of 10

The headline gap law is a Taylor identity read off the same entropy function whose fitted curvature and measured shortfall are then called a prediction; the independent content is limited to the causal η-sweep.

  1. self definitional [Section 3, Eqs. (3)-(4)]
    "H(σ(c)) = H ⋆ − 1 2 κ(c−c ⋆)2 +O((c−c ⋆)3). ... A point on the segment at entropy shortfall δ := H ⋆ −H (σ(c)) therefore satisfies 1 2 κ(c−c ⋆)2 ≈δ , i.e. gap :=|c−c ⋆| ≈ p 2δ/κ . (4)"

    Eq. (4) is a rearrangement of the Taylor expansion (3), with κ defined as the second derivative of H at c⋆ and δ defined as the entropy gap from H⋆. For any point on the Nash segment, δ = (1/2)κ·gap² up to cubic terms, so the two quantities are not independent. 'Verifying' the gap law across games is therefore a tautological check of the quadratic approximation rather than an empirical discovery. The paper itself concedes this in Section 6, calling it 'a local geometric identity—a second-order Taylor readout,' yet the abstract and Section 5 present the agreement as a prediction holding across five games.

  2. fitted input called prediction [Section 4 (curvature estimation) and Section 5.2 (Table 2)]
    "For the curvature we evaluate H(σ(c)) on an 801-point uniform grid over the coordinate range and estimate κ by a least-squares quadratic fit to H on the window |c−c ⋆| ≤0.10 (range)."

    The 'predicted' Kuhn gap √(2δ/κ) = 0.0204 is computed using κ = 4.0 obtained by fitting the same entropy landscape H(σ(c)) from which δ = 8.3×10⁻⁴ is measured. The agreement with the observed gap 0.0205 to within 2×10⁻⁴ therefore only confirms that the quadratic fit is a good local approximation of H; it is not an independent prediction of the solver's landing point. The fitted κ carries no information beyond the shape of the landscape, so this is a self-consistency check, not a test of the removal hypothesis.

1 more flagged steps
  1. self citation load bearing [Section 2, Conjecture 1; Section 1; Conclusion]
    "Conjecture 1 (I-projection selection; Conjecture 1 of Leal, 2026). R-N aDinitialized and referenced at the uniform strategy selects the I-projection of the uniform reference onto NE, i.e. the maximum-entropy equilibrium."

    The theoretical account that the paper aims to rescue—'the I-projection account is upheld'—is imported as Conjecture 1 from the same author's prior work (Leal, 2026). The paper's four matrix games provide some independent support for the conjecture in those cases, but the general maximum-entropy selection claim, and the framing of Kuhn as an 'apparent failure' requiring removal, rest on this self-citation. The conclusion therefore partly states that the author's own conjecture is consistent with the new Kuhn data rather than being derived from an independent source.

full rationale

The central quantitative result, gap ≈ √(2δ/κ), is not an empirical law but a second-order Taylor identity: κ is defined as −H''(c⋆), δ is defined as H⋆ − H(σ(c)), and the relation follows by rearranging the quadratic expansion. Consequently, the five-game 'verification' is a consistency check of the quadratic approximation, and the Kuhn 'prediction' 0.0204 vs 0.0205 is a self-consistency readout from the same entropy function whose fitted κ and measured δ are the inputs. The paper transparently acknowledges this in Section 6, which lowers the severity of the circularity but does not remove the fact that the headline law is constructional rather than predictive. The genuinely independent evidence is the η-sweep: varying the magnet drives δ toward zero and the landed coordinate toward c⋆, with the gap shrinking along the predicted curve until a stability floor. That causal sweep does discriminate removability from a fixed bias, though the claimed exponent 1/2 is itself the generic Taylor prediction and cannot by itself separate H-flat from a small residual bias over the accessible δ range. A minor self-citation (Conjecture 1 from Leal, 2026) is load-bearing for the interpretation but is partially counterbalanced by the paper's own matrix-game exact matches. Overall, partial circularity: the quantitative 'prediction' reduces by construction, while the qualitative removal claim retains independent experimental content.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central relation needs no free physical parameters, but it is a local identity on the same H that supplies δ, κ, and gap, so the ledger is dominated by estimation choices and the inherited entropy-objective assumption rather than by new fitted quantities.

free parameters (2)
  • Peak curvature κ of the entropy landscape = kuhn: 4.01, asym safe: 20.08, pennies safe: 2.27, two safe: 2.01, dup action: 4.02
    Estimated by least-squares quadratic fit to H(σ(c)) on an 801-point grid over window |c−c⋆|≤0.10 (Section 5.1). The central prediction gap≈sqrt(2δ/κ) depends on κ; a 3% intercept-recovery cross-check gives κ=3.88.
  • Curvature-fit window width = 0.10 (coordinate units)
    Hand-chosen estimation window; varying it from ±5% to ±20% changes κ across 3.99–4.09, which the paper uses as the error bar on the predicted gap.
assumptions (4)
  • standard math H(σ(c)) is C² in a neighborhood of c⋆ with H′(c⋆)=0 and κ>0 (interior maximum).
    Used in Section 3 to expand H and derive gap≈sqrt(2δ/κ); verified for all five games by checking c⋆ is strictly interior and by the quadratic fit.
  • domain assumption The Nash equilibrium set of each panel game is a one-dimensional segment parameterized by c, and R-NaD's converged profile lies exactly on that segment.
    Required for the scalar gap law to apply; verified via explicit families σ(c) with NashConv≤1.1e-16 and on-manifold coincidence ~1e-11 (Section 5.2).
  • domain assumption The selection target is the unweighted mean Shannon entropy (I-projection of a uniform reference).
    Inherited from Conjecture 1 (Leal, 2026); the paper does not prove it and notes objective-dependence in Section 7 and Figure 5. If the true target were reach-weighted or Tsallis entropy, δ would not be a true shortfall.
  • ad hoc to paper The η→0 limit is the right causal handle and the stability floor at η≈0.15 does not conceal a fixed bias.
    The removability conclusion extrapolates the magnet sweep across a stability floor; the paper explicitly calls the zero-gap result a limit statement, not an observed endpoint (Section 7).

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact." pith.science (2026). https://pith.science/paper/DO4I4WLG

@misc{pith2026260717543,
  author       = {Pith},
  title        = {Pith review of: The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DO4I4WLG}},
  note         = {Machine review of arXiv:2607.17543}
}
abstract

In two-player zero-sum games whose Nash equilibria form a convex set, regularized solvers such as Regularized Nash Dynamics (R-NaD) empirically select the maximum-entropy member: the information projection (I-projection) of a uniform reference onto the Nash set. On a panel of small games this match is exact, with one apparent exception: in Kuhn poker R-NaD lands at bluff coordinate 0.180 while the maximum-entropy member sits at 0.201, a coordinate gap of about 0.021, even though R-NaD attains 99.7 percent of the maximum entropy. We ask whether this gap is a genuine selection bias or an artifact, and answer it quantitatively. We show that for selection on a one-dimensional Nash manifold the coordinate gap factorizes as $\mathrm{gap} \approx \sqrt{2\delta/\kappa}$, where $\delta$ is the entropy shortfall of the solver and $\kappa$ is the curvature of the entropy landscape at its peak. Across five games this relation holds to within $2 \times 10^{-4}$ (under 1 percent relative error). The four matrix games have $\delta \approx 0$ (R-NaD reaches the maximum-entropy member exactly) and therefore no gap regardless of curvature; only the sequential game (Kuhn) has $\delta > 0$. A causal sweep of the magnet strength drives $\delta \to 0$ and the gap toward zero along the predicted curve (fitted scaling exponent 0.50, $R^2 > 0.999999$, against the exact prediction of 1/2), until the dynamics destabilize at a stability floor: behavior consistent with a removable shortfall and inconsistent with a fixed bias. We quantify the curvature half of the law from measured curvatures and flag a moving-target pitfall in the natural Tsallis-entropy experiment. The Kuhn gap is thus the curvature shadow of a small, removable entropy shortfall on an unusually flat peak; the I-projection account is upheld up to a flatness-limited residual.

Figures

Figures reproduced from arXiv: 2607.17543 by the authors.

Figure 1
Figure 1. Mean-entropy landscapes H(σ(c)) over each Nash segment. Dotted line: the maximum-entropy coordinate c ⋆ . Shaded: the band of coordinates within 99.7% of H⋆ . Flatter peaks (smaller κ) yield wider bands; Kuhn (κ= 4.0) is flatter than the sharp matrix games but not the flattest overall. of coordinates within 99.7% of H⋆ . The band is wide where the peak is flat and narrow where it is sharp: a fixed entropy tolerance … view at source ↗
Figure 2
Figure 2. Measured coordinate gap vs. the prediction p 2δ/κ across the five games. Points lie on the diagonal; the matrix games cluster at the origin (δ ≈ 0), Kuhn alone has a nonzero gap fully accounted for by its shortfall and curvature. 5.3 Causal test on the shortfall: weakening the magnet H-flat predicts the gap is removable: drive δ → 0 and the gap should follow p 2δ/κ → 0. We sweep 15 magnet strengths η ∈ [0.10, 0.80] … view at source ↗
Figure 3
Figure 3. Kuhn magnet sweep in the (δ, gap) plane. Dashed: the prediction p 2δ/κ for κ = 4.0. Stable runs (blue) lie on the curve and move toward the origin as η decreases; the unstable run (red ×) leaves the curve once the dynamics limit-cycle. Comparing such a solver against the Shannon maximum-entropy coordinate would conflate this moving target with the curvature effect and produce an uninterpretable trend; any Tsallis cu… view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: The selection target moves with the en￾tropy index. Maximum-Tsallis-q coordinate c ⋆ q over the fixed Kuhn family vs. q; dotted line is the Shannon target. A Tsallis-regularized solver must be compared to c ⋆ q , not the Shannon c ⋆ . Why does the shortfall appear only…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

8 extracted references · 2 linked inside Pith

  1. [1]

    Csisz\'ar

    I. Csisz\'ar. I-divergence geometry of probability distributions and minimization problems. The Annals of Probability, 3(1):146--158, 1975. doi:10.1214/aop/1176996454 https://doi.org/10.1214/aop/1176996454

  2. [2]

    H. W. Kuhn. A simplified two-person poker. In Contributions to the Theory of Games, vol. 1, pp. 97--103. Princeton University Press, 1950. doi:10.1515/9781400881727-010 https://doi.org/10.1515/9781400881727-010

  3. [3]

    L. Leal. Which Nash equilibrium? Solver-dependent selection on zero-sum Nash polytopes. arXiv:2606.28308, 2026

  4. [4]

    R. D. McKelvey and T. R. Palfrey. Quantal response equilibria for normal form games. Games and Economic Behavior, 10(1):6--38, 1995. doi:10.1006/game.1995.1023 https://doi.org/10.1006/game.1995.1023

  5. [5]

    Perolat, R

    J. Perolat, R. Munos, J.-B. Lespiau, S. Omidshafiei, M. Rowland, P. Ortega, N. Burch, T. Anthony, D. Balduzzi, B. De Vylder, G. Piliouras, M. Lanctot, and K. Tuyls. From Poincar\'e recurrence to convergence in imperfect information games: Finding equilibrium via regularization. In ICML, pp. 8525--8535, 2021

  6. [6]

    Perolat, B

    J. Perolat, B. De Vylder, D. Hennes, E. Tarassov, F. Strub, et al. Mastering the game of Stratego with model-free multiagent reinforcement learning. Science, 378(6623):990--996, 2022. doi:10.1126/science.add4679 https://doi.org/10.1126/science.add4679

  7. [7]

    Sokota, R

    S. Sokota, R. D'Orazio, J. Z. Kolter, N. Loizou, M. Lanctot, I. Mitliagkas, N. Brown, and C. Kroer. A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games. In ICLR, 2023. doi:10.48550/arxiv.2206.05825 https://doi.org/10.48550/arxiv.2206.05825

  8. [8]

    Zinkevich, M

    M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione. Regret minimization in games with incomplete information. In NeurIPS, 2007

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.