Pith. sign in

REVIEW 1 minor 39 references

Coarse Q-learning: Indifference, Indeterminacy, and Instability

T0 review · 0 major / 1 minor · reviewed 2026-05-23 · grok-4.3

Pith's one-line read Coarse Q-learning can lead to multiple stable equilibria, class indifference, or limit cycles in the high payoff-sensitivity limit.

desk verdict CQL adds class-level pooling to Q-learning and derives new high-sensitivity outcomes like class indifference or limit cycles that the per-alternative benchmark lacks. read the letter →

arxiv 2412.09321 v6 submitted 2024-12-12 econ.TH cs.GT

classification econ.THcs.GT
keywords coarseq-learningreinforcementlearningbanditproblemsmean-fielddynamicsvaluationequilibrialimitcyclespayoffsensitivityindifference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper develops Coarse Q-learning for bandit problems where alternatives are grouped into similarity classes and feedback is pooled at the class level. Using stochastic approximation, it derives mean-field dynamics and characterizes steady states. In the high payoff-sensitivity limit, the dynamics can result in multiple stable strict equilibria, a globally stable mixed equilibrium with indifference, or a stable limit cycle with no equilibrium. These behaviors stem from the coarse aggregation and contrast with standard Q-learning at the alternative level, suggesting that categorization affects long-run outcomes in learning and choice.

What carries the argument

Mean-field dynamics derived from stochastic approximation of class-level Q-learning updates with multinomial logit choice probabilities over class valuations.

What would settle it

An experiment tracking whether agents update valuations from class-pooled feedback versus individual alternatives and checking if choice dynamics exhibit multiple equilibria, indifference, or limit cycles as payoff sensitivity increases.

Watch

Extended reading notes

Core claim

Coarse Q-learning pools feedback within exogenously given similarity classes to form class valuations, which guide multinomial logit choices and update according to Q-learning rules; the resulting mean-field dynamics have steady states that are smooth versions of Valuation Equilibria, and in the high payoff-sensitivity limit yield multiple stable strict equilibria, a unique globally stable mixed equilibrium featuring indifference across classes, or convergence to a stable limit cycle, all driven by the coarseness and absent from the alternative-level benchmark.

Load-bearing premise

The similarity classes are fixed and exogenous, with feedback always pooled within each class rather than tracked individually.

Editorial extensions

If this is right

  • Multiple stable strict equilibria can exist depending on the environment.
  • A unique globally stable mixed equilibrium with indifference across classes can occur.
  • Valuations and choice probabilities can converge to a stable limit cycle with no stable equilibrium.
  • These phenomena are specific to coarse aggregation and do not appear in standard alternative-level Q-learning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Real-world agents using coarse categories might experience persistent instability or cycles in preferences rather than settling to fixed choices.
  • The results imply that interventions changing how options are categorized could alter market stability or convergence.
  • The model could be extended to settings where class partitions evolve over time or based on experience.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 1 minor

Summary. The paper introduces Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, with feedback pooled within classes into class-level valuations. Choices follow multinomial logit over class valuations, and valuations update toward realized payoffs as in Q-learning. Using stochastic approximation, the paper derives the mean-field dynamics and characterizes the steady states as smooth analogues of Valuation Equilibria. In the high payoff-sensitivity limit, CQL exhibits novel long-run phenomena depending on the environment: multiple stable strict equilibria, a unique globally stable mixed equilibrium with indifference across classes, or no stable equilibrium with convergence to a stable limit cycle. These outcomes are driven by coarse aggregation and do not arise in the standard alternative-level benchmark.

Significance. If the derivations hold, the paper offers a significant contribution to modeling coarse thinking in reinforcement learning and its implications for long-run choice dynamics. The stochastic approximation approach to obtain mean-field dynamics is a methodological strength that enables the characterization of equilibria and the identification of phenomena (indifference, indeterminacy, instability) absent from the alternative-level benchmark. The contrast with standard Q-learning highlights the role of class-level aggregation.

minor comments (1)
  1. The abstract refers to 'smooth analogues of Valuation Equilibria' without defining the precise sense in which the steady states are smooth or how they relate to the original Valuation Equilibria concept; a brief clarification in the introduction would aid readability.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their careful reading and for the accurate summary of the paper's contributions. We are encouraged by the positive assessment of the methodological approach and the identification of novel phenomena driven by coarse aggregation. No specific major comments were provided in the report, so we have no points to address point-by-point at this stage. We remain available to clarify any aspects of the derivations or results if the editor or referee requests further details.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity identified

full rationale

The paper introduces Coarse Q-learning via an exogenous partition into similarity classes with pooled class-level valuations, multinomial logit choice, and standard Q-learning updates. It applies stochastic approximation to derive mean-field dynamics and characterizes steady states as smooth analogues of Valuation Equilibria. The long-run phenomena (multiple strict equilibria, globally stable mixed equilibrium with class indifference, or stable limit cycles) in the high-sensitivity limit are presented as direct consequences of the coarse aggregation structure, explicitly contrasted with the alternative-level benchmark that does not produce them. No step reduces by construction to a fitted parameter, self-citation chain, or definitional equivalence; the derivations remain self-contained against the model definition.

Assumptions & free parameters 2 free parameters · 2 assumptions · 0 invented entities

The model rests on an exogenous fixed partition into similarity classes and standard stochastic approximation results; no new entities are postulated.

free parameters (2)
  • exogenous similarity class partition
    The grouping of alternatives into classes is chosen outside the model and directly determines the pooling of feedback.
  • payoff-sensitivity parameter (beta)
    Analysis focuses on the limit beta to infinity; the value is not fitted but the limit behavior depends on taking this limit.
assumptions (2)
  • standard math Stochastic approximation theorems convert the discrete learning process into mean-field ODEs
    Invoked to derive the continuous-time dynamics whose steady states are then characterized.
  • domain assumption Choice probabilities follow multinomial logit over class valuations
    Modeling assumption stated in the abstract for how agents select among classes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Coarse Q-learning: Indifference, Indeterminacy, and Instability." pith.science (2026). https://pith.science/paper/2412.09321

@misc{pith2026241209321,
  author       = {Pith},
  title        = {Pith review of: Coarse Q-learning: Indifference, Indeterminacy, and Instability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2412.09321}},
  note         = {Machine review of arXiv:2412.09321}
}
read the original abstract

We introduce Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, and feedback from sampled alternatives is pooled within classes into class-level valuations. Choices follow multinomial logit over class valuations, and valuations update toward realized payoffs as in Q-learning. Using stochastic approximation, we derive the mean-field dynamics and characterize the steady states as smooth analogues of Valuation Equilibria. The model yields novel long-run phenomena in the high payoff-sensitivity limit: depending on the environment, CQL may exhibit multiple stable strict equilibria, a unique globally stable mixed equilibrium with indifference across classes, or no stable equilibrium at all, with valuations and choice probabilities converging instead to a stable limit cycle. These outcomes are driven by coarse aggregation and do not arise in the standard alternative-level benchmark.

Figures

Figures reproduced from arXiv: 2412.09321 by the authors.

Figure 1
Figure 1. Bob’s original decision tree r ′ ω1 (2) apples (3) citrus 2 3 ω2 (2) citrus 1 3 [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Bob’s simplified decision tree tions determine her strategy in each stage.10 Once an alternative is chosen, its corresponding payoff is observed, and the valuation of the similarity class containing the chosen alternative is updated based on the observed payoff. More precisely, Alice encounters a stage decision tree T ′ repeated infinitely. At each stage k ∈ N ∪ {0}, nature presents Alice with a choice problem ω ∈ Ω… view at source ↗
Figure 3
Figure 3. Example of a Decision Tree with Two Similarity Classes [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Stable Strict Pure SVE at (1.0, 0.0); β = 50 [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 6
Figure 6. Figure 6: Stable Unique Mixed SVE at (2.577, 2.577); β = 50 [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 8
Figure 8. Figure 8: Decision Tree T ′ 2 with Two Similarity Classes In this section, we restrict our attention to decision trees with generic payoffs where Alice has two similarity classes available to her.20 We imagine a general decision tree T ′ 2 with generic payoffs (depicted in [PIT…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

39 extracted references · 39 canonical work pages

  1. [1]

    Anderson, S., De Palma, A., and Thisse, J. (1992). Discrete Choice Theory of Product Differentiation . Mit Press. MIT Press

  2. [2]

    Bena \" m, M. (1999). Dynamics of stochastic approximation algorithms. S\'eminaire de probabilit\'es de Strasbourg , 33:1--68

  3. [3]

    Bena\" m, M., Hofbauer, J., and Sorin, S. (2005). Stochastic approximations and differential inclusions. SIAM Journal on Control and Optimization , 44(1):328--348

  4. [4]

    Bordalo, P., Gennaioli, N., and Shleifer, A. (2012). Salience Theory of Choice Under Risk . The Quarterly Journal of Economics , 127(3):1243--1285

  5. [5]

    Bordalo, P., Gennaioli, N., and Shleifer, A. (2013). Salience and consumer choice. Journal of Political Economy , 121(5):803--843

  6. [6]

    Brown, G. W. (1951). Iterative solution of games by fictitious play. In Koopmans, T. C., editor, Activity Analysis of Production and Allocation , pages 374--376. Wiley, New York

  7. [7]

    and Sarin, R

    Börgers, T. and Sarin, R. (1997). Learning through reinforcement and replicator dynamics. Journal of Economic Theory , 77(1):1--14

  8. [8]

    Cominetti, R., Melo, E., and Sorin, S. (2010). A payoff-based learning procedure and its application to traffic games. Games and Economic Behavior , 70(1):71--83

Show all 39 references
  1. [9]

    Conley, C. C. (1978). Isolated invariant sets and the Morse index / Charles Conley. Regional conference series in mathematics ; no. 38. Published for the Conference Board of the Mathematical Sciences by the American Mathematical Society, Providence

  2. [10]

    DellaVigna, S. (2009). Psychology and economics: Evidence from the field. Journal of Economic Literature , 47(2):315–72

  3. [11]

    and Roth, A

    Erev, I. and Roth, A. E. (1998). Predicting how people play games: Reinforcement learning in experimental games with unique, mixed strategy equilibria. The American Economic Review , 88(4):848--881

  4. [12]

    and Kreps, D

    Fudenberg, D. and Kreps, D. M. (1993). Learning mixed equilibria. Games and Economic Behavior , 5(3):320--367

  5. [13]

    and Levine, D

    Fudenberg, D. and Levine, D. K. (1998). The theory of learning in games , volume 2. MIT press

  6. [14]

    Gittins, J. C. (1979). Bandit processes and dynamic allocation indices. Journal of the Royal Statistical Society: Series B (Methodological) , 41(2):148--164

  7. [15]

    K., HOLT, C

    Goeree, J. K., HOLT, C. A., and PALFREY, T. R. (2016). Quantal Response Equilibrium: A Stochastic Theory of Games . Princeton University Press

  8. [16]

    and McFadden, D

    Hausman, J. and McFadden, D. (1984). Specification tests for the multinomial logit model. Econometrica , 52(5):1219--1240

  9. [17]

    W., Smith, H

    Hirsch, M. W., Smith, H. L., and Zhao, X.-Q. (2001). Chain transitivity, attractivity, and strong repellors for semidynamical systems. Journal of Dynamics and Differential Equations , 13:107--131

  10. [18]

    and Sandholm, W

    Hofbauer, J. and Sandholm, W. H. (2002). On the global convergence of stochastic fictitious play. Econometrica , 70(6):2265--2294

  11. [19]

    Iyengar, S. S. and Lepper, M. R. (2000). When choice is demotivating: Can one desire too much of a good thing? Journal of personality and social psychology , 79(6):995

  12. [20]

    Jehiel, P. (2005). Analogy-based expectation equilibrium. Journal of Economic theory , 123(2):81--104

  13. [21]

    and Samet, D

    Jehiel, P. and Samet, D. (2007). Valuation equilibrium. Theoretical Economics , 2(2):163--185

  14. [22]

    and Singh, J

    Jehiel, P. and Singh, J. (2021). Multi-state choices with aggregate feedback on unfamiliar alternatives. Games and Economic Behavior , 130:1--24

  15. [23]

    and Yin, G

    Kushner, H. and Yin, G. (2003). Stochastic Approximation and Recursive Algorithms and Applications . Stochastic Modelling and Applied Probability. Springer New York

  16. [24]

    McKelvey, R. D. and Palfrey, T. R. (1995). Quantal response equilibria for normal form games. Games and economic behavior , 10(1):6--38

  17. [25]

    A., Veness, J., Bellemare, M

    Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015). Human-level control through deep reinforcement learning. nature , 518(7540):529--533

  18. [26]

    and Shapley, L

    Monderer, D. and Shapley, L. S. (1996a). Fictitious play property for games with identical interests. Journal of Economic Theory , 68(1):258--265

  19. [27]

    and Shapley, L

    Monderer, D. and Shapley, L. S. (1996b). Potential games. Games and Economic Behavior , 14(1):124--143

  20. [28]

    Nachbar, J. H. (1990). Evolutionary selection dynamics in games: Convergence and limit properties. International Journal of Game Theory , 19(1):59--89

  21. [29]

    Pemantle, R. (1990). Nonconvergence to Unstable Points in Urn Models and Stochastic Approximations . The Annals of Probability , 18(2):698 -- 712

  22. [30]

    Robinson, J. (1951). An iterative method of solving a game. Annals of Mathematics , 54(2):296--301

  23. [31]

    and Lloyd, B

    Rosch, E. and Lloyd, B. B. (1978). Principles of categorization . MIT press

  24. [32]

    Roth, A. E. and Erev, I. (1995). Learning in extensive-form games: Experimental data and simple dynamic models in the intermediate term. Games and Economic Behavior , 8(1):164--212

  25. [33]

    Rustichini, A., Soukupova, M., and Palminteri, S. (2023). Adaptive coding is optimal in reinforcement learning. Available at SSRN 4320894

  26. [34]

    and Vahid, F

    Sarin, R. and Vahid, F. (1999). Payoff assessments without probabilities: A simple dynamic model of choice. Games and Economic Behavior , 28(2):294--309

  27. [35]

    Shapley, L. S. (1964). 1. Some Topics in Two-Person Games , pages 1--28. Princeton University Press, Princeton

  28. [36]

    Smith, H. L. (1995). Monotone dynamical systems: an introduction to the theory of competitive and cooperative systems . American Mathematical Soc

  29. [37]

    Sutton, R. S. and Barto, A. G. (2018). Reinforcement Learning: An Introduction . The MIT Press, second edition

  30. [38]

    Tversky, A. (1972). Elimination by aspects: a theory of choice. Psychological Review , 79:281--299

  31. [39]

    Watkins, C. J. C. H. and Dayan, P. (1992). Q-learning. Machine Learning , 8(3):279--292

Pith tools

Reviewed May 23, 2026 · model on record in the stance chip above.