REVIEW 1 minor 39 references
Coarse Q-learning: Indifference, Indeterminacy, and Instability
T0 review · 0 major / 1 minor · reviewed 2026-05-23 · grok-4.3
Pith's one-line read Coarse Q-learning can lead to multiple stable equilibria, class indifference, or limit cycles in the high payoff-sensitivity limit.
desk verdict CQL adds class-level pooling to Q-learning and derives new high-sensitivity outcomes like class indifference or limit cycles that the per-alternative benchmark lacks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Mean-field dynamics derived from stochastic approximation of class-level Q-learning updates with multinomial logit choice probabilities over class valuations.
What would settle it
An experiment tracking whether agents update valuations from class-pooled feedback versus individual alternatives and checking if choice dynamics exhibit multiple equilibria, indifference, or limit cycles as payoff sensitivity increases.
Extended reading notes
Core claim
Coarse Q-learning pools feedback within exogenously given similarity classes to form class valuations, which guide multinomial logit choices and update according to Q-learning rules; the resulting mean-field dynamics have steady states that are smooth versions of Valuation Equilibria, and in the high payoff-sensitivity limit yield multiple stable strict equilibria, a unique globally stable mixed equilibrium featuring indifference across classes, or convergence to a stable limit cycle, all driven by the coarseness and absent from the alternative-level benchmark.
Load-bearing premise
The similarity classes are fixed and exogenous, with feedback always pooled within each class rather than tracked individually.
Editorial extensions
If this is right
- Multiple stable strict equilibria can exist depending on the environment.
- A unique globally stable mixed equilibrium with indifference across classes can occur.
- Valuations and choice probabilities can converge to a stable limit cycle with no stable equilibrium.
- These phenomena are specific to coarse aggregation and do not appear in standard alternative-level Q-learning.
Reading between the lines
- Real-world agents using coarse categories might experience persistent instability or cycles in preferences rather than settling to fixed choices.
- The results imply that interventions changing how options are categorized could alter market stability or convergence.
- The model could be extended to settings where class partitions evolve over time or based on experience.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, with feedback pooled within classes into class-level valuations. Choices follow multinomial logit over class valuations, and valuations update toward realized payoffs as in Q-learning. Using stochastic approximation, the paper derives the mean-field dynamics and characterizes the steady states as smooth analogues of Valuation Equilibria. In the high payoff-sensitivity limit, CQL exhibits novel long-run phenomena depending on the environment: multiple stable strict equilibria, a unique globally stable mixed equilibrium with indifference across classes, or no stable equilibrium with convergence to a stable limit cycle. These outcomes are driven by coarse aggregation and do not arise in the standard alternative-level benchmark.
Significance. If the derivations hold, the paper offers a significant contribution to modeling coarse thinking in reinforcement learning and its implications for long-run choice dynamics. The stochastic approximation approach to obtain mean-field dynamics is a methodological strength that enables the characterization of equilibria and the identification of phenomena (indifference, indeterminacy, instability) absent from the alternative-level benchmark. The contrast with standard Q-learning highlights the role of class-level aggregation.
minor comments (1)
- The abstract refers to 'smooth analogues of Valuation Equilibria' without defining the precise sense in which the steady states are smooth or how they relate to the original Valuation Equilibria concept; a brief clarification in the introduction would aid readability.
Simulated Author's Rebuttal
We thank the referee for their careful reading and for the accurate summary of the paper's contributions. We are encouraged by the positive assessment of the methodological approach and the identification of novel phenomena driven by coarse aggregation. No specific major comments were provided in the report, so we have no points to address point-by-point at this stage. We remain available to clarify any aspects of the derivations or results if the editor or referee requests further details.
Circularity Check
No significant circularity identified
full rationale
The paper introduces Coarse Q-learning via an exogenous partition into similarity classes with pooled class-level valuations, multinomial logit choice, and standard Q-learning updates. It applies stochastic approximation to derive mean-field dynamics and characterizes steady states as smooth analogues of Valuation Equilibria. The long-run phenomena (multiple strict equilibria, globally stable mixed equilibrium with class indifference, or stable limit cycles) in the high-sensitivity limit are presented as direct consequences of the coarse aggregation structure, explicitly contrasted with the alternative-level benchmark that does not produce them. No step reduces by construction to a fitted parameter, self-citation chain, or definitional equivalence; the derivations remain self-contained against the model definition.
Assumptions & free parameters
free parameters (2)
- exogenous similarity class partition
- payoff-sensitivity parameter (beta)
assumptions (2)
- standard math Stochastic approximation theorems convert the discrete learning process into mean-field ODEs
- domain assumption Choice probabilities follow multinomial logit over class valuations
Cite this review
Pith. "Pith review of Coarse Q-learning: Indifference, Indeterminacy, and Instability." pith.science (2026). https://pith.science/paper/2412.09321
@misc{pith2026241209321,
author = {Pith},
title = {Pith review of: Coarse Q-learning: Indifference, Indeterminacy, and Instability},
year = {2026},
howpublished = {\url{https://pith.science/paper/2412.09321}},
note = {Machine review of arXiv:2412.09321}
}
read the original abstract
We introduce Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, and feedback from sampled alternatives is pooled within classes into class-level valuations. Choices follow multinomial logit over class valuations, and valuations update toward realized payoffs as in Q-learning. Using stochastic approximation, we derive the mean-field dynamics and characterize the steady states as smooth analogues of Valuation Equilibria. The model yields novel long-run phenomena in the high payoff-sensitivity limit: depending on the environment, CQL may exhibit multiple stable strict equilibria, a unique globally stable mixed equilibrium with indifference across classes, or no stable equilibrium at all, with valuations and choice probabilities converging instead to a stable limit cycle. These outcomes are driven by coarse aggregation and do not arise in the standard alternative-level benchmark.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Anderson, S., De Palma, A., and Thisse, J. (1992). Discrete Choice Theory of Product Differentiation . Mit Press. MIT Press
work page 1992
-
[2]
Bena \" m, M. (1999). Dynamics of stochastic approximation algorithms. S\'eminaire de probabilit\'es de Strasbourg , 33:1--68
work page 1999
-
[3]
Bena\" m, M., Hofbauer, J., and Sorin, S. (2005). Stochastic approximations and differential inclusions. SIAM Journal on Control and Optimization , 44(1):328--348
work page 2005
-
[4]
Bordalo, P., Gennaioli, N., and Shleifer, A. (2012). Salience Theory of Choice Under Risk . The Quarterly Journal of Economics , 127(3):1243--1285
work page 2012
-
[5]
Bordalo, P., Gennaioli, N., and Shleifer, A. (2013). Salience and consumer choice. Journal of Political Economy , 121(5):803--843
work page 2013
-
[6]
Brown, G. W. (1951). Iterative solution of games by fictitious play. In Koopmans, T. C., editor, Activity Analysis of Production and Allocation , pages 374--376. Wiley, New York
work page 1951
-
[7]
Börgers, T. and Sarin, R. (1997). Learning through reinforcement and replicator dynamics. Journal of Economic Theory , 77(1):1--14
work page 1997
-
[8]
Cominetti, R., Melo, E., and Sorin, S. (2010). A payoff-based learning procedure and its application to traffic games. Games and Economic Behavior , 70(1):71--83
work page 2010
Show all 39 references
-
[9]
Conley, C. C. (1978). Isolated invariant sets and the Morse index / Charles Conley. Regional conference series in mathematics ; no. 38. Published for the Conference Board of the Mathematical Sciences by the American Mathematical Society, Providence
1978
-
[10]
DellaVigna, S. (2009). Psychology and economics: Evidence from the field. Journal of Economic Literature , 47(2):315–72
2009
-
[11]
and Roth, A
Erev, I. and Roth, A. E. (1998). Predicting how people play games: Reinforcement learning in experimental games with unique, mixed strategy equilibria. The American Economic Review , 88(4):848--881
1998
-
[12]
and Kreps, D
Fudenberg, D. and Kreps, D. M. (1993). Learning mixed equilibria. Games and Economic Behavior , 5(3):320--367
1993
-
[13]
and Levine, D
Fudenberg, D. and Levine, D. K. (1998). The theory of learning in games , volume 2. MIT press
1998
-
[14]
Gittins, J. C. (1979). Bandit processes and dynamic allocation indices. Journal of the Royal Statistical Society: Series B (Methodological) , 41(2):148--164
1979
-
[15]
K., HOLT, C
Goeree, J. K., HOLT, C. A., and PALFREY, T. R. (2016). Quantal Response Equilibrium: A Stochastic Theory of Games . Princeton University Press
2016
-
[16]
and McFadden, D
Hausman, J. and McFadden, D. (1984). Specification tests for the multinomial logit model. Econometrica , 52(5):1219--1240
1984
-
[17]
W., Smith, H
Hirsch, M. W., Smith, H. L., and Zhao, X.-Q. (2001). Chain transitivity, attractivity, and strong repellors for semidynamical systems. Journal of Dynamics and Differential Equations , 13:107--131
2001
-
[18]
and Sandholm, W
Hofbauer, J. and Sandholm, W. H. (2002). On the global convergence of stochastic fictitious play. Econometrica , 70(6):2265--2294
2002
-
[19]
Iyengar, S. S. and Lepper, M. R. (2000). When choice is demotivating: Can one desire too much of a good thing? Journal of personality and social psychology , 79(6):995
2000
-
[20]
Jehiel, P. (2005). Analogy-based expectation equilibrium. Journal of Economic theory , 123(2):81--104
2005
-
[21]
and Samet, D
Jehiel, P. and Samet, D. (2007). Valuation equilibrium. Theoretical Economics , 2(2):163--185
2007
-
[22]
and Singh, J
Jehiel, P. and Singh, J. (2021). Multi-state choices with aggregate feedback on unfamiliar alternatives. Games and Economic Behavior , 130:1--24
2021
-
[23]
and Yin, G
Kushner, H. and Yin, G. (2003). Stochastic Approximation and Recursive Algorithms and Applications . Stochastic Modelling and Applied Probability. Springer New York
2003
-
[24]
McKelvey, R. D. and Palfrey, T. R. (1995). Quantal response equilibria for normal form games. Games and economic behavior , 10(1):6--38
1995
-
[25]
A., Veness, J., Bellemare, M
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015). Human-level control through deep reinforcement learning. nature , 518(7540):529--533
2015
-
[26]
and Shapley, L
Monderer, D. and Shapley, L. S. (1996a). Fictitious play property for games with identical interests. Journal of Economic Theory , 68(1):258--265
-
[27]
and Shapley, L
Monderer, D. and Shapley, L. S. (1996b). Potential games. Games and Economic Behavior , 14(1):124--143
-
[28]
Nachbar, J. H. (1990). Evolutionary selection dynamics in games: Convergence and limit properties. International Journal of Game Theory , 19(1):59--89
1990
-
[29]
Pemantle, R. (1990). Nonconvergence to Unstable Points in Urn Models and Stochastic Approximations . The Annals of Probability , 18(2):698 -- 712
1990
-
[30]
Robinson, J. (1951). An iterative method of solving a game. Annals of Mathematics , 54(2):296--301
1951
-
[31]
and Lloyd, B
Rosch, E. and Lloyd, B. B. (1978). Principles of categorization . MIT press
1978
-
[32]
Roth, A. E. and Erev, I. (1995). Learning in extensive-form games: Experimental data and simple dynamic models in the intermediate term. Games and Economic Behavior , 8(1):164--212
1995
-
[33]
Rustichini, A., Soukupova, M., and Palminteri, S. (2023). Adaptive coding is optimal in reinforcement learning. Available at SSRN 4320894
2023
-
[34]
and Vahid, F
Sarin, R. and Vahid, F. (1999). Payoff assessments without probabilities: A simple dynamic model of choice. Games and Economic Behavior , 28(2):294--309
1999
-
[35]
Shapley, L. S. (1964). 1. Some Topics in Two-Person Games , pages 1--28. Princeton University Press, Princeton
1964
-
[36]
Smith, H. L. (1995). Monotone dynamical systems: an introduction to the theory of competitive and cooperative systems . American Mathematical Soc
1995
-
[37]
Sutton, R. S. and Barto, A. G. (2018). Reinforcement Learning: An Introduction . The MIT Press, second edition
2018
-
[38]
Tversky, A. (1972). Elimination by aspects: a theory of choice. Psychological Review , 79:281--299
1972
-
[39]
Watkins, C. J. C. H. and Dayan, P. (1992). Q-learning. Machine Learning , 8(3):279--292
1992
Reviewed May 23, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.