Pith. sign in

REVIEW 2 major objections 7 minor 23 references

A model of discrete choice based on reinforcement learning under short-term memory

T0 review · 2 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Finite-memory reinforcement learning produces choice probabilities that violate Luce's choice axiom, even when initial biases satisfy it.

desk verdict A genuinely new RL-based choice model that derives real anomalies from short-memory learning, but the advertised large-memory recovery of Luce's axiom is asserted rather than proved. read the letter →

arxiv 1908.06133 v1 pith:T6IBDPUK submitted 2019-08-16 econ.EM

classification econ.EM MSC 91B0691B1691B30
keywords discretechoicemodelsLuce'saxiomreinforcementlearningexpectedutilityintransitivityindependenceframingeffectinsurancedemand
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper constructs a family of discrete choice models in which a subject's response strengths are updated by reinforcement learning with a short, k-term memory span, and choice probabilities are read off from the stationary distribution of the resulting Markov chain. The paper's central claim is that these equilibrium probabilities deviate from Luce's choice axiom even when the initial learning bias satisfies it, and that the deviations disappear as the memory span becomes large. Using the shortest-memory case k=1 with a logit scaling function, the paper shows that the binary preferences derived from these probabilities are intransitive, violate the independence axiom in an Allais-style experiment, and exhibit framing: risk aversion for gains and risk seeking for losses. The paper argues these results matter because they reproduce canonical empirical anomalies of choice from a single behavioral mechanism, finite memory, without additional psychological assumptions.

What carries the argument

The central object is the RL(k) learning model: a Markov chain on the set of the last k selected alternatives, where the response strength to alternative i is \(U_i^n = U_0^i\) plus the average of the response values \(u\) of the reinforcements received for i over the last k periods in which i was chosen. Choice probabilities are \(\Phi(U_i)\) divided by the sum of \(\Phi\) over all alternatives. The equilibrium of this chain is the set of model choice probabilities. The load-bearing identity is formula (15), which gives the stationary probability of each alternative in the k=1 case; the two-alternative version (16) yields the binary preference rule (7) used in all applications.

What would settle it

Run the k=1 learning process with known reinforcement distributions and compare the observed stationary choice frequencies with the closed-form probabilities from formula (15); a systematic mismatch would refute the model. Or present subjects with the three two-outcome lotteries described in Lemma 3 (suitably instantiated) and check whether pairwise choices form the predicted cycle; transitive choices would falsify the claimed intransitivity.

Watch

Extended reading notes

Core claim

The discovery is that a subject who learns by reinforcement but remembers only the last k outcomes will, in the long run, choose according to probabilities that are not of the Luce form, even if the initial biases are. For k=1 with \(\Phi(u)=$e^{{u/\beta}}$\), response \(u(s)=s\), and priors \(U_0^i=E[R_i]\), the equilibrium choice probabilities satisfy an explicit formula, and the binary trace relation becomes the inequality in (7). The paper proves that this relation is generically intransitive and violates the independence axiom, and that its certainty equivalent lies below the expected payoff for gains and above it for losses; it also respects first-order stochastic dominance. Thus finite memory alone can produce the main empirical anomalies that motivated prospect theory and other departures from expected utility.

Load-bearing premise

The paper assumes that the initial learning bias of an alternative equals its expected utility and that this bias does not depend on the set of alternatives offered to the subject.

Editorial extensions

If this is right

  • A subject whose initial biases satisfy Luce's choice axiom will, after learning with finite memory, choose in a way that violates the independence of irrelevant alternatives.
  • The binary preferences from the k=1 model are generically intransitive and violate the independence axiom, providing a learning-based account of Allais-type paradoxes.
  • The model produces framing effects: risk aversion for gains and risk seeking for losses, with certainty equivalents below expected payoff for gains and above for losses.
  • The binary preferences respect first-order stochastic dominance, so persistently better lotteries are always preferred.
  • The parametric extension (8)-(10) applied to insurance demand predicts that optimal coverage can jump between no insurance, full insurance, and overinsurance as income and loss probability vary.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If finite memory is the source of choice anomalies, then experiments that shorten effective memory (e.g., cognitive load) should strengthen violations of Luce's axiom, while practice that lengthens memory should weaken them.
  • The long-memory recovery of Luce's axiom is argued informally in Section 2.3; a rigorous quantitative bound on how violations shrink as k grows would let experiments estimate a subject's effective memory span.
  • The binary preference rule (7) depends nonlinearly on the expectation \(E[X]\), which suggests a broader class of expectation-dependent preferences that could be tested against reference-dependent theories.
  • The insurance model's predicted phase transitions in coverage (jumps between a=-1, 1, and 2) are a sharp, testable signature distinguishing the learning model from standard expected-utility choice.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The paper introduces RL(k), a family of discrete choice models in which choice probabilities are the stationary probabilities of a Markov chain where response strengths are updated by averaging reinforcements over the last k trials. It derives exact formulas for k=1 (Eqs. (7), (15), (16)) and shows that, for exponential scale Φ(u)=e^{u/β} and initial bias U0_i=E[R_i], the induced binary preferences violate transitivity (Appendix 5.2) and independence (Section 3.2), display a gain/loss asymmetry with risk aversion for gains and risk seeking for losses (Lemma 1), and satisfy first-order stochastic dominance (Lemma 2). A parametric extension with response Ui = E[u(X_i)] + αu(X_i) (Section 4) is proposed as a model of deviations from expected utility and applied to insurance demand. The paper also claims that Luce's choice axiom is recovered as memory span k→∞ (Section 2.3).

Significance. If the results are correct, the paper offers a parsimonious mechanism: finite-memory reinforcement learning generates the standard anomalies of choice (intransitivity, IIA violation, framing effects) without adding psychological parameters beyond β, α, k, and the reference point. The derivations of the stationary probabilities for k=1 are algebraically sound, and Lemma 1 is rigorously proved. The model is tractable and the application to insurance demand produces a phase diagram that can be tested. The main weakness is that the large-k recovery of Luce's axiom is only sketched informally; because this claim appears in the abstract, the paper as written does not yet fully establish one of its advertised properties.

major comments (2)
  1. [Section 2.3] The paragraph 'We will proceed informally...' asserts that the averages in Eq. (4) converge by the law of large numbers to E[u(R_i)] and hence that choice probabilities reduce to LCA values as k→∞. This is not a proof: the Markov chain on k-tuples has a state-dependent stationary distribution, the counts N(n,i) in Eq. (4) are endogenous, and the exchange of the stationary limit with the limit k→∞ is not justified. One must show that every alternative's count diverges in the stationary chain and that the transition probabilities converge uniformly to the LCA vector; this is a nontrivial mean-field problem. The authors should either provide a rigorous convergence theorem, or explicitly label the recovery of Luce's axiom as a conjecture and soften the abstract accordingly.
  2. [Section 4, after Eq. (9)] The statement that in the limit α,β→∞ with α/β=1 the model 'is also described by another EU principle based on the minimization of E[1/(1+e^{u(R_i)})]' is exact only for q=2. Using Eq. (9), the limiting stationary probability is proportional to (E[1/(q-1+e^{u(R_i)})])^{-1}, because K0→q and e^{U0_i/β}→1. For q>2 the objective depends on the size of the choice set and is not an EU principle in the usual sense. Please correct this claim or state explicitly that the statement holds only for binary choice sets.
minor comments (7)
  1. [Abstract and Introduction] There are typos: 'we shown' should be 'we show', 'Kaheman and Trversky' should be 'Kahneman and Tversky', and 'by me ans' should be 'by means'.
  2. [Section 2.3] 'last k trails' should be 'last k trials'.
  3. [Appendix 5.2] The proof of Lemma 3 implicitly sets β=1; the authors should state that this is without loss of generality or present the β-dependent formulas.
  4. [Appendix 5.2] The sentence 'lotteries can be suitably perturbed to show that ≻ is not transitive as well' is asserted without justification; a brief perturbation argument would make the proof complete.
  5. [Section 3.2] The Allais-type example relies entirely on Figure 1; the numerical parameters used (β=0.1, U0_i=x) should be stated in the text so the reader can reproduce the figure without reading the caption.
  6. [Section 4.1] Figures 4 and 5 would benefit from explicit parameter values (e.g., the exact grid of a, y, p, q, and the values of α, β) so the phase diagrams are reproducible.
  7. [Global notation] Equation (9) uses q for the number of alternatives while Appendix 5.1 uses n; please align the notation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's claimed anomalies are mathematical consequences of the model equations, not fitted or self-referential.

full rationale

The paper derives choice probabilities from an explicit reinforcement-learning process (Eqs. 4–6) and then proves properties such as LCA violation, intransitivity (Lemma 3), independence violation (Sec. 3.2), and gain/loss asymmetry (Lemma 1) from those equations. No parameter is fitted to the target phenomena; the modeling choices U0_i = E[u(R_i)] and Phi(u)=e^{u/beta} are stated assumptions, not data-driven calibrations, so the subsequent deductions are not equivalent to their inputs. The only flagged concern is Sec. 2.3, where the claim that Luce's axiom is recovered as k grows is supported by an informal law-of-large-numbers argument ('We will proceed informally... converge by the law of large numbers to the mean') rather than a rigorous proof of the k-to-infinity limit for the stationary distribution. That is a correctness/rigor gap in an auxiliary claim, not circularity: the limit is not assumed as an input, and the finite-k deviations are independently proven. There are also no load-bearing self-citations; the paper's references are standard external sources (Luce, Fudenberg-Levine, Kahneman-Tversky) and do not supply the derivation. Hence no circular step can be exhibited under the stated standards.

Assumptions & free parameters 4 free parameters · 8 assumptions · 0 invented entities

All results are derived from the stated behavioral laws, not from data. The model's free parameters, beta, alpha, k, and the reference point u0, are fixed by hand for each example; none are estimated. The most consequential identifying assumptions are the initial bias U0_i=E[u(R_i)] and the max-probability selection rule in Section 4. No new physical or economic entities are postulated.

free parameters (4)
  • beta (response scale parameter) = 0.1 in Figures 1-2, 0.2 in Figure 3, 1 in Figure 5
    Scale parameter in the logit map Phi(u)=e^{u/beta}; set by hand per example, not estimated.
  • alpha (deviation from EU parameter) = 0.4 in Figures 4-5; also alpha/beta=0.4 with alpha,beta -> infinity
    Parameter in equation (8) that controls the weight of one-period experience; chosen to illustrate the model, not fitted.
  • memory span k = 1
    The paper sets k=1 for all applications and shows the large-k limit only informally.
  • reference point u0, or s0 = s0=2 in Figure 3; reference income 4 in the insurance example
    Frames gains and losses in equation (12); the value is chosen by hand to illustrate the framing effect.
assumptions (8)
  • domain assumption Reinforcement-to-response law is additive over the last k reinforcements, equation (4).
    Defines the model; without this law the stationary probabilities and all subsequent results do not follow.
  • domain assumption Choice probability is proportional to a positive, non-decreasing scale Phi of response strengths, equation (5).
    The logit form is a special case; all applications use Phi(u)=e^{u/beta}.
  • domain assumption Initial priors U0_i are independent of the offered subset S, stated in Section 2.1.
    This gives the base probabilities Luce's axiom before learning; if false, the statement 'even if initial bias adheres to LCA' is vacuous.
  • ad hoc to paper In applications, initial bias U0_i equals expected utility E[u(R_i)], stated in Section 3.
    This identifying assumption is not derived from the learning process; it makes formula (7) and the subsequent behavioral results tractable.
  • domain assumption Reinforcements are independent across periods and alternatives, stated in Section 2.1.
    Needed for the Markov property and for the law of large numbers used in the long-memory limit.
  • standard math The learning Markov chain is irreducible and ergodic, so a unique stationary distribution exists, used in Appendix 5.1.
    Relies on positivity of transition probabilities (13); citation to Feller is given.
  • domain assumption Binary preferences are derived by the trace relation: x is preferred to y iff the stationary probability of x exceeds 1/2, stated in Section 3.
    This link from probabilistic choice to deterministic preference is standard but not forced.
  • ad hoc to paper In the parametric model, the subject chooses the alternative with maximum stationary probability, equation (10).
    A new selection principle introduced in Section 4; the insurance demand results depend on it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A model of discrete choice based on reinforcement learning under short-term memory." pith.science (2026). https://pith.science/paper/T6IBDPUK

@misc{pith2026190806133,
  author       = {Pith},
  title        = {Pith review of: A model of discrete choice based on reinforcement learning under short-term memory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T6IBDPUK}},
  note         = {Machine review of arXiv:1908.06133}
}
read the original abstract

A family of models of individual discrete choice are constructed by means of statistical averaging of choices made by a subject in a reinforcement learning process, where the subject has short, k-term memory span. The choice probabilities in these models combine in a non-trivial, non-linear way the initial learning bias and the experience gained through learning. The properties of such models are discussed and, in particular, it is shown that probabilities deviate from Luce's Choice Axiom, even if the initial bias adheres to it. Moreover, we shown that the latter property is recovered as the memory span becomes large. Two applications in utility theory are considered. In the first, we use the discrete choice model to generate binary preference relation on simple lotteries. We show that the preferences violate transitivity and independence axioms of expected utility theory. Furthermore, we establish the dependence of the preferences on frames, with risk aversion for gains, and risk seeking for losses. Based on these findings we propose next a parametric model of choice based on the probability maximization principle, as a model for deviations from expected utility principle. To illustrate the approach we apply it to the classical problem of demand for insurance.

Figures

Figures reproduced from arXiv: 1908.06133 by the authors.

Figure 1
Figure 1. figure 1. The independence axiom requires the same preference be [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 1
Figure 1. Allais Paradox. Red line is the certainty equivalent for lotter [PITH_FULL_IMAGE:figures/full_fig_p022_1.png] view at source ↗
Figure 2
Figure 2. Certainty equivalent curves for simple lotteries for losses [PITH_FULL_IMAGE:figures/full_fig_p023_2.png] view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: Certainty equivalent curves for simple lotteries for losses [PITH_FULL_IMAGE:figures/full_fig_p024_3.png]
Figure 4
Figure 4. Figure 4: Demand for insurance I. The figure shows fraction [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]
Figure 5
Figure 5. Figure 5: Demand for insurance II. The figure shows fraction [PITH_FULL_IMAGE:figures/full_fig_p026_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 23 canonical work pages

  1. [1]

    (1953) Les comportement de l’homme rationnal devant le risque: critique des postulates and axioms de l’ecole americaine

    Allais, M. (1953) Les comportement de l’homme rationnal devant le risque: critique des postulates and axioms de l’ecole americaine. Econometrica, 21

  2. [2]

    Arrow, K. (1951). Social Choice and Individual Values. Wiley & Son s, New York, NY

  3. [3]

    Bush, R.R., and Mosteller, F. (1951). A mathematical model for s im- ple learning. Psychological Review 58, 313–323. 19

  4. [4]

    and Mosteller, F

    Bush, R.R. and Mosteller, F. (1955). Stochastic models for learn ing. Wiley & Sons, New York, NY

  5. [5]

    and Roth, A

    Erev, I. and Roth, A. E. (1996). On the need of low rationality co gni- tive game theory: reinforcement learning in experimental games wit h unique mixed equilibria. Mimeo. University of Pittsburgh

  6. [6]

    and Roth, A

    Erev, I. and Roth, A. E. (1998). Predicting how people play game s: reinforcement learning in experimental games with unique, mixed strategy equilibrium. American Econ. Review 88, 848–881

  7. [7]

    Feller, W. (1957). An Introduction to Probability Theory and Its Applications. Vol II. John Wiley & Sons, New York, NY

  8. [8]

    Fudenberg, D., and Levine, D. (1998). The theory of learning in games. MIT Press, Cambridge, MA., London, England

Show all 23 references
  1. [9]

    Harley, C.B. (1981). Learning the Evolutionary Stable Strategy . J. theor. Biol. 89, 611–633

  2. [10]

    Kahneman, D., and Tversky, A. (1979). Prospect Theory: An Anal- ysis of Decision under Risk, Econometrica, XVLII 263–291

  3. [11]

    Kahneman, D., and Tversky, A. (1984). Choices, Values and Fr ames. American Psychologist 39, 341–350

  4. [12]

    Marschak, J. (1960). Binary choice constraints on random ut ility in- dicators. in K. Arrow (ed.), Stanford Symposium on Mathematical Models in the Social Sciences, Stanford University Press, Stanfor d, CA

  5. [13]

    Luce, R. D. (1959). Individual Choice Behavior. Wiley &Sons, Ne w York, NY

  6. [14]

    Luce, R. D. (1977). The choice axiom after twenty years. Jou rnal of Math. Psych. 15, 215–233

  7. [15]

    Machina, M.J. (1982). Expected utility analysis without the indep en- dence axiom”. Econometrica. 50 (2), 277–323

  8. [16]

    Quiggin, J. (1982). A theory of anticipated utility. Journal of E co- nomic Behavior and Organization 3(4), 323–343. 20

  9. [17]

    Quiggin, J. (1993). Generalized Expected Utility Theory. The Ra nk- Dependent Model. Kluwer Academic Publishers, Boston, MA

  10. [18]

    Pleskac, T. (2015). Decision and Choice: Luce’s Choice Axiom. in P. Bona, (ed.) International Encyclopedia of the Social & Behavior al Sciences

  11. [19]

    Roth, A.E., and Erev, I. (1995). Learning in extensive-form ga mes: experimental data and simple dynamics models in the intermediate term. Games and Economic Behavior, 8 164–212

  12. [20]

    and Barto, A.G

    Sutton, R.S. and Barto, A.G. (1998). Reinforcement learning: an introduction. MIT Press, Cambridge MA

  13. [21]

    Tversky, A., Kahneman, D. (1992). Advances in prospect the ory: Cumulative representation of uncertainty. Journal of Risk and Un - certainty, 5:297–323

  14. [22]

    and Morgenstern, O

    Von Neumann, J. and Morgenstern, O. (1947). Theory of Gam es and Economic Behavior. Princeton, NJ: Princeton University Press

  15. [23]

    Yaari, M. (1987). The dual theory of choice under risk. Econo metrica, 55, 95–115. 21 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 certainty equivalent, c 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 lottery weight, x Figure 1: Allais Paradox. Red line is the certainty equivalent for lo...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.