Pith. sign in

REVIEW 3 major objections 4 minor 56 references

History-Dependent Recursive Preferences in Markov Decision Processes

T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This paper proves that, for finite-horizon Markov decision problems with history-dependent preferences, the whole history can be quotiented to a canonical preference-augmented state that is the coarsest reachable recursive factorization, an

desk verdict A coherent state-reduction theory for history-dependent recursive preferences, but the central representation theorem leans on a certainty-equivalent solvability assumption that standard preferences need not satisfy. read the letter →

arxiv 2607.16538 v1 pith:6UJL2HJK submitted 2026-07-17 math.OC

classification math.OC MSC 90C4091B16
keywords history-dependentpreferencesrecursiveutilityMarkovdecisionprocessesstateaggregationpreference-augmentedtimeandriskaggregatorsBellmanrecursionbelieftasteseparation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes a behavioral state-reduction theory for finite-horizon Markov decision problems in which preferences can depend on the entire realized history, not just the current physical state. Its central claim is that, under six behavioral axioms plus a certainty-equivalent richness condition, any compatible utility system for such preferences admits a recursive representation built from time aggregators and risk aggregators. It then constructs a canonical preference-augmented (PA) state by identifying histories that no continuation plan can distinguish and that remain indistinguishable after every common one-step extension, and proves this quotient is the coarsest reachable recursive factorization of the preferences. On this reduced state, a Bellman recursion holds and any Bellman selector induces a history-dependent optimal policy. Under extra rectangularity and exhaustiveness assumptions, the preference memory splits into separate belief and taste coordinates, yielding a separated representation and Bellman recursion. A sympathetic reader cares because this turns an apparently intractable full-history problem into ordinary dynamic programming on a state derived endogenously from behavior.

What carries the argument

The load-bearing object is the canonical preference-augmented (PA) state: a quotient of the history space by an equivalence relation that (i) preserves the current physical Markov state, (ii) identifies histories that are indifferent under every common continuation plan, and (iii) is forward-stable, meaning equivalent histories stay equivalent after any common one-step extension. This equivalence is built from a fixed compatible utility system and enforces that the quotient has continuous, well-defined transitions and compact metrizable state spaces. The recursive representation itself is carried by two named objects: time aggregators and risk aggregators, which the axioms separate out of th

What would settle it

Construct a finite-horizon preference satisfying Axioms 1–6 whose attainable continuation-utility sets are disconnected or otherwise shaped so that no continuous, internal, normalized certainty-equivalent completion exists. If such a preference still admitted the paper's recursive representation, Assumption 3.6 would be shown unnecessary; conversely, exhibiting one such preference with no recursive representation would confirm the assumption is doing the work.

Watch

Extended reading notes

Core claim

The paper's core discovery is that full-history recursive preferences are not inherently high-dimensional: the behavior itself determines a canonical quotient state. Define two histories equivalent when they share the current physical state, assign the same utility to every common continuation plan, and remain equivalent after every common one-step extension. The paper proves that this quotient is a compact metrizable state space with continuous transitions, that the fixed utility system factorizes through it, and that any other reachable recursive factorization refines it. It then shows that the optimal value functions satisfy a Bellman recursion on this PA state and that a Bellman selector

Load-bearing premise

The load-bearing premise is the certainty-equivalent richness assumption: for every period there exists a jointly continuous, monotone, normalized, internal certainty-equivalent functional on the attainable continuation utilities, with a deterministic solvability property; this is not derived from the more primitive behavioral axioms but assumed for the fixed utility system, and without it the recursive representation, the PA quotient, and all downstream Bellman results lose

Editorial extensions

If this is right

  • If the central claims are correct, history-dependent MDPs can be solved by backward induction on the canonical PA state, whose dimension is a behavioral property of preferences rather than an ad hoc modeling choice.
  • The minimality result implies that no reachable recursive factorization of the same utility system can use a strictly coarser state than the canonical quotient; any other memory augmentation contains at least as much state as the PA state.
  • Bellman selectors computed on the PA state induce policies that are optimal among all history-dependent policies, so the reduction is lossless for planning.
  • When rectangularity and exhaustiveness hold, the SPA separation into belief and taste coordinates means the same data can support identification of subjective beliefs or ambiguity on one side and time preference on the other, with a Bellman recursion in which the two aggregators do not cross-depend.
  • The framework encompasses a wide taxonomy of known models—discounted expected utility, Epstein–Zin, smooth ambiguity, habit and wealth effects, multiplier preferences—as special cases of PA or SPA constructions, unifying them under one reduction argument.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • [Editorial inference] The canonical PA quotient is conceptually a behavioral version of an information state or bisimulation: histories are merged exactly when no continuation plan and no future extension can separate them, suggesting a direct bridge to state-minimization algorithms for controlled Markov processes with history-dependent rewards.
  • [Editorial inference] Because the PA state is constructed from a fixed compatible utility system and the global aggregator extensions are non-unique, the belief/taste split in Section 6 is relative to that choice; a natural extension would be to quantify how sensitive the SPA coordinates are to alternative compatible representations of the same preferences.
  • [Editorial inference] The theory suggests practical state-compression algorithms: given a discretized utility system, compute histories indistinguishable under a finite collection of continuation plans and one-step extensions; the resulting quotient should approximate the canonical PA state and can be used for memory design in reinforcement learning for history-dependent environments.
  • [Editorial inference] Certainty-equivalent richness is doing most of the work; if it fails in a given application, the recursive decomposition may still hold on the attainable domain but without continuous global aggregators, leaving open a purely effective-domain Bellman theory.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies finite-horizon Markov decision processes with history-dependent preferences and asks when the full history can be replaced by a smaller, endogenously constructed preference-augmented (PA) state. Under weak order, continuity, dynamic consistency, terminal compatibility, compensated monotonicity, and weak separability (Axioms 1–6), together with Assumption 3.6 (certainty-equivalent richness for a fixed compatible utility system), Theorem 3.7 establishes a full-history recursive representation through time and risk aggregators. Section 4 defines a canonical PA-state quotient by identifying histories with the same physical state, the same utility under every common continuation plan, and forward stability under common extensions; Theorem 4.12 gives a recursive factorization on this quotient, and Proposition 4.15 asserts minimality among reachable recursive factorizations. Section 5 proves a Bellman recursion and a verification theorem on the PA state under Markov feasibility and DP regularity. Section 6 imposes rectangularity and aggregator exhaustiveness to separate the preference-memory state into belief and taste coordinates, with a separated recursive representation and Bellman recursion. A taxonomy of examples illustrates history-independent, pure-belief, pure-taste, decoupled, coupled, and entangled models.

Significance. If the results hold as stated, the paper makes a substantial contribution: it provides an endogenous, preference-based method for reducing history-dependent MDPs to a compact state, with a minimality result relative to recursive factorizations, and it unifies several strands of recursive preferences, risk-sensitive MDPs, and information-state theory. The proof strategy is careful and technically interesting: the effective-domain aggregators are built by backward induction, separated by behavioral axioms, and then extended by a fiberwise monotone extension theorem. The quotient topology arguments are rigorous, and the examples provide a useful taxonomy. The main theorems are conditional on explicit assumptions, and the paper does not claim empirical identification. However, the central recursive-representation theorem rests on Assumption 3.6, which is strong, is not derived from Axioms 1–6, and is not verified in the examples. The canonical/minimality claims are also relative to a fixed compatible utility system and to a specific class of factorizations. These caveats do not destroy the paper's value as a conditional theory, but they need to be addressed before the results can be ad

major comments (3)
  1. [§3.3, Assumption 3.6(iii) and Proposition B.15] Deterministic solvability is load-bearing: Proposition B.15 constructs W^0_t by choosing a plan that realizes the constant certainty equivalent v, and then uses M^0_t(u)=M^0_t(v) to assert indifference. Without (iii), W^0_t is undefined and Theorem B.22 cannot be established. The assumption is not implied by Axioms 1–6. Counterexample: T=2, S={0,1}, C=[0,1], V_3(h_2,c_2)=c_2 if s_2=1 and 0.5 c_2 if s_2=0, U_1(h_1,(c,f_+))=c+β E_s V_3(...). Then Λ^U_1(h_1,c)={(0.5 a,b): a,b∈[0,1]}, whose only constant vector is 0. The induced risk ranking is expectation, so any normalized, monotone M^0_1 representing it assigns v>0 to u=(0.5,1), and no attainable constant vector equals v. All Axioms 1–6 hold. The paper should either derive (iii) from more basic axioms or explicitly frame the main theorem as conditional on this substantive CE-richness property and verify it in the examples.
  2. [§3.1–3.3, Assumption 3.6 and Lemma B.9] Assumption 3.6 is stated for a fixed compatible utility system U, while Lemma B.9 only proves existence of some compatible system. The attainable sets Λ^U_t, the effective aggregators, and the canonical quotient in Definition 4.6 are all U-relative. The paper acknowledges this in words, but the abstract and introduction present the result as a property of preferences. If Assumption 3.6 holds for one U and fails for another ordinally equivalent representation, then the phrase 'behavioral axioms and a certainty-equivalent richness condition' overstates the contribution. The manuscript should either prove the invariance of Assumption 3.6 and the quotient under the non-unique choice of U, or provide a canonical cardinal normalization, or explicitly state the theorem as representation-relative at the level of the main claims.
  3. [§6.3–6.4, Definition 6.8 and Theorem 6.14] The belief/taste separation is not canonical in the preference sense: Definition 6.8 defines belief and taste equivalence relations using the selected global aggregators M^*_t and W^*_t from a chosen PA factorization, and the paper itself notes in Remark B.24 that these global extensions are not unique. Thus the spaces Y_t and Z_t, and therefore the SPA representation and the separated Bellman recursion of Corollary 6.19, depend on the choice of extension. Since the separation of beliefs and tastes is one of the paper's headline contributions, this dependence needs to be stated in the main text and either proven to be extension-invariant or illustrated with a concrete example where different extensions yield different separated coordinates.
minor comments (4)
  1. [Assumption 3.6(iii)] The statement 'v = \tilde u_{t+1}(h_t,(c,g_+))' conflates a scalar v with a vector in L. The intended meaning is that the constant vector v belongs to Λ^U_t(h_t,c). Please clarify the notation, since the equality as written is not well-typed.
  2. [§4.4, Proposition 4.15] The minimality claim is weaker than the word 'canonical' may suggest: it is minimality among factorizations that preserve the fixed utility system U and satisfy exact transition consistency. Since ⊙_t is defined by indifference under all common continuation plans, any such factorization automatically refines ⊙_t by the proof's Step 2. This is a useful consistency result, but it should be described as a relative minimality theorem, not as a fully preference-based coarsest-state theorem independent of the selected U and the functional form of the aggregators.
  3. [§7, examples] None of the examples explicitly verifies Assumption 3.6(iii), the deterministic-solvability clause. Since the taxonomy is meant to illustrate the scope of the framework, at least one representative example from each group should state how the attainable constant vectors cover the certainty equivalents, or otherwise show that the assumption is satisfied. This would materially help readers assess how restrictive the assumption really is.
  4. [§6.3, Definition 6.10 and Theorem 6.14] The dependence on the selected rectangularization (Assumption 6.2) should be emphasized in the main text. Definition 6.1 makes clear that the fiber identifications X_t(s) ≅ M_t are a choice, and the subsequent belief/taste separation is relative to that choice. This is stated in passing before Definition 6.8, but it deserves a more prominent caveat given the 'canonical' language.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the main theorems are conditional constructions; some minimality/separation steps are close to definitional but transparent, and no prediction reduces to a fit or self-citation.

full rationale

The paper's central recursive representation (Theorem 3.7) is not circular: Assumption 3.6 assumes existence of a certainty-equivalent functional M0 on the effective domain, and the proof constructs W0 from U and M0, extends both, and verifies by backward induction that the generated utility equals U. The assumption is powerful and not implied by Axioms 1-6 (the deterministic-solvability clause (iii) can fail, e.g., in a two-state example where the certainty equivalent is not an attainable constant vector), but that is an assumption-strength/robustness issue, not a reduction of the conclusion to its input. The PA minimality claim (Proposition 4.15) is largely a consequence of the definitions: Definition 4.6(ii) declares histories equivalent when they give the same evaluation under every common continuation plan, and Definition 4.5(i) says a factorization must preserve utilities, so any factorization must refine the canonical equivalence. This is definitional but explicitly constructed, not hidden. The SPA separation (Theorem 6.14) similarly forms belief/taste quotients from the kernels of the selected aggregators (Definition 6.8), and factorization follows from the universal property of quotients; the paper itself warns (Remark B.24) that this separation depends on global extension choices. No self-citations, fitted parameters, or imported uniqueness theorems are load-bearing. The fragile Assumption 3.6 and the extension-dependence of SPA states are genuine limitations and are flagged in the paper, but they do not amount to circularity; score 2 reflects only the mild definitional flavor of the quotient/minimality and SPA steps.

Assumptions & free parameters 0 free parameters · 15 assumptions · 2 invented entities

The central theorems rest on behavioral axioms, the strong certainty-equivalent richness assumption, a fixed cardinal utility scale, and several topological/order-theoretic background results. No free parameters are fitted to data; the parameters in the Section 7 examples are illustrative model choices and are not load-bearing for the main theorems.

assumptions (15)
  • domain assumption Axiom 1 (Weak order): ⪰(t) is complete and transitive on D_t.
    Standard behavioral assumption; used in Lemma B.9 to obtain a utility representation. Location: Section 3.1.
  • domain assumption Axiom 2 (Joint continuity on grand domain D_t).
    Assumes closed upper/lower contour sets on D_t, stronger than history-independent continuity. Location: Section 3.1.
  • domain assumption Axiom 3 (Dynamic consistency across one-step extensions).
    Requires the decision maker to agree with future selves after observing the next state; used to reduce plans to continuation-utility vectors. Location: Section 3.1.
  • domain assumption Axiom 4 (Terminal compatibility with V_{T+1}).
    Fixes terminal preferences to an exogenous value function. Location: Section 3.1.
  • domain assumption Axiom 5 (Compensated consumption monotonicity).
    More current consumption is strictly better when continuation utilities are equalized; used in Proposition B.15 to obtain strict monotonicity of the time aggregator. Location: Section 3.1.
  • domain assumption Axiom 6 (Weak separability of consumption and risk ranking).
    Ensures the induced risk comparison over continuation-utility vectors is independent of current consumption; used in Lemma B.13. Location: Section 3.1.
  • ad hoc to paper Assumption 3.6 (Certainty-equivalent richness).
    Postulates the existence of continuous, internal, monotone certainty-equivalent completions on every period's attainable continuation-utility domain. This is the main load-bearing premise of the recursive representation theorem. Location: Section 3.3.
  • domain assumption A fixed compatible utility system U is chosen as the cardinal scale.
    All derived objects (attainable utility sets, effective aggregators, quotient utilities) are relative to this fixed U; no behavioral invariance under increasing transformations is established. Location: Definition 2.4 and Remark B.24.
  • domain assumption Assumption 5.1 (Markov feasibility with respect to the PA state).
    Feasible consumption sets must depend only on the PA state; without it the Bellman state must be enlarged. Location: Section 5.
  • domain assumption Assumption 5.3 (DP regularity: continuous compact-valued feasible-consumption correspondence).
    Topological regularity needed for Berge's maximum theorem in the Bellman recursion. Location: Section 5.
  • ad hoc to paper Assumption 6.2 (Rectangularity of the canonical PA state).
    Requires the canonical quotient to split homeomorphically into physical state × memory space; needed before belief/taste coordinates can be defined. Location: Section 6.1.
  • ad hoc to paper Assumption 6.11 (Aggregator exhaustiveness).
    Requires that any two memory states differing in both belief and taste equivalence classes coincide; ensures the SPA triple is a lossless reparameterization. Location: Section 6.3.
  • standard math Debreu's continuous utility representation theorem.
    Used in Lemma B.9 to represent weak orders on second-countable grand domains. Location: Appendix A.1 and Lemma B.9.
  • standard math Closed-equivalence quotient compactness/metrizability theorem.
    Ensures the canonical quotient spaces are compact and metrizable. Location: Theorem A.3 and Lemma 4.9.
  • standard math Fiber-isotone continuous extension theorem (Minguzzi).
    Provides the nonstandard extension result used to turn effective time/risk aggregators into global continuous monotone aggregators. This is a critical external tool. Location: Theorem B.19.
invented entities (2)
  • Preference-memory state m_t (and canonical PA state x_t)
    purpose: Summarize the history dependence of preferences so dynamic programming can run on a smaller state.
    Derived as a quotient of histories; it has no measurement handle outside the model and exists only under the paper's axioms. It is a mathematical construct, not an independently observable quantity.
  • Separated belief and taste coordinates (y_t, z_t)
    purpose: Split preference memory into risk-evaluation memory and utility/time-preference memory.
    Constructed from a chosen rectangularization and selected global aggregator extensions; the paper itself notes the result is not invariant to those choices (Remark B.24), so it carries no external falsifiable handle.

how reviews work

0 comments
Cite this review

Pith. "Pith review of History-Dependent Recursive Preferences in Markov Decision Processes." pith.science (2026). https://pith.science/paper/6UJL2HJK

@misc{pith2026260716538,
  author       = {Pith},
  title        = {Pith review of: History-Dependent Recursive Preferences in Markov Decision Processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6UJL2HJK}},
  note         = {Machine review of arXiv:2607.16538}
}
read the original abstract

In finite horizon dynamic programming with history-dependent preferences, the relevant state may be the entire realized history, even when the physical state is Markov. This paper develops a behavioral state-reduction theory for such Markov decision processes. Under behavioral axioms and a certainty-equivalent richness condition, the full-history problem admits a recursive representation composed of time and risk aggregators. We then derive a canonical preference-augmented (PA) state by quotienting histories that have the same current physical Markov state, are indifferent under every common continuation plan, and remain equivalent after every common one-step extension. This canonical PA state is minimal among reachable recursive factorizations of the underlying preferences. Under Markov feasibility and standard dynamic-programming regularity, a PA Bellman selector induces an optimal full-history policy. With additional rectangularity and exhaustiveness conditions, we reparameterize the preference memory into distinct belief and taste coordinates, and obtain a separated representation and Bellman recursion. We give a taxonomy of examples to illustrate the scope of our framework.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 3 linked inside Pith

  1. [1]

    Aliprantis and Kim C

    Charalambos D. Aliprantis and Kim C. Border.Infinite Dimensional Analysis: A Hitchhiker’s Guide. Springer, Berlin, 3 edition, 2006

  2. [2]

    Optimal control of Markov processes with incomplete state information

    Karl Johan ˚Astr¨om. Optimal control of Markov processes with incomplete state information. Journal of Mathematical Analysis and Applications, 10(1):174–205, 1965. 28

  3. [3]

    Exotic preferences for macroe- conomists.NBER Macroeconomics Annual, 19:319–390, 2004

    David K Backus, Bryan R Routledge, and Stanley E Zin. Exotic preferences for macroe- conomists.NBER Macroeconomics Annual, 19:319–390, 2004

  4. [4]

    Markov decision processes with risk-sensitive criteria: an overview.Mathematical Methods of Operations Research, 99:141–178, 2024

    Nicole B ¨auerle and Anna Ja´skiewicz. Markov decision processes with risk-sensitive criteria: an overview.Mathematical Methods of Operations Research, 99:141–178, 2024

  5. [5]

    Markov decision processes with average-value-at-risk crite- ria.Mathematical Methods of Operations Research, 74(3):361–379, 2011

    Nicole B ¨auerle and Jonathan Ott. Markov decision processes with average-value-at-risk crite- ria.Mathematical Methods of Operations Research, 74(3):361–379, 2011

  6. [6]

    Oliver & Boyd, Edinburgh, 1963

    Claude Berge.Topological Spaces: Including a Treatment of Multi-Valued Functions, Vector Spaces and Convexity. Oliver & Boyd, Edinburgh, 1963. Translated by E. M. Patterson

  7. [7]

    Athena Scientific, 4th edition, 2012

    Dimitri P Bertsekas.Dynamic Programming and Optimal Control, volume 2. Athena Scientific, 4th edition, 2012

  8. [8]

    A survey of time consistency of dynamic risk measures and dynamic performance measures in discrete time: Lm-measure perspective

    Tomasz R Bielecki, Igor Cialenco, and Marcin Pitera. A survey of time consistency of dynamic risk measures and dynamic performance measures in discrete time: Lm-measure perspective. Probability, Uncertainty and Quantitative Risk, 2:1–52, 2017

Show all 56 references
  1. [9]

    On monotone recursive preferences

    Antoine Bommier, Asen Kochov, and Franc ¸ois Le Grand. On monotone recursive preferences. Econometrica, 85(5):1433–1466, 2017

  2. [10]

    By force of habit: A consumption-based explanation of aggregate stock market behavior.Journal of political Economy, 107(2):205–251, 1999

    John Y Campbell and John H Cochrane. By force of habit: A consumption-based explanation of aggregate stock market behavior.Journal of political Economy, 107(2):205–251, 1999

  3. [11]

    Ambiguity aversion and wealth effects.Journal of Economic Theory, 199:104898, 2022

    Simone Cerreia-Vioglio, Fabio Maccheroni, and Massimo Marinacci. Ambiguity aversion and wealth effects.Journal of Economic Theory, 199:104898, 2022

  4. [12]

    Recursive utility under uncertainty

    Soo H Chew and Larry G Epstein. Recursive utility under uncertainty . InEquilibrium theory in infinite dimensional spaces, pages 352–369. Springer, 1991

  5. [13]

    Risk-sensitive and robust decision-making: a cvar optimization approach.arXiv preprint arXiv:1506.02188, 2015

    Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone. Risk-sensitive and robust decision-making: a cvar optimization approach.arXiv preprint arXiv:1506.02188, 2015

  6. [14]

    Representation of a preference ordering by a numerical function

    Gerard Debreu. Representation of a preference ordering by a numerical function. In Robert M. Thrall, Clyde H. Coombs, and Robert L. Davis, editors,Decision Processes, pages 159–165. Wiley , New York, 1954

  7. [15]

    Springer Nature, 2024

    Darinka Dentcheva and Andrzej P Ruszczy ´nski.Risk-Averse Optimization and Control: Theory and Methods. Springer Nature, 2024

  8. [16]

    History-dependent risk attitude.Journal of Economic Theory, 157:445–477, 2015

    David Dillenberger and Kareen Rozen. History-dependent risk attitude.Journal of Economic Theory, 157:445–477, 2015

  9. [17]

    Helder- mann Verlag, Berlin, 1989

    Ryszard Engelking.General Topology, volume 6 ofSigma Series in Pure Mathematics. Helder- mann Verlag, Berlin, 1989

  10. [18]

    Recursive multiple-priors.Journal of Economic Theory, 113(1):1–31, 2003

    Larry G Epstein and Martin Schneider. Recursive multiple-priors.Journal of Economic Theory, 113(1):1–31, 2003

  11. [19]

    Substitution, risk aversion, and the temporal behav- ior of consumption and asset returns: A theoretical framework.Econometrica (1986-1998), 57(4):937, 1989

    Larry G Epstein and Stanley E Zin. Substitution, risk aversion, and the temporal behav- ior of consumption and asset returns: A theoretical framework.Econometrica (1986-1998), 57(4):937, 1989. 29

  12. [20]

    A recursive formulation for repeated agency with history dependence.Journal of Economic Theory, 91(2):223–247, 2000

    Ana Fernandes and Christopher Phelan. A recursive formulation for repeated agency with history dependence.Journal of Economic Theory, 91(2):223–247, 2000

  13. [21]

    Dynamic random utility .Econometrica, 87(6):1941–2002, 2019

    Mira Frick, Ryota Iijima, and Tomasz Strzalecki. Dynamic random utility .Econometrica, 87(6):1941–2002, 2019

  14. [22]

    Equivalence notions and model minimiza- tion in Markov decision processes.Artificial Intelligence, 147(1-2):163–223, 2003

    Robert Givan, Thomas Dean, and Matthew Greig. Equivalence notions and model minimiza- tion in Markov decision processes.Artificial Intelligence, 147(1-2):163–223, 2003

  15. [23]

    Robust control and model uncertainty .American Economic Review, 91(2):60–66, 2001

    Lars Peter Hansen and Thomas J Sargent. Robust control and model uncertainty .American Economic Review, 91(2):60–66, 2001

  16. [24]

    Intertemporal substitution, risk aversion and ambiguity aversion.Economic Theory, 25(4):933–956, 2005

    Takashi Hayashi. Intertemporal substitution, risk aversion and ambiguity aversion.Economic Theory, 25(4):933–956, 2005

  17. [25]

    Risk-sensitive Markov decision processes.Manage- ment Science, 18(7):356–369, 1972

    Ronald A Howard and James E Matheson. Risk-sensitive Markov decision processes.Manage- ment Science, 18(7):356–369, 1972

  18. [26]

    Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005

    Garud N Iyengar. Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005

  19. [27]

    Ambiguity , learning, and asset returns.Econometrica, 80(2):559–591, 2012

    Nengjiu Ju and Jianjun Miao. Ambiguity , learning, and asset returns.Econometrica, 80(2):559–591, 2012

  20. [28]

    On state dependent preferences and subjective probabilities.Econometrica, 51(4):1021–1031, 1983

    Edi Karni, David Schmeidler, and Karl Vind. On state dependent preferences and subjective probabilities.Econometrica, 51(4):1021–1031, 1983

  21. [29]

    Recursive smooth ambiguity prefer- ences.Journal of Economic Theory, 144(3):930–976, 2009

    Peter Klibanoff, Massimo Marinacci, and Sujoy Mukerji. Recursive smooth ambiguity prefer- ences.Journal of Economic Theory, 144(3):930–976, 2009

  22. [30]

    Stationary ordinal utility and impatience.Econometrica: Journal of the Econometric Society, pages 287–309, 1960

    Tjalling C Koopmans. Stationary ordinal utility and impatience.Econometrica: Journal of the Econometric Society, pages 287–309, 1960

  23. [31]

    Temporal resolution of uncertainty and dynamic choice theory .Econometrica: journal of the Econometric Society, pages 185–200, 1978

    David M Kreps and Evan L Porteus. Temporal resolution of uncertainty and dynamic choice theory .Econometrica: journal of the Econometric Society, pages 185–200, 1978

  24. [32]

    Quantile Markov decision process

    Xiaocheng Li, Huaiyang Zhong, and Margaret L Brandeau. Quantile Markov decision process. arXiv preprint arXiv:1711.05788, 2017

  25. [33]

    Dynamic variational preferences

    Fabio Maccheroni, Massimo Marinacci, and Aldo Rustichini. Dynamic variational preferences. Journal of Economic Theory, 128(1):4–44, 2006

  26. [34]

    Robust MDPs with k-rectangular uncertainty .Mathe- matics of Operations Research, 41(4):1484–1509, 2016

    Shie Mannor, Ofir Mebel, and Huan Xu. Robust MDPs with k-rectangular uncertainty .Mathe- matics of Operations Research, 41(4):1484–1509, 2016

  27. [35]

    Recursive contracts.Econometrica, 87(5):1589–1631, 2019

    Albert Marcet and Ramon Marimon. Recursive contracts.Econometrica, 87(5):1589–1631, 2019

  28. [36]

    Recursive preferences and ambiguity attitudes, 2026

    Massimo Marinacci, Giulio Principi, and Lorenzo Stanca. Recursive preferences and ambiguity attitudes, 2026

  29. [37]

    Normally preordered spaces and utilities.Order, 30(1):137–150, 2013

    Ettore Minguzzi. Normally preordered spaces and utilities.Order, 30(1):137–150, 2013

  30. [38]

    Munkres.Topology

    James R. Munkres.Topology. Prentice Hall, Upper Saddle River, NJ, 2 edition, 2000. 30

  31. [39]

    Robust control of Markov decision processes with uncer- tain transition matrices.Operations Research, 53(5):780–798, 2005

    Arnab Nilim and Laurent El Ghaoui. Robust control of Markov decision processes with uncer- tain transition matrices.Operations Research, 53(5):780–798, 2005

  32. [40]

    Iterated risk measures for risk-sensitive Markov decision processes with discounted cost

    Takayuki Osogami. Iterated risk measures for risk-sensitive Markov decision processes with discounted cost. InProceedings of the 27th Conference on Uncertainty in Artificial Intelligence (UAI 2011), pages 567–574, 2011

  33. [41]

    Time-consistent decisions and temporal decomposition of coherent risk functionals.Mathematics of Operations Research, 41(2):682–699, 2016

    Georg Ch Pflug and Alois Pichler. Time-consistent decisions and temporal decomposition of coherent risk functionals.Mathematics of Operations Research, 41(2):682–699, 2016

  34. [42]

    Dynamic programming with recursive preferences: Op- timality and applications.arXiv preprint arXiv:1812.05748, 2018

    Guanlong Ren and John Stachurski. Dynamic programming with recursive preferences: Op- timality and applications.arXiv preprint arXiv:1812.05748, 2018

  35. [43]

    Foundations of intrinsic habit formation.Econometrica, 78(4):1341–1373, 2010

    Kareen Rozen. Foundations of intrinsic habit formation.Econometrica, 78(4):1341–1373, 2010

  36. [44]

    Risk-averse dynamic programming for Markov decision processes

    Andrzej Ruszczy ´nski. Risk-averse dynamic programming for Markov decision processes. Mathematical programming, 125(2):235–261, 2010

  37. [45]

    Conditional risk mappings.Mathematics of operations research, 31(3):544–561, 2006

    Andrzej Ruszczy ´nski and Alexander Shapiro. Conditional risk mappings.Mathematics of operations research, 31(3):544–561, 2006

  38. [46]

    Completely abstract dynamic programming.arXiv preprint arXiv:2308.02148, 2023

    Thomas J Sargent and John Stachurski. Completely abstract dynamic programming.arXiv preprint arXiv:2308.02148, 2023

  39. [47]

    Dynamic mixture-averse preferences.Econometrica, 86(4):1347–1382, 2018

    Todd Sarver. Dynamic mixture-averse preferences.Econometrica, 86(4):1347–1382, 2018

  40. [48]

    An isomorphism between asset pricing models with and without linear habit formation.The Review of Financial Studies, 15(4):1189–1221, 2002

    Mark Schroder and Costis Skiadas. An isomorphism between asset pricing models with and without linear habit formation.The Review of Financial Studies, 15(4):1189–1221, 2002

  41. [49]

    SIAM, 2021

    Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczynski.Lectures on stochastic pro- gramming: modeling and theory. SIAM, 2021

  42. [50]

    Dynamic choice under ambiguity .Theoretical Economics, 6(3):379–421, 2011

    Marciano Siniscalchi. Dynamic choice under ambiguity .Theoretical Economics, 6(3):379–421, 2011

  43. [51]

    The optimal control of partially observable Markov processes over a finite horizon.Operations Research, 21(5):1071–1088, 1973

    Richard D Smallwood and Edward J Sondik. The optimal control of partially observable Markov processes over a finite horizon.Operations Research, 21(5):1071–1088, 1973

  44. [52]

    Dynamic programming with state-dependent discount- ing.Journal of Economic Theory, 192:105190, 2021

    John Stachurski and Junnan Zhang. Dynamic programming with state-dependent discount- ing.Journal of Economic Theory, 192:105190, 2021

  45. [53]

    Axiomatic foundations of multiplier preferences.Econometrica, 79(1):47– 73, 2011

    Tomasz Strzalecki. Axiomatic foundations of multiplier preferences.Econometrica, 79(1):47– 73, 2011

  46. [54]

    Temporal resolution of uncertainty and recursive models of ambiguity aversion.Econometrica, 81(3):1039–1074, 2013

    Tomasz Strzalecki. Temporal resolution of uncertainty and recursive models of ambiguity aversion.Econometrica, 81(3):1039–1074, 2013

  47. [55]

    Approximate in- formation state for approximate planning and reinforcement learning in partially observed systems.J

    Jayakumar Subramanian, Amit Sinha, Raihan Seraj, and Aditya Mahajan. Approximate in- formation state for approximate planning and reinforcement learning in partially observed systems.J. Mach. Learn. Res., 23:12–1, 2022

  48. [56]

    History-dependent risk aversion, the reinforcement effect, and dynamic monotonicity

    Gerelt Tserenjigmid. History-dependent risk aversion, the reinforcement effect, and dynamic monotonicity . Technical report, working paper, 2019. 31 A Mathematical Background This section collects supporting results for ease of reference. A.1 Continuous Utility Representation ...

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.