{"id":"9caa2c31-7839-4a6c-8af2-86e1e5aabd19","arxiv_id":"2607.16538","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"History-dependent recursive preferences have a canonical minimal preference-augmented state and Bellman recursion when certainty-equivalent richness and separability axioms hold.","lead":"This paper develops a theory for compressing history-dependent preferences in Markov decision problems into a small 'preference state,' and proves that optimal policies can be computed on that reduced state. It matters because recursive preferences over habits, sentiment, or wealth usually suffer from state explosion, and this framework gives conditions under which tractable dynamic programming is justified.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 3.6(iii) is the load-bearing premise: certainty equivalents must be attainable as constant continuation vectors, which Axioms 1–6 do not imply; a two-state recursive-utility example violates it.","rationale":"The reader's weakest-assumption analysis identifies Assumption 3.6, and my read agrees; I sharpen it to the deterministic-solvability subclause. I traced the proof chain: Theorem 3.7 is a direct corollary of the abstract Theorem B.22; Proposition B.15 depends on deterministic solvability to define W0_t; Corollaries B.20 and B.21 supply global extensions but do not repair a failure of (iii). The concrete counterexample is a standard recursive expected-utility preference with state-dependent terminal scaling: it satisfies Axioms 1–6, has a continuous compatible utility system, yet Assumption 3.6(iii) fails because the certainty equivalent of an attainable vector lies outside the set of attainable constant vectors. This shows the assumption is a genuine extra structural condition, not a harmless regularity assumption, and it is not derived from the behavioral axioms. It also confirms the reader's statement that Assumption 3.6 already contains much of the content of the decomposition. The paper is internally consistent and honest about using Assumption 3.6, so the concern does not invalidate the conditional theorems; it only reinforces that the central claim is conditional on a nontrivial, non-behavioral hypothesis. No change to the CONDITIONAL verdict is needed.","tokens_in":62159,"tokens_out":13563,"duration_ms":156689,"concrete_test":"Analytically check the two-period counterexample: set T=2, S={0,1}, C=[0,1], V_3(h_2,c_2)=c_2 for s_2=1 and 0.5 c_2 for s_2=0, with U_2=V_3 and U_1(h_1,(c,f_+))=c+β(f_+(0)+f_+(1))/2. Verify Axioms 1–6 hold. Compute Λ^U_1(h_1,c)={(0.5 a, b): a,b∈[0,1]}. Choose f_+ with f_+(0)=1, f_+(1)=1, giving u=(0.5,1). If Assumption 3.6(iii) held, the certainty equivalent v=M0(u) would need to be an attainable constant vector, so v=(v,v)∈Λ and v∈[0,0.5]. But representing the expectation ranking forces v=E[u]=0.75, contradiction. This directly tests whether Assumption 3.6 can fail for a natural recursive preference satisfying all behavioral axioms.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest point is not the existence/continuity part of Assumption 3.6 but its deterministic-solvability clause (iii). Axioms 1–6 only ensure a continuous compatible utility system (Lemma B.9) and an induced risk ranking; they do not guarantee that for every f the certainty equivalent v=M0_t(h_t,u_{t+1}(h_t,f)) is itself an attainable constant vector in Λ^U_t(h_t,c). Proposition B.15 needs exactly this: it constructs W0_t by picking a plan realizing constant v and uses M0_t(u)=M0_t(v) to assert indifference. If (iii) fails, W0_t is undefined and Theorem B.22 — hence Theorem 3.7 and all downstream PA/SPA results — collapses. This is not an idle technicality. Take T=2, S={0,1}, C=[0,1], terminal value V_3(h_2,c_2)=c_2 when s_2=1 and V_3=0.5 c_2 when s_2=0, and let U_1(h_1,(c,f_+))=c+β E_s U_2(ι_1(h_1,c,s), f_+(s)). Then Λ^U_1(h_1,c)={(0.5 a, b): a,b∈[0,1]}. For u=(0.5,1), any normalized M0 representing the expectation risk ranking must satisfy M0(u)=E[u]=0.75, but the constant vector 0.75 is not in Λ^U_1(h_1,c). All Axioms 1–6 hold, yet no M0 can simultaneously represent the ranking and satisfy deterministic solvability. Since Assumption 3.6 is stated relative to a fixed U, the theorem's applicability depends on an arbitrary cardinal choice and on a richness property of the attainable-utility set that the behavioral axioms do not pin down.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies finite-horizon Markov decision processes with history-dependent preferences and asks when the full history can be replaced by a smaller, endogenously constructed preference-augmented (PA) state. Under weak order, continuity, dynamic consistency, terminal compatibility, compensated monotonicity, and weak separability (Axioms 1–6), together with Assumption 3.6 (certainty-equivalent richness for a fixed compatible utility system), Theorem 3.7 establishes a full-history recursive representation through time and risk aggregators. Section 4 defines a canonical PA-state quotient by identifying histories with the same physical state, the same utility under every common continuation plan, and forward stability under common extensions; Theorem 4.12 gives a recursive factorization on this quotient, and Proposition 4.15 asserts minimality among reachable recursive factorizations. Section 5 proves a Bellman recursion and a verification theorem on the PA state under Markov feasibility and DP regularity. Section 6 imposes rectangularity and aggregator exhaustiveness to separate the preference-memory state into belief and taste coordinates, with a separated recursive representation and Bellman recursion. A taxonomy of examples illustrates history-independent, pure-belief, pure-taste, decoupled, coupled, and entangled models.","tokens_in":62601,"tokens_out":8963,"duration_ms":99696,"significance":"If the results hold as stated, the paper makes a substantial contribution: it provides an endogenous, preference-based method for reducing history-dependent MDPs to a compact state, with a minimality result relative to recursive factorizations, and it unifies several strands of recursive preferences, risk-sensitive MDPs, and information-state theory. The proof strategy is careful and technically interesting: the effective-domain aggregators are built by backward induction, separated by behavioral axioms, and then extended by a fiberwise monotone extension theorem. The quotient topology arguments are rigorous, and the examples provide a useful taxonomy. The main theorems are conditional on explicit assumptions, and the paper does not claim empirical identification. However, the central recursive-representation theorem rests on Assumption 3.6, which is strong, is not derived from Axioms 1–6, and is not verified in the examples. The canonical/minimality claims are also relative to a fixed compatible utility system and to a specific class of factorizations. These caveats do not destroy the paper's value as a conditional theory, but they need to be addressed before the results can be ad","major_comments":[{"comment":"Deterministic solvability is load-bearing: Proposition B.15 constructs W^0_t by choosing a plan that realizes the constant certainty equivalent v, and then uses M^0_t(u)=M^0_t(v) to assert indifference. Without (iii), W^0_t is undefined and Theorem B.22 cannot be established. The assumption is not implied by Axioms 1–6. Counterexample: T=2, S={0,1}, C=[0,1], V_3(h_2,c_2)=c_2 if s_2=1 and 0.5 c_2 if s_2=0, U_1(h_1,(c,f_+))=c+β E_s V_3(...). Then Λ^U_1(h_1,c)={(0.5 a,b): a,b∈[0,1]}, whose only constant vector is 0. The induced risk ranking is expectation, so any normalized, monotone M^0_1 representing it assigns v>0 to u=(0.5,1), and no attainable constant vector equals v. All Axioms 1–6 hold. The paper should either derive (iii) from more basic axioms or explicitly frame the main theorem as conditional on this substantive CE-richness property and verify it in the examples.","section":"§3.3, Assumption 3.6(iii) and Proposition B.15"},{"comment":"Assumption 3.6 is stated for a fixed compatible utility system U, while Lemma B.9 only proves existence of some compatible system. The attainable sets Λ^U_t, the effective aggregators, and the canonical quotient in Definition 4.6 are all U-relative. The paper acknowledges this in words, but the abstract and introduction present the result as a property of preferences. If Assumption 3.6 holds for one U and fails for another ordinally equivalent representation, then the phrase 'behavioral axioms and a certainty-equivalent richness condition' overstates the contribution. The manuscript should either prove the invariance of Assumption 3.6 and the quotient under the non-unique choice of U, or provide a canonical cardinal normalization, or explicitly state the theorem as representation-relative at the level of the main claims.","section":"§3.1–3.3, Assumption 3.6 and Lemma B.9"},{"comment":"The belief/taste separation is not canonical in the preference sense: Definition 6.8 defines belief and taste equivalence relations using the selected global aggregators M^*_t and W^*_t from a chosen PA factorization, and the paper itself notes in Remark B.24 that these global extensions are not unique. Thus the spaces Y_t and Z_t, and therefore the SPA representation and the separated Bellman recursion of Corollary 6.19, depend on the choice of extension. Since the separation of beliefs and tastes is one of the paper's headline contributions, this dependence needs to be stated in the main text and either proven to be extension-invariant or illustrated with a concrete example where different extensions yield different separated coordinates.","section":"§6.3–6.4, Definition 6.8 and Theorem 6.14"}],"minor_comments":[{"comment":"The statement 'v = \\tilde u_{t+1}(h_t,(c,g_+))' conflates a scalar v with a vector in L. The intended meaning is that the constant vector v belongs to Λ^U_t(h_t,c). Please clarify the notation, since the equality as written is not well-typed.","section":"Assumption 3.6(iii)"},{"comment":"The minimality claim is weaker than the word 'canonical' may suggest: it is minimality among factorizations that preserve the fixed utility system U and satisfy exact transition consistency. Since ⊙_t is defined by indifference under all common continuation plans, any such factorization automatically refines ⊙_t by the proof's Step 2. This is a useful consistency result, but it should be described as a relative minimality theorem, not as a fully preference-based coarsest-state theorem independent of the selected U and the functional form of the aggregators.","section":"§4.4, Proposition 4.15"},{"comment":"None of the examples explicitly verifies Assumption 3.6(iii), the deterministic-solvability clause. Since the taxonomy is meant to illustrate the scope of the framework, at least one representative example from each group should state how the attainable constant vectors cover the certainty equivalents, or otherwise show that the assumption is satisfied. This would materially help readers assess how restrictive the assumption really is.","section":"§7, examples"},{"comment":"The dependence on the selected rectangularization (Assumption 6.2) should be emphasized in the main text. Definition 6.1 makes clear that the fiber identifications X_t(s) ≅ M_t are a choice, and the subsequent belief/taste separation is relative to that choice. This is stated in passing before Definition 6.8, but it deserves a more prominent caveat given the 'canonical' language.","section":"§6.3, Definition 6.10 and Theorem 6.14"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically serious and the conditional theorems appear to be proved carefully. My main concern is the gap between the axiomatic framing and Assumption 3.6, which is a strong, representation-relative richness condition that is not derived from the behavioral axioms and can fail in simple recursive-utility examples. The authors should either derive the condition from more primitive axioms, substantially soften the claimed behavioral foundation, or add explicit verification of the condition in the examples. If they do so, I would support acceptance; in its current form, the central representation theorem is a conditional existence result whose behavioral scope is substantially narrower than the abstract suggests."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Core: for a theory paper, this is a solid piece of work. It defines a canonical preference-augmented quotient, proves minimality among reachable recursive factorizations, and gives Bellman and verification theorems on the reduced state. The belief/taste separation section is a genuine organizational contribution, and the taxonomy of examples (habit, wealth, sentiment, ambiguity) shows the framework has real scope. The proofs are detailed and internally coherent, and the paper is honest about its dependence on global extensions in Remark B.24.\n\nThe main soft spot is Theorem 3.7. The theorem is only as good as Assumption 3.6, and clause (iii) — deterministic solvability — is doing most of the work. It requires that every certainty equivalent be attainable as a constant continuation vector. That is not a consequence of Axioms 1–6. A simple two-state example with terminal value c in state 1 and 0.5c in state 0 gives attainable continuation-utility set {(0.5a, b) : a,b in [0,1]}. For u = (0.5, 1), any normalized certainty equivalent representing the expectation ranking must assign 0.75, but (0.75, 0.75) is not in the attainable set. So Assumption 3.6(iii) fails; the recursive representation theorem, and everything downstream that uses it, loses its foundation in that case. This is not a minor technicality, and it is not flagged in the body of the paper. The abstract and Section 3 do say \"certainty-equivalent richness condition,\" but a reader could easily take that to be a mild regularity condition rather than a substantive richness assumption that is essentially the separation result in disguise.\n\nOther concerns are smaller. Theorem B.19 is imported from an external order-topology source; the proof is not machine-checked. The SPA belief/taste split depends on the selected global extension of the aggregators, so it is not fully preference-based — the paper admits this in Remark B.24, which is to its credit, but it does limit the economic interpretation.\n\nIf Assumption 3.6 is accepted as a primitive or supplied with sufficient conditions for natural preference classes, the rest of the architecture holds together. The paper is honest, well-structured, and the conditional theorems are new and meaningful. It deserves a serious referee; the main task for the referee is to determine how restrictive Assumption 3.6 really is, and whether it can be weakened or derived for standard recursive-utility families. I would send it out.","headline":"A coherent state-reduction theory for history-dependent recursive preferences, but the central representation theorem leans on a certainty-equivalent solvability assumption that standard preferences need not satisfy.","tokens_in":63051,"tokens_out":3023,"would_cite":true,"duration_ms":35524,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C40","91B16"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that, for finite-horizon Markov decision problems with history-dependent preferences, the whole history can be quotiented to a canonical preference-augmented state that is the coarsest reachable recursive factorization, an","keywords":["history-dependent preferences","recursive utility","Markov decision processes","state aggregation","preference-augmented state","time and risk aggregators","Bellman recursion","belief and taste separation"],"falsifier":"Construct a finite-horizon preference satisfying Axioms 1–6 whose attainable continuation-utility sets are disconnected or otherwise shaped so that no continuous, internal, normalized certainty-equivalent completion exists. If such a preference still admitted the paper's recursive representation, Assumption 3.6 would be shown unnecessary; conversely, exhibiting one such preference with no recursive representation would confirm the assumption is doing the work.","tokens_in":61974,"feed_emoji":"","tokens_out":4496,"duration_ms":49122,"temperature":0.7,"pith_summary":"The paper establishes a behavioral state-reduction theory for finite-horizon Markov decision problems in which preferences can depend on the entire realized history, not just the current physical state. Its central claim is that, under six behavioral axioms plus a certainty-equivalent richness condition, any compatible utility system for such preferences admits a recursive representation built from time aggregators and risk aggregators. It then constructs a canonical preference-augmented (PA) state by identifying histories that no continuation plan can distinguish and that remain indistinguishable after every common one-step extension, and proves this quotient is the coarsest reachable recursive factorization of the preferences. On this reduced state, a Bellman recursion holds and any Bellman selector induces a history-dependent optimal policy. Under extra rectangularity and exhaustiveness assumptions, the preference memory splits into separate belief and taste coordinates, yielding a separated representation and Bellman recursion. A sympathetic reader cares because this turns an apparently intractable full-history problem into ordinary dynamic programming on a state derived endogenously from behavior.","feed_headline":"History-dependent choices reduce to a minimal Bellman state","feed_subtitle":"Quotienting histories no continuation plan can tell apart gives the coarsest state that still supports optimal planning.","key_machinery":"The load-bearing object is the canonical preference-augmented (PA) state: a quotient of the history space by an equivalence relation that (i) preserves the current physical Markov state, (ii) identifies histories that are indifferent under every common continuation plan, and (iii) is forward-stable, meaning equivalent histories stay equivalent after any common one-step extension. This equivalence is built from a fixed compatible utility system and enforces that the quotient has continuous, well-defined transitions and compact metrizable state spaces. The recursive representation itself is carried by two named objects: time aggregators and risk aggregators, which the axioms separate out of th","core_discovery":"The paper's core discovery is that full-history recursive preferences are not inherently high-dimensional: the behavior itself determines a canonical quotient state. Define two histories equivalent when they share the current physical state, assign the same utility to every common continuation plan, and remain equivalent after every common one-step extension. The paper proves that this quotient is a compact metrizable state space with continuous transitions, that the fixed utility system factorizes through it, and that any other reachable recursive factorization refines it. It then shows that the optimal value functions satisfy a Bellman recursion on this PA state and that a Bellman selector","pith_inferences":["[Editorial inference] The canonical PA quotient is conceptually a behavioral version of an information state or bisimulation: histories are merged exactly when no continuation plan and no future extension can separate them, suggesting a direct bridge to state-minimization algorithms for controlled Markov processes with history-dependent rewards.","[Editorial inference] Because the PA state is constructed from a fixed compatible utility system and the global aggregator extensions are non-unique, the belief/taste split in Section 6 is relative to that choice; a natural extension would be to quantify how sensitive the SPA coordinates are to alternative compatible representations of the same preferences.","[Editorial inference] The theory suggests practical state-compression algorithms: given a discretized utility system, compute histories indistinguishable under a finite collection of continuation plans and one-step extensions; the resulting quotient should approximate the canonical PA state and can be used for memory design in reinforcement learning for history-dependent environments.","[Editorial inference] Certainty-equivalent richness is doing most of the work; if it fails in a given application, the recursive decomposition may still hold on the attainable domain but without continuous global aggregators, leaving open a purely effective-domain Bellman theory."],"forward_implications":["If the central claims are correct, history-dependent MDPs can be solved by backward induction on the canonical PA state, whose dimension is a behavioral property of preferences rather than an ad hoc modeling choice.","The minimality result implies that no reachable recursive factorization of the same utility system can use a strictly coarser state than the canonical quotient; any other memory augmentation contains at least as much state as the PA state.","Bellman selectors computed on the PA state induce policies that are optimal among all history-dependent policies, so the reduction is lossless for planning.","When rectangularity and exhaustiveness hold, the SPA separation into belief and taste coordinates means the same data can support identification of subjective beliefs or ambiguity on one side and time preference on the other, with a Bellman recursion in which the two aggregators do not cross-depend.","The framework encompasses a wide taxonomy of known models—discounted expected utility, Epstein–Zin, smooth ambiguity, habit and wealth effects, multiplier preferences—as special cases of PA or SPA constructions, unifying them under one reduction argument."],"fun_headline_variants":["Full-history preferences reduce to a minimal Bellman state","Quotient histories: a coarsest state for optimal planning","History-dependent choice collapses to a canonical reduced state","Minimal PA state supports Bellman recursion for full history"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is the certainty-equivalent richness assumption: for every period there exists a jointly continuous, monotone, normalized, internal certainty-equivalent functional on the attainable continuation utilities, with a deterministic solvability property; this is not derived from the more primitive behavioral axioms but assumed for the fixed utility system, and without it the recursive representation, the PA quotient, and all downstream Bellman results lose","fun_headline_variants_meta":{"raw":{"variants":["Full-history preferences reduce to a minimal Bellman state","Quotient histories: a coarsest state for optimal planning","History-dependent choice collapses to a canonical reduced state","Minimal PA state supports Bellman recursion for full history"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1018,"prompt_tokens":675,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":419,"completion_tokens_details":{"reasoning_tokens":277}},"tokens_in":419,"tokens_out":343,"duration_ms":4439,"temperature":1.0,"reasoning_tokens":277,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T20:40:53.848926+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a finite-horizon preference satisfying Axioms 1–6 whose attainable continuation-utility sets are disconnected or otherwise shaped so that no continuous, internal, normalized certainty-equivalent completion exists. If such a preference still admitted the paper's recursive representation, Assumption 3.6 would be shown unnecessary; conversely, exhibiting one such preference with no recursive representation would confirm the assumption is doing the work.","supporting_citations":[],"review_version":1}