{"id":"eb00e7be-b203-4832-a16e-d6ee7af093e3","arxiv_id":"2608.06152","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A continuous-time stochastic choice model with distribution-dependent preferences is developed, behaviorally characterized, shown to be identified only under strong conditions, and separated from dynamic random utility.","lead":"This paper builds a model where people's preferences evolve based on the conditional distribution of latent preferences inferred from observed choices, and shows such feedback behavior can be told apart from standard models where preferences evolve exogenously. It also proves the underlying stochastic system has a unique solution under stated conditions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strict behavioral separation (Theorem 19) is proven only for counterfactual arrays over arbitrary latent distributions and measure flows; it does not follow for realized stochastic-choice data, so the central empirical claim is under-supported.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the behavioral separation theorem is formulated over a counterfactual observation domain, not over realized stochastic-choice data. I checked the internal logic of Theorem 19 and its proof in Supplementary Appendix SB.3. Given the full array O0∪O1, the proof is coherent: if a DDU representation were observationally equivalent to a DRU representation on this full domain, the DRU representation's distributional invariance would transfer to the DDU representation, contradicting behavioral distributional feedback. The issue is therefore not an internal inconsistency but an overstatement of what 'observable' means. In standard stochastic-choice theory, the analyst sees choices from menus conditional on histories, not conditional choice probabilities under arbitrary latent-state distributions ν, distributional states μ, initial laws λ, or counterfactual measure flows m. Definitions 6–9 and the observation map in Section 3.3 require exactly those counterfactual variations. Consequently, the impossibility theorem establishes a separation between two model classes at the level of their full counterfactual kernels, but it does not establish that a DDU array with behavioral distributional feedback is distinguishable from every DRU array on the realized histories alone. The paper's probabilistic contributions—existence and weak uniqueness for the conditional McKean-Vlasov system—are plausible and largely independent of this concern, and the behavioral characterization is internally consistent once the counterfactual domain is granted. But the headline claim that endogenous preference evolution is behaviorally distinct from exogenous dynamics needs the realized-data version to be proven separately. Since the reader's conditional verdict already reflects this gap, my stress-test does not change the verdict.","tokens_in":54432,"tokens_out":8056,"duration_ms":96718,"concrete_test":"Replace O0∪O1 in Section 3.3 by the realized-data domain O* consisting only of (s,A,z,μ_s^θ,μ_s^θ) for histories generated by θ and of continuation probabilities under the equilibrium flow m=μ^θ; re-run the proof of Theorem 19 on this restricted observation map. Settle the question with a two-state binary-menu construction: take u(s,a,x,μ)=x, u(s,b,x,μ)=φ(μ) with φ nonconstant, X_t a one-dimensional OU process, Y_t the choice/signal process, and ask whether any DDU specification with φ nonconstant on the reachable μ-set has a DRU representation matching all O* probabilities. If such a pair exists, the strict enlargement fails under realized data; if no such pair exists for this class, the counterfactual domain is precisely the boundary of the theorem.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central impossibility theorem is valid as a statement about the model's counterfactual array, but not as stated about 'stochastic-choice arrays' in the usual observational sense. Definitions 6–9 define behavioral distributional feedback through kernels C_s(z;A|ν,μ) evaluated at arbitrary ν while holding μ, and through D_{s,r}(z;A|λ,m) at arbitrary externally specified measure flows m,m'. Section 3.3 then declares the 'observable stochastic-choice array' Q_θ to be the full family indexed by O0∪O1, which contains arbitrary latent-state distributions ν, distributional states μ, initial laws λ, and measure flows m. Theorem 19's proof transfers DRU's distributional invariance to the DDU representation only because T(θ)=T(θ~) is assumed to hold on all of O0∪O1. If the analyst records only realized histories and the equilibrium diagonal C_s(z;A|μ_s,μ_s), the off-diagonal comparisons in Definitions 6 and 8 are not observed and cannot be used to reject a DRU alternative. Thus R_DDU strictly enlarges R_DRU only relative to a counterfactual observation protocol; the claim that endogenous preference evolution is behaviorally distinct from exogenous dynamics at the level of realized stochastic choice is not established. The paper's frequent qualifier 'on the reachable domain' does not solve this: reachability is defined through the model, whereas the analyst's data do not include the counterfactual ν and m variations used in the feedback definition.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a continuous-time stochastic choice model, DDU, in which the analyst's conditional distribution of latent preferences (the CPD) enters both the current felicity index and the law of motion of latent preferences. It proves existence and weak uniqueness for the resulting conditional McKean-Vlasov system (Theorem 3), defines a behavioral representation Φ=(Φ_u,Φ_P), proves identification of Φ from the observable array (Theorem 22), characterizes when DDU reduces to DRU (Theorem 18), and claims a behavioral impossibility theorem: any array exhibiting behavioral distributional feedback lies outside the DRU class (Theorem 19). The main text contains detailed proof sketches, with the probabilistic proofs in a Supplementary Appendix.","tokens_in":54812,"tokens_out":5112,"duration_ms":49947,"significance":"The probabilistic well-posedness of the conditional McKean-Vlasov system is a valuable technical contribution, and the conceptual decomposition of feedback into felicity and preference channels is useful. The paper is ambitious in linking stochastic choice, filtering, and mean-field dynamics. However, the claimed behavioral separation and identification results are established for an observation domain that includes counterfactual latent-state distributions and measure flows; for realized stochastic-choice data the central separation claim is not demonstrated. With appropriate qualifications the structural results may be publishable, but as stated the behavioral theorems overreach.","major_comments":[{"comment":"The impossibility theorem is proven for the counterfactual array Qθ defined on O0∪O1, which includes arbitrary ν, μ, λ, and m. In actual stochastic-choice data the analyst observes only the equilibrium diagonal C_s(z;A|μ_s,μ_s) and realized histories. The proof of Theorem 19 transfers DRU's distributional invariance to DDU only because T(θ)=T(θ~) is assumed on all of O0∪O1; without observing the off-diagonal comparisons in Definitions 6 and 8, a DRU alternative cannot be rejected from realized data. Therefore the conclusion that R_DDU strictly enlarges R_DRU holds only relative to a counterfactual observation protocol, not for stochastic-choice arrays in the usual observational sense.","section":"Section 3.3 (observable domain) and Theorem 19"},{"comment":"Theorem 22 (Behavioral identification) relies on Assumption 16 (injectivity of μ↦C_s and ν↦C_r) and Assumption 21 (the choice-test class H_r separates and is measure-determining on reachable laws P_r). These assumptions are imposed on the full counterfactual domain, and the proof uses equality of integrals against all ν∈N_s(μ) and all reachable ρ. Since the observation domain already contains all such ν and ρ, the identification is essentially an injectivity assumption on the observation map. The abstract's claim that the representation 'is identified from stochastic choice' would require identification from the diagonal C_s(·;·|μ_s,μ_s) alone; that claim is not established.","section":"Section 3.3, Theorem 22"},{"comment":"Behavioral distributional feedback is defined through C_s(z;A|ν,μ) at arbitrary ν holding μ fixed, and through D_{s,r}(z;A|λ,m) at externally specified measure flows m,m'. Reachability (Assumptions 8, 13, 16) is defined through the model, but the analyst's data do not include these counterfactual variations. Consequently, the 'observable restrictions' on stochastic choice highlighted in the introduction and abstract are not restrictions on realized choice frequencies; they are restrictions on a hypothetical array that varies latent distributions and measure flows. This is the load-bearing issue for the paper's central empirical claims.","section":"Section 3.1, Definitions 6–9"}],"minor_comments":[{"comment":"The labeling of proofs is inconsistent with the main text: SA.3.1 proves 'Proposition 26' and SA.3.2 proves 'Theorem 27', while the main text numbers these as Lemma 26 and Theorem 28; the numbering should be harmonized throughout.","section":"Supplementary Appendix SA.3.1–SA.3.3"},{"comment":"There are several typos: 'therefor,e' in Supplement SA.3.3; a repeated 'decision' in the Introduction; '∂ εΓε̸=0' with missing spaces in Section 4.4; and the tuple in Remark 8 lists u twice.","section":"Throughout"},{"comment":"The paper claims to be the 'first' continuous-time theory of endogenous preference evolution. Given the proximity to conditional McKean-Vlasov models (Carmona et al.; Buckdahn et al.), the novelty claim should be stated more cautiously, e.g., 'first to our knowledge' or with a more precise delineation from existing conditional-law frameworks.","section":"Section 1.1 and 5.1"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claim—that endogenous preference evolution is behaviorally distinct from DRU—depends on an unusual definition of 'observable' that includes counterfactual variations of latent distributions and measure flows. If the authors are willing to recast the results as statements about a hypothetical experiment or about structural equivalence classes, the paper could be acceptable. I would not recommend rejection because the probabilistic core appears sound and the counterfactual interpretation can be stated honestly, but as written the separation and identification theorems overstate their empirical reach."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a serious theoretical paper, but its headline separation claim—endogenous preference evolution is behaviorally distinct from exogenous DRU—is proven over a counterfactual observation domain, not over realized stochastic-choice data. That gap should not get it desk-rejected; it should get it a referee who will push on what 'observable' means.\n\nWhat is genuinely new: the conditional-McKean-Vlasov feedback loop, where the analyst's posterior over the latent preference state (μ_s) enters both the felicity index and the drift/diffusion of the latent process. That is a real extension of Frick-Iijima-Strzalecki, and the behavioral representation Φ = (Φ_u, Φ_P) splitting contemporaneous choice from continuation behavior is clean. Theorem 19 is logically correct given the definitions: if the full counterfactual array T(θ) shows behavioral distributional feedback, no DRU representation can match it. The probabilistic side—existence and weak uniqueness for the conditional MVSDE—uses the standard frozen-flow fixed-point machinery from Buckdahn-Li-Ma and Carmona-Delarue-Lacker, and it looks plausible. The citation pattern is appropriate; the self-citations are marginal and not load-bearing.\n\nThe main soft spot, as the stress-test note says, is the observational domain. Section 3.3 defines the 'observable' array Q_θ as the full family over O0 ∪ O1, including arbitrary latent compositions ν, distributional states μ, and externally specified measure flows m. The impossibility theorem only works because T(θ) = T(θ̃) is required on all of that domain. An analyst who sees realized histories and the equilibrium diagonal C_s(z;A|μ_s, μ_s) gets none of the off-diagonal variation in Definitions 6 and 8. This is not the counterfactual-menu variation standard in the DRU literature: ν and m are model-internal objects, not things an experimenter can offer or induce. The 'reachable domain' qualifiers do not fix it, because reachability is defined through the model. So the abstract's empirical framing is oversold; the separation is a model-class result, not a testable-data result.\n\nSecondary issue: the identification theorems stack injectivity and measure-determining assumptions (8, 13, 16, 21). They are flagged, honestly, but not motivated. I would want a section that either says what data could identify Φ or retreats to the cleaner conceptual claim.\n\nWho this is for: decision theorists and people at the boundary of stochastic choice and mean-field economics. It deserves a serious referee. My recommendation: send it out, with a referee who knows both the DRU literature and the conditional McKean-Vlasov SDE literature. The machinery is worth having; the observational claims need scaling back.","headline":"Genuinely new conditional-McKean-Vlasov machinery and a sound impossibility theorem, but the separation from dynamic random utility is proven over a counterfactual observation domain, not over realized choice data—worth referee time, with the empirical claims needing to be scaled back.","tokens_in":55208,"tokens_out":7770,"would_cite":true,"duration_ms":77266,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91B06","91B16","60H10","60G35"],"pacs":[],"model":"deepseek-v4-flash","headline":"A stochastic-choice array generated by endogenous preference feedback through the conditional distribution of latent states can never be represented by dynamic random utility.","keywords":["stochastic choice","distribution-dependent utility","dynamic random utility","endogenous preference evolution","conditional McKean-Vlasov dynamics","behavioral identification","behavioral impossibility","conditional preference distribution"],"falsifier":"Take a two-period binary-choice panel in which an exogenous public signal shifts the analyst's posterior over latent preferences without changing the decision maker's fundamentals. Under DRU, the equilibrium choice kernels must satisfy the distributional invariance identities $C_s(z;A|\\nu,\\mu)=C_s(z;A|\\nu,\\mu')$ and $D_{s,r}^{\\bar\\mu}(z;A|\\lambda,m)=D_{s,r}^{\\bar\\mu}(z;A|\\lambda,m')$ for admissible flows. If estimated choice probabilities from such an experiment reject these identities, Theorem 19 predicts the data lie outside $R_{\\mathrm{DRU}}$; if the identities always hold, the behavioral separation would fail. A concrete implementation is to compare choices after different public information histories that induce the same posterior $\\mu_s$ but different counterfactual flows.","tokens_in":54198,"feed_emoji":"🎲","tokens_out":8761,"duration_ms":77852,"temperature":0.7,"pith_summary":"This paper tries to establish that preferences can evolve endogenously: the analyst's conditional distribution of latent preference states, inferred from observed choices, feeds back into current utility and into the future dynamics of preferences. The author contrasts this distribution-dependent utility (DDU) with dynamic random utility (DRU), where latent preferences evolve exogenously and history matters only through learning. The central claim is that behavioral distributional feedback is observable and places the induced stochastic-choice array outside the DRU class, so endogenous preference evolution is a genuinely new behavioral phenomenon rather than a reparameterization. The paper also argues that stochastic choice identifies a two-part behavioral representation, contemporaneous choice and continuation behavior, and that DDU reduces to DRU if and only if both parts are invariant. A sympathetic reader would care because, if true, choice data can in principle detect whether preferences are being changed by what people have already chosen.","feed_headline":"Choices that reshape preferences break dynamic random utility","feed_subtitle":"Endogenous preference feedback is behaviorally distinct from exogenous dynamics, and choice data can tell them apart.","key_machinery":"The central object is the conditional preference distribution (CPD) $\\mu_s = L(X_s \\mid F^Y_s)$, the analyst's posterior over the latent preference state given observed behavior; the paper treats this distribution as an endogenous state rather than a passive posterior. The behavioral representation $\\Phi=(\\Phi_u,\\Phi_P)$ decomposes observable implications into contemporaneous choice rankings $\\Phi_u$ and continuation transition operators $\\Phi_P$. The argument is carried by the requirement that the conditional-law operator $\\Gamma$ on $C([0,t];\\mathcal{P}_2(\\mathbb{R}^d))$ has a fixed point $\\mu=\\Gamma(\\mu)$, where the coefficients of the dynamics depend on $(s,X_s,Y_s,\\mu_s)$. Well-posedness comes from Lipschitz and linear-growth conditions in the 2-Wasserstein metric, a Schauder-Tychonoff fixed-point argument for existence, and a Gronwall-style filter-stability condition for weak uniqueness. The impossibility theorem builds on the equivalence between DRU reducibility and behavioral distributional invariance.","core_discovery":"Distribution-dependent utility makes the conditional preference distribution $\\mu_s = L(X_s \\mid F^Y_s)$ an endogenous state variable: $\\mu_s$ enters the felicity index $u(s,\\cdot,X_s,\\mu_s)$ and the coefficients $(b,\\sigma,\\sigma_0,h,\\Sigma_Y)$ of the latent-observable system, generating the feedback $X \\to Y \\to \\mu \\to X$. The paper's main theorem states that if the stochastic-choice array generated by a DDU representation exhibits behavioral distributional feedback---either choice-relevant felicity feedback or behaviorally relevant preference feedback---then the array admits no dynamic random utility representation. Equivalently, $R_{\\mathrm{DDU}}$ strictly enlarges $R_{\\mathrm{DRU}}$. The same behavioral apparatus yields a characterization of reducibility: DDU reduces to DRU exactly when both the contemporaneous choice map $\\mu \\mapsto C_s$ and the continuation transition map $m \\mapsto D_{s,r}$ are constant on the reachable domain. On the probabilistic side, the paper establishes existence and weak uniqueness of the underlying conditional McKean-Vlasov system with conditional-law feedback, so the behavioral objects are generated by a well-posed fixed-point problem.","pith_inferences":["Editorial inference: the impossibility theorem is stated for full counterfactual arrays; with equilibrium panel data alone, detecting feedback may require exogenous variation in information or menus that changes the analyst's posterior while holding the decision maker's payoff-relevant state fixed.","Editorial inference: the same feedback mechanism could be imported into dynamic discrete choice estimation, where conditional choice probability inversion typically assumes state evolution is exogenous; Theorem 19 gives a formal sense in which such estimates are misspecified when distributional feedback operates.","Editorial inference: a direct empirical strategy is to test whether public signals that move the conditional preference distribution shift future choice probabilities even after controlling for the private latent state; DDU predicts yes, DRU predicts no.","Editorial inference: welfare and policy exercises should be designed around the transition operator and its fixed-point manifold, since the paper's rigidity results imply interventions that leave it unchanged are behaviorally null."],"forward_implications":["If behavioral distributional feedback is present, any fitted dynamic random utility model will misattribute endogenous preference change to Bayesian learning about exogenous preferences.","Stochastic-choice data identify the behavioral image $(\\Phi_u,\\Phi_P)$ but not the primitive coefficient tuple, so structural parameters are only identified up to observational equivalence.","Distribution-dependent utility reduces to dynamic random utility if and only if both contemporaneous and continuation choice maps are invariant, which is a testable condition for when exogenous preference dynamics suffice.","The conditional McKean-Vlasov system is well posed, so the model yields coherent equilibrium preference dynamics rather than an arbitrary feedback loop.","Comparative statics and policy analysis should target the behavioral quotient and the conditional-law operator rather than individual coefficients."],"supporting_citations":[{"why":"Supplies the continuous-time dynamic random utility benchmark that DDU is designed to enlarge and to separate from behaviorally.","marker":"Frick et al., 2019"},{"why":"Provides the conditional McKean-Vlasov well-posedness framework used for existence and weak uniqueness of system (3).","marker":"Buckdahn et al., 2023"},{"why":"Grounds the distribution-dependent dynamics and mean-field equilibrium structure that the conditional-law feedback extends.","marker":"Carmona et al., 2018"},{"why":"Supplies the common-noise conditional-law techniques used in the coupled X-Y system with conditional distribution feedback.","marker":"Carmona et al., 2016"},{"why":"Originates the McKean-Vlasov class of Markov processes with nonlinear distribution-dependent coefficients.","marker":"McKean Jr, 1966"},{"why":"Supplies the nonparametric identification and testing of random utility from stochastic-choice data that the paper's identification results extend.","marker":"Kitamura and Stoye, 2018"}],"fun_headline_variants":["Preference feedback breaks dynamic random utility","Distribution-dependent utility escapes dynamic random utility","Endogenous preference dynamics defy dynamic random utility","Choices reshaping preferences elude random utility"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The identification and impossibility results assume the analyst observes the full counterfactual stochastic-choice array on the domain $O_0 \\cup O_1$---including choices under arbitrary latent-state distributions, distributional states, initial laws, and measure flows---whereas real choice data contain only realized histories and equilibrium-path choice kernels. If that counterfactual domain is not granted, the empirical separation of DDU from DRU loses its observational grounding.","fun_headline_variants_meta":{"raw":{"variants":["Preference feedback breaks dynamic random utility","Distribution-dependent utility escapes dynamic random utility","Endogenous preference dynamics defy dynamic random utility","Choices reshaping preferences elude random utility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3221,"prompt_tokens":933,"completion_tokens":2288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":2247}},"tokens_in":549,"tokens_out":2288,"duration_ms":18443,"temperature":1.0,"reasoning_tokens":2247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:53:09.988208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a two-period binary-choice panel in which an exogenous public signal shifts the analyst's posterior over latent preferences without changing the decision maker's fundamentals. Under DRU, the equilibrium choice kernels must satisfy the distributional invariance identities $C_s(z;A|\\nu,\\mu)=C_s(z;A|\\nu,\\mu')$ and $D_{s,r}^{\\bar\\mu}(z;A|\\lambda,m)=D_{s,r}^{\\bar\\mu}(z;A|\\lambda,m')$ for admissible flows. If estimated choice probabilities from such an experiment reject these identities, Theorem 19 predicts the data lie outside $R_{\\mathrm{DRU}}$; if the identities always hold, the behavioral separation would fail. A concrete implementation is to compare choices after different public information histories that induce the same posterior $\\mu_s$ but different counterfactual flows.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the continuous-time dynamic random utility benchmark that DDU is designed to enlarge and to separate from behaviorally."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the conditional McKean-Vlasov well-posedness framework used for existence and weak uniqueness of system (3)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the distribution-dependent dynamics and mean-field equilibrium structure that the conditional-law feedback extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the common-noise conditional-law techniques used in the coupled X-Y system with conditional distribution feedback."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Originates the McKean-Vlasov class of Markov processes with nonlinear distribution-dependent coefficients."}],"review_version":1}