{"id":"3be83614-d5b0-496e-8865-f6755d6e9001","arxiv_id":"2501.08865","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Boundedly rational moral choice is modeled as maximizing expected utility minus a divergence penalty, with the penalty interpreted as deontology and the coupling constant left to human authority.","lead":"An information-theoretic framework treats deontology as a penalty term that pulls decisions toward a prior, balanced against expected utility. The paper applies this to constitutional rights and leaves the balance parameter to a legal authority.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 7's legal example reverses the sign of the deontic term: with D(d)=dmax−DKL, Eq. (77) is equivalent to maximizing E[U]+β⁻¹DKL, the opposite of Eq. (78).","rationale":"The reader's weakest assumption was that the legal disutility model is an ungrounded interpretive choice. My concern is more specific and more damaging to the internal argument: the concrete legal model in Section 7, using D(d)=dmax−DKL, reverses the sign of the deontic penalty. This is not merely a matter of calibration or external validity; it is a consistency check that fails. If the sign is taken literally, the 'deontic' term pushes the optimizing policy toward large KL divergence, exactly the opposite of the support-restricting role assigned to R in Eq. (78). The core variational mathematics in Sections 5–6 appears coherent, and the KKT derivation of the Gibbs/Boltzmann distribution is standard and correctly executed. A repaired version might define D as an increasing divergence cost, e.g., D=DKL, or explicitly explain the sign convention, and then the legal example could instantiate (78). Therefore I do not recommend rejection: the issue is localized and fixable, and the reader's CONDITIONAL verdict already captures the need for revision. I nevertheless flag this as the single most load-bearing technical check, because without a sign-consistent legal example the paper's headline moral claim lacks its only concrete application. The garbled inserted passage and the several unproved geometric assertions, noted by the reader, are secondary; they do not by themselves determine the central claim as sharply as the Section 7 sign error does.","tokens_in":81512,"tokens_out":5955,"duration_ms":72128,"concrete_test":"Independently substitute D(d(p,q)) := dmax − DKL(p||q) into Eq. (77) and reduce the objective to E_p[U] + (1/β)DKL(p||q) − dmax/β. Then evaluate that objective at p=q and at a vertex p=δ_j for a fixed q, U, and β; if the vertex with larger DKL gives a larger value, the term is a pro-divergence reward. Repeat with D=+DKL; if the maximizers change, the sign of Section 7's deontic term is decisive. This directly settles whether (77) is an instance of (78).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that deontology is a regularizer in max_p E_p[U] − (1/β) R(p,q), as stated in Eq. (78) and in the multiplier-robust problem (25) where the KL term is subtracted. The legal application in Section 7 is the only place where this identification is supposed to be instantiated for constitutional doctrine. But as written, it does the opposite. The text defines D(d(p,q)) := dmax − DKL(p||q) and then uses (77): max_p E_p[U] − (1/β) D(d(p,q)). Substituting gives E_p[U] − (1/β)(dmax − DKL(p||q)) = E_p[U] − dmax/β + (1/β)DKL(p||q). Since dmax/β is constant, the optimization is equivalent to maximizing E_p[U] + (1/β)DKL(p||q). The KL term is now a reward for moving away from the prior, not a penalty that confines decisions to the support of q. This contradicts the role asserted for R in Eq. (78), where R is a cut-off that incorporates deontological protections, and also contradicts the sign of the cost term in Eq. (25). The chosen D is also not strictly convex in the required sense: D(d)=dmax−d is affine, and the composed expression is concave after the minus sign. Thus the paper's only concrete legal model does not support the deontology-as-regularization claim; it is internally inconsistent with the central variational objective unless some unstated sign convention or different interpretation of D is supplied.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes that bounded-rational moral and legal decision making is captured by a variational objective in which expected utility plays the role of utilitarianism and a divergence/regularization term plays the role of deontology. After introducing Markov kernels, co-sheaves, and information-geometric tools, the paper derives the Boltzmann-Gibbs solution for the multiplier-robust control problem, studies a rate-utility version of the constraint robust-control problem, and applies the framework to a model of constitutional-right restrictions. The central claim is Eq. (78): the optimal policy maximizes E_p[U] minus (1/beta)R(p,q), where R is interpreted as a deontic regularizer and beta is a free coupling constant to be fixed by a legislative or judicial authority.","tokens_in":81888,"tokens_out":8543,"duration_ms":94571,"significance":"If the moral identification in Eq. (78) were established, the paper would offer a compact mathematical bridge between information-theoretic bounded rationality, moral theory, and legal doctrine. The variational derivations in Section 5 are standard and mostly correct: the KKT conditions yield the Gibbs distribution, and the convexity/concavity statements about the rate-utility function in Section 6 follow familiar arguments. The paper is also honest that beta is not determined internally. However, the central 'deontology as regularization' claim is not derived but assumed by labelling R as deontic, the legal model in Section 7 contains a sign inconsistency, and the paper provides no falsifiable prediction or independent constraint on q, R, or beta. The mathematical framework is therefore not yet the moral theory the title promises.","major_comments":[{"comment":"The legal application has a sign inconsistency with the main objective. With D(d(p,q)) = d_max - D_KL(p||q), the objective in (77) becomes E_p[U] - (1/beta)(d_max - D_KL(p||q)) = E_p[U] - d_max/beta + (1/beta)D_KL(p||q). Since the constant term does not affect the argmax, the optimization is equivalent to maximizing E_p[U] + (1/beta)D_KL(p||q), so the KL term rewards departure from the prior instead of penalizing it. This contradicts the sign of the deontic term in Eqs. (25) and (78), where R is subtracted as a cut-off protecting the support of q. In addition, D(d) = d_max - d is affine rather than strictly convex, contrary to the paper's own condition that the disutility function be a strictly decreasing convex function of d. As written, Section 7 does not instantiate 'deontology as regularization'; it instantiates the opposite of Eq. (78).","section":"§7, Eq. (77) and the definition of D(d)"},{"comment":"The identification of the divergence penalty with deontology is an interpretive assumption rather than a derived result. The label 'deontic' is attached to R(p,q) in Eq. (78), and the support of q is described as a deontological cut-off, but no argument from deontological ethics, legal texts, or behavioral data establishes that a moral system's content is represented by a prior q and a regularizer R. Section 7 selects the disutility function from three ad-hoc 'basic types (not to scale)' without doctrinal derivation. The paper's own statement that beta 'will remain a free parameter to be determined by the competent legislative or judicial authority' concedes that the framework has no independent predictive content for the central mapping. Unless the mapping is constrained by a separate theory or by testable predictions, Eq. (78) is a relabeling of a known bounded-rationality optimization problem, not a discovery about deontology.","section":"§7 and Eq. (78)"},{"comment":"The transition from the stated optimization problem (47) to the minimization in (48) and (52) is not an equivalence. Problem (47) maximizes over the prior kernel κ, while (48) fixes K and minimizes D_KL(P⋊K || P⋊κ) over κ; these are different variational problems. For a fixed K, the minimizer κ* = K, or q* = K_*P in the constant-kernel case, does not in general maximize the free-energy expression in (47), which is driven toward priors concentrated on high-utility actions. Consequently the rate-utility problem (53) and the subsequent concavity analysis are not consequences of the stated starting point. This does not invalidate the Section 5 derivation of the Gibbs solution, but it undermines the claimed generalization in Section 6 unless an additional equivalence argument is supplied.","section":"§6, Eqs. (47)–(53)"}],"minor_comments":[{"comment":"The displayed objective in Eq. (1) uses a minimization with a plus sign in front of (1/beta)D_KL(p||q), whereas Eqs. (25) and (78) use maximization with a minus sign; this inconsistency should be corrected.","section":"§1, Eq. (1)"},{"comment":"The text says 'The expression is referred to as multiplier robust-control problem []' with an empty citation; the reference should be filled in.","section":"§1, after Eq. (1)"},{"comment":"The existence and uniqueness of the utility expansion path satisfying (74) and the contraction path satisfying (76) are not proved; the claimed disjointness and reflection symmetry of the two paths should be stated as a proposition with explicit hypotheses.","section":"§6.3, Definitions 6.1–6.2"},{"comment":"The caption says the z-axis shows F_beta[p] as a function of temperature 1/beta, but also lists a prior q = (0.7,0.2,0.1) in Delta_2 and a utility vector U = (7,5); the dimensional mismatch and the precise role of the displayed curve should be clarified.","section":"Figure 9 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is a broad interdisciplinary synthesis whose main formal results are standard within decision theory. The sign error in Section 7 is fixable, but the more serious issue is that the moral and legal content is put in by definition, and the paper explicitly leaves beta as an external free parameter. If the journal values conceptual cross-disciplinary papers, a major revision that replaces the interpretive claims with a clearly labeled modeling proposal and repairs the Section 6 equivalence would be appropriate. The manuscript would also benefit from a more rigorous treatment of the legal doctrine or from an explicit statement that the legal section is only an illustrative analogy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: the paper does not deliver a new mathematical result. The variational objective in Eq. (78) is the standard bounded rationality / free-energy objective, and the Gibbs solution is textbook. The new part is the interpretive claim that deontology is the KL-regularizer and utilitarianism is the expected utility. That is a framing, not a discovery, and the paper's Eq. (78) builds the deontic term in as the regularizer and then reads it back out as a finding.\n\nWhat the paper does well: Sections 5 and 6 are a clean, correct account of the multiplier-robust and constraint-robust control problems. The KKT derivation, the convexity/concavity of the rate-utility function, and the geometric remarks about m-/e-geodesics are all sound. The author is honest about provenance, citing Ortega and Braun, Ortega and Stocker, Mattsson and Weibull, and the rate-distortion literature. The utility expansion and contraction paths (gamma+ and gamma-) are a reasonable geometric addition, though not deeply developed.\n\nThe soft spots are real. First, the moral mapping is an assumption. The prior q and the regularizer R are free inputs, and beta is explicitly a free parameter left to the legislator. So there is no independent prediction that could validate the deontology-as-regularization identification. Second, the legal section contains a load-bearing sign error. The text defines D(d) = dmax − DKL(p||q) and then maximizes Ep[U] − (1/beta)D(d). That is equivalent to maximizing Ep[U] + (1/beta)DKL(p||q), which rewards moving away from the prior—exactly the opposite of the deontic cut-off in Eq. (78). The chosen D is also affine, not strictly convex as the text claims. This is the only worked legal application, so the central claim is not actually instantiated. Third, several geometric claims (the reflection symmetry of the expansion and contraction paths, for instance) are stated without proof. Fourth, there is an unedited garbled paragraph near the start of Section 3.2 that should have been removed.\n\nThere is enough coherent mathematics here that a referee could help the author fix the sign error and tighten the framing. As it stands, the paper is a conceptually suggestive, mathematically standard sketch with a flawed application. I would not cite it yet.\n\nWho is this for? Readers interested in information-theoretic formulations of moral psychology or AI-agent constraints might find the framing worth a look, but they should treat the legal claims with suspicion. I would send it to peer review because the issues are fixable and the topic deserves engagement, but I would not accept it without major revision.","headline":"A clean recapitulation of bounded rationality dressed as a moral theory, with a legal example that has a sign error making the deontic term a reward rather than a penalty.","tokens_in":82403,"tokens_out":3054,"would_cite":false,"duration_ms":32841,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A17","94A34","62B11","91B06"],"pacs":[],"model":"deepseek-v4-flash","headline":"Moral and legal decision making, the paper claims, is one variational problem whose two terms are utilitarianism, $E_p[U]$, and deontology, the penalty $\\frac{1}{\\beta} R(p,q)$.","keywords":["bounded rationality","deontology","utilitarianism","rate distortion","information geometry","Boltzmann-Gibbs distribution","Markov kernels","constitutional rights"],"falsifier":"The framework predicts that, for fixed utility assignments, observed choices follow the Boltzmann-Gibbs form $p(i) \\propto q_i e^{\\beta u_i}$ for some prior $q$ and weight $\\beta \\geq 0$. That prediction is falsifiable by behavioral data: fit $\\beta$ and $q$ to the response frequencies of a large set of moral-dilemma decisions with known utilities, and the claim fails if the residuals are systematic — for example, if choices persistently violate the ratio property of Luce's choice axiom, or if no single fitted pair $(q, \\beta)$ reproduces the frequencies across variations of the dilemma. In the legal domain, the applied model predicts discontinuous switches between permitted restrictions as $\\beta$ crosses thresholds; observing smooth, history-dependent restriction patterns that no piecewise-constant $\\beta$ can reproduce would falsify the application.","tokens_in":81239,"feed_emoji":"⚖️","tokens_out":16188,"duration_ms":142093,"temperature":0.7,"pith_summary":"The paper tries to establish that resource-bounded moral and legal decision making is a single optimization problem, not a clash between two incompatible theories. The decision maker chooses the policy $p$ that maximizes expected utility $E_p[U]$ minus the scaled divergence penalty $\\frac{1}{\\beta} R(p,q)$, where the expected-utility term is the utilitarian component and the penalty, anchored to a prior $q$ over the permitted actions, is the deontological component. The optimum is always a Boltzmann-Gibbs distribution, interpolating from pure rule-following at $\\beta \\to 0$ to pure consequence-maximizing at $\\beta \\to \\infty$, and the coupling constant $\\beta$ is declared a free parameter to be fixed by the legislator or the court. This would matter because it reduces an old philosophical dichotomy to a concrete quantitative trade-off and gives one formal object, Eq. (78), in which the restriction of constitutional rights appears as a constrained second-best optimization with the authority's discretion built in as the free parameter.","feed_headline":"Reduce moral choice to one formula: utility minus a rules penalty","feed_subtitle":"Rule-following enters as a penalty term, and the weight of rules vs. results is left to judges and legislators.","key_machinery":"The central object is the variational objective of Eq. (78), $\\max_p \\left( E_p[U] - \\frac{1}{\\beta} R(p,q) \\right)$, in which $q$ is a prior over actions, $R$ is a divergence-based regularizer, and $\\beta$ is the inverse temperature. That objective carries the argument because its support condition makes the prior the encoding of the legal or moral code — only actions in the support of $q$ are eligible — and the divergence term makes rule-following a graded pull rather than a hard constraint. Solving it yields the Boltzmann-Gibbs distribution, an exponential family on the probability simplex, and varying $\\beta$ moves the solution along the $e$-geodesic through $q$. The constraint-robust variant replaces the fixed prior with a source distribution and organizes the same trade-off through the rate-utility function, whose slope is $1/\\beta$ and whose tangency points with constant-mutual-information surfaces define the utility expansion path.","core_discovery":"The central claim is that a resource-bounded decision maker resolves the conflict between deontology and utilitarianism inside a single variational objective, Eq. (78): the optimal policy maximizes expected utility $E_p[U]$ minus the scaled penalty $\\frac{1}{\\beta} R(p,q)$, where the expected-utility term is the utilitarian component and the regularizer $R$, anchored to a prior $q$ whose support is the set of permitted actions, is the deontological component. The solution is the Boltzmann-Gibbs distribution, $p^*_\\beta(i) \\propto q_i e^{\\beta u_i}$, an exponential-family weighting of actions by utility that interpolates between pure rule-following at $\\beta \\to 0$ and unrestricted utility maximization at $\\beta \\to \\infty$, tracing an $e$-geodesic through the prior as the weight is swept. In the constraint-robust version with a source distribution over world states, the same trade-off is organized by a rate-utility function whose slope at every point is $1/\\beta$, and the optimal kernels solve self-consistent equations of rate-distortion type. The author's stated point is that neither moral theory determines the coupling constant: it remains a free parameter for the legislative or judicial authority, and that free parameter is the formal locus of the margin of discretion.","pith_inferences":["Editorial extension: with $\\beta$ treated as a quantity to be fitted rather than legislated, the framework becomes behaviorally testable — one could estimate $\\beta$ and $q$ from observed choice frequencies in moral dilemmas and ask whether a single pair transfers across situations, a check the paper does not perform.","Editorial extension: because the constraint-robust problem is formally a rate-distortion problem, moral or legal vagueness can be read as a compression phenomenon — coarse representations of a situation cost less information, so an 'optimal vagueness' would trade decision accuracy against coding cost, a consequence the paper leaves implicit.","Editorial consequence of the mapping: if deontology is a regularizer, then disagreements between rule-based and consequentialist moral theories are disagreements about the support of the prior and the value of the coupling constant — a philosophical dispute relocated onto two numbers that the paper hands to the legislative authority."],"forward_implications":["If the framework is right, every moral choice problem is specified by the prior $q$, the regularizer $R$, and the coupling $\\beta$; no separate moral theory is needed beyond these ingredients.","The optimal moral policy is never a pure rule or a pure maximizer in the interior regime: it is the Boltzmann-Gibbs distribution, and the two classical theories are recovered only in the limits $\\beta \\to 0$ and $\\beta \\to \\infty$.","In the legal application, a restriction of a fundamental right is justified exactly when the public utility gain exceeds the disutility of the restriction, and the model predicts that the chosen restriction switches discontinuously as the authority's weight $\\beta$ crosses critical thresholds.","Because the constraint-robust problem is formally a rate-distortion problem, coarse-graining the space of states is governed by the data-processing inequality: abstraction cannot increase the information a decision can carry, so hierarchical and simplified moral reasoning follows from the same objective.","Autonomous agents implementing Eq. (78) inherit a tunable deontology: their rule-following behavior is controlled by a single externally supplied constant rather than by a hard-coded rule list."],"supporting_citations":[{"why":"supplies the probabilistic bounded-rationality model whose optimal solution is the Gibbs distribution, the starting point of the multiplier-robust control problem","marker":"[27]"},{"why":"provides the thermodynamic reading of bounded rationality as free-energy differences that the paper adopts for the decision objective","marker":"[31]"},{"why":"sets out the free-energy objective for hierarchical bounded-rational decision making that the constraint-robust problem generalizes","marker":"[19]"},{"why":"introduces the Information Bottleneck method, the conceptual template for the constraint robust-control problem","marker":"[37]"},{"why":"frames bounded rationality as a regularization phenomenon, the perspective the paper transfers to deontology","marker":"[32]"},{"why":"supplies rate-distortion theory, including the structure of the optimality equations and the analogue of the rate-utility function","marker":"[4]"},{"why":"provides the information-theoretic facts used throughout: KL convexity, the data-processing inequality, and rate-distortion duality","marker":"[13]"},{"why":"grounds the information geometry of exponential families, Bregman divergences, and m- and e-projections used to describe the solutions","marker":"[2]"},{"why":"supplies the multiplier-versus-constraint robust-control classification used to organize the two optimization problems","marker":"[22]"}],"fun_headline_variants":["One formula for moral dilemmas: utility minus rules penalty","Moral decisions as exponential weighting of allowed actions","Deontology is a penalty term in utility maximization","The temperature of morality: trading rules for utility"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands on the interpretive identification of rule-following (deontology) with a divergence penalty against a prior over permitted actions — a modeling choice that the paper asserts rather than derives from moral theory, legal texts, or data — together with the external supply of the penalty weight $\\beta$.","fun_headline_variants_meta":{"raw":{"variants":["One formula for moral dilemmas: utility minus rules penalty","Moral decisions as exponential weighting of allowed actions","Deontology is a penalty term in utility maximization","The temperature of morality: trading rules for utility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000386,"raw_usage":{"total_tokens":2035,"prompt_tokens":935,"completion_tokens":1100,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":1039}},"tokens_in":551,"tokens_out":1100,"duration_ms":9371,"temperature":1.0,"reasoning_tokens":1039,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:16:36.562830+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The framework predicts that, for fixed utility assignments, observed choices follow the Boltzmann-Gibbs form $p(i) \\propto q_i e^{\\beta u_i}$ for some prior $q$ and weight $\\beta \\geq 0$. That prediction is falsifiable by behavioral data: fit $\\beta$ and $q$ to the response frequencies of a large set of moral-dilemma decisions with known utilities, and the claim fails if the residuals are systematic — for example, if choices persistently violate the ratio property of Luce's choice axiom, or if no single fitted pair $(q, \\beta)$ reproduces the frequencies across variations of the dilemma. In the legal domain, the applied model predicts discontinuous switches between permitted restrictions as $\\beta$ crosses thresholds; observing smooth, history-dependent restriction patterns that no piecewise-constant $\\beta$ can reproduce would falsify the application.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the probabilistic bounded-rationality model whose optimal solution is the Gibbs distribution, the starting point of the multiplier-robust control problem"},{"cited_title":"A., and Braun, D","cited_arxiv_id":null,"evidence_quote":"provides the thermodynamic reading of bounded rationality as free-energy differences that the paper adopts for the decision objective"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"sets out the free-energy objective for hierarchical bounded-rational decision making that the constraint-robust problem generalizes"},{"cited_title":"A., and Stocker, A","cited_arxiv_id":null,"evidence_quote":"frames bounded rationality as a regularization phenomenon, the perspective the paper transfers to deontology"},{"cited_title":"Rate distortion theory : a mathematical basis for data compression","cited_arxiv_id":null,"evidence_quote":"supplies rate-distortion theory, including the structure of the optimality equations and the analogue of the rate-utility function"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the information-theoretic facts used throughout: KL convexity, the data-processing inequality, and rate-distortion duality"},{"cited_title":"Information geometry and its applications , vol","cited_arxiv_id":null,"evidence_quote":"grounds the information geometry of exponential families, Bregman divergences, and m- and e-projections used to describe the solutions"},{"cited_title":"P., and Sargent, T","cited_arxiv_id":null,"evidence_quote":"supplies the multiplier-versus-constraint robust-control classification used to organize the two optimization problems"}],"review_version":1}