REVIEW 1 major objections 4 minor 46 references
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
T0 review · 1 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves that a computable near-optimal social-welfare policy exists whenever a predictive world model is accurate enough, and that any black-box policy can be made safe by blocking destructive actions.
desk verdict A worthwhile SMDP/PAA framework whose all-q existence theorem is undercut by a false concentration lemma for q<0 and mismatched Gamma constants for q=0,1. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the sparse-sampling $Q$-estimator $\hat Q^h$ defined recursively in Eq. (6), where the reward is the power-mean welfare of a random subset of assessors and the next-state expectations are taken under the approximate model. The proof skeleton follows the classic sparse-sampling planner, but the approximation machinery is new: Lemma 5 gives a Hoeffding--Serfling concentration bound for the power mean, Lemma 6 bounds the reward error from approximating $p$ by $\hat p$, and Lemma 7 bounds the value loss incurred by acting greedily with respect to an approximate $Q$-function. The planner's parameters $K$, $C$, $n$, and $H$ are chosen so that all six error terms fit inside the approximation budget determined by $\epsilon$ and the model error.
What would settle it
Evaluate the claimed inequality used in Lemma 5 with $q=-1$ and $x=0.1$: it asserts $(1+x)^q \le 1-(1-2q)x$, i.e. $0.909 \le 0.7$, which is false; that single failure means the concentration bound for $q<0$ is not established, and the theorem's negative-$q$ branch collapses unless a different bound is found.
Extended reading notes
Core claim
The central claim is Theorem 2: for any social MDP with power-mean welfare $W_q$ and any tolerances $\epsilon>0$, $\delta\in[0,1)$, if there exists an approximate dynamics model $\hat p$ with $\sup_{(s,a)} D_{\mathrm{KL}}(p(\cdot|s,a)\,\|\,\hat p(\cdot|s,a)) < \epsilon^2(1-\gamma)^4 / (8\,\Delta W^2)$, then a computable $\delta$-$\epsilon$-PAA policy exists. The policy is the greedy action on recursively estimated $Q$-values from a sparse-sampling planner that simulates transitions with $\hat p$ and estimates rewards from a finite panel of sampled utilities. The proof decomposes the error into six terms and bounds each with concentration inequalities: Lemma 5 controls sampling error of the power mean, Lemma 6 controls model mismatch through KL divergence, and Lemma 7 converts approximate $Q$-values into a value-function loss. Consequently, near-optimal alignment can be certified from model accuracy and finite utility feedback without ever observing the agent's true objective.
Load-bearing premise
The existence theorem depends on a concentration bound for how well a random sample of citizens estimates the social-welfare average, and the proof of that bound for negative welfare exponents uses an algebraic inequality that does not hold, so the theorem's coverage of that range is not established.
Editorial extensions
If this is right
- For any $\epsilon>0$ and $\delta\in[0,1)$, a computable $\delta$-$\epsilon$-PAA policy exists whenever the world model's worst-case KL error is below the stated threshold, with explicit sample sizes and horizon.
- Because the bounds do not depend on the size of the state space, the existence result applies to infinite state spaces as long as the action space is finite.
- Any black-box policy can be converted into a $\delta$-$\omega$-safe policy by action masking, at the cost of refusing actions whose estimated continuation value is too low.
- The safe-policy result does not require the world model to meet the PAA accuracy threshold; lower model accuracy only shrinks the set of verifiably safe actions.
- Alignment becomes an a priori, quantitative property: a society can verify the guarantee from the world model and a finite utility sample rather than from observed behavior.
Reading between the lines
- If the theorem survives the $q<0$ gap, alignment audits could shift from inspecting an agent's objective to validating the predictive model, because the guarantee is driven entirely by world-model accuracy and sampled utilities.
- The safe-policy wrapper suggests a practical deployment path: keep a black-box policy for its competence, compute $\hat Q^H$ from a learned simulator, and veto any action whose estimated continuation value falls below the safe threshold; this is testable in any simulator with known ground truth.
- Because the reward class is tied to the power mean, the framework's applicability depends on which informational basis a society adopts for comparing utilities; changing that basis changes the concentration constants and hence the required sample sizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defines a Social Markov Decision Process whose reward is the expected future discounted power-mean social welfare W_q of the members of a society, and defines a policy to be δ-ε-PAA if, with probability at least 1−δ, its expected social welfare is within ε of the optimal policy. The central theoretical contribution is Theorem 2: under a uniform KL-divergence bound on an approximate world model, a computable δ-ε-PAA policy exists for every q∈R. The proof adapts the sparse-sampling planner of Kearns et al., bounding separately the error from sampling a subset of assessors, from the approximate dynamics model, and from Monte Carlo rollouts. The paper also introduces δ-ω-safe policies and proves in Theorem 8 that any black-box policy can be restricted to a safe action set while preserving a desired safety level.
Significance. If the technical gaps are repaired, the paper offers a useful formal bridge between social choice and MDP planning: alignment is quantified as ε-δ near-optimality in a constructed social MDP, and the sufficient condition on model accuracy is explicit and checkable in principle. The proof architecture is transparent and largely faithful to sparse sampling, and the claimed sample complexities are independent of the number of states. I found no circularity: PAA is defined independently as near-optimality in the constructed MDP, and the theorems are derived from standard concentration and planning results rather than by fitting the definitions to the conclusions. The main blocker is a false concentration lemma used for negative power-mean social welfare functions.
major comments (1)
- [Appendix A.2.1, Lemma 5 (Eqs. (8)–(9))] The q<0 branch of Lemma 5 is invalid. The proof asserts the inequality (1+x)^q ≤ 1−(1−2q)x for 0<x≤1 and q<0, which is false; for q=−1 and x=0.1 it gives 0.909 ≤ 0.7. The stated bound is not merely unproved but contradicted by a finite example: take q=−1, X={1,2}, N=2, n=1, ε=0.6. The population power mean is 4/3, the sample is 1 or 2 with equal probability, so P(|S−4/3|≥0.6)=0.5. With Γ(ε,a,b,q)=(1−2q)^2 b^{2q−2}/(a^q−b^q)^2=9/4, Eq. (8) gives 2 exp(−2·1·0.36·2.25/((1−1/2)(1+1/1)))=2e^{−1.62}≈0.396<0.5. Since Theorem 3 uses Lemma 5 to bound the term Z1 for every q∈R, and Theorem 2 claims existence of PAA policies for every q∈R, the existence theorem is unproved for negative-power-mean social welfare functions, including the harmonic and near-egalitarian cases the paper explicitly claims to cover. Theorem 8 inherits the same gap through Γmax in the definition of α. A valid substitute concentration bound for q<0, or a restriction of the theorems to q≥0, is needed.
minor comments (4)
- [Appendix A.2.1, Lemma 6] The lemma statement contains a dangling fragment: "such that D_KL(p∥p̂) ≤ d ∈ R and ." This should be completed or the stray "and ." removed.
- [Section 2.1.1 and Eq. (9)] The text says Umin=0 is allowed in specific cases, but the q=0 branch of Lemma 5 requires log a to be defined and hence requires Umin>0 for the geometric mean. Please clarify which values of q are admissible when Umin=0.
- [Theorem 3 proof, near Eq. (18)] The bookkeeping of the γ^k term is hard to follow because the displayed expression and the definition of β contain unbalanced parentheses and a conditional formatting artifact. Rewriting that line with explicit parentheses would improve readability.
- [Eq. (3)] The notation \hat E^K_{s'∼\hat p} is used in Eq. (3) before it is defined in the surrounding text; please define the empirical expectation operator before first use.
Circularity Check
No significant circularity: the PAA existence theorem is a constructive MDP planning result whose alignment objective is explicitly defined as social welfare; no fitted quantity is renamed as a prediction.
full rationale
The central chain (Section 3, Theorem 2 then Theorem 3) proves that if a world model p_hat satisfies the KL condition (Eq. 5), then the sparse-sampling policy of Eq. (7) is epsilon-delta optimal with respect to the expected future discounted social welfare W^pi. The alignment definition in Section 2.1.3 and the SMDP reward r_I in Eq. (2) use the same social-welfare functional W_q; this is an explicit modeling commitment, not an input secretly equal to the output. The parameters K, C, n, H are chosen in Theorem 3 to meet an epsilon-delta budget, not fitted to any data and then reported as predictions. The external ingredients (Kearns et al. sparse sampling, Hoeffding-Serfling concentration, social-choice representation theorems, Roberts/Cousins) are independent citations and are not self-citations by the authors. Hence no derivation step reduces by construction to its own input. Separately, and without changing the circularity verdict, the q<0 branch of Lemma 5 in Appendix A.2.1 appears mathematically unsupported: the proof's inequality (1+x)^q <= 1-(1-2q)x is false for q<0, so the all-q statement of Theorem 3 has a correctness risk; that is a proof flaw, not circularity.
Assumptions & free parameters
assumptions (6)
- standard math Debreu's representation theorem: complete, continuous, transitive preferences admit continuous utility functions.
- domain assumption The social welfare function is a power mean W_q, derived from axioms (U), (XI), (IIA), (WP), (A) and an informational basis.
- domain assumption Utilities are bounded, time-invariant, and interpersonally comparable at the level required for W_q.
- domain assumption The approximate world model p_hat satisfies the stated sup-KL condition.
- domain assumption Sampled assessors report truthful utilities and can evaluate complete social states.
- standard math Hoeffding-Serfling concentration inequalities and Pinsker/Bretagnolle-Huber inequalities hold as cited.
Cite this review
Pith. "Pith review of Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies." pith.science (2026). https://pith.science/paper/N3BKQSBE
@misc{pith2026241200033,
author = {Pith},
title = {Pith review of: Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies},
year = {2026},
howpublished = {\url{https://pith.science/paper/N3BKQSBE}},
note = {Machine review of arXiv:2412.00033}
}
read the original abstract
While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications such as social decision processes. More importantly, existing alignment methods provide no formal guarantees on the safety of such models. Drawing from utility and social choice theory, we provide a novel quantitative definition of alignment in the context of social decision-making. Building on this definition, we introduce probably approximately aligned (i.e., near-optimal) policies, and we derive a sufficient condition for their existence. Lastly, recognizing the practical difficulty of satisfying this condition, we introduce the relaxed concept of safe (i.e., nondestructive) policies, and we propose a simple yet robust method to safeguard the black-box policy of any autonomous agent, ensuring all its actions are verifiably safe for the society.
Figures
Reference graph
Works this paper leans on
-
[1]
Constrained Markov Decision Processes
Eitan Altman. Constrained Markov Decision Processes. Routledge, 2021
work page 2021
-
[2]
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. Concrete Problems in AI Safety. arXiv preprint arXiv:1606.06565, 2016
arXiv 2016
-
[3]
Social Choice and Individual Values
Kenneth Joseph Arrow. Social Choice and Individual Values. Wiley: New York, 1951
work page 1951
-
[4]
Concentration Inequalities for Sampling without Replace- ment
Rémi Bardenet and Odalric-Ambrym Maillard. Concentration Inequalities for Sampling without Replace- ment. Bernoulli, 21(3):1361 – 1385, 2015
work page 2015
-
[5]
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile Lukosuite, Amanda Askell, Andy Jones, Anna Chen, et al. Measuring Progress on Scalable Oversight for Large Language Models. arXiv preprint arXiv:2211.03540, 2022
arXiv 2022
-
[6]
Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei
Paul F. Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems 30 (NIPS), pages 4299–4307, 2017
work page 2017
-
[7]
An Axiomatic Theory of Provably-Fair Welfare-Centric Machine Learning
Cyrus Cousins. An Axiomatic Theory of Provably-Fair Welfare-Centric Machine Learning. In Advances in Neural Information Processing Systems 34 (NeurIPS), pages 16610–16621, 2021
work page 2021
-
[8]
Rethinking the Maturity of Artificial Intelligence in Safety-Critical Settings
Mary Cummings. Rethinking the Maturity of Artificial Intelligence in Safety-Critical Settings . AI Magazine, 42(1):6–15, 2021
work page 2021
Show all 46 references
-
[9]
Equity and the Informational Basis of Collective Choice
Claude D’Aspremont and Louis Gevers. Equity and the Informational Basis of Collective Choice. The Review of Economic Studies, 44(2):199–209, 1977
1977
-
[10]
Representation of a preference ordering by a numerical func- tion
Gerard Debreu and Werner Hildenbrand. Representation of a preference ordering by a numerical func- tion. In Mathematical Economics: Twenty Papers of Gerard Debreu, Econometric Society Monographs, chapter 6, page 105–110. Cambridge University Press, 1983
-
[11]
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment. arXiv preprint arXiv:2304.06767, 2023
2023 arXiv
-
[12]
Feinberg
Eugene A. Feinberg. Total Expected Discounted Reward MDPS: Existence of Optimal Policies. Wiley Encyclopedia of Operations Research and Management Science. John Wiley & Sons, Ltd, 2011
2011
-
[13]
Artificial intelligence, values, and alignment
Iason Gabriel. Artificial intelligence, values, and alignment. Minds and machines, 30(3):411–437, 2020
2020
-
[14]
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv preprint arXiv:2209.07858, 2022
2022 arXiv
-
[15]
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll. A Review of Safe Reinforcement Learning: Methods, Theory and Applications. arXiv preprint arXiv:2205.10330, 2023
2023 arXiv
-
[16]
Hicks and Roy G
John R. Hicks and Roy G. D. Allen. A Reconsideration of the Theory of Value. Part I. Economica, 1(1): 52–76, 1934
1934
-
[17]
AI Alignment: A Comprehensive Survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhong- hao He, Jiayi Zhou, Zhaowei Zhang, et al. AI Alignment: A Comprehensive Survey. arXiv preprint arXiv:2310.19852, 2023
2023 arXiv
-
[18]
On the Sample Complexity of Reinforcement Learning
Sham Machandranath Kakade. On the Sample Complexity of Reinforcement Learning. Doctoral dissertation, University College London (United Kingdom), 2003
2003
-
[19]
A Sparse Sampling Algorithm for Near-optimal Planning in Large Markov Decision Processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng. A Sparse Sampling Algorithm for Near-optimal Planning in Large Markov Decision Processes. Machine learning, 49:193–208, 2002. 10
2002
-
[20]
Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking
Hanna Krasowski, Jakob Thumm, Marlon Müller, Lukas Schäfer, Xiao Wang, and Matthias Althoff. Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking. arXiv preprint arXiv:2205.06750, 2023
2023 arXiv
-
[21]
Social Choice with Changing Preferences: Representation Theorems and Long-run Policies
Kshitij Kulkarni and Sven Neth. Social Choice with Changing Preferences: Representation Theorems and Long-run Policies. arXiv preprint arXiv:2011.02544, 2020
2011 arXiv
-
[22]
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications
Abhilash Mishra. AI Alignment and Social Choice: Fundamental Limitations and Policy Implications. arXiv preprint arXiv:2310.16048, 2023
2023 arXiv
-
[23]
Value-Guided Synthesis of Parametric Normative Systems
Nieves Montes and Carles Sierra. Value-Guided Synthesis of Parametric Normative Systems. In Proceed- ings of the 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), page 907–915, 2021
2021
-
[24]
The Alignment Problem from a Deep Learning Perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann. The Alignment Problem from a Deep Learning Perspective. arXiv preprint arXiv:2209.00626, 2022
2022 arXiv
-
[25]
Accountability in Artificial Intelligence: What it is and how it works
Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi. Accountability in Artificial Intelligence: What it is and how it works. AI & SOCIETY, pages 1–12, 2023
2023
-
[26]
Christiano, Jan Leike, and Ryan Lowe
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, ...
2022
-
[27]
Manuale di Economia Politica con una Introduzione alla Scienza Sociale
Vilfredo Pareto. Manuale di Economia Politica con una Introduzione alla Scienza Sociale . Piccola biblioteca scientifica. Societa editrice libraria, 1906
1906
-
[28]
Dynamic Social Choice with Evolving Preferences
David Parkes and Ariel Procaccia. Dynamic Social Choice with Evolving Preferences. In Proceedings of the 27th AAAI conference on artificial intelligence, pages 767–773, 2013
2013
-
[29]
Kevin W. S. Roberts. Interpersonal Comparability and Social Choice Theory. The Review of Economic Studies, 47(2):421–439, 1980
1980
-
[30]
Research Priorities for Robust and Beneficial Artificial Intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark. Research Priorities for Robust and Beneficial Artificial Intelligence. AI magazine, 36(4):105–114, 2015
2015
-
[31]
Schwartz
Shalom H. Schwartz. Universals in the Content and Structure of Values: Theoretical Advances and Empirical Tests in 20 Countries. In Mark P. Zanna, editor, Advances in Experimental Social Psychology, volume 25, pages 1–65. Academic Press, 1992
1992
-
[32]
On Weights and Measures: Informational Constraints in Social Welfare Analysis
Amartya Sen. On Weights and Measures: Informational Constraints in Social Welfare Analysis. Econo- metrica, 45(7):1539–1572, 1977
1977
-
[33]
R. J. Serfling. Probability Inequalities for the Sum in Sampling without Replacement. The Annals of Statistics, 2(1):39 – 48, 1974
1974
-
[34]
Value Alignment: A Formal Approach
Carles Sierra, Nardine Osman, Pablo Noriega, Jordi Sabater-Mir, and Antoni Perelló. Value Alignment: A Formal Approach. arXiv preprint arXiv:2110.09240, 2021
2021 arXiv
-
[35]
Singh and Richard C
Satinder P. Singh and Richard C. Yee. An Upper Bound on the Loss from Approximate Optimal-Value Functions. Machine Learning, 16(3):227–233, 1994
1994
-
[36]
Robert H. Strotz. Cardinal Utility. The American Economic Review, 43(2):384–397, 1953
1953
-
[37]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto. Reinforcement learning: An introduction. MIT press, 2018
2018
-
[38]
A Theory of the Learnable
Leslie Gabriel Valiant. A Theory of the Learnable. Communications of the ACM, 27(11):1134–1142, 1984
1984
-
[39]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y . Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V . Le. Finetuned Language Models are Zero-Shot Learners. In Proceedings of the 10th International Conference on Learning Representations (ICLR), 2022
2022
-
[40]
A Survey of Preference-Based Reinforcement Learning Methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz. A Survey of Preference-Based Reinforcement Learning Methods. Journal of Machine Learning Research, 18(136):1–46, 2017
2017
-
[41]
Consequences of Misaligned AI
Simon Zhuang and Dylan Hadfield-Menell. Consequences of Misaligned AI. In Advances in Neural Information Processing Systems 33 (NeurIPS), pages 15763–15773, 2020
2020
-
[42]
Bob isx times more satisfied than Alice
Åström, Karl Johan. Optimal Control of Markov Processes with Incomplete State Information I. Journal of Mathematical Analysis and Applications, 10:174–205, 1965. 11 A Appendix A.1 Utility and Social Choice Theory A.1.1 Measuring and comparing utilities Measurability of u ∈ Uca...
1965
-
[43]
V ∗(s) − V π ˆQ (s) ≤ 2ε 1 − γ + γk(Vmax − Vmin) with probability at least 1 − 2kδ, ∀k ∈ N∗,
-
[44]
∞X i=0 γir(si, ai) # − Eτ ∼pτ (·|π ˆQ,s0=s)
V ∗(s) − V π ˆQ (s) ≤ 2ε + 2δ(Vmax − Vmin) 1 − γ almost surely. Proof. First, note that if |Q∗(s, a) − ˆQ(s, a)| ≤ε with probability at least 1 − δ for all state-action pairs (s, a), then Q∗(s, π∗(s)) − Q∗(s, πˆQ(s)) ≤ 2ε with probability at least 1 − 2δ since Q∗(s, π∗(s)) ≤ ˆ...
-
[45]
P 1 n nX j=1 Xj − µ ≥ ε ≤ exp − 2ε2n (1 − n N )(1 + 1 n )(b − a)2 ,
-
[46]
approximation budget
P 1 n nX j=1 Xj − µ ≤ −ε ≤ exp ( − 2ε2n (1 − n N )(1 + 1 N −n )(b − a)2 ) . The proof of 1) is similar to the one proposed in [4]: Let Zn = 1 n Pn j=1 Xj − µ. We have for any λ >0: P [Zn ≥ ε] = P eλnZn ≥ eλnε ≤ E[eλnZn ] eλnε ≤ exp 1 8 (b − a)2λ2(n + 1) 1 − n N − λnε ,...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.