Pith. sign in

REVIEW 1 major objections 4 minor 46 references

Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies

T0 review · 1 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper proves that a computable near-optimal social-welfare policy exists whenever a predictive world model is accurate enough, and that any black-box policy can be made safe by blocking destructive actions.

desk verdict A worthwhile SMDP/PAA framework whose all-q existence theorem is undercut by a false concentration lemma for q<0 and mismatched Gamma constants for q=0,1. read the letter →

arxiv 2412.00033 v1 pith:N3BKQSBE submitted 2024-11-21 cs.AI cs.CY

classification cs.AIcs.CY MSC 91B1468Q32
keywords probablyapproximatelyalignedsocialwelfarepowermeanMarkovdecisionprocessAIalignmentsafepoliciessamplecomplexitychoice
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether an autonomous AI agent can be trusted to make social decisions and answers with a formal conditional yes. It defines social welfare in a Markov decision process as the power-mean aggregation of individual utilities, and calls a policy probably approximately aligned (PAA) if it achieves near-optimal expected discounted social welfare with high probability. The paper proves that a computable PAA policy exists whenever the predictive world model's worst-case KL divergence from the true dynamics is below an explicit threshold. The proof also specifies how many sampled utility reports, how many model calls, and what planning depth suffice. A separate result shows that any black-box policy can be wrapped into a safe policy that verifiably avoids destructive states.

What carries the argument

The carrying object is the sparse-sampling $Q$-estimator $\hat Q^h$ defined recursively in Eq. (6), where the reward is the power-mean welfare of a random subset of assessors and the next-state expectations are taken under the approximate model. The proof skeleton follows the classic sparse-sampling planner, but the approximation machinery is new: Lemma 5 gives a Hoeffding--Serfling concentration bound for the power mean, Lemma 6 bounds the reward error from approximating $p$ by $\hat p$, and Lemma 7 bounds the value loss incurred by acting greedily with respect to an approximate $Q$-function. The planner's parameters $K$, $C$, $n$, and $H$ are chosen so that all six error terms fit inside the approximation budget determined by $\epsilon$ and the model error.

What would settle it

Evaluate the claimed inequality used in Lemma 5 with $q=-1$ and $x=0.1$: it asserts $(1+x)^q \le 1-(1-2q)x$, i.e. $0.909 \le 0.7$, which is false; that single failure means the concentration bound for $q<0$ is not established, and the theorem's negative-$q$ branch collapses unless a different bound is found.

Watch

Extended reading notes

Core claim

The central claim is Theorem 2: for any social MDP with power-mean welfare $W_q$ and any tolerances $\epsilon>0$, $\delta\in[0,1)$, if there exists an approximate dynamics model $\hat p$ with $\sup_{(s,a)} D_{\mathrm{KL}}(p(\cdot|s,a)\,\|\,\hat p(\cdot|s,a)) < \epsilon^2(1-\gamma)^4 / (8\,\Delta W^2)$, then a computable $\delta$-$\epsilon$-PAA policy exists. The policy is the greedy action on recursively estimated $Q$-values from a sparse-sampling planner that simulates transitions with $\hat p$ and estimates rewards from a finite panel of sampled utilities. The proof decomposes the error into six terms and bounds each with concentration inequalities: Lemma 5 controls sampling error of the power mean, Lemma 6 controls model mismatch through KL divergence, and Lemma 7 converts approximate $Q$-values into a value-function loss. Consequently, near-optimal alignment can be certified from model accuracy and finite utility feedback without ever observing the agent's true objective.

Load-bearing premise

The existence theorem depends on a concentration bound for how well a random sample of citizens estimates the social-welfare average, and the proof of that bound for negative welfare exponents uses an algebraic inequality that does not hold, so the theorem's coverage of that range is not established.

Editorial extensions

If this is right

  • For any $\epsilon>0$ and $\delta\in[0,1)$, a computable $\delta$-$\epsilon$-PAA policy exists whenever the world model's worst-case KL error is below the stated threshold, with explicit sample sizes and horizon.
  • Because the bounds do not depend on the size of the state space, the existence result applies to infinite state spaces as long as the action space is finite.
  • Any black-box policy can be converted into a $\delta$-$\omega$-safe policy by action masking, at the cost of refusing actions whose estimated continuation value is too low.
  • The safe-policy result does not require the world model to meet the PAA accuracy threshold; lower model accuracy only shrinks the set of verifiably safe actions.
  • Alignment becomes an a priori, quantitative property: a society can verify the guarantee from the world model and a finite utility sample rather than from observed behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the theorem survives the $q<0$ gap, alignment audits could shift from inspecting an agent's objective to validating the predictive model, because the guarantee is driven entirely by world-model accuracy and sampled utilities.
  • The safe-policy wrapper suggests a practical deployment path: keep a black-box policy for its competence, compute $\hat Q^H$ from a learned simulator, and veto any action whose estimated continuation value falls below the safe threshold; this is testable in any simulator with known ground truth.
  • Because the reward class is tied to the power mean, the framework's applicability depends on which informational basis a society adopts for comparing utilities; changing that basis changes the concentration constants and hence the required sample sizes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper defines a Social Markov Decision Process whose reward is the expected future discounted power-mean social welfare W_q of the members of a society, and defines a policy to be δ-ε-PAA if, with probability at least 1−δ, its expected social welfare is within ε of the optimal policy. The central theoretical contribution is Theorem 2: under a uniform KL-divergence bound on an approximate world model, a computable δ-ε-PAA policy exists for every q∈R. The proof adapts the sparse-sampling planner of Kearns et al., bounding separately the error from sampling a subset of assessors, from the approximate dynamics model, and from Monte Carlo rollouts. The paper also introduces δ-ω-safe policies and proves in Theorem 8 that any black-box policy can be restricted to a safe action set while preserving a desired safety level.

Significance. If the technical gaps are repaired, the paper offers a useful formal bridge between social choice and MDP planning: alignment is quantified as ε-δ near-optimality in a constructed social MDP, and the sufficient condition on model accuracy is explicit and checkable in principle. The proof architecture is transparent and largely faithful to sparse sampling, and the claimed sample complexities are independent of the number of states. I found no circularity: PAA is defined independently as near-optimality in the constructed MDP, and the theorems are derived from standard concentration and planning results rather than by fitting the definitions to the conclusions. The main blocker is a false concentration lemma used for negative power-mean social welfare functions.

major comments (1)
  1. [Appendix A.2.1, Lemma 5 (Eqs. (8)–(9))] The q<0 branch of Lemma 5 is invalid. The proof asserts the inequality (1+x)^q ≤ 1−(1−2q)x for 0<x≤1 and q<0, which is false; for q=−1 and x=0.1 it gives 0.909 ≤ 0.7. The stated bound is not merely unproved but contradicted by a finite example: take q=−1, X={1,2}, N=2, n=1, ε=0.6. The population power mean is 4/3, the sample is 1 or 2 with equal probability, so P(|S−4/3|≥0.6)=0.5. With Γ(ε,a,b,q)=(1−2q)^2 b^{2q−2}/(a^q−b^q)^2=9/4, Eq. (8) gives 2 exp(−2·1·0.36·2.25/((1−1/2)(1+1/1)))=2e^{−1.62}≈0.396<0.5. Since Theorem 3 uses Lemma 5 to bound the term Z1 for every q∈R, and Theorem 2 claims existence of PAA policies for every q∈R, the existence theorem is unproved for negative-power-mean social welfare functions, including the harmonic and near-egalitarian cases the paper explicitly claims to cover. Theorem 8 inherits the same gap through Γmax in the definition of α. A valid substitute concentration bound for q<0, or a restriction of the theorems to q≥0, is needed.
minor comments (4)
  1. [Appendix A.2.1, Lemma 6] The lemma statement contains a dangling fragment: "such that D_KL(p∥p̂) ≤ d ∈ R and ." This should be completed or the stray "and ." removed.
  2. [Section 2.1.1 and Eq. (9)] The text says Umin=0 is allowed in specific cases, but the q=0 branch of Lemma 5 requires log a to be defined and hence requires Umin>0 for the geometric mean. Please clarify which values of q are admissible when Umin=0.
  3. [Theorem 3 proof, near Eq. (18)] The bookkeeping of the γ^k term is hard to follow because the displayed expression and the definition of β contain unbalanced parentheses and a conditional formatting artifact. Rewriting that line with explicit parentheses would improve readability.
  4. [Eq. (3)] The notation \hat E^K_{s'∼\hat p} is used in Eq. (3) before it is defined in the surrounding text; please define the empirical expectation operator before first use.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the PAA existence theorem is a constructive MDP planning result whose alignment objective is explicitly defined as social welfare; no fitted quantity is renamed as a prediction.

full rationale

The central chain (Section 3, Theorem 2 then Theorem 3) proves that if a world model p_hat satisfies the KL condition (Eq. 5), then the sparse-sampling policy of Eq. (7) is epsilon-delta optimal with respect to the expected future discounted social welfare W^pi. The alignment definition in Section 2.1.3 and the SMDP reward r_I in Eq. (2) use the same social-welfare functional W_q; this is an explicit modeling commitment, not an input secretly equal to the output. The parameters K, C, n, H are chosen in Theorem 3 to meet an epsilon-delta budget, not fitted to any data and then reported as predictions. The external ingredients (Kearns et al. sparse sampling, Hoeffding-Serfling concentration, social-choice representation theorems, Roberts/Cousins) are independent citations and are not self-citations by the authors. Hence no derivation step reduces by construction to its own input. Separately, and without changing the circularity verdict, the q<0 branch of Lemma 5 in Appendix A.2.1 appears mathematically unsupported: the proof's inequality (1+x)^q <= 1-(1-2q)x is false for q<0, so the all-q statement of Theorem 3 has a correctness risk; that is a proof flaw, not circularity.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No numbers are fitted to data: the algorithm parameters H, K, C, n are chosen from closed-form formulas, and the KL error d and power-mean exponent q are inputs, not fitted values. The paper introduces mathematical objects (SMDP, PAA, safe policy), not natural-kind entities with independent empirical handles, so the invented-entity ledger is empty.

assumptions (6)
  • standard math Debreu's representation theorem: complete, continuous, transitive preferences admit continuous utility functions.
    Invoked in Section 2.1.1 to justify the existence of utility functions representing individual preferences.
  • domain assumption The social welfare function is a power mean W_q, derived from axioms (U), (XI), (IIA), (WP), (A) and an informational basis.
    Section 2.1.2; this is a substantive normative choice about how to aggregate utilities, not a purely mathematical fact.
  • domain assumption Utilities are bounded, time-invariant, and interpersonally comparable at the level required for W_q.
    Section 2.1.3 and limitations; required for the PAA definition, for Lemma 5, and for the discounted welfare calculation.
  • domain assumption The approximate world model p_hat satisfies the stated sup-KL condition.
    Theorems 2 and 3; the condition cannot be verified exactly when the true model p is unknown, as the paper notes.
  • domain assumption Sampled assessors report truthful utilities and can evaluate complete social states.
    Section 2.2.3; the limitations section acknowledges that full-state observability and truthful reporting are assumed.
  • standard math Hoeffding-Serfling concentration inequalities and Pinsker/Bretagnolle-Huber inequalities hold as cited.
    Used in Lemma 5, Lemma 6, and the main proof in Appendix A.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies." pith.science (2026). https://pith.science/paper/N3BKQSBE

@misc{pith2026241200033,
  author       = {Pith},
  title        = {Pith review of: Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N3BKQSBE}},
  note         = {Machine review of arXiv:2412.00033}
}
read the original abstract

While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications such as social decision processes. More importantly, existing alignment methods provide no formal guarantees on the safety of such models. Drawing from utility and social choice theory, we provide a novel quantitative definition of alignment in the context of social decision-making. Building on this definition, we introduce probably approximately aligned (i.e., near-optimal) policies, and we derive a sufficient condition for their existence. Lastly, recognizing the practical difficulty of satisfying this condition, we introduce the relaxed concept of safe (i.e., nondestructive) policies, and we propose a simple yet robust method to safeguard the black-box policy of any autonomous agent, ensuring all its actions are verifiably safe for the society.

Figures

Figures reproduced from arXiv: 2412.00033 by the authors.

Figure 1
Figure 1. Democratic (left) vs autonomous (right) governments. Transparent elections must be [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 34 canonical work pages

  1. [1]

    Constrained Markov Decision Processes

    Eitan Altman. Constrained Markov Decision Processes. Routledge, 2021

  2. [2]

    Concrete Problems in AI Safety

    Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. Concrete Problems in AI Safety. arXiv preprint arXiv:1606.06565, 2016

  3. [3]

    Social Choice and Individual Values

    Kenneth Joseph Arrow. Social Choice and Individual Values. Wiley: New York, 1951

  4. [4]

    Concentration Inequalities for Sampling without Replace- ment

    Rémi Bardenet and Odalric-Ambrym Maillard. Concentration Inequalities for Sampling without Replace- ment. Bernoulli, 21(3):1361 – 1385, 2015

  5. [5]

    Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile Lukosuite, Amanda Askell, Andy Jones, Anna Chen, et al

    Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile Lukosuite, Amanda Askell, Andy Jones, Anna Chen, et al. Measuring Progress on Scalable Oversight for Large Language Models. arXiv preprint arXiv:2211.03540, 2022

  6. [6]

    Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei

    Paul F. Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems 30 (NIPS), pages 4299–4307, 2017

  7. [7]

    An Axiomatic Theory of Provably-Fair Welfare-Centric Machine Learning

    Cyrus Cousins. An Axiomatic Theory of Provably-Fair Welfare-Centric Machine Learning. In Advances in Neural Information Processing Systems 34 (NeurIPS), pages 16610–16621, 2021

  8. [8]

    Rethinking the Maturity of Artificial Intelligence in Safety-Critical Settings

    Mary Cummings. Rethinking the Maturity of Artificial Intelligence in Safety-Critical Settings . AI Magazine, 42(1):6–15, 2021

Show all 46 references
  1. [9]

    Equity and the Informational Basis of Collective Choice

    Claude D’Aspremont and Louis Gevers. Equity and the Informational Basis of Collective Choice. The Review of Economic Studies, 44(2):199–209, 1977

  2. [10]

    Representation of a preference ordering by a numerical func- tion

    Gerard Debreu and Werner Hildenbrand. Representation of a preference ordering by a numerical func- tion. In Mathematical Economics: Twenty Papers of Gerard Debreu, Econometric Society Monographs, chapter 6, page 105–110. Cambridge University Press, 1983

  3. [11]

    RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

    Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment. arXiv preprint arXiv:2304.06767, 2023

  4. [12]

    Feinberg

    Eugene A. Feinberg. Total Expected Discounted Reward MDPS: Existence of Optimal Policies. Wiley Encyclopedia of Operations Research and Management Science. John Wiley & Sons, Ltd, 2011

  5. [13]

    Artificial intelligence, values, and alignment

    Iason Gabriel. Artificial intelligence, values, and alignment. Minds and machines, 30(3):411–437, 2020

  6. [14]

    Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

    Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv preprint arXiv:2209.07858, 2022

  7. [15]

    A Review of Safe Reinforcement Learning: Methods, Theory and Applications

    Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll. A Review of Safe Reinforcement Learning: Methods, Theory and Applications. arXiv preprint arXiv:2205.10330, 2023

  8. [16]

    Hicks and Roy G

    John R. Hicks and Roy G. D. Allen. A Reconsideration of the Theory of Value. Part I. Economica, 1(1): 52–76, 1934

  9. [17]

    AI Alignment: A Comprehensive Survey

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhong- hao He, Jiayi Zhou, Zhaowei Zhang, et al. AI Alignment: A Comprehensive Survey. arXiv preprint arXiv:2310.19852, 2023

  10. [18]

    On the Sample Complexity of Reinforcement Learning

    Sham Machandranath Kakade. On the Sample Complexity of Reinforcement Learning. Doctoral dissertation, University College London (United Kingdom), 2003

  11. [19]

    A Sparse Sampling Algorithm for Near-optimal Planning in Large Markov Decision Processes

    Michael Kearns, Yishay Mansour, and Andrew Y Ng. A Sparse Sampling Algorithm for Near-optimal Planning in Large Markov Decision Processes. Machine learning, 49:193–208, 2002. 10

  12. [20]

    Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking

    Hanna Krasowski, Jakob Thumm, Marlon Müller, Lukas Schäfer, Xiao Wang, and Matthias Althoff. Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking. arXiv preprint arXiv:2205.06750, 2023

  13. [21]

    Social Choice with Changing Preferences: Representation Theorems and Long-run Policies

    Kshitij Kulkarni and Sven Neth. Social Choice with Changing Preferences: Representation Theorems and Long-run Policies. arXiv preprint arXiv:2011.02544, 2020

  14. [22]

    AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

    Abhilash Mishra. AI Alignment and Social Choice: Fundamental Limitations and Policy Implications. arXiv preprint arXiv:2310.16048, 2023

  15. [23]

    Value-Guided Synthesis of Parametric Normative Systems

    Nieves Montes and Carles Sierra. Value-Guided Synthesis of Parametric Normative Systems. In Proceed- ings of the 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), page 907–915, 2021

  16. [24]

    The Alignment Problem from a Deep Learning Perspective

    Richard Ngo, Lawrence Chan, and Sören Mindermann. The Alignment Problem from a Deep Learning Perspective. arXiv preprint arXiv:2209.00626, 2022

  17. [25]

    Accountability in Artificial Intelligence: What it is and how it works

    Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi. Accountability in Artificial Intelligence: What it is and how it works. AI & SOCIETY, pages 1–12, 2023

  18. [26]

    Christiano, Jan Leike, and Ryan Lowe

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, ...

  19. [27]

    Manuale di Economia Politica con una Introduzione alla Scienza Sociale

    Vilfredo Pareto. Manuale di Economia Politica con una Introduzione alla Scienza Sociale . Piccola biblioteca scientifica. Societa editrice libraria, 1906

  20. [28]

    Dynamic Social Choice with Evolving Preferences

    David Parkes and Ariel Procaccia. Dynamic Social Choice with Evolving Preferences. In Proceedings of the 27th AAAI conference on artificial intelligence, pages 767–773, 2013

  21. [29]

    Kevin W. S. Roberts. Interpersonal Comparability and Social Choice Theory. The Review of Economic Studies, 47(2):421–439, 1980

  22. [30]

    Research Priorities for Robust and Beneficial Artificial Intelligence

    Stuart Russell, Daniel Dewey, and Max Tegmark. Research Priorities for Robust and Beneficial Artificial Intelligence. AI magazine, 36(4):105–114, 2015

  23. [31]

    Schwartz

    Shalom H. Schwartz. Universals in the Content and Structure of Values: Theoretical Advances and Empirical Tests in 20 Countries. In Mark P. Zanna, editor, Advances in Experimental Social Psychology, volume 25, pages 1–65. Academic Press, 1992

  24. [32]

    On Weights and Measures: Informational Constraints in Social Welfare Analysis

    Amartya Sen. On Weights and Measures: Informational Constraints in Social Welfare Analysis. Econo- metrica, 45(7):1539–1572, 1977

  25. [33]

    R. J. Serfling. Probability Inequalities for the Sum in Sampling without Replacement. The Annals of Statistics, 2(1):39 – 48, 1974

  26. [34]

    Value Alignment: A Formal Approach

    Carles Sierra, Nardine Osman, Pablo Noriega, Jordi Sabater-Mir, and Antoni Perelló. Value Alignment: A Formal Approach. arXiv preprint arXiv:2110.09240, 2021

  27. [35]

    Singh and Richard C

    Satinder P. Singh and Richard C. Yee. An Upper Bound on the Loss from Approximate Optimal-Value Functions. Machine Learning, 16(3):227–233, 1994

  28. [36]

    Robert H. Strotz. Cardinal Utility. The American Economic Review, 43(2):384–397, 1953

  29. [37]

    Sutton and Andrew G

    Richard S. Sutton and Andrew G. Barto. Reinforcement learning: An introduction. MIT press, 2018

  30. [38]

    A Theory of the Learnable

    Leslie Gabriel Valiant. A Theory of the Learnable. Communications of the ACM, 27(11):1134–1142, 1984

  31. [39]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y . Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V . Le. Finetuned Language Models are Zero-Shot Learners. In Proceedings of the 10th International Conference on Learning Representations (ICLR), 2022

  32. [40]

    A Survey of Preference-Based Reinforcement Learning Methods

    Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz. A Survey of Preference-Based Reinforcement Learning Methods. Journal of Machine Learning Research, 18(136):1–46, 2017

  33. [41]

    Consequences of Misaligned AI

    Simon Zhuang and Dylan Hadfield-Menell. Consequences of Misaligned AI. In Advances in Neural Information Processing Systems 33 (NeurIPS), pages 15763–15773, 2020

  34. [42]

    Bob isx times more satisfied than Alice

    Åström, Karl Johan. Optimal Control of Markov Processes with Incomplete State Information I. Journal of Mathematical Analysis and Applications, 10:174–205, 1965. 11 A Appendix A.1 Utility and Social Choice Theory A.1.1 Measuring and comparing utilities Measurability of u ∈ Uca...

  35. [43]

    V ∗(s) − V π ˆQ (s) ≤ 2ε 1 − γ + γk(Vmax − Vmin) with probability at least 1 − 2kδ, ∀k ∈ N∗,

  36. [44]

    ∞X i=0 γir(si, ai) # − Eτ ∼pτ (·|π ˆQ,s0=s)

    V ∗(s) − V π ˆQ (s) ≤ 2ε + 2δ(Vmax − Vmin) 1 − γ almost surely. Proof. First, note that if |Q∗(s, a) − ˆQ(s, a)| ≤ε with probability at least 1 − δ for all state-action pairs (s, a), then Q∗(s, π∗(s)) − Q∗(s, πˆQ(s)) ≤ 2ε with probability at least 1 − 2δ since Q∗(s, π∗(s)) ≤ ˆ...

  37. [45]

    P   1 n nX j=1 Xj − µ ≥ ε   ≤ exp − 2ε2n (1 − n N )(1 + 1 n )(b − a)2 ,

  38. [46]

    approximation budget

    P   1 n nX j=1 Xj − µ ≤ −ε   ≤ exp ( − 2ε2n (1 − n N )(1 + 1 N −n )(b − a)2 ) . The proof of 1) is similar to the one proposed in [4]: Let Zn = 1 n Pn j=1 Xj − µ. We have for any λ >0: P [Zn ≥ ε] = P eλnZn ≥ eλnε ≤ E[eλnZn ] eλnε ≤ exp 1 8 (b − a)2λ2(n + 1) 1 − n N − λnε ,...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.