REVIEW 3 major objections 4 minor 51 references
A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A set of fairness axioms forces the form of an AI objective that measures human power, and an AI that softly maximizes that objective will ask for confirmation, make commitments, and follow social norms.
desk verdict A solid axiomatic derivation of a human-empowerment objective whose behavioral claims depend on an underspecified goal-set input; worth reviewing, with conditions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing structure is the four-level hierarchy of metrics listed in Table 1: $C$ (goal-attainment capability), $I$ (individual power), $P$ (present aggregate power), and $L$ (long-term aggregate power), with the robot policy $\pi_r(s)(a_r) \propto E_{s'\sim s,a_r,\pi_H}[2^{-L(s')}]^{-\beta_r}$. Each level is pinned down by a representation theorem: separability axioms combined with Pigou–Dalton-style inequality aversion, stationarity/impatience axioms, and translation-invariance arguments force the logarithmic-exponential functional forms and restrict parameter ranges ($\zeta>1$, $\xi, \eta, \rho>0$, $\gamma_h, \gamma_r\in(0,1)$). The human behavior model $\pi_H(s,g)$ and the goal set $G_h$—required to cover every state (G1)—enter only structurally, so the robot never needs to know a human's actual goal. The 'power' metric is thus a function of the effective number of goals humans can bring about, measured in bits.
What would settle it
Implement the two-armed confirmation game from Proposition 8 with a simulated human whose error rate $\epsilon$ is known and the robot's discount factor $\gamma_h$ set to 0.99; the paper predicts the robot asks for confirmation exactly twice ($k^* = 2$). If across a sweep of $\epsilon$ and $\gamma_h$ the $L$-maximizing robot's number of confirmation requests deviates from the formula $k^* \approx \ln(1-\gamma_h)/\ln \epsilon$, the derived behavioral consequences are refuted. A second, more direct test: construct any finite acyclic game form satisfying axioms C1–C6, I1–I7, P0–P8, T1–T8, L1–L5 whose optimal policy cannot be expressed through Table 1 metrics; that would falsify the representation theorems.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that the long-term aggregate human power metric $L(s) = -\log_2 E_{s_{\ge t}\sim s_t,\pi}\left[\left(\sum_{u\ge t} \gamma_r^{u-t} 2^{-\eta P(s_u)}\right)^\rho\right]$ is not an arbitrary design choice but the unique form enforced by the stated desiderata C1–C6, I1–I7, P0–P8, T1–T8, L1–L5. The same axioms fix the individual-level pieces: goal-attainment capability $C$ is a truncated Bellman equation (a discounted probability), individual power is $I_h = \log_2 \sum_{g\in G_h} C^{\zeta}$ with $\zeta>1$, and present aggregate power is $P = -\log_2 \sum_h 2^{-\xi I_h}$ with $\xi>0$. The paper further claims that a robot using the Boltzmann policy over $L$ (with soft maximization degree $\beta_r$) will, in stylized models, choose non-overwhelming menus, ask for confirmation a finite number of times, make transparent commitments, allow pausing but disable destruction, and follow social norms. These behaviors are presented not as added constraints but as emergent from a single objective.
Load-bearing premise
The entire objective is conditional on the AI having a correct and structurally complete world model—including the set of possible human goals covering every state (axiom G1) and a faithful goal-conditioned model of human behavior—so if either is inaccurate, every 'power' metric is measuring a fiction; the paper itself concedes this risk as 'wishful thinking' in its conclusion.
Editorial extensions
If this is right
- A robot following the derived objective will present humans with a menu of options that is large but not overwhelming; the optimal menu size is approximately $(e^{\beta_h}-1)/(\zeta-1)$ for a Boltzmann-rational human.
- It will ask for confirmation before irreversible or mistake-prone commands—a finite, computable number of times that grows with the human's discount factor $\gamma_h$—and will eventually obey.
- It will voluntarily make binding commitments that restrict its future behavior because predictable robots increase human goal-attainment capability.
- It will tend to follow social norms and to allocate scarce resources equally (unless power is very convex in resources), because doing so raises the inequality-averse aggregate power $P$.
- It will allow itself to be paused when humans can cope without it, but will disable a destroy button when destruction would permanently end its empowering assistance.
Reading between the lines
- The paper's goal-agnostic premise invites a concrete halfway design it only sketches: replace the flat sum over possible goals with a slowly changing Bayesian prior, and the same formulas interpolate between a pure empowering agent and a cooperative inverse-RL assistant.
- The confirmation-count formula from the paper's toy model translates directly into a human-factor design rule: how many times a system should double-check a destructive or irreversible action can be computed from the user's measured error rate and discount factor, before any training.
- A telling boundary case the paper leaves open is that the robot's own power never enters $P$, so an agent that accumulates capabilities while boosting human power could concentrate power over the long run; the authors list this as a possibly undesirable effect, and its resolution would require a second fairness axiom over the human–AI distribution.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops an axiomatic foundation for an AI objective that preserves and fairly distributes human power. In a finite acyclic stochastic game form with a given set of possible goals G_h and a goal-conditioned human policy pi_H, the authors derive representation theorems (Propositions 1-6) pinning the functional forms of goal-attainment capability C, individual power I_h, present aggregate power P, and long-term aggregate power L to a small set of normative parameters (gamma_h, gamma_r, zeta, xi, eta, rho; Table 1). A Boltzmann robot policy over L is then proposed (Eq. 8), and Section 3 analyzes toy models claiming that softly maximizing L yields a finite optimal menu size (Prop. 7), confirmation-seeking before irreversible action (Prop. 8), commitment-making (Prop. 9), and selective enabling of pause/destroy buttons (Prop. 10), followed by more speculative emergent behaviors (norms, fair allocation, manipulation of expectations). The paper positions this objective against channel-capacity empowerment and extrinsic-reward maximization.
Significance. If the derivations hold, this is a substantial theory contribution: it delivers a decomposable, parameter-transparent objective for human-empowerment-preserving AI, with proofs in the appendix, explicit normative parameters with interpretable behavioral trade-offs, a formal comparison with Klyubin's empowerment metric (Prop. 3), and concrete falsifiable predictions (menu size, confirmation count) that could be tested in simulation. The authors deserve credit for shipping the full proof apparatus and for stating limitations candidly (Section 4 on 'wishful thinking'; Appendix E.2). The central qualification is that the Section 3 behavioral predictions are consequences of the axioms plus the non-canonical goal set G_h, not of the axioms alone: the axioms fix functional forms but not the granularity of G_h, and changing granularity changes or annuls predictions such as Prop. 7's optimal menu size. The metric-design contribution stands independently, but the behavioral claims need either a canonical G_h construction or an explicit scope restriction.
major comments (3)
- [Section 3.1, Prop. 7; Appendix E.2] The optimal-menu-size prediction is not invariant under the choice of G_h, and axiom G1 (coverage) does not fix granularity. Under the stated assumption G_h = {{s'}: s' in S_top}, Prop. 7 yields k* approx (e^{beta_h}-1)/(zeta-1); however, if G_h is the single covering goal {S_top}, then C identically equals 1 and I_h identically equals 0 for every menu size, so the robot is indifferent among all k and the predicted finite optimum disappears; if terminal states are instead paired into goals, the same dynamics give a different I_h(s_k) and a different optimal k*. Because I_h(s) = log2 sum_g C_g^zeta is not invariant under refining or merging goals with positive attainment capability, and because Appendix E.2's canonical construction G_h = {{s}: s in S} is justified only for tree-shaped transition graphs, the claim that 'r will likely choose k approx ...' (and the related 'will not overwhelm humans with too many options' conclusions in Sections 3.2 and 4) is conditional on an arbitrary modeling choice. Section 4's own warning about 'convenient but inaccurate world models and goal sets' causing 'wishful thinking' applies directly here. Please extend the canonical goal-set construction to general state spaces or restrict the menu-size claim to the canonical setting and add a sensitivity analysis identifying which Section 3 predictions survive changes of goal granularity.
- [Section 2 vs. Prop. 8] Proposition 8 uses a game outside the stated framework. Section 2 assumes the game form is 'finite and acyclic,' but the confirmation game in Prop. 8 returns to state s_k whenever the final confirmation fails, and the proof sums an infinite geometric series over repeated rounds, i.e., it is a discounted infinite-horizon game. Relatedly, the self-referential definition of the robot policy in Eq. (8) (pi_r depends on L, which depends on the trajectory distribution under pi_r) is benign in the finite-acyclic case because it can be resolved by backward induction, but the paper never states this, and the Prop. 8 setting requires an existence/uniqueness argument for the fixed point of (8) in cyclic discounted games. Please either extend the framework to discounted infinite-horizon games with a well-definedness statement for (8), or rework Prop. 8 to fit the acyclic framework (e.g., a bounded-horizon approximation).
- [Section 3.2; Section 4] Several 'likely behavioral consequences' are already encoded in the axioms or parameter choices and should be labeled as such. The equal-split resource allocation claim in Section 3.2 follows directly from the Pigou-Dalton axiom (P6) together with the functional form of P in Prop. 4, and the choices eta > 0 and rho > 0 are explicitly selected in Section 2.2 to incentivize intertemporal equality and uncertainty reduction. Presenting these in Section 3 as consequences of 'softly maximizing L' overstates what is emergent; the Section 4 remark distinguishing properties 'directly baked into the metric' from 'emergent' ones is appropriate and should be applied consistently in Section 3, so that the genuinely emergent claims (commitment-making, confirmation-seeking, norm-following) can be identified and tested separately.
minor comments (4)
- [Appendix A.3, proof of Prop. 7] The displayed formula for C writes the zeta-th power into the definition of C; the subsequent I_h formula is correct, but the display should read C = e^{beta_h}/(e^{beta_h}+k-1) with I_h = log2[k C^zeta].
- [Prop. 3, proof paragraph] The claim that I_h and E_zeta 'share their range [-log2 k, log2 k]' is only correct for zeta = 2; the common range is [-(zeta-1) log2 k, log2 k], attained at fully deterministic and fully uniform probability matrices.
- [Figure 2 caption] The parameter values are printed as '= 1, = 0.99' with the symbols (presumably eta and gamma) missing; please regenerate the caption with the symbols labeled.
- [Section 3.1] The qualifier 'will likely choose' in Props. 7 and 8 should state the exact optimality criterion (beta_r -> infinity, tie-breaking, and the chosen objective Q_r(s, a_r)), since the appendix already contains the formal versions.
Circularity Check
No significant circularity; the power-metric forms follow from explicit axioms, and behavioral examples are conditional on declared inputs rather than fitted predictions.
full rationale
The central derivation (Section 2) is not circular: C, I_h, P, T, and L are obtained from the stated axioms (C1-C6, I1-I7, P0-P8, T1-T8, L1-L5) via classical representation theorems (Debreu, Dasgupta et al., Koopmans, Pfanzagl, Roberts), with no fitted parameter renamed as a prediction. The goal set G_h and behavior model pi_H are explicitly treated as given inputs ('We are only assuming the AI has (i) a model of how humans behave conditional on their goals, and (ii) a structural understanding of possibly dynamics, interactions, and transition probabilities'), and the paper concedes that 'convenient but inaccurate world models and goal sets' can produce 'wishful thinking' (Section 4). The behavioral examples in Section 3 are conditional consequences of these inputs and of the chosen parameter cases (e.g., eta>0, rho>0), not hidden re-uses of the conclusions; the paper itself distinguishes 'some directly baked into the metric, others emergent.' The two self-citations (Heitzig 2026, Potham and Harms 2025) are peripheral and not load-bearing. Any concern about goal-individuation sensitivity (Appendix E.2) is an assumption-robustness limitation, not a circular derivation.
Assumptions & free parameters
free parameters (7)
- gamma_h (human goal discounting) =
(0,1), not fixed
- gamma_r (robot discounting) =
(0,1), not fixed
- zeta (reliability preference) =
>1; set to 2 in Appendix C
- xi (inequality aversion) =
>0, not fixed
- eta (intertemporal inequality aversion) =
>0, chosen to incentivize reduced intertemporal inequality
- rho (uncertainty aversion) =
>0, chosen; rho=1 recommended for recursivity
- beta_r (soft-max degree of maximization) =
>0, not fixed
assumptions (7)
- standard math Goal-attainment representation axioms (C1-C6)
- domain assumption World-model inputs: complete goal sets G_h and goal-conditioned human policy pi_H
- domain assumption Individual power axioms (I1-I7)
- domain assumption Population axioms (P0-P8)
- domain assumption Temporal axioms (T1-T8)
- domain assumption Uncertainty axioms (L1-L5)
- domain assumption Finite acyclic game form
Cite this review
Pith. "Pith review of A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences." pith.science (2026). https://pith.science/paper/KGA72MP5
@misc{pith2026260808240,
author = {Pith},
title = {Pith review of: A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences},
year = {2026},
howpublished = {\url{https://pith.science/paper/KGA72MP5}},
note = {Machine review of arXiv:2608.08240}
}
read the original abstract
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.
Figures
Reference graph
Works this paper leans on
-
[1]
Bayesian theory of mind: Modeling joint belief-desire attribution
[Bakeret al., 2011 ] Chris Baker, Rebecca Saxe, and Joshua Tenenbaum. Bayesian theory of mind: Modeling joint belief-desire attribution. InProceedings of the annual meeting of the cognitive science society, volume 33,
work page 2011
-
[2]
α, C=−x α/ϵ. Then solving the system of linear equations for the Markov chains ofΓ 2 states resulting from the three possible robot policies gives W0 =A 0/ϵ,(26) W1 = δA1 +γpB 1 ϵ ,(27) W2 = δA2 +γpB 2 +γqC 1−γ(1−q) (28) If bothx, y≫1andd=y−x, we have approximately W1 > W0 ⇔pγd <1,(29) W2 > W0 ⇔qγd < ϵ(2−γpd),(30) W2 > W1 ⇔qγ(1 +dδ)< ϵ.(31) Hence, enablin...
work page 2024
-
[3]
[Bengioet al., 2025 ] Yoshua Bengio, Michael Cohen, Dami- ano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, S ¨oren Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, et al. Superintelligent agents pose catastrophic risks: Can Scientist AI offer a safer path?arXiv preprint arXiv:2502.15657,
arXiv 2025
-
[9]
Ai deskilling is a struc- tural problem.AI & SOCIETY, pages 1–13,
[Ferdman, 2025] Avigail Ferdman. Ai deskilling is a struc- tural problem.AI & SOCIETY, pages 1–13,
work page 2025
-
[10]
The effect of model- ing human rationality level on learning rewards from mul- tiple feedback types
[Ghosalet al., 2023 ] Gaurav R Ghosal, Matthew Zurek, Daniel S Brown, and Anca D Dragan. The effect of model- ing human rationality level on learning rewards from mul- tiple feedback types. InProceedings of the AAAI Con- ference on Artificial Intelligence, volume 37, pages 5983– 5992,
work page 2023
-
[14]
[Hill Jr, 2002] Thomas E Hill Jr.Human welfare and moral worth: Kantian perspectives. Clarendon Press,
work page 2002
-
[18]
Position: Humanity faces existential risk from grad- ual disempowerment
[Kulveitet al., 2025 ] Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger, and David Duve- naud. Position: Humanity faces existential risk from grad- ual disempowerment. InForty-second International Con- ference on Machine Learning Position Paper Track,
work page 2025
-
[19]
A path towards autonomous machine intelligence version 0.9
[LeCun, 2022] Yann LeCun. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27.Open Re- view, 62(1):1–62,
work page 2022
Show all 51 references
-
[22]
Bakker, Fazl Barez, Matija Franklin, Andreas Haupt, Jobst Heitzig, Wesley H
[Loweet al., 2025 ] Ryan Lowe, Joe Edelman, Tan Zhi- Xuan, Oliver Klingefjord, Ellie Hain, Vincent Wang, Atrisha Sarkar, Michiel A. Bakker, Fazl Barez, Matija Franklin, Andreas Haupt, Jobst Heitzig, Wesley H. Hol- liday, Julian Jara-Ettinger, Atoosa Kasirzadeh, Ryan Oth- niel ...
2025
-
[23]
Bloomsbury Publishing,
[Lukes, 2026] Steven Lukes.Power: A radical view. Bloomsbury Publishing,
2026
-
[24]
Performative reinforcement learning
[Mandalet al., 2023 ] Debmalya Mandal, Stelios Triantafyl- lou, and Goran Radanovic. Performative reinforcement learning. InInternational Conference on Machine Learn- ing, pages 23642–23680. PMLR,
2023
-
[25]
Learning to assist humans without inferring rewards.Advances in Neural Information Processing Systems, 37:71540–71567,
[Myerset al., 2024 ] Vivek Myers, Evan Ellis, Sergey Levine, Benjamin Eysenbach, and Anca Dragan. Learning to assist humans without inferring rewards.Advances in Neural Information Processing Systems, 37:71540–71567,
2024
-
[26]
Yale University Press,
[Nordhaus, 2008] William Nordhaus.A question of balance: Weighing the options on global warming policies. Yale University Press,
2008
-
[31]
Corrigibility as a singular target: A vision for in- herently reliable foundation models.arXiv preprint arXiv:2506.03056,
[Potham and Harms, 2025] Ram Potham and Max Harms. Corrigibility as a singular target: A vision for in- herently reliable foundation models.arXiv preprint arXiv:2506.03056,
2025 arXiv
-
[32]
A theory of anticipated utility.Journal of economic behavior & organization, 3(4):323–343,
[Quiggin, 1982] John Quiggin. A theory of anticipated utility.Journal of economic behavior & organization, 3(4):323–343,
1982
-
[35]
Possibility theorems with interpersonally comparable welfare levels.The Re- view of Economic Studies, 47(2):409–420,
[Roberts, 1980] Kevin WS Roberts. Possibility theorems with interpersonally comparable welfare levels.The Re- view of Economic Studies, 47(2):409–420,
1980
-
[38]
Empowerment as replacement for the three laws of robotics.Frontiers in Robotics and AI, 4:260425,
[Salge and Polani, 2017] Christoph Salge and Daniel Polani. Empowerment as replacement for the three laws of robotics.Frontiers in Robotics and AI, 4:260425,
2017
-
[41]
Conservative agency via attainable utility preservation
[Turneret al., 2020 ] Alexander Matt Turner, Dylan Hadfield-Menell, and Prasad Tadepalli. Conservative agency via attainable utility preservation. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 385–391,
2020
-
[42]
Moral deskilling and up- skilling in a new machine age: Reflections on the am- biguous future of character.Philosophy & Technology, 28(1):107–124,
[Vallor, 2015] Shannon Vallor. Moral deskilling and up- skilling in a new machine age: Reflections on the am- biguous future of character.Philosophy & Technology, 28(1):107–124,
2015
-
[44]
Consequences of misaligned AI.Advances in Neural Information Processing Systems, 33:15763–15773,
[Zhuang and Hadfield-Menell, 2020] Simon Zhuang and Dylan Hadfield-Menell. Consequences of misaligned AI.Advances in Neural Information Processing Systems, 33:15763–15773,
2020
-
[45]
This is the Pigou–Dalton axiom in its traditional form
+ ˜f(y ′ 1). This is the Pigou–Dalton axiom in its traditional form. As this must hold for ally∈ [−1,0], another classical result by Dasguptaet al.[1973] implies that ˜fmust be strictly concave, hencefis strictly convex as claimed. (I6) now implies: If P i f(x i)< P i f(x ′ i)...
1973
-
[46]
non- expected
Strict convexity then impliesζ >1as claimed. Now assume (I7). From the previous proposition and the in- dependence ofΓ 1 andΓ 2, we know thatC g1 h×g2 h h =C g1 h h C g2 h h , hence F X g1 h X g2 h f(C g1 h h C g2 h h ) (14) =F X g1 h f(C g1 h h ) +F X g2 h f(...
1982
-
[47]
Also (P5) can be treated very similarly, only using ˜f(y) = f(log 2 y)instead off
Let us abbreviate these commonf I,F I by justf,F. Also (P5) can be treated very similarly, only using ˜f(y) = f(log 2 y)instead off. This is becauseI h is a logarithmic quantity, so then allC gh h get multiplied byγ∈(0,1),I h gets addedδ:=ζlog 2 γ <0. In other words, (P5) impl...
1973
-
[48]
k,I h0 (s′) = 0, andI h′(s′) =I h′(s)for all h /∈ {h0, . . . , hk}. Then−2 −P(s) =−(k+ 1)2 −ξ×1 +c≥ −2ξ2−ξ +c=−1+cand−2 −P(s ′) <−2 −ξ×0 +c=−1+c, wherec= P h′ 2−ξIh′ (s) represents the remaining humans. HenceP(s)> P(s ′)as claimed. Aggregation along time: trajectory-specific a...
1959
-
[49]
not committed
Let us abbreviate these commonf I i ,F I by justf i,F. (T5) implies that there are transformationsψ n withPn i=1 fi+1(xi) =ψ n(Pn i=1 fi(xi))for allnand all x1, . . . , xn. A classical results from decision theory shows that then there is a real valueγ r and constantsc i so th...
1960
-
[51]
intrinsic reward
E Alternative metrics E.1 Alternative aggregation order Instead of aggregatingI h →P→T, one could instead first aggregateI h intoindividual lifetime powerL h and then into aggregate long-term human powerL ′ via Lh(st) =−log 2 E s≥t∼st,π P u≥t γu−t rh 2−ηIh(su) ρ ,(32) L′(s) =−...
1928
-
[1928]
First contact: Unsupervised human- machine co-adaptation via mutual information maximiza- tion.Advances in Neural Information Processing Systems, 35:31542–31556,
[Reddyet al., 2022 ] Siddharth Reddy, Sergey Levine, and Anca Dragan. First contact: Unsupervised human- machine co-adaptation via mutual information maximiza- tion.Advances in Neural Information Processing Systems, 35:31542–31556,
2022
-
[1959]
A survey of world models for autonomous driving.arXiv preprint arXiv:2501.11260,
[Fenget al., 2025 ] Tuo Feng, Wenguan Wang, and Yi Yang. A survey of world models for autonomous driving.arXiv preprint arXiv:2501.11260,
2025 arXiv
-
[1960]
Measuring and avoid- ing side effects using relative reachability.arXiv preprint arXiv:1806.01186,
[Krakovnaet al., 2018 ] Victoria Krakovna, Laurent Orseau, Miljan Martic, and Shane Legg. Measuring and avoid- ing side effects using relative reachability.arXiv preprint arXiv:1806.01186,
2018 arXiv
-
[1973]
Topological methods in cardinal utility theory.Mathematical Methods in the Social Sci- ences, page 16,
[Debreu, 1959] G Debreu. Topological methods in cardinal utility theory.Mathematical Methods in the Social Sci- ences, page 16,
1959
-
[1980]
The capability approach in practice.Journal of political philosophy, 14(3),
[Robeyns, 2006] Ingrid Robeyns. The capability approach in practice.Journal of political philosophy, 14(3),
2006
-
[1982]
A mathematical theory of saving.The economic journal, 38(152):543–559,
[Ramsey, 1928] Frank Plumpton Ramsey. A mathematical theory of saving.The economic journal, 38(152):543–559,
1928
-
[1996]
Performative pre- diction
[Perdomoet al., 2020 ] Juan Perdomo, Tijana Zrnic, Celes- tine Mendler-D¨unner, and Moritz Hardt. Performative pre- diction. InInternational conference on machine learning, pages 7599–7609. PMLR,
2020
-
[2002]
Empowerment: A universal agent-centric measure of control
[Klyubinet al., 2005 ] Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv. Empowerment: A universal agent-centric measure of control. In2005 ieee congress on evolutionary computation, volume 1, pages 128–135. IEEE,
2005
-
[2005]
Stationary ordi- nal utility and impatience.Econometrica: Journal of the Econometric Society, pages 287–309,
[Koopmans, 1960] Tjalling C Koopmans. Stationary ordi- nal utility and impatience.Econometrica: Journal of the Econometric Society, pages 287–309,
1960
-
[2006]
Classification of mental workload using brain connectivity and machine learning on electroencephalogram data.Scientific Reports, 14(1):9153,
[Safariet al., 2024 ] MohammadReza Safari, Reza Shalbaf, Sara Bagherzadeh, and Ahmad Shalbaf. Classification of mental workload using brain connectivity and machine learning on electroencephalogram data.Scientific Reports, 14(1):9153,
2024
-
[2008]
Aristotelian social democracy
[Nussbaum, 2019] Martha Nussbaum. Aristotelian social democracy. InLiberalism and the Good, pages 203–252. Routledge,
2019
-
[2011]
Public Affairs,
[Banerjee and Duflo, 2011] Abhijit V Banerjee and Esther Duflo.Poor economics: A radical rethinking of the way to fight global poverty. Public Affairs,
2011
-
[2014]
The concordia contest: Advanc- ing the cooperative intelligence of language agents
[Smithet al., 2024 ] Chandler Smith, Rakshit Trivedi, Jesse Clifton, Lewis Hammond, Akbir Khan, Sasha Vezhnevets, John P Agapiou, Edgar A Du ´e˜nez-Guzm´an, Jayd Matyas, Danny Karmon, et al. The concordia contest: Advanc- ing the cooperative intelligence of language agents. In...
2024
-
[2015]
Is sora a world simulator? a comprehensive survey on general world mod- els and beyond.arXiv preprint arXiv:2405.03520,
[Zhuet al., 2024 ] Zheng Zhu, Xiaofeng Wang, Wangbo Zhao, Chen Min, Nianchen Deng, Min Dou, Yuqi Wang, Botian Shi, Kai Wang, Chi Zhang, et al. Is sora a world simulator? a comprehensive survey on general world mod- els and beyond.arXiv preprint arXiv:2405.03520,
2024
-
[2016]
Unbiased canonical set- valued oracles via lattice theory.arXiv preprint arXiv:2606.26418,
[Heitzig, 2026] Jobst Heitzig. Unbiased canonical set- valued oracles via lattice theory.arXiv preprint arXiv:2606.26418,
2026 arXiv
-
[2017]
Development as freedom (1999)
[Sen, 2014] Amartya Sen. Development as freedom (1999). The globalization and development reader: Perspectives on development and global change, 525,
1999
-
[2018]
Cooper- ative inverse reinforcement learning.Advances in neural information processing systems, 29,
[Hadfield-Menellet al., 2016 ] Dylan Hadfield-Menell, Stu- art J Russell, Pieter Abbeel, and Anca Dragan. Cooper- ative inverse reinforcement learning.Advances in neural information processing systems, 29,
2016
-
[2019]
Individual rights and social evalua- tion: a conceptual framework.Oxford Economic Papers, 48(2):194–212,
[Pattanaik and Suzumura, 1996] Prasanta K Pattanaik and Kotaro Suzumura. Individual rights and social evalua- tion: a conceptual framework.Oxford Economic Papers, 48(2):194–212,
1996
-
[2020]
A general theory of mea- surement applications to utility.Naval research logistics quarterly, 6(4):283–294,
[Pfanzagl, 1959] Johann Pfanzagl. A general theory of mea- surement applications to utility.Naval research logistics quarterly, 6(4):283–294,
1959
-
[2021]
Notes on the measurement of inequality
[Dasguptaet al., 1973 ] Partha Dasgupta, Amartya Sen, and David Starrett. Notes on the measurement of inequality. Journal of economic theory, 6(2):180–187,
1973
-
[2022]
A theory of appropriateness with applications to generative artificial intelligence.arXiv preprint arXiv:2412.19010,
[Leiboet al., 2024 ] Joel Z Leibo, Alexander Sasha Vezhn- evets, Manfred Diaz, John P Agapiou, William A Cun- ningham, Peter Sunehag, Julia Haas, Raphael Koster, Edgar A Du´e˜nez-Guzm´an, William S Isaac, et al. A theory of appropriateness with applications to generative artif...
2024 arXiv
-
[2023]
World models.arXiv preprint arXiv:1803.10122,
[Ha and Schmidhuber, 2018] David Ha and J ¨urgen Schmid- huber. World models.arXiv preprint arXiv:1803.10122,
2018 arXiv
-
[2024]
Beneficent intelligence: a capability approach to modeling benefit, assistance, and associated moral failures through ai systems.Minds and Machines, 34(4):41,
[London and Heidari, 2024] Alex John London and Hoda Heidari. Beneficent intelligence: a capability approach to modeling benefit, assistance, and associated moral failures through ai systems.Minds and Machines, 34(4):41,
2024
-
[2025]
Safety from honesty in a disinter- ested ai predictor.arXiv preprint arXiv:2606.29657,
[Bengioet al., 2026 ] Yoshua Bengio, Oliver Richardson, Tom´aˇs Gavenˇciak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, et al. Safety from honesty in a disinter- ested ai predictor.arXiv preprint arXiv:2606.29657,
2026 arXiv
-
[2026]
Identifiability in inverse reinforcement learn- ing.Advances in Neural Information Processing Systems, 34:12362–12373,
[Caoet al., 2021 ] Haoyang Cao, Samuel Cohen, and Lukasz Szpruch. Identifiability in inverse reinforcement learn- ing.Advances in Neural Information Processing Systems, 34:12362–12373,
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.