{"id":"92242c3f-daba-4d02-ab06-ef2ffc5fc5a6","arxiv_id":"2411.17749","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Under partial observability, optimal human-AI play can include the AI disabling its off switch, and more human information or more communication can reduce AI deference.","lead":"This paper introduces a game-theoretic model of the AI shutdown problem in which the human and the AI each see only part of the world, and proves that even perfectly rational human-AI teams sometimes optimally let the AI bypass its off switch. It also shows that giving the human more information, giving the AI less, or adding communication can backfire and reduce how often the AI defers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 5.7's proof misses an equally optimal OPP in which A waits on all three observations, so the claim that expanding A's message set makes A defer strictly less in optimal policy pairs is not established as stated.","rationale":"The paper's core Section 4 examples survive direct checking: Example 4.1 and the unique-OPP enumerations behind Propositions 4.9 and 4.11 are sound modulo the Appendix A.4 typos the reader already flagged. The more serious internal issue is in the communication half. Proposition 5.7 is one of the two supports for the abstract's claim that bounded communication can make the AI defer less, but its proof is not exhaustive and the proposition overstates the conclusion. The alternative OPP I identified is not exotic; it is simply a different partition of the three assistant observations into two messages, and it has the same payoff. This means the nonmonotonicity claim survives only as an existence claim under a tie-breaking convention, which the paper does not state. Because the Section 4 results are not affected and the Section 5 issue is repairable, the reader's CONDITIONAL verdict is unchanged.","tokens_in":29572,"tokens_out":28423,"duration_ms":266490,"concrete_test":"Enumerate all deterministic policy pairs for the Appendix B.1 game with |MA| = 2, allowing arbitrary message mappings, including non-injective ones. In particular evaluate pi_A(A1) = m1, pi_A(A2) = m2, pi_A(A3) = m2 with all wait actions and H best-responding as above; it scores 39/4, tying the paper's claimed optimum. If full enumeration or an upper-bound argument confirms no policy exceeds 39/4, then Proposition 5.7 must be weakened to an existential claim, or the construction perturbed to break the tie.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Appendix B.1, the proof of Proposition 5.7 claims that when |MA| = 2, the optimal policy is to defer on A1 and A2 and play a on A3, with payoff 39/4, and that this follows 'by exhaustion.' The enumeration is incomplete: it only considers policies in which deferring observations get distinct messages. The policy pi_A(A1) = (m1, w(a)), pi_A(A2) = (m2, w(a)), pi_A(A3) = (m2, w(a)), with H best-responding by playing ON on {B1, B4} after m1 and on {B2, B4} after m2, gives the same expected payoff. In probability-weighted units, the three assistant observations contribute v1 = (5, -10, -10, 10), v2 = (-10, 5, -10, 10), and v3 = (-1, -1, 1, 10) over B1..B4. Partitioning as {A1}, {A2, A3} yields (1/4)[(5+10) + (4+20)] = 39/4. Thus A can wait on all three observations in an OPP, not only on A1 and A2. Consequently the proposition's universal reading fails, and the 'arbitrarily low probability' claim requires an unstated tie-breaking convention. This weakens the abstract's bounded-communication nonmonotonicity result, though the Section 4 existence examples appear intact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Partially Observable Off-Switch Game (PO-OSG), a single-round, common-payoff Bayesian game in which a human H and an AI assistant A receive correlated but private observations of a state, and A chooses to act, wait, or shut down; if A waits, H approves or shuts down. The authors show that, unlike in the fully observable off-switch game, optimal play can involve A acting unilaterally even when H is rational (Example 4.1), and that making H better informed or A worse informed can reduce the set of observations on which A waits in optimal policy pairs (Propositions 4.9 and 4.11). The model is extended to cheap-talk communication, with a Blackwell-type value-of-information theorem (Theorem 4.7), a monotonicity result for richer message sets (Theorem 5.4), and examples where more communication for A or less communication for H decreases deference (Propositions 5.7 and 5.8). An A-unaware human variant is analyzed, and finding an optimal policy pair is shown NP-hard by a reduction from MAX-CUT (Appendix D).","tokens_in":29823,"tokens_out":22107,"duration_ms":218489,"significance":"If the main claims are correct, this is a useful contribution to the AI-safety and shutdown-problem literature: it shows that partial observability changes the qualitative conclusion of Hadfield-Menell et al., and that naive informational interventions can backfire. The paper's explicit finite examples are a genuine strength, since they are small enough to be checked by hand, and the NP-hardness reduction ties the model to a standard computational problem. The paper is also careful to distinguish coordinated from arbitrary garblings and to separate the A-aware and A-unaware human models. However, the headline bounded-communication result (Proposition 5.7) needs a quantifier clarification: as proved, it establishes the existence of an optimal policy with lower deferral, while the statement and proof suggest a stronger uniqueness or universal reading. This should be fixed before publication.","major_comments":[{"comment":"The proof of Proposition 5.7 contains an incomplete enumeration for the case |M_A| = 2. In addition to the policy 'm1 in A1, m2 in A2, a in A3', the policy pi_A(A1) = (m1, w(a)), pi_A(A2) = (m2, w(a)), pi_A(A3) = (m2, w(a)), with H responding ON on {B1,B4} after m1 and ON on {B2,B4} after m2, gives the same expected payoff of 39/4. Thus there is an optimal policy pair in which A waits on all three observations, not only on A1 and A2. Consequently, the proposition's claim that expanding M_A makes A play w(a) strictly less often in optimal policy pairs is false under the natural universal reading, and the statement that A 'defers with probability 2p' requires an unstated tie-breaking convention. The result should be reformulated existentially, with an explicit quantifier over optimal policy pairs, and the proof should enumerate message assignments even for observations where A acts or shuts down, since messages are sent before A's action and can affect H's beliefs.","section":"Appendix B.1 / Proposition 5.7"},{"comment":"Theorem 4.7 states an if-and-only-if: O1 is better in optimal play than O2 iff O1 is more informative than O2. The 'if' direction is immediate and is the direction used later. The 'only if' direction is attributed to Theorem 3.5 of Lehrer, Rosenberg, and Shmaya (2010), but their theorem concerns all common-interest Bayesian games, whereas 'better in optimal play' here is quantified only over PO-OSGs, a restricted class with fixed action sets {a, w(a), OFF} and {ON, OFF} and payoffs that depend only on whether the action is taken. It is not obvious that this restricted class is rich enough to detect every violation of the coordinated-garbling informativeness order, so the converse does not follow from the cited theorem as written. Please either prove the converse for the PO-OSG class or state Theorem 4.7 only in the direction actually needed.","section":"Section 4.3 / Theorem 4.7"},{"comment":"The paper defines 'A plays w(a) strictly less often' for a pair of policies, but the propositions in Sections 4 and 5 use the phrase 'in optimal policy pairs' without specifying whether the comparison is over all optimal policy pairs, some optimal policy pair, or a set-valued relation. This ambiguity is hidden in Section 4, where the examples have unique OPPs, but it becomes load-bearing in Proposition 5.7, where multiple OPPs exist. The statements should either introduce a set-valued comparison of waiting sets for OPPs or explicitly say 'there exists an optimal policy pair in which A plays w(a) strictly less often,' and the proofs should match that statement.","section":"Definition 4.8 and Section 5 statements"}],"minor_comments":[{"comment":"In the proof of Proposition 4.9, the expected payoff of acting when H observes 1.x is -3/4, not -1/2; the summary policy line swaps ON and OFF for the observations 1.x and 2.x; and the stated expected payoff 2/3 should be 1. These are local inconsistencies with the surrounding figures and do not change the qualitative conclusion, but they make the formal example harder to verify.","section":"Appendix A.4"},{"comment":"Definition 6.2 says 'A policy pair is an A-aware optimal policy pair' but the context and surrounding text describe an A-unaware human; this appears to be a typo and should be corrected.","section":"Definition 6.2"},{"comment":"In the NP-hardness reduction, the claim that there is never an incentive for A to play a is stated too tersely: the argument should condition on A's observation OA and bound the expected payoff by (1/n)(-n^4) + (deg(v)/n)n^2 < 0. Also, the verification paragraph contains a garbled sentence ('given an optimal policy pair to determine if the optimal policy pair has expected payoff bigger than k') that should be rewritten.","section":"Appendix D"},{"comment":"In the enumerated policies for |M_A| = 2, the entry 'Playing m1 in A1, m2 in A2, a in A3' does not specify which message A sends when it plays a on A3. Since H conditions on the message, a full deterministic policy must include a message for every observation, including those where A acts or shuts down; this omission is part of why the later enumeration is incomplete.","section":"Appendix B.1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid core of checkable examples and appears worth publishing after revision. The main issue is that Section 5's headline nonmonotonicity result is stated more strongly than the proof supports: the constructed game has multiple OPPs, and the low-deference behavior is only one of them. The authors should either weaken the proposition to an existential statement or add a tie-breaking rule that selects the less-deferring OPP. I would also ask them to clarify the scope of Theorem 4.7, since the converse direction as stated does not obviously follow from the cited LRS theorem for the restricted PO-OSG class."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know before reading. The model is a genuine extension of Hadfield-Menell et al.: partial observability, common payoff, finite states, with optimal policy pairs as the solution concept. The Section 4 examples are explicit and checkable, and they make the point convincingly: an AI assisting a rational human can optimally act unilaterally (File Deletion Game), and giving the human strictly more information can make the AI wait strictly less (Example 4.10). Those proofs are basically enumeration, and the enumerations are complete. Good work.\n\nThe soft spot is Section 5. Proposition 5.7 claims that expanding A's message set makes A play w(a) strictly less often in optimal policy pairs. The appendix proof says 'by exhaustion' but the exhaustion is incomplete. In the |M_A|=2 case, there is an equally good policy where A waits on A1, A2, and A3: send message m1 on A1, message m2 on A2 and A3, and have H play ON on {B1,B4} after m1 and {B2,B4} after m2. That yields the same 39/4 payoff as the paper's preferred OPP. So 'strictly less often in optimal policy pairs' is false on the universal reading. It is true only if you mean 'there exists an OPP,' and the authors do not say that. This is a correctable gap—add a tie-breaking convention or modify the game to make the OPP unique—but it sits on one of the abstract's headline claims.\n\nTwo smaller items. Theorem 4.7 states an iff between 'better in optimal play' and 'more informative,' citing Lehrer-Rosenberg-Shmaya; the 'only if' direction looks stronger than that theorem alone supports, and it is not needed for the rest of the paper, so weakening it would be cheap. And in Appendix A.4 there is a typo in the displayed policy for π_A (o_H instead of o_A) that will trip up anyone verifying Example 4.10.\n\nWhat is solid: the NP-hardness reduction from MAX-CUT is clean, the A-unaware human variant is a thoughtful robustness check, and the paper openly states its limiting assumptions. The citations to prior off-switch and assistance-game work are appropriate.\n\nWho this is for: anyone working on corrigibility, assistance games, or shutdown incentives. It deserves a real referee, but the referee should require the Section 5 claim to be restated or fixed before the paper appears.\n\nRecommendation: send to peer review, conditional on the authors fixing the tie-breaking in Proposition 5.7 and toning down the iff in Theorem 4.7.","headline":"The PO-OSG is a genuine contribution and the Section 4 examples hold up; Section 5's Proposition 5.7 overreaches because the proof misses equally optimal policies that defer everywhere.","tokens_in":30475,"tokens_out":8271,"would_cite":true,"duration_ms":65153,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A27","91A80"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under partial observability, optimal AI policies can rationally avoid shutdown, and information changes can make the AI defer less often, even with a perfectly rational human.","keywords":["partial observability","off-switch game","corrigibility","shutdown problem","asymmetric information","assistance games","cheap talk","optimal policy pairs"],"falsifier":"Enumerate all deterministic policy pairs — finitely many — for the two observation structures of Example 4.10 and compute their expected payoffs. The paper claims that when H observes only the first digit of the version number, the unique optimal policy pair has A wait under both code types, and when H observes the full version number, the unique optimal policy pair has A wait only under code type B and act unilaterally under code type A; because every optimal stochastic policy pair is a mixture of deterministic optimal policy pairs, finding any deterministic policy pair in either game with a strictly higher expected payoff would refute Proposition 4.9's non-monotonicity claim.","tokens_in":29311,"feed_emoji":"🔌","tokens_out":12260,"duration_ms":98596,"temperature":0.7,"pith_summary":"The paper introduces the Partially Observable Off-Switch Game (PO-OSG), a two-player common-payoff model in which a human and an AI assistant each observe only part of the state and the AI can act, wait for the human's approval, or turn itself off. Its central finding is that partial observability alone — with a perfectly rational human, common payoffs, and costless feedback — makes it optimal for the AI to sometimes disable its own off switch and act unilaterally. It then proves that common-sense information interventions can backfire: making the human's observations strictly more informative, or the AI's strictly less informative, can make the AI defer strictly less often in optimal policy pairs, and bounded communication can do the same. These results matter because they overturn the intuition that transparency or bandwidth can monotonically restore shutdown incentives, and because safe-agent design must then weigh payoff maximization against maintaining deference. If the paper's examples generalize, the off-switch guarantee that holds when the human knows everything the AI knows breaks precisely in the high-information-asymmetry settings that motivate shutdown concerns.","feed_headline":"Better-informed humans can make AIs wait less","feed_subtitle":"In optimal play of off-switch games, adding human information or communication can backfire and reduce AI deference.","key_machinery":"The central object is the Partially Observable Off-Switch Game (PO-OSG): a dynamic Bayesian game in which an action either goes through or not; if it goes through both players receive $u_a(S)$, otherwise $u_o(S)$, with H observing $O_H$ and A observing $O_A$ before choosing. A can take the action, wait, or shut down; if it waits, H chooses ON or OFF. The load-bearing result that lets the paper separate payoff effects from deference effects is Theorem 4.7, which says an observation structure on both players is better in optimal play than another exactly when it is more informative in the sense of coordinated garblings — a multi-agent analogue of the nonnegative value of information. The mechanism behind the counterintuitive deference results is 'deferral as implicit communication': because H knows A's policy, A's choice to wait reveals a bit of A's private observation, so giving either player new information can change which subsets of states are jointly selectable, and the newly optimal subset may be achievable only by waiting on fewer observations. In the communication extension (PO-OSG-C), messages are cheap talk and unbounded communication by either player makes the other player's observation redundant, restoring the classic results.","core_discovery":"Under full observability, the classic off-switch game has a clean result: an AI that shares the human's payoff never has a reason to avoid shutdown. The paper's central discovery is that this clean result fails once information is asymmetric. In a PO-OSG, there are optimal policy pairs — policy pairs maximizing the common expected payoff over all policy pairs — in which the AI plays the action unilaterally instead of deferring, even when the human is perfectly rational and feedback is free; the File Deletion Game (Example 4.1) has a unique optimal policy pair in which the AI acts directly on one of its observations. The paper further proves that the deference choice is not monotone in information: there are games where making the human's observations strictly more informative (Proposition 4.9) or the AI's observations strictly less informative (Proposition 4.11) makes the AI wait strictly less often in every optimal policy pair, and the same holds for expanding the AI's message set (Proposition 5.7) or shrinking the human's (Proposition 5.8). At the extremes, the classic result survives: if one player's observations are redundant given the other's, the better-informed player can be given sole decision-making power (Proposition 4.3), and unbounded communication by either player has the same effect (Corollary 5.6). These results persist when the human is A-unaware and cannot use the AI's deferral as a signal, though in that case giving the human more information can strictly reduce the expected common payoff (Proposition 6.3(b)).","pith_inferences":["This suggests a design caution beyond the paper's examples: any proposal to restore corrigibility by giving humans more transparency or by restricting AI sensing must be checked against the optimal policy pair of the resulting game, since the map from relative information to deference is provably non-monotone.","A testable extension is to replace the implicit-communication channel used by the wait action with explicit cheap-talk messages of the same bandwidth and measure whether optimal deference changes; the paper's Propositions 5.7 and 5.8 predict it can, which would show that the channel, not its implicitness, drives the backfire.","One open question the paper flags is whether the constructed payoff tables are natural; a concrete way to test this is to search real human-AI assistance datasets for payoff structures shaped like the File Deletion table, where a better-informed human makes unilateral AI action optimal."],"forward_implications":["When one player's observations are redundant given the other's, there is an optimal policy pair giving sole decision-making power to the better-informed player: A can always defer if H is at least as informed, and never needs to defer if A is at least as informed (Proposition 4.3).","Giving either player more informative observations never decreases the expected common payoff in optimal play, so payoff losses cannot justify interventions that restrict information (Theorem 4.7).","Increasing the human's information, decreasing the AI's information, expanding the AI's message set, or shrinking the human's message set can each make the AI defer strictly less often in optimal policy pairs (Propositions 4.9, 4.11, 5.7, 5.8).","If either player can communicate their full observation, the other player can be given sole decision-making power, so unbounded communication restores the classic off-switch-game dichotomy of full deference or no deference (Corollary 5.6).","All of the counterintuitive deference findings survive when the human is A-unaware and cannot use deferral as a signal, and in that setting more human information can strictly lower the common expected payoff; solving for optimal policy pairs is NP-hard while A-unaware policies are polynomial-time computable (Section 6 and Appendix D)."],"supporting_citations":[{"why":"The Off-Switch Game that PO-OSGs generalize; supplies the baseline result that a fully informed rational human makes A always defer in optimal play.","marker":"Hadfield-Menell et al. 2017"},{"why":"Provides the characterization (used as Theorem 4.7) that an observation structure is better in optimal play iff it is more informative under coordinated garblings; the backbone of the payoff and information results.","marker":"Lehrer, Rosenberg, and Shmaya 2010"},{"why":"Defines assistance games, the common-payoff framework in which PO-OSGs are situated (Appendix E).","marker":"Shah et al. 2020"},{"why":"Supplies the motivating 'you can't fetch the coffee if you're dead' framing of the shutdown problem.","marker":"Russell 2019"},{"why":"Earlier work showing partial observability can incentivize observation tampering in assistance games; the catalogue of concerning behaviors this paper extends to shutdown-avoidance.","marker":"Emmons et al. 2024"},{"why":"Shows the OSG's main results require costless human feedback; the paper inherits that assumption and cites it.","marker":"Freedman and Gleave 2022"},{"why":"The cheap-talk model used to define communication in PO-OSG-Cs (Section 5).","marker":"Crawford and Sobel 1982"},{"why":"Decentralized POMDP complexity results that situate the NP-hardness of finding optimal policy pairs in PO-OSGs (Appendix D).","marker":"Bernstein et al. 2002"},{"why":"The classic experiment-comparison theorem used in the proof that more informative A-observations improve A-unaware optimal play (Proposition 6.3(a)).","marker":"Blackwell 1951"}],"fun_headline_variants":["Asymmetric info lets optimal AIs avoid shutdown","More human information can make AIs defer less","Off-switch game: hidden info breaks the rational-AI rule","Communication can backfire: AIs defer less with more info"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline results are proved for optimal policy pairs in which both agents coordinate on a policy pair maximizing a common payoff, with costless human feedback, a single interaction round, and a payoff function that exactly captures the human's preferences; if real humans cannot infer and follow the AI's equilibrium policy, or the payoff misses what they truly want, the predicted deference behavior need not occur.","fun_headline_variants_meta":{"raw":{"variants":["Asymmetric info lets optimal AIs avoid shutdown","More human information can make AIs defer less","Off-switch game: hidden info breaks the rational-AI rule","Communication can backfire: AIs defer less with more info"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000633,"raw_usage":{"total_tokens":2994,"prompt_tokens":1093,"completion_tokens":1901,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":1847}},"tokens_in":709,"tokens_out":1901,"duration_ms":21098,"temperature":1.0,"reasoning_tokens":1847,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:10:38.544423+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Enumerate all deterministic policy pairs — finitely many — for the two observation structures of Example 4.10 and compute their expected payoffs. The paper claims that when H observes only the first digit of the version number, the unique optimal policy pair has A wait under both code types, and when H observes the full version number, the unique optimal policy pair has A wait only under code type B and act unilaterally under code type A; because every optimal stochastic policy pair is a mixture of deterministic optimal policy pairs, finding any deterministic policy pair in either game with a strictly higher expected payoff would refute Proposition 4.9's non-monotonicity claim.","supporting_citations":[],"review_version":1}