{"id":"2e4bed96-e780-4f51-9352-447af77e3e54","arxiv_id":"2506.02923","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Observed behavior only weakly constrains an intentional agent's choices under distribution shift, and its perceived fairness and harm cannot be identified from behavior alone.","lead":"This paper derives mathematical bounds on how much an AI's decisions in new environments can be predicted from its past behavior, assuming the AI acts through a causal world model. The bounds show which actions can be ruled out as suboptimal and which cannot, with direct consequences for AI safety and fairness evaluation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the central conditional claim is coherent, and the main caveat (grounding) is explicitly scoped.","rationale":"The reader's weakest assumption (grounding) is indeed the point where the practical guarantee is least secure, so I partially agree. However, grounding is an explicit premise of the theorem, not a hidden or circular step; the paper states it as Definition 3, uses it consistently, and offers an approximate-grounding relaxation in Section 5.1. I checked the main proof steps: the denominator bound for P_{z,d}(C=c) is correct for general C once the c' shorthand is expanded, and the tightness constructions can be combined into a single SCM by letting structural equations depend on D (with off-diagonal values of f_Y unconstrained by observed data and free to realize the bounds). The remaining typesetting errors (e.g., missing parentheses in Theorem 1's second numerator, the mixed denominators in the appendix proof of Corollary 1) are presentation issues, not load-bearing. A verification of the general denominator inequality would be a useful belt-and-braces check, but I do not believe it will overturn the central claim.","tokens_in":43371,"tokens_out":32920,"duration_ms":380209,"concrete_test":"Independently re-derive Theorem 1's two-sided bound for a multi-valued context C, replacing the Appendix C c' notation by the general inequality P_{z,d}(R=r) ≤ P_d(R=r,Z=z)+1-P_d(Z=z); if the derivation fails at any step, the characterization is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim is a conditional characterization: for an agent exactly grounded in M and with access to P_d(V) for every decision d, Theorem 1's inequality is necessary and sufficient for weak predictability. The lower/upper bound derivation in Appendix C is valid (the c' shorthand conceals a legitimate partition of C, and the 'summands>0' steps have the correct direction), and the tightness construction can in principle be made simultaneous by defining f_Y(D,Z,U) piecewise on the diagonal/off-diagonal of Z relative to f_Z(D,U). The strongest dependency is Definition 3: if the internal model is not exactly grounded, the inequalities do not constrain out-of-distribution choices. This is an explicit, unscored assumption rather than an internal inconsistency, and Section 5.1 partially addresses it. I do not see a load-bearing flaw in the theorem as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how much of an AI agent's future out-of-distribution decisions can be predicted from its observed behavior, under the assumptions that the agent is an expected-utility maximizer over an internal structural causal model (SCM) and that the agent is 'grounded' in its training domain, i.e. its internal interventional distributions coincide with the real environment's for every decision. The central result (Theorem 1) is a necessary-and-sufficient inequality, computable from observed interventional distributions P_d(V), for the existence of a decision that can be ruled out as suboptimal under a hard intervention do(z). Section 4 also derives bounds for multiple environments, general shifts, and counterfactual fairness and harm gaps; Section 5 relaxes exact grounding, exact expected-utility maximization, and exact observation of the utility; the appendices contain proofs, canonical-model constructions, and a limitations discussion.","tokens_in":43489,"tokens_out":23595,"duration_ms":231400,"significance":"If correct, Theorem 1 is a clean, parameter-free characterization: either behavioral data certifies that a grounded agent will not take a given action in a novel environment, or there exist two equally data-compatible internal models with different optimal actions. The paper is explicit that the grounding assumption is load-bearing, and Section 5.1 extends it approximately; the limitations appendix (B.5) honestly acknowledges the finite-sample, acyclicity, and verbal-behavior caveats. The bounds follow the Balke-Pearl style and are derived rather than assumed, and the multiple-environments and fairness/harm results are natural and falsifiable. However, the tightness proofs as printed are flawed, which affects the central if-and-only-if claim; the issues appear repairable with the piecewise diagonal construction, so the contribution remains of high potential value.","major_comments":[{"comment":"The SCM displayed for the lower-bound tightness construction branches on 'f_Z(u)=z' and sets Y=0 for d1 and Y=1 for d0 whenever f_Z(u) is not equal to the intervention value z. But in the observed regime the actual value of Z is f_Z(u), so this branch forces P_d1(Z≠z,Y=1)=0 and P_d0(Z≠z,Y=0)=0 irrespective of the observed distribution. For Example 1, Table 1 gives P_d1(Z=0,Y=1)=0.2, so the displayed construction cannot generate the observed P_d1. The correct construction should branch on the equality between the actual value of Z and f_Z(u): define f_Y(d,c,Z,u) to use the observed response when Z=f_Z(u) and to take the extreme values only on the off-diagonal Z≠f_Z(u) in the intervened regime. Because this construction is the only demonstration of tightness for Theorem 1, the if-and-only-if claim is not proved as written.","section":"Appendix C, Eq. (79)"},{"comment":"Weak predictability is defined by the existence of d* such that for every valid SCM some alternative d (possibly depending on the SCM) is preferred to d*. Theorem 1, however, characterizes weak predictability by the existence of a single pair d,d* with min_{cM} Δ_{d≻d*} > 0. This is a quantifier exchange: from 'for every cM there exists d with Δ_d(cM)>0' it does not in general follow that 'there exists d with Δ_d(cM)>0 for every cM'. The proof only constructs tightness examples separately for each pair, so no argument is given that when all pairwise minima are non-positive, a single SCM exists in which d* is optimal against all alternatives simultaneously. A simultaneous construction is needed, or the theorem should be restated as a sufficiency result.","section":"Theorem 1"},{"comment":"The tightness construction for the multiple-environments bound suffers from the same off-diagonal branching problem as Theorem 1: the branch on f_S(u)=s forces outcomes in the observed regime that need not match the data. In addition, the proof of tightness is only given for the special case of two environments with Z=R1∪R2; the general statement for arbitrary R_i and k>2 is asserted without a construction. The theorem should either be proved in full generality or restricted to the case for which the tightness argument is supplied.","section":"Theorem 2"}],"minor_comments":[{"comment":"The numerator of the second term in the displayed inequality is typeset as a separate line without a clear fraction; it should read [E[Y|c,z]P(c,z)+1-P(z)] / [P(c,z)+1-P(z)], as in Eq. (77) of Appendix C.","section":"Theorem 1"},{"comment":"The displayed condition is split across two fractions with denominator P_{σ,d*}(c); it would be much clearer to write the numerator as a single bracketed term, e.g. 1 - [2 + E[Y|c]P(c) - E[Y|c]P(c) - 2P(z) + P(c)] / P_{σ,d*}(c) > 0.","section":"Theorem 4"},{"comment":"The statement and proof use a single 'min over bP' for expressions involving both bP_d and bP_d*, which is ambiguous about whether the two decision-indexed distributions are optimized independently or jointly as marginals of the same internal model; the proof's denominators are also inconsistent (P_d versus bP_d). Please clarify the feasible set and the optimization.","section":"Corollary 1"},{"comment":"In the displayed SCMs, the branch for C sets C to '1 otherwise'; the subsequent denominator computations require P(C=c | f_Z(u)≠z)=1, so the constant should be c (the context value) rather than the number 1.","section":"Appendix C"},{"comment":"The expression 'argmax_π E[Y|do(π)]' is used without defining the intervened SCM induced by a policy π; the paper should define do(π) or replace it by an explicit policy-intervention notation.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The central idea is strong and the bounds are likely correct, but the tightness proofs in the appendix are not valid as written, and the if-and-only-if claim needs either a simultaneous construction or a careful restatement. These are fixable within the manuscript's scope, so I do not recommend rejection; the paper should be reconsidered after a revision that repairs the canonical constructions and clarifies the quantifier issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a careful, largely successful formal treatment of what behavior can tell us about an AI's choices in new environments. The main result, Theorem 1, is an iff condition for when observed decisions P_d(V) rule out a decision d* as suboptimal under an intervention do(z). That's genuinely new: it connects partial identification bounds (Manski, Pearl, Tian) to intentional-agent prediction. The fairness impossibility (Theorem 5) is also a clean new application: the counterfactual fairness gap is essentially unconstrained by behavior, width 1, so you can never infer 'intends to be fair' from data alone. Harm bounds (Theorem 6) are a legitimate extension of probability-of-causation bounds.\n\nThe proofs are real proofs, with explicit SCM constructions showing tightness of each bound. I checked the lower/upper bound derivation in Appendix C and the c' shorthand is a legitimate partition of C; the 'summands>0' steps have correct direction. The bounds are parameter-free, so circularity burden is low. The paper also openly states its limitations: infinite sample, no feedback, grounding assumption, no exploitation of verbal behavior. Section 5.1 partially addresses grounding with approximate grounding.\n\nSoft spots: The load-bearing premise is grounding (Def. 3): internal model must assign same intervention probabilities as real environment. If the AI's model is misspecified, the bounds don't constrain. The authors acknowledge this but the exact theorems depend on it. That's an assumption, not a flaw, but it limits direct application. Some equations in the main text are garbled (Theorem 4 and Corollary 2 are hard to parse), though the appendix restates them correctly. Finite-sample analysis is missing; they flag it, but it matters for safety use-cases. Theorem 3 and 4 have a slightly awkward boundary: under completely unknown shifts no predictability; with covariate data you get non-tight bound. That's honest but leaves a gap.\n\nMy verdict: this deserves a serious referee. The central conditional claim holds up as stated. It's not a breakthrough, but it's a solid, citable formal result. I'd send it to review (maybe a causality or AI-safety venue) and ask for a fix of the typesetting and a short discussion of finite-sample implications.\n\nFor reading group: maybe, yes if people are into causal bounds. I'd cite it if I work on agent behavior prediction.\n\nRecommendation: accept with minor revisions after a technical check of the formulas.","headline":"A genuinely new and correct conditional characterization of when behavioral data can rule out an agent's decisions under intervention, with the caveat that everything rests on the grounding assumption.","tokens_in":43992,"tokens_out":3053,"would_cite":true,"duration_ms":27372,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Behavioral data alone can guarantee which actions an AI will not take in a novel environment exactly when a data-computable inequality is positive.","keywords":["AI safety","behavioural prediction","structural causal models","weak predictability","causal bounds","counterfactual fairness","counterfactual harm","distribution shift"],"falsifier":"Construct a grounded AI whose internal structural causal model is fully known, record its behavior P_d(V), compute the Theorem 1 expression for a candidate unsafe decision d* under an intervention do(z), and then actually deploy the AI under do(z). If the expression is positive and the AI nevertheless chooses d*, the theorem's sufficiency claim is false; if the expression is negative but d* is never chosen across many runs with different internal models, the necessity claim needs qualification. The cleanest single check is to build two SCMs with identical P_d(V) but opposite optimal decisions under do(z) while the Theorem 1 expression is positive, which the theorem says is impossible.","tokens_in":43169,"feed_emoji":"🤖","tokens_out":6819,"duration_ms":68486,"temperature":0.7,"pith_summary":"This paper asks what an outside observer can guarantee about an AI's choices in a new environment after watching only its behavior — the decisions it took, the contexts, and the resulting rewards — in a training environment. Treating the AI as an expected-utility maximizer whose internal world model is a structural causal model, the paper proves that a decision can be ruled out under an intervention exactly when a data-computable inequality is positive. If the inequality fails, no decision can be certified as sub-optimal: two different internal models, both fully compatible with the observed behavior, will favor different actions. The result is an if-and-only-if boundary: it tells practitioners when behavioral audits can support safety claims and when they fundamentally cannot. The paper also shows that an AI's counterfactual fairness gap cannot be identified from behavior alone and derives tight bounds for its perceived counterfactual harm.","feed_headline":"A single inequality sets the limit of predicting AI behavior","feed_subtitle":"When it holds, observers can rule out an unsafe action in a new environment; when it fails, behavior data cannot certify any choice.","key_machinery":"Structural causal models (SCMs) serve as the common language for the AI's internal world model and the observer's data, with grounding as the bridge between them: the AI's internal model probabilities P_hat_d(V) equal the observed behavior probabilities P_d(V) for every decision d. The carrying object is the preference gap Δ_{d≻d*} = E_{P_hat}[Y | do(σ,d), c] − E_{P_hat}[Y | do(σ,d*), c]; weak predictability of d* holds exactly when some alternative d has a positive lower bound on Δ. The lower bound is derived by manipulating counterfactual probabilities using the axioms of counterfactuals and the σ-calculus, replacing the AI's internal expectations with observable P_d terms through grounding, and tightness is shown by explicit SCM constructions that attain the bound.","core_discovery":"An AI that is grounded in a domain M, meaning its internal world model assigns the same intervention probabilities as the observed environment for every decision, is weakly predictable under a shift do(z) in a context c if and only if there is a decision d* such that for some alternative d the lower bound on the preference gap Δ_{d≻d*} is strictly positive. The preference gap is the difference in the AI's expected utility between the two decisions under the shift. The theorem gives a closed-form expression in terms of the observable behavior distribution P_d(V) and proves the bound is tight: whenever the expression is not positive, one can construct two structural causal models generating exactly the same behavior but making d and d* respectively optimal, so behavior alone cannot rule out either action. The same machinery yields tighter bounds when behavior is observed in multiple interventional domains, a negative result when the shift is entirely unspecified, partial predictability when covariate data from the shifted environment is available, and tight bounds on the AI's perceived counterfactual fairness gap and counterfactual harm gap.","pith_inferences":["A practical audit protocol follows directly: before deploying under an intervention, compute the inequality on logged behavior; a negative result should push evaluators toward model inspection or interventional data collection rather than more observational data.","For large language models, the grounding assumption could be tested by comparing predicted optimal decisions under prompt-level interventions against the bounds computed from a behavior corpus, though assigning variables and decisions to text is nontrivial.","The paper's own limitation notes point to natural extensions: relaxing acyclicity, incorporating verbal self-reports from the AI, and moving from exact to high-probability bounds by parameterizing the SCM family.","When fairness and harm are defined counterfactually, audits that rely only on behavioral logs can certify harm bounds but not fairness intentions, shifting the burden of evidence toward model-based or interventional sources."],"forward_implications":["Behavioral audits can certify that an AI will avoid specific unsafe actions under well-defined interventions, without inspecting its internal weights, whenever the Theorem 1 inequality is positive.","Because the characterization is if-and-only-if, a failed inequality is a genuine epistemic limit: more behavior data from the same domain cannot rule out any action.","Observations from several interventional domains tighten the bounds, so multi-environment evaluations can strictly expand the set of provably avoided actions.","If the deployment shift is completely unspecified, behavior alone provides no action guarantees at all; knowing the shifted environment's covariate distribution restores partial predictability.","Counterfactual-fairness claims cannot be certified from behavior alone — the fairness gap is compatible with any observed behavior — while the counterfactual harm gap has tight, informative bounds.","The paper states its guarantees in the infinite-sample limit, so finite-sample application requires additional statistical treatment that the paper does not provide."],"supporting_citations":[{"why":"Supplies the structural causal model formalism, the axioms of counterfactuals, and the causal-hierarchy perspective that ground all the paper's derivations.","marker":"Pearl (2009)"},{"why":"Provides the σ-calculus inference rules used to move between policy-induced behavior distributions and intervention distributions.","marker":"Correa and Bareinboim (2020a)"},{"why":"Gives the probability-of-causation bounds that Theorem 6 extends to an AI's perceived counterfactual harm.","marker":"Tian and Pearl (2000)"},{"why":"Supplies the canonical model constructions used in the proof of Theorem 3 to show that under a vague shift the preference gap is unconstrained.","marker":"Jalaldoust et al. (2024)"},{"why":"Supports the representation result connecting rational agent behavior to causal world models, motivating the SCM treatment of AI beliefs.","marker":"Halpern and Piermont (2024)"},{"why":"Supports the assumption that agents robust across environments behave as if they operate with a causal world model.","marker":"Richens and Everitt (2024)"}],"fun_headline_variants":["An inequality determines when AI behavior is predictable","Tight bounds on predicting AI agents from behavior","Theoretical limit on inferring AI goals from actions","Predicting AI actions: a single inequality decides","When behavior data can't certify an AI's next move"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument rests on grounding: the AI's internal world model must assign exactly the same intervention probabilities as the real training environment for every decision; if the AI learned a distorted causal model, the observed-data inequalities no longer constrain its out-of-distribution choices.","fun_headline_variants_meta":{"raw":{"variants":["An inequality determines when AI behavior is predictable","Tight bounds on predicting AI agents from behavior","Theoretical limit on inferring AI goals from actions","Predicting AI actions: a single inequality decides","When behavior data can't certify an AI's next move"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1285,"prompt_tokens":914,"completion_tokens":371,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":298}},"tokens_in":530,"tokens_out":371,"duration_ms":4032,"temperature":1.0,"reasoning_tokens":298,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:12:28.730799+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a grounded AI whose internal structural causal model is fully known, record its behavior P_d(V), compute the Theorem 1 expression for a candidate unsafe decision d* under an intervention do(z), and then actually deploy the AI under do(z). If the expression is positive and the AI nevertheless chooses d*, the theorem's sufficiency claim is false; if the expression is negative but d* is never chosen across many runs with different internal models, the necessity claim needs qualification. The cleanest single check is to build two SCMs with identical P_d(V) but opposite optimal decisions under do(z) while the Theorem 1 expression is positive, which the theorem says is impossible.","supporting_citations":[],"review_version":1}