{"id":"16e33e8e-61c0-45ab-9288-25e2b4c288c6","arxiv_id":"2411.15951","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Standard inverse reinforcement learning models identify rewards only up to reward shaping and redistribution, and are not robust to even small misspecifications of transition dynamics or discount factors.","lead":"This paper maps, mathematically, when reward functions can and cannot be recovered from observed behavior in inverse reinforcement learning. It shows that the three standard behavioral models are ambiguous in predictable ways and are highly sensitive to small errors in assumed environment parameters.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the negative results are correctly scoped to the unrestricted reward space, and the main caveat is explicitly discussed in Appendix A.5.","rationale":"The reader's weakest-assumption analysis correctly identifies the unrestricted reward space as the main scope condition of the impossibility results. That condition is load-bearing in the sense that the negative theorems rely on it, and Appendix A.5 explicitly demonstrates that common restricted classes can evade the constructions. However, the paper does not hide this: the theorems are stated over all of R, and the appendix warns that the negative results may fail under restrictions. This is a limitation of the practical interpretation, not a flaw in the formal claims. The proof structure is extensive and the main theorems are precisely stated with proofs deferred to Appendix C. The reader's ACCEPT verdict with moderate confidence is appropriate, and no correction to the verdict is needed. I would only encourage readers to carry the Appendix A.5 caveat into any citation of the 'even slight misspecification' slogan.","tokens_in":54580,"tokens_out":18387,"duration_ms":184066,"concrete_test":"Re-derive the counterexample behind Theorem 74 for a concrete MDP, e.g., |S|=2, |A|=2, a deterministic non-trivial transition function, γ1=0.9 and γ2=0.91, by explicitly solving the Boltzmann fixed-point equations and checking that within one b_{τ,γ1,β}-equivalence class there exist R1,R2 with dSTARC_{τ,γ2}(R1,R2) > 0.98 (i.e., > 2ε for ε=0.49). If the construction fails in this instance, the theorem's proof has a gap; if it succeeds, the unrestricted-R impossibility is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I read the paper as establishing conditional impossibility results: for the standard behavioural models over the full reward space R = R^{S×A×S}, even arbitrarily small misspecification of τ or γ leaves the reward unidentifiable up to STARC distance at least 1, hence no ε<0.5 robustness (Theorems 73 and 74). The load-bearing modelling choices are the asymptotic convergence model, the unrestricted space R, and worst-case quantification over all compatible rewards. All three are stated in Section 3, and Appendices A.3–A.5 examine their consequences. In particular, Appendix A.5 shows that restricting R to state-action rewards or single-transition rewards can remove the S′-redistribution and potential-shaping constructions that drive Theorems 52–55 and 73–74, so the impossibility does not automatically transfer to common restricted reward parameterizations. This is a genuine scope limitation, but it is not a hidden assumption or an internal inconsistency: the theorems are formulated over the full space, and the paper does not claim the impossibility holds for every restricted class. The central mathematical claim is therefore supported, and the practical message should be read as a worst-case caution rather than a universal empirical statement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a formal theory of partial identifiability and misspecification robustness for inverse reinforcement learning (IRL). It introduces reward objects, invariance partitions, and two notions of misspecification robustness, one based on equivalence relations and one based on pseudometrics. For the three standard behavioural models (optimal policies, Boltzmann-rational policies, and maximal causal entropy policies), it characterizes the ambiguity of the inferred reward exactly: the first two determine the reward up to potential shaping and S'-redistribution, while the optimality model determines it up to optimality-preserving transformations. The paper also introduces STARC metrics, proves they are sound and complete (inducing both upper and lower regret bounds) and unique up to bilipschitz equivalence, and uses them to quantify ambiguity diameters and misspecification sensitivity. The main negative results (Theorems 52-55, 62-63, 73-74) show that any behavioural model invariant to S'-redistribution or potential shaping cannot guarantee transfer to a different transition function or discount factor, and is not epsilon-robust to arbitrarily small misspecification of those parameters for epsilon < 0.5 under the STARC metric. Appendices A.3-A.5 explicitly examine how the results change under inductive-bias assumptions, prior assumptions about the true reward, and restrictions of the reward space.","tokens_in":54771,"tokens_out":7803,"duration_ms":71387,"significance":"If correct, this is a substantial theoretical contribution. The STARC framework provides a canonical, regret-grounded way to compare reward functions, and the invariance theorems give exact characterizations of partial identifiability under the most common IRL behavioural models. The negative results are sharp conditional impossibilities rather than vague warnings, and the paper is unusually transparent about its scope: the asymptotic convergence model, the unrestricted reward space R, and worst-case quantification are all stated in Section 3, and Appendix A.5 explicitly shows that the transfer and misspecification impossibility results can fail when the reward space is restricted to state-action rewards or single-transition rewards. The proofs are detailed and parameter-free, and the appendices provide reusable tools for analysing new behavioural models. The heavy reliance on the authors' own prior conference papers is acknowledged in Section 1.2, and the novel parts (Section 6.2, Appendix A, and parts of Section 5) are clearly identified.","major_comments":[],"minor_comments":[{"comment":"The abstract's 'comprehensive mathematical analysis' should be qualified by the unrestricted-reward-space assumption. Appendix A.5 shows that Theorems 52-55, 62-63, and 73-74 can fail when the reward space is restricted, for example to state-action rewards or single-transition rewards, so the reader should be told up front that the impossibility results are worst-case over R.","section":"Abstract and Section 1"},{"comment":"The paragraph beginning 'Do do this, we will fist give' contains typos; it should read 'To do this, we will first give'.","section":"Section 3.2"},{"comment":"The stated condition 'g ≠ c_{τ,γ,ψ}' uses an undefined ψ; it should presumably read 'g ≠ c_{τ,γ,α}'.","section":"Theorem 59"},{"comment":"The first clause 'If f_τ is invariant to S′-redistribution with τ' does not quantify τ; the intended reading is that each f_τ is invariant to S′-redistribution with its own τ, and this should be stated explicitly.","section":"Theorem 62"},{"comment":"The soundness inequality places the normalizing range term on the right-hand side to avoid division by zero; a parenthetical remark clarifying this would help readers who might otherwise misread the bound as depending on the scale of R1.","section":"Definition 28"}],"recommendation":"accept","confidential_remarks":"This manuscript reuses substantial material from the authors' own prior conference papers (Skalse et al. 2023, 2024; Skalse and Abate 2023a, 2024), and some central theorems are adapted from those works rather than proved from scratch here. This is not a fatal issue for a synthesis-and-extension paper, but the editor may wish to confirm that the journal's policy on self-reuse is satisfied and that the provenance of each theorem is clear. The fit with a theory-oriented machine learning venue is strong, and the practical message is correctly scoped as a worst-case caution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a careful, mostly consolidating theory paper that makes precise an important negative message for IRL — the standard behavioral models cannot guarantee accurate reward inference under even slight misspecification of transition dynamics or discount factor when reward functions are unrestricted. The paper is honest about what is inherited from earlier conference papers and what is new. That transparency earns real credit.\n\nWhat is new: Section 6.2 (continuity/surjectivity arguments about which behavioral models can tolerate misspecification), parts of Section 5, and Appendix A including Theorems 77-83. The STARC framework and the main identifiability characterizations (potential shaping plus S'-redistribution) are restated from prior papers, but they are stated cleanly and the proofs are in the appendix. The paper ships no fitted parameters; it's all derivation, so the central claims are either right or wrong on the page.\n\nThe soft spots: the headline negative theorems (Theorems 73-74) quantify over the full reward space R = R^{S×A×S} and use worst-case compatible rewards under an asymptotic convergence model. Those three modeling choices are stated up front, and Appendix A.5 shows they matter: restrict R to state-action rewards or single-transition rewards and the S'-redistribution/potential-shaping constructions that drive the impossibility can disappear. So the practical message is conditional — worst-case caution, not a universal impossibility for every parameterization. That is a genuine scope limit, but not a hidden one. Theorem 72's perturbation non-robustness is similarly built on rewards near the zero reward; the paper says so. The proof appendix is not machine-checked; I didn't find an error in the readable core, but it's long and the edge cases are exactly where mistakes hide.\n\nWho it's for: theorists working on IRL identifiability, reward learning safety, and anyone building on STARC metrics. It deserves a serious referee and probably will get one.\n\nRecommendation: send to peer review. The caveats should be pressed in revision: make the scope of the negative results prominent in the abstract or main text, not just the appendix.","headline":"A transparent consolidation of prior IRL theory that makes the negative transfer and misspecification results precise, with the main caveat being that the impossibility lives in the unrestricted reward space.","tokens_in":55327,"tokens_out":2117,"would_cite":true,"duration_ms":18996,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T05","90C40"],"pacs":[],"model":"deepseek-v4-flash","headline":"Standard inverse reinforcement learning models recover rewards only up to shaping transformations, and they are not robust to even slight misspecification of the environment dynamics or discount factor.","keywords":["inverse reinforcement learning","partial identifiability","misspecification robustness","reward ambiguity","potential shaping","Boltzmann rationality","maximal causal entropy","STARC metrics"],"falsifier":"On a small MDP with a non-trivial transition function, compute the diameter, under the STARC metric $d^{\\mathrm{STARC}}_{\\tau,\\gamma_2}$, of the equivalence class of an arbitrary reward under potential shaping with discount $\\gamma_1$; Theorem 55 predicts this diameter is exactly 1, so any computed diameter below 1 would refute the transfer-impossibility result.","tokens_in":54338,"feed_emoji":"🤖","tokens_out":9945,"duration_ms":77736,"temperature":0.7,"pith_summary":"Inverse reinforcement learning (IRL) aims to infer a reward function from a demonstrated policy, but the same policy can be optimal, Boltzmann-rational, or entropy-maximising for many different rewards. This paper proves that, within the training environment, the Boltzmann-rational and maximal-causal-entropy behavioural models pin the reward down exactly up to potential shaping and S'-redistribution, while the optimality model pins it down only up to optimality-preserving transformations. It then asks how much misspecification these models tolerate, and answers: for the transition function and the discount factor, essentially none. The paper shows that any behavioural model invariant to these transformations is not $\\epsilon$-robust for any $\\epsilon < 0.5$ under the STARC metric to misspecified $\\tau$ or $\\gamma$, and that no continuous behavioural model is robust to arbitrarily small perturbations of the observed policy. If correct, this means that a small error in the environment model used by an IRL algorithm can produce a reward function that is nearly orthogonal to the true reward.","feed_headline":"Slight model errors can make IRL rewards completely wrong","feed_subtitle":"Optimal, Boltzmann-rational, and MCE models fail robustness tests for transition and discount misspecification.","key_machinery":"The machinery has two parts. First, reward transformations: potential shaping (adding $\\gamma\\Phi(s') - \\Phi(s)$), S'-redistribution (changing $R(s,a,s')$ without changing its expectation under $\\tau$), and optimality-preserving transformations. These transformations generate the invariance partitions of the standard behavioural models, and the paper proves that Boltzmann-rational and MCE policies are invariant exactly to potential shaping composed with S'-redistribution. Second, STARC (Standardised Reward Comparison) metrics, constructed by canonicalising away potential shaping and S'-redistribution, normalising away positive linear scaling, and then measuring distance; the paper shows these metrics are sound and complete, meaning small STARC distance is necessary and sufficient for low worst-case regret, and that any metric with this property is bilipschitz equivalent to them. The negative misspecification theorems work by showing that invariance to a transformation forces the diameter of the invariance partition to be 1 under the STARC metric of the misspecified environment.","core_discovery":"The central discovery is a complete characterisation of reward ambiguity and misspecification tolerance for the three standard behavioural models in IRL. A behavioural model is treated as a function $f : \\mathcal{R} \\to \\Pi$, and its invariance partition $\\mathcal{Am}(f)$ describes which reward functions are indistinguishable from a given policy. The paper shows that Boltzmann-rational and maximal-causal-entropy policies have the same invariance partition: rewards are identified only up to potential shaping (with discount $\\gamma$) and S'-redistribution (with transition $\\tau$), so under the same environment the learned reward has the same policy ordering as the true reward and zero STARC distance. The optimality model is more ambiguous: it determines the reward only up to optimality-preserving transformations, which preserve optimal policies but not the full policy ordering. The main negative results are that any model invariant to S'-redistribution is not robust to misspecification of $\\tau$, and any model invariant to potential shaping is not robust to misspecification of $\\gamma$, for any $\\epsilon < 0.5$ under a STARC metric, even when the misspecification is arbitrarily small. The paper also proves that continuous behavioural models are not $\\epsilon/\\delta$-separating, hence not robust to arbitrarily small perturbations of the observed policy.","pith_inferences":["Beyond the paper, these results imply that an IRL pipeline that estimates the discount factor or transition dynamics from data should treat those estimates as safety-critical: a small estimation error can yield a recovered reward that is nearly orthogonal to the true one under the STARC metric.","A testable extension would be to compute, on small MDPs, the diameter of potential-shaping equivalence classes under a slightly different discount factor; Theorem 55 predicts the diameter is 1 even when the discount error is arbitrarily small.","The framework also suggests a design principle for robust reward learning: explicitly breaking invariance to potential shaping or S'-redistribution, for example by anchoring rewards to a fixed reference transition or combining data from multiple environments with different dynamics, may evade the negative theorems."],"forward_implications":["Within the training environment, Boltzmann-rational and MCE IRL recover a reward with the same policy ordering as the true reward, so the ambiguity is harmless if the learned reward is deployed in the same MDP.","The optimality model preserves optimal policies but not the full policy ordering, so its ambiguity has positive upper diameter under any sound and complete metric.","Any behavioural model invariant to S'-redistribution is not $\\epsilon$-robust to a misspecified transition function for any $\\epsilon < 0.5$ under a STARC metric.","Any behavioural model invariant to potential shaping is not $\\epsilon$-robust to a misspecified discount factor for any $\\epsilon < 0.5$ under a STARC metric, for any non-trivial transition function.","If the true reward is known to lie in a restricted class, such as rewards depending only on state and action, the negative transfer and misspecification results can fail, as shown in Appendix A.5."],"supporting_citations":[{"why":"Introduces the IRL problem and its first ambiguity characterisation for optimal policies, the baseline this paper extends.","marker":"(Ng and Russell, 2000)"},{"why":"Introduces potential shaping and proves it preserves optimal policies, a pivotal transformation for the invariance theorems.","marker":"(Ng et al., 1999)"},{"why":"Introduces the Boltzmann-rational behavioural model whose partial identifiability and robustness the paper characterises.","marker":"(Ramachandran and Amir, 2007)"},{"why":"Introduces maximal causal entropy IRL, the other stochastic behavioural model the paper characterises.","marker":"(Ziebart, 2010)"},{"why":"Supplies the soft Q-function recursion used to express the MCE policy, which the paper's MCE theorems rely on.","marker":"(Haarnoja et al., 2017)"},{"why":"Prior work on invariance in policy optimisation and partial identifiability that Section 5 builds on.","marker":"(Skalse et al., 2023)"},{"why":"Prior work on misspecification in IRL that grounds the equivalence-relation robustness results of Section 6.","marker":"(Skalse and Abate, 2023a)"},{"why":"Introduces STARC metrics, the reward-distance tool used for the quantitative robustness bounds.","marker":"(Skalse et al., 2024)"},{"why":"Prior work quantifying sensitivity to misspecification that Section 7 extends.","marker":"(Skalse and Abate, 2024)"}],"fun_headline_variants":["IRL rewards are fragile: tiny model errors, big mistakes","Partial identifiability dooms IRL to ambiguity and fragility","IRL reward inference breaks with slight model misspecification","Even small transition or discount errors wreck IRL rewards","Ambiguity and fragility: the twin perils of IRL reward learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that IRL algorithms are analysed in the asymptotic limit in which they may converge to any reward function in the unrestricted space of all possible rewards; if the true reward is known to lie in a restricted class, such as state-action rewards, the negative results can fail, as the paper notes in Appendix A.5.","fun_headline_variants_meta":{"raw":{"variants":["IRL rewards are fragile: tiny model errors, big mistakes","Partial identifiability dooms IRL to ambiguity and fragility","IRL reward inference breaks with slight model misspecification","Even small transition or discount errors wreck IRL rewards","Ambiguity and fragility: the twin perils of IRL reward learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000786,"raw_usage":{"total_tokens":3548,"prompt_tokens":1104,"completion_tokens":2444,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":2358}},"tokens_in":720,"tokens_out":2444,"duration_ms":17373,"temperature":1.0,"reasoning_tokens":2358,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:42:22.388968+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a small MDP with a non-trivial transition function, compute the diameter, under the STARC metric $d^{\\mathrm{STARC}}_{\\tau,\\gamma_2}$, of the equivalence class of an arbitrary reward under potential shaping with discount $\\gamma_1$; Theorem 55 predicts this diameter is exactly 1, so any computed diameter below 1 would refute the transfer-impossibility result.","supporting_citations":[],"review_version":1}