{"id":"9e5c73fb-e62b-43c5-97ac-3fe6ae244698","arxiv_id":"2505.23445","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Optimising a proxy can still maximise the true goal when both are light tailed, but a heavy-tailed discrepancy makes the goal collapse at a rate set by the tail shape.","lead":"This paper formalises Goodhart's law without assuming the goal and the proxy metric are independent. It shows that tail thickness, not dependence alone, determines whether optimising a proxy helps or harms the true goal.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Light-tailed claim is overgeneralized: the paper's own Gaussian formulas show dependence can flip benign to strong Goodhart when Var(G) < Var(ξ).","rationale":"The reader's weakest_assumption concerns the tail-conditioning model of optimization (Section 3). That is an external-validity worry: if real optimizers do not select extreme tails of M, the theorems may not transfer. I share that worry but do not consider it the most load-bearing issue for the paper's central claim. The more direct threat is internal: the abstract's headline assertion about light-tailed goal and light-tailed discrepancy is overgeneralized and, as shown by the Gaussian formulas in Lemma 4.1, actually false in a parameter regime the paper itself permits (Var(G) < Var(ξ) with sufficiently negative covariance). This is not a matter of empirical calibration; it is a mismatch between the paper's advertised contribution and its own derived asymptotics. The main heavy-tailed theorem (Theorem 4.1) appears correct and is the paper's strongest contribution, so the right response is a revision that restricts or removes the unqualified light-tailed claim and qualifies the abstract with the variance condition used in Section 4.3. That is consistent with the reader's CONDITIONAL verdict, so no verdict change is needed.","tokens_in":30728,"tokens_out":21962,"duration_ms":208853,"concrete_test":"Plug the valid Gaussian parameters a=1, b=2, c=-1.4 (det Σ = 0.04 > 0) into Lemma 4.1's asymptotic: E[G|M>m] ~ ((1-1.4)/(1+2-2.8)) m = -2m. This tends to -∞ (strong Goodhart), while c=0 gives +m/3 (benign). This single parameter check uses only the paper's own formula and shows the abstract's unqualified light-tailed claim is false.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims that 'in the case of light tailed goal and light tailed discrepancy, dependence does not change the nature of Goodhart's effect.' This is the advertised independence-free selling point, but it is not established and is contradicted by the paper's own Section 4.3 Gaussian analysis. Lemma 4.1 gives E[G|M>m] ~ (a+c)/(a+b+2c) m in the Gaussian case with (G,ξ) ~ N(0,Σ), a=Var(G), b=Var(ξ), c=Cov(G,ξ). The benign outcome (E[G|M>m]→+∞, Corr→0) holds only when a+c>0, which is guaranteed by the stated condition Var(G)>Var(ξ) but is not generally true for light-tailed G and ξ. If a<b, one can choose c ∈ (-√ab, -a) (e.g., a=1, b=2, c=-1.4); then the coefficient is negative, so E[G|M>m] → -∞, which is strong Goodhart by the paper's own Table 1, whereas for c=0 the coefficient is positive (benign). Dependence therefore changes the nature of Goodhart's law in the light-tailed case, contrary to the abstract. The paper's Section 5 even concedes that dependence matters when 'relative tail thickness are not favorable,' so the abstract is internally inconsistent with the body's results.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalises Goodhart's law in the paradigm-agnostic framework G = M + ξ, where optimisation of the proxy metric is modelled as conditioning on M > m with m → S(M). It defines four outcomes (no, benign, weak, strong Goodhart) and studies how the dependence between the goal G and the discrepancy ξ changes the outcome. The main results are: (i) Theorem 4.1, asserting that when G is heavy-tailed and ξ is two-sided light-tailed, E[G|M>m] ≥ m + o(m) regardless of dependence; (ii) a Gaussian analysis (Lemma 4.1, Theorem 4.2) giving the conditional expectation and correlation asymptotics and identifying a 'benign Goodhart' regime under Var(G) > Var(ξ); and (iii) an exponential-goal / conditionally heavy-tailed discrepancy example (Lemmas 4.2 and 4.3) producing strong Goodhart with rate η^{b-1}/m^{b-1}. The appendix contains detailed computations, several of which are verified with Sympy.","tokens_in":31046,"tokens_out":11394,"duration_ms":119479,"significance":"If the claims are appropriately qualified, the paper makes a useful contribution: it extends the independence-based analysis of El-Mhamdi and Hoang to dependent couplings, provides explicit and checkable Gaussian formulas, and exhibits a concrete heavy-tailed coupling that turns weak into strong Goodhart. The machine-checked symbolic computations are a genuine strength, as is the effort to state tail conditions explicitly. However, the headline claim that dependence does not change the nature of Goodhart's law in the light-tailed case is not supported by the paper's own Gaussian formulas, and one of the heavy-tailed lemmas is stated outside its domain of validity. With corrections and a more accurate abstract, the core material can be made publishable.","major_comments":[{"comment":"The abstract claims that in the light-tailed-goal/light-tailed-discrepancy case 'dependence does not change the nature of Goodhart's effect.' This is contradicted by Lemma 4.1, which gives E[G|M>m] ∼ (a+c)/(a+b+2c) m. For admissible parameters with Var(G)<Var(ξ), for example a=1, b=2, c=−1.4 (satisfying |c|<√(ab)=1.414), the coefficient a+c is negative, so E[G|M>m]→−∞, which is strong Goodhart by Table 1, whereas under independence (c=0) the same a,b give a positive coefficient and E[G|M>m]→+∞, i.e. benign Goodhart. Dependence therefore flips the nature of the effect within the paper's own definitions. Section 5's admission that dependence matters when 'relative tail thickness are not favorable' is inconsistent with the abstract. The abstract and Section 4.1 overview must be revised to state the condition a+c>0 (e.g. Var(G)>Var(ξ)) under which the benign conclusion holds, and Theorem 4.2's 'no matter the correlation' wording should be qualified because decorrelation alone does not imply benign Goodhart.","section":"Abstract and §4.3 (Lemma 4.1)"},{"comment":"Lemma 4.2 states E[ξ|M>m] ∼ m(b−1)/(b−2) under the setup of Section 6.4, where b is only assumed to lie in (1,∞). This asymptotic is valid only for b>2. The proof in Section 6.7 contains the term (b−1)η^{b−1}∫_m^∞ x^{−(b−1)}dx, which diverges for 1<b<2, and Lemma 6.7 itself explicitly restricts to b>2. For 1<b<2 the printed leading coefficient (b−1)/(b−2) is negative and the conditional expectation of ξ is not covered by the stated expansion. The domain restriction must be imposed in Lemma 4.2 and in the Section 4.4 claims that the discrepancy is maximised; as written, the strong-Goodhart example is stated beyond its range of validity.","section":"§4.4, Lemma 4.2 and §6.4"},{"comment":"The proof of Theorem 4.1 is not complete as written. In bounding I1 and E1, the text invokes Slutsky's theorem to pass from e^{cm}P(ξ>m|G)→0 almost surely to e^{cm}E[|G|1{G<0}P(ξ>m|G)]→0, but the random variable |G| is unbounded and, for heavy-tailed G with γ≥1, may have infinite mean; no uniform integrability or dominating function is supplied, and convergence in probability does not imply convergence of the expectation. The earlier bound P(ξ>m−g)≤P(ξ>m) used for negative g is similarly insufficient when E|G| is infinite. The theorem may be true, but the stated lower bound E[G|M>m]≥m+o(m) is not established for the full class of heavy-tailed G admitted by the statement, and the theorem should either add conditions ensuring the conditional expectation is finite or provide a rigorous truncation argument.","section":"§4.2 / Appendix 6.8 (Theorem 4.1)"}],"minor_comments":[{"comment":"The heading 'Proof of lemma 3.3' should read 'Proof of Lemma 4.3'; the lemma is numbered 4.3 in the main text.","section":"§6.6"},{"comment":"In the proof of Lemma 4.2, the displayed asymptotic for αE[ξ|G+ξ>m] omits the leading term (b−1)η^{b−1}/((b−2)m^{b−2}); as printed, multiplying by Δ does not produce the stated equivalent and the proof appears internally inconsistent.","section":"§6.7"},{"comment":"The error term o(1/m^{1−4b}) in Lemma 6.15 should presumably be o(1/m^{4b−1}); the current exponent changes sign for non-integer b and is inconsistent with the expansion it follows.","section":"§6.5, Lemma 6.15"},{"comment":"Table 3 labels the effect of c>0 on E[G|M>m] as negative, but under the stated condition Var(G)>Var(ξ) the conditional expectation still diverges to +∞; the sign refers to the coefficient in the asymptotic equivalent, not to the direction of divergence, and the table should say so explicitly.","section":"§4.3, Table 3"},{"comment":"The overview of the exponential case writes the conditional density as exp(G((x/η)^{b−1}−1)x^{b−2}; the typesetting loses the minus sign in the exponential and the intended exponent, making the density ambiguous compared with the correct definition in Section 6.4.","section":"§4.1"},{"comment":"The identification of 'optimising the proxy' with conditioning on M>m for m→S(M) is a strong mechanism-agnostic modelling choice that is not empirically justified; the paper should state more prominently that its results describe the tail of M under this conditioning model and may not transfer to specific optimisation algorithms.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The paper's most prominently advertised claim is not supported by its own Gaussian analysis, and the heavy-tailed example has a domain-of-validity problem. These are fixable by qualification and restriction, and the remaining material (the Gaussian expansions, the heavy-tailed lower-bound theorem once the proof gap is closed, and the exponential-heavy-tail example for b>2) is a solid contribution. No concerns about citation practice; the self-citation to El-Mhamdi and Hoang (2024) is appropriate given the direct extension."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new result here is Theorem 4.1: with a heavy-tailed goal and a light-tailed discrepancy, E[G|M>m] grows at least like m + o(m) regardless of the dependence structure. That is a clean, valuable extension of the independence-based results in EH24 and KTG24, and it deserves to be cited. The Gaussian analysis in Section 4.3 is also self-contained and correct as far as I checked, and the exponential-goal/heavy-tailed-discrepancy construction is a useful worked example of strong Goodhart emerging from coupling.\n\nThe soft spots are real, and one is load-bearing. The abstract claims that in the light-tailed/light-tailed case, dependence does not change the nature of Goodhart's effect. The paper's own Lemma 4.1 contradicts this. For (G,ξ) ~ N(0,Σ), E[G|M>m] ~ (a+c)/(a+b+2c) m. If a<b, choose c in (-√(ab), -a) (e.g. a=1, b=2, c=-1.4) and the coefficient is negative, so E[G|M>m] → -∞, which is strong Goodhart by the paper's own Table 1. With c=0, the same a,b give a positive coefficient and benign Goodhart. Dependence therefore flips the outcome. This is not a minor wording slip; it undercuts the advertised independence-free selling point for the light-tailed case. Section 5's discussion of 'relative tail thickness' reads like an implicit retreat from the abstract's stronger statement, but the abstract needs to be fixed directly.\n\nThere are two smaller issues. Lemma 4.2's asymptotic (b-1)/(b-2) requires b>2, while the setup allows b in (1,∞); for 1<b<2 the formula gives a negative expectation, which is impossible. The missing domain restriction is easy to repair but should not be left as is. Also, the paper promises Sympy verification code in the supplementary materials, but I could not find the code in the text. That reproducibility claim should be either substantiated or removed.\n\nThe heavy-tailed theorem is solid enough to carry the paper. The Gaussian counterexample means the paper needs major revision before publication, but it is worth refereeing.\n\nMy recommendation: send to peer review, but require the authors to fix the abstract, correct the b>2 restriction, and either provide the Sympy code or drop the claim.","headline":"Useful core theorem on heavy-tailed goals without independence, but the abstract's light-tailed claim is contradicted by the paper's own Gaussian formulas.","tokens_in":31540,"tokens_out":1456,"would_cite":true,"duration_ms":18017,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60E05","62E20","62H20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Proxy-goal dependence, not tail weight alone, settles which Goodhart regime applies.","keywords":["Goodhart's law","proxy metric optimisation","heavy-tailed distributions","light-tailed distributions","dependence structure","tail thickness","benign Goodhart","reward hacking"],"falsifier":"Simulate a heavy-tailed $G$ (for instance Pareto with tail index $1/\\gamma$) and a two-sided light-tailed $\\xi$ with an adversarially chosen dependence (for example $\\xi$ almost surely negative when $G$ is large), set $M=G+\\xi$, and compute $\\mathbb{E}[G\\mid M>m]$ for increasing $m$; Theorem 4.1 predicts this conditional mean is at least $m+o(m)$ for every coupling, so a bounded or decreasing value at large $m$ would refute it. For the Gaussian claim, fit a bivariate normal with $\\mathrm{Var}(G)>\\mathrm{Var}(\\xi)$ and check whether $\\log\\mathrm{Corr}(G,M\\mid M>m)$ versus $\\log m$ has slope $-1$; any different slope would refute the claimed benign decorrelation.","tokens_in":30505,"feed_emoji":"🎯","tokens_out":13004,"duration_ms":120771,"temperature":0.7,"pith_summary":"This paper formalises Goodhart's law—'when a measure becomes a target, it ceases to be a good measure'—in a setting where a proxy metric $M$ is optimised by conditioning on $M>m$ with $m\\to\\infty$, for an intended goal $G$ and a discrepancy $\\xi$ linked by $M=G+\\xi$. Earlier formalisations assumed $G$ and $\\xi$ are independent; this work removes that assumption and asks what tail thickness and coupling jointly imply. It proves that if $G$ is heavy-tailed with a regularly varying survival function and $\\xi$ is light-tailed on both sides, then $\\mathbb{E}[G\\mid M>m]\\ge m+o(m)$ under any dependence structure, so weak and strong Goodhart are impossible. In the light-tailed Gaussian case, dependence produces what the paper calls benign Goodhart: the goal still grows linearly in $m$, but the correlation between goal and proxy decays like $1/m$. In an exponential-goal/heavy-tailed-discrepancy example, dependence turns optimisation into strong Goodhart, with $\\mathbb{E}[G\\mid M>m]\\sim \\eta^{b-1}/m^{b-1}\\to 0$, and a lighter discrepancy tail makes the collapse faster.","feed_headline":"Heavy-tailed goals block weak and strong Goodhart","feed_subtitle":"Even with arbitrary coupling, expected goal grows at least linearly when the discrepancy is light-tailed.","key_machinery":"The carrying object is the conditional expectation $\\mathbb{E}[G\\mid M>m]$ together with the conditional correlation $\\mathrm{Corr}(G,M\\mid M>m)$, evaluated as $m$ approaches the upper support of $M$. Optimisation is modelled mechanism-agnostically as conditioning on the extreme tail event $M>m$. For the heavy-tailed-goal theorem, the proof machinery is a lemma showing that for a two-sided light-tailed variable $\\xi$, the conditional tail $\\mathbb{P}(\\xi>t\\mid G)$ decays almost surely faster than any $e^{-ct}$; this lets the proof cut the integral defining $\\mathbb{E}[G\\mid M>m]$ and show that the slow tail of $G$ dominates. In the Gaussian case, exact Gaussian-tail expansions and conditional moment formulas give explicit asymptotics for the conditional mean, variance, covariance, and hence correlation. In the exponential-goal case, an integration-by-parts expansion for integrals of the form $P(g)\\exp(-g((m-g)/\\eta)^{b-1})$ yields the leading terms.","core_discovery":"The central claim is a tail-thickness trichotomy that survives dependence, plus two dependent examples showing that coupling controls the outcome. Formally, Theorem 4.1 says that whenever the goal's survival function is regularly varying of index $-1/\\gamma$ and the discrepancy is exponentially light on both sides, conditioning on $M>m$ forces $\\mathbb{E}[G\\mid M>m]$ to grow at least as $m+o(m)$, regardless of how $G$ and $\\xi$ are coupled. Theorem 4.2 and Lemma 4.1 show that for Gaussian $(G,\\xi)$ with $\\mathrm{Var}(G)>\\mathrm{Var}(\\xi)$, $\\mathbb{E}[G\\mid M>m]\\sim \\frac{a+c}{a+b+2c}m$ while $\\mathrm{Corr}(G,M\\mid M>m)\\sim \\frac{(a+c)\\sqrt{a+b+2c}}{m\\sqrt{ab-c^2}}\\to 0$: the benign case where the proxy becomes uninformative before the goal stops improving. Lemma 4.3 shows that when $G\\sim\\mathrm{Exp}(1)$ and $\\xi\\mid G$ has conditional density $G\\exp(-G((x/\\eta)^{b-1}-1))x^{b-2}(b-1)/\\eta^{b-1}$ for $x>\\eta$, the same conditioning sends $\\mathbb{E}[G\\mid M>m]$ to zero at rate $m^{1-b}$, with larger $b$ (lighter discrepancy tail) meaning faster collapse. Together these results motivate the paper's formal four-way classification: no, benign, weak, and strong Goodhart.","pith_inferences":["In deployed systems where the discrepancy between a learned reward and the true reward is conditionally heavier when the true reward is small, independence-based Goodhart analyses may be systematically optimistic; measuring the conditional law of $\\xi\\mid G$, not just its marginal tail, would indicate whether collapse is faster than the independent-case formulas predict.","The benign Gaussian case carries a practical warning: using the correlation between proxy and goal as an early-stopping or 'proxy still valid' signal can advise stopping just as the goal is still improving, since the correlation decays like $1/m$.","A natural next test is to take real proxy-goal pairs (for example a learned reward model and a gold-standard reward), estimate the tail index of $G$ and the conditional tail of $\\xi\\mid G$, and compare the measured $\\mathbb{E}[G\\mid M>m]$ along an actual optimisation trajectory to the three asymptotic regimes.","The dependence-free lower bound suggests a design heuristic the paper does not itself propose: making the discrepancy two-sided light-tailed, or ensuring the goal's tail dominates the discrepancy's, would rule out the worst Goodhart outcomes no matter how the two are coupled."],"forward_implications":["If the goal is heavy-tailed with regularly varying survival and the discrepancy is light-tailed on both sides, no dependence structure can create weak or strong Goodhart: the expected goal must grow at least linearly with the optimisation threshold.","In the Gaussian light-tailed regime with $\\mathrm{Var}(G)>\\mathrm{Var}(\\xi)$, a monitor that tracks only the proxy-goal correlation will see it decay toward zero while the goal continues to grow, so correlation collapse is not by itself evidence of goal failure.","In the exponential-goal/heavy-tailed-discrepancy regime, a coupling that associates large discrepancies with small goal values turns an independent-case weak Goodhart situation into strong Goodhart, and a lighter discrepancy tail makes the expected goal fall to zero faster.","The paper's formal classification (no, benign, weak, strong) gives empirical auditors a concrete checklist: estimate the tail classes of goal and discrepancy, estimate their copula, and read off the predicted regime from the asymptotics."],"supporting_citations":[{"why":"Supplies the conditioning-on-$M>m$ formalisation of proxy optimisation and the independence-assumption results this paper relaxes.","marker":"[EH24]"},{"why":"Provides the prior heavy-tailed-reward Goodhart analysis that the paper's heavy-tailed-discrepancy example extends and contrasts with.","marker":"[KTG24]"},{"why":"Contributes the taxonomy of Goodhart variants that motivates the formal no/benign/weak/strong classification.","marker":"[MG18]"},{"why":"Originates the Goodhart statement that the paper formalises.","marker":"[Goo75]"}],"fun_headline_variants":["Goodhart's law survives arbitrary coupling, tails set the strength","Independence-free Goodhart: tail thickness decides benign, weak, strong","Heavy-tailed discrepancy accelerates Goodhart's over-optimization","Goodhart's law formalized without independence or paradigm assumptions","Benign, weak, strong: Goodhart's law depends only on tail thicknesses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper treats optimising a proxy as conditioning on the proxy exceeding a very high threshold $m$; if real optimisation procedures do not drive the proxy into its extreme upper tail this way, the predicted Goodhart regimes need not transfer to practice.","fun_headline_variants_meta":{"raw":{"variants":["Goodhart's law survives arbitrary coupling, tails set the strength","Independence-free Goodhart: tail thickness decides benign, weak, strong","Heavy-tailed discrepancy accelerates Goodhart's over-optimization","Goodhart's law formalized without independence or paradigm assumptions","Benign, weak, strong: Goodhart's law depends only on tail thicknesses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000306,"raw_usage":{"total_tokens":1828,"prompt_tokens":1091,"completion_tokens":737,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":707,"completion_tokens_details":{"reasoning_tokens":645}},"tokens_in":707,"tokens_out":737,"duration_ms":7471,"temperature":1.0,"reasoning_tokens":645,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:45:52.667005+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a heavy-tailed $G$ (for instance Pareto with tail index $1/\\gamma$) and a two-sided light-tailed $\\xi$ with an adversarially chosen dependence (for example $\\xi$ almost surely negative when $G$ is large), set $M=G+\\xi$, and compute $\\mathbb{E}[G\\mid M>m]$ for increasing $m$; Theorem 4.1 predicts this conditional mean is at least $m+o(m)$ for every coupling, so a bounded or decreasing value at large $m$ would refute it. For the Gaussian claim, fit a bivariate normal with $\\mathrm{Var}(G)>\\mathrm{Var}(\\xi)$ and check whether $\\log\\mathrm{Corr}(G,M\\mid M>m)$ versus $\\log m$ has slope $-1$; any different slope would refute the claimed benign decorrelation.","supporting_citations":[],"review_version":1}