{"id":"532908d2-649f-4f4d-bb04-dbf78820b330","arxiv_id":"2608.08158","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A unified framework and review of dynamic reward shaping, with a new taxonomy over twelve method families and a gap analysis of optimality guarantees.","lead":"This paper proposes a unified taxonomy for dynamic reward shaping, separating additive shaping from reward replacement and other adaptive guidance, and reviewing twelve method families along temporal, informational, and theoretical axes. It consolidates existing theoretical guarantees and highlights that no performance-driven method in the surveyed set preserves exact optimality.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"G1 classification for value-derived T2 entries is not reconciled with the paper's own Remark 1(iii), so the 'G1 only under T2' observation may rest on mislabeled guarantee classes.","rationale":"The reader identified candidate-set completeness as the weakest assumption, and that is indeed a genuine limitation. However, the more immediate threat to the central claim is internal: the paper's own audit standard in Remark 1(iii) says that potentials depending on value-function parameters do not imply original-MDP invariance, yet Table 3 labels two such T2 entries as G1. This is not a matter of searching more broadly; it is a matter of whether the reviewed evidence, as classified, supports the Figure 5 pattern. The paper's careful hedging about the T3-G1 cell being a property of the reviewed set does not repair the classification if the G1 labels themselves are inconsistent with the stated theoretical framework. The recommended check is therefore a focused re-audit of the two value-derived G1 entries, followed by a recomputation of the T2-G1 count. If the audit confirms the labels, the reader's ACCEPT stands; if not, the central observation should be revised or the entries downgraded, which would make acceptance conditional on that correction.","tokens_in":44916,"tokens_out":17607,"duration_ms":176142,"concrete_test":"Obtain Grześ and Kudenko (2010) and Adamczyk et al. (2025); for each, determine whether the invariance or convergence theorem is stated for the original environment MDP over S or only for an augmented process that includes the learner's value estimates or abstract-value parameters. If the latter, reclassify the entries from G1 to G3 (or mark ‡) under Section 3.4's standard, then recompute Figure 5's T2-G1 row. If the row's remaining G1 entries are only PBIM/ADOPS/BAMDP-meta/potential-based difference rewards, the claim survives in weaker form; if the row becomes empty, the claim 'exact optimality preservation occurs only under experience-driven adaptation' is not supported by the reviewed evidence as labelled.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2.4, Remark 1 case (iii) states that when a shaping potential depends on learner-internal quantities (value-function parameters, replay contents, current policy), invariance in the original stationary MDP over S does not follow and should not be claimed; the guarantee is only over an augmented learner-environment process. Table 2's Theorem 2 row echoes this by restricting original-S guarantees to potentials depending solely on k or a Markov sufficient statistic. Yet Table 3 assigns G1 to two T2/I3 C1 entries whose potentials are exactly learner-internal: 'Potential from an abstract-state value function' (Online shaping-reward learning) and 'Potential set to the current V estimate' (Bootstrapped shaping). Section 4.2 reports only 'convergence results for the tabular setting' for the latter, and no separate theorem is stated that would bridge Remark 1(iii). If these G1 assignments are not backed by an original-MDP proof in the cited papers, the Figure 5 T2-G1 count is inflated and the central observation that exact optimality preservation occurs only under experience-driven adaptation is weakened.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This review paper proposes a unified analytical framework for dynamic reward shaping and neighbouring adaptive reward mechanisms in reinforcement learning. It formalises dynamic shaping as an information-conditioned process (Eqs. 2-4), separates parametric revision of a shaping rule from state-dependent variation of a fixed rule, and introduces four mechanism classes (C1 additive shaping, C2 reward replacement, C3 redistribution/relabelling, C4 reward-adjacent guidance), three taxonomy dimensions (temporal signatures T1-T5 plus S, information sources I1-I9, guarantee classes G1-G5), and a classification of twelve method families in Table 3. The central findings, explicitly framed as observations about the reviewed set rather than theorems, are that among reviewed C1 methods exact optimality preservation (G1) arises only under experience-driven adaptation (T2), that no performance-driven (T3) method attains G1 (Figure 5), and that agent-internal estimates (I3) are the most frequent information source. The paper also provides an implementation audit of replay, bootstrapped critics, multi-step returns, terminal conditions, and reward normalisation (Section 5.4), a practical decision guide (Table 5), evaluation reporting standards (Table 7), and thirteen open research directions.","tokens_in":45178,"tokens_out":33015,"duration_ms":296230,"significance":"Should the classification be accepted, the framework is a genuinely useful organising device for a fragmented literature, and the paper's self-auditing machinery is exemplary in conception: Table 2 records exact conclusions and information-timing assumptions per theorem; unverifiable claims are marked with a double dagger and excluded from Figure 5's counts; the DPBA entry is reclassified to G5† with its documented refutation; Remark 2 and Figure 9 give a clean separation of index pairing, faithful reconstruction, and target stationarity under replay; and the search protocol's limits are disclosed in unusual detail (Section 1.1, Appendix A.1), with all counts explicitly scoped to the reviewed candidate pool. The taxonomy's empty cells are stated as falsifiable claims about that pool, with the broader compliance audit explicitly delegated to future work, and the historical narrative of Section 5.1 is honestly hedged as a plausible reading.","major_comments":[{"comment":"Table 3 and Remark 1(iii) are in direct tension. Table 3 assigns G1 to “Online shaping-reward learning” (Grześ and Kudenko 2010) and “Bootstrapped shaping” (Adamczyk et al. 2025), both C1/T2/I3, and Figure 5 counts them among the six T2-G1 entries this review can stand behind. Both potentials, however, are read from the learner's own concurrently learned value estimates, which is exactly the learner-internal case Remark 1(iii) excludes: “Invariance in the original stationary MDP over S does not follow and should not be claimed.” Table 2's Theorem 2 row likewise restricts original-S guarantees to potentials depending solely on k or a Markov sufficient statistic, and the paper states that establishing stationarity of a learner-augmented process is “a further, non-trivial step that this review does not take.” The two entries are absent from Table 2, so the audit trail promised by that table's header (“a G1-G3 assignment in Table 3 can be checked against its source”) does not exist for precisely the assignments at issue. Section 4.2 reports for Adamczyk et al. only “convergence results for the tabular setting,” which under the paper's own definitions (G1 structural preservation, G2 exact recovery in the limit, G3 other formal result) is a G2/G3-type claim rather than a G1 claim. Unless the cited papers prove an original-MDP invariance theorem for the online update process, the G1 labels are earned only by the frozen-potential sub-claim (a trivial application of Theorem 1), which does not certify the dynamic process to which the T2 label refers. There is also an asymmetry of audit standards: GRM and VLM-guided potentials are excluded from Figure 5 because their G1 claims could not be independently verified, whereas Remark 1(iii) affirmatively says these two claims should not be made, which is a stronger reason for exclusion. The authors should either supply the missing bridging theorem, or reclassify the entries (e.g., as G2 for the tabular convergence result) and propagate the change to Figure 5's counts, Section 3.5's first pattern, and Table 5's safety-critical row, which currently recommends value-derived potentials as a “G1 (structural)” route.","section":"§3.4/Table 3 vs. §2.4 Remark 1(iii)"},{"comment":"The G1 category is applied over different decision processes without Table 3 recording which. For the static and time-indexed PBRS rows, G1 is structural preservation over the original MDP S. For BAMDP shaping, the G1 meta-RL row is a claim over the belief/history-augmented BAMDP (“necessary and sufficient… for the shaped BAMDP's optimal algorithm to remain Bayes-optimal”), as Table 2 correctly documents, while ADOPS's G1 is conditional on behaviourally enforced tie/dominance conditions computed from the learner's own estimates (Eq. 9). Section 3.5's sentence “every G1 entry adapts through experience-driven read-outs (T2)” therefore aggregates guarantees stated at different levels, and Section 3.4's “Basis of the audited G3 and G4 assignments” paragraph supplies the evidentiary basis for G3/G4 but not for the G1 rows. I recommend (i) adding a column or footnote to Table 3 stating the decision process over which each G1/G2/G3 claim is made (original S, augmented (s,k), belief/history, BAMDP, product MDP), and (ii) adding the remaining G1 rows (Grześ and Kudenko 2010; Adamczyk et al. 2025; Devlin et al. 2014) to Table 2, so that the pattern statements of Section 3.5 and the “guarantee reachable” column of Table 5 are checkable from the paper's own audit tables.","section":"§3.3-3.5, Table 3"}],"minor_comments":[{"comment":"The phrase “The empty T1 column” is imprecise: Figure 5's T1 row contains one G3 entry (heuristic-guided RL), so what is empty is the T1-G1 cell, not the whole column; rewording would prevent a misreading of the count.","section":"§3.5"},{"comment":"The inclusion criterion “sharpens a class boundary in Section 2.2” is partly circular, since the C1-C4 classes are the framework's own contribution; restating it operationally (for example, as a method that the literature describes as shaping but that classifies differently under Section 2.2) would make the protocol more reproducible.","section":"§1.1, App. A.1"},{"comment":"The replay-recomputation distinction (mismatched pairing versus consistent current-potential relabelling) is explained three times: Remark 2(ii), Section 4.11, and Section 5.4/Figure 9. A forward cross-reference from Section 4.11 to Section 5.4 would shorten the repeated discussion without loss of precision.","section":"§4.11"}],"recommendation":"major_revision","confidential_remarks":"This is a well-disciplined survey whose principal risk is the internal inconsistency described in Major Comment 1; because the authors' own audit apparatus (Table 2, the dagger/double-dagger conventions, the “Basis of the audited G3 and G4 assignments” paragraph) provides a clear template for the fix, I expect the revision to be tractable. The novelty disclosure is handled properly: the taxonomy is explicitly labelled as the review's own apparatus rather than borrowed terminology, and the two self-citations appear only as application examples in Section 6. Fit with the journal's scope is good. If the G1 assignments are reclassified or backed by the missing theorems, and the counts in Figure 5 and Section 3.5 are updated accordingly, I would support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it twice. The taxonomy is a genuine contribution: the C1-C4 mechanism classes, the T1-T5 temporal signatures, the I1-I9 information sources, and the G1-G5 guarantee ladder give the field a shared vocabulary it has been missing. The distinction between parametric revision of a shaping rule and mere state-dependent evaluation of a fixed rule is the right axis, and the paper applies it consistently. The explicit marking of unverified claims with double-dagger symbols, the assumption audit in Table 2, and the replay analysis in Figure 9 are all done carefully. Credit where earned: this is one of the more self-aware reviews I have seen in the shaping literature.\n\nBut the stress-test note is right, and it lands on the paper's own central empirical observation. Section 2.4, Remark 1(iii) says that when a potential depends on value-function parameters, replay contents, or the current policy, invariance in the original stationary MDP over S does not follow and should not be claimed. The G1 definition in Section 3.4 says G1 means a Theorem 1 or Theorem 2-style transformation preserves the optimal-policy set, and Table 2 restricts original-S Theorem 2 guarantees to potentials depending only on k or a Markov sufficient statistic. Yet Table 3 assigns G1 to Online shaping-reward learning and Bootstrapped shaping, both of which use learner-internal potentials. The paper cites only tabular convergence for Adamczyk et al. and offers no separate theorem bridging Remark 1(iii). So two or three entries in the T2-G1 cell of Figure 5 are not backed by the paper's own standard. That does not destroy the taxonomy, but it does inflate the headline pattern and makes the 'exact optimality only under experience-driven adaptation' claim look stronger than the evidence.\n\nThe literature-selection limitation is real but disclosed and not fatal. The empty T3-G1 cell is an observation about a curated candidate pool, and the paper says so. The self-citations appear only as application examples in Section 6 and do not support the framework, so no issue there.\n\nWho is this for? Researchers working in reward shaping, intrinsic motivation, or reward-model alignment will get a useful map and a consistent vocabulary. The paper deserves a serious referee, but not as-is. I would send it to review with a clear instruction to check the consistency between Remark 1(iii), the G1 definition, and the Table 3 entries, and to either provide the missing proofs, reclassify the affected entries, or restrict the claims to the augmented learner-environment process. With that fixed, I would be happy to see it published.","headline":"A useful, carefully hedged taxonomy with a real internal inconsistency: the G1 labels for value-derived T2 entries contradict the paper's own Remark 1(iii), and the central 'G1 only under T2' claim is weaker than the table suggests.","tokens_in":45627,"tokens_out":2819,"would_cite":true,"duration_ms":33459,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a unified taxonomy for dynamic reward shaping and shows that, across the methods it reviews, exact optimality preservation arises only when the shaping signal is revised by accumulated experience, never when it is tuned…","keywords":["dynamic reward shaping","reward shaping taxonomy","potential-based reward shaping","policy invariance","optimality preservation","experience-driven adaptation","performance-driven adaptation","reinforcement learning"],"falsifier":"Run a systematic, query-logged database search for peer-reviewed additive reward-shaping methods whose signal is revised by an outer loop against measured task return and whose paper proves preservation of the exact optimal-policy set; finding one, or constructing a tabular MDP with a known optimal policy in which a performance-driven update rule provably satisfies the chronological potential pairing at every update, would refute the observed emptiness of the T3–G1 cell.","tokens_in":44737,"feed_emoji":"🧭","tokens_out":13004,"duration_ms":111590,"temperature":0.7,"pith_summary":"Classical reward shaping is provably safe only when the shaping signal is a fixed potential, yet contemporary reinforcement-learning systems revise their guidance as training proceeds, and no shared vocabulary existed for comparing those adaptive methods. This review supplies one: it classifies each mechanism by how it relates to the task reward (additive shaping, replacement, redistribution, or reward-adjacent), how it changes over training, what information drives the change, and what guarantee it actually carries. Applying the taxonomy to twelve method families yields the central observation: among the additive-shaping methods this review can stand behind, exact preservation of the optimal policy occurs only under experience-driven adaptation, while no performance-driven method achieves it. The framework matters because it separates genuine structural guarantees from objective-aware heuristics, and it pinpoints where guarantees survive, or quietly break, in modern deep-RL pipelines with replay buffers, bootstrapped critics, and reward normalisation. The paper's open challenge to the field is a theory of how fast a shaping signal may change before learning becomes unstable.","feed_headline":"Provably safe reward shaping is always experience-driven","feed_subtitle":"A twelve-family taxonomy shows where exact optimality guarantees survive in deep RL and where they break","key_machinery":"The framework's load-bearing object is a four-part taxonomy. Mechanism classes C1–C4 separate additive dynamic shaping (the only class that inherits the invariance theorems) from reward replacement, reward redistribution or relabelling, and reward-adjacent guidance; temporal signatures T1–T5 plus S separate schedule-, experience-, performance-, structure-, and interaction-driven revision from mere state dependence; information sources I1–I9 name what drives the revision; and guarantee classes G1–G5 rank what is actually proved, from exact structural preservation down to purely empirical reports. The identity carrying the theoretical content is the telescoping pairing of time-indexed potentials, $F_k(s,a,s') = \\gamma\\Phi_{k+1}(s') - \\Phi_k(s)$, whose discounted sum collapses to a boundary term that is the same constant for every policy, provided the potential version at departure and at arrival are chronological. On the implementation side the machinery is a trilemma: replay storage cannot simultaneously keep the index pairing intact, preserve the reward the transition was actually given, and give the critic a stationary regression target; the review catalogues storage policies and which property each forfeits.","core_discovery":"The paper claims that training-time revision of reward guidance is a distinct, organising phenomenon worth its own analytical apparatus, and that a single information-conditioned picture captures it: the learner consumes $R'_k = R + F_{\\theta_k}$ while an update rule $\\theta_{k+1} = U(\\theta_k, \\mathcal{I}_{k+1})$ revises the shaping term as information accumulates. Within that picture the decisive distinction is parametric revision, the rule itself changes, versus state-dependent variation, a fixed rule evaluated on a moving argument such as a belief or a plan index; only the former counts as dynamic. The review then consolidates the theory governing safety: the time-indexed potential form $F_k(s,a,s') = \\gamma\\Phi_{k+1}(s') - \\Phi_k(s)$ preserves the optimal-policy set when the paired potentials are chronological, and that guarantee belongs only to additive dynamic shaping proper (class C1). Its measured finding, over the twenty-one audited C1 methods, is that every exact-preservation (G1) entry adapts through experience (T2), such as value estimates, visitation statistics, or correction terms, while the cell pairing performance-driven adaptation (T3) with G1 is empty; the paper is explicit that this describes its reviewed set, not a proof of impossibility. It further argues that in deep-RL practice the guarantee is conditional: replay recomputation, bootstrapped critics fed by the same value function that supplies the potential, terminal potentials, and reward normalisation each create conditions the classical theorems do not cover.","pith_inferences":["If the empty T3–G1 cell survives a broader systematic audit, it would amount to a design law for the field: a performance-driven outer loop either forgoes exact preservation or must impose an additional structural constraint, such as forcing every update to land on a valid potential, a step the paper leaves as an open question rather than a theorem.","The reward-versus-advantage argument, made in the paper for human feedback, plausibly extends to intrinsic-motivation and value-derived signals, which are also defined relative to the learner's current behaviour; whether treating them as advantages rather than rewards explains observed instabilities is a testable hypothesis.","The replay trilemma predicts a concrete experimental outcome no reviewed study isolates: storing paired potentials with each transition versus recomputing both terms under a current potential should differ measurably in off-policy deep RL, since the two policies trade the chronological identity against static structure in different ways.","A rate-of-change theory connecting potential-update frequency to critic stability could turn the central tuning knob of dynamic shaping into an analyzable quantity; a natural conjecture is an analogue of target-network lag, bounding how many critic updates may separate consecutive potential versions."],"forward_implications":["A designer who needs exact optimality preservation should derive the shaping signal from the agent's own experience, such as value estimates, visitations, or prediction errors, and cast it as a consistently paired potential; an outer loop that tunes the signal against measured return buys measured performance but, in the reviewed record, not preservation.","The empty T3–G1 cell points to the route most likely to combine rich information sources with exact guarantees: constrain the learned object to an invariant form and let any information source supply its content, learning the potential rather than the reward.","Implementation is where guarantees die: recomputing only one term of a replayed transition from a later potential breaks the telescoping identity, while consistently recomputing both terms under one current potential replaces the dynamic construction by a sequence of static problems.","Claims of a benefit from dynamism should rest on a frozen-adaptation baseline alongside an unshaped baseline and the best static potential; without the frozen baseline, an apparent gain may come from the content of the shaping signal rather than from adaptation itself.","The taxonomy converts method families into a design space where sparse cells are informative: interaction-driven adaptation appears in the reviewed set only as reward replacement or policy-level guidance, never as additive shaping."],"supporting_citations":[{"why":"Establishes the baseline theorem: fixed potential-based shaping $F(s,a,s') = \\gamma\\Phi(s') - \\Phi(s)$ preserves the optimal-policy set, the safety standard every dynamic method is measured against.","marker":"[Ng et al., 1999]"},{"why":"Extends invariance to time-indexed potentials revised once per step; the construction behind the G1 assignments for experience-driven dynamic shaping.","marker":"[Devlin and Kudenko, 2012]"},{"why":"Bootstrapped shaping uses the current value estimate directly as the potential; the archetypal T2/G1 entry and the taxonomy's worked example.","marker":"[Adamczyk et al., 2025]"},{"why":"PBIM converts intrinsic-motivation bonuses into potential form, showing experience-driven adaptation can be made exactly optimality-preserving, and supplies the gridworld evidence of uncorrected bonuses changing the optimum.","marker":"[Forbes et al., 2024a]"},{"why":"ADOPS relaxes the future-agnostic condition with an online correction from the learner's own estimates, extending G1 to action-dependent intrinsic returns.","marker":"[Forbes et al., 2025]"},{"why":"Refutes the claimed G1 invariance of dynamic potential-based advice and supplies the PIES replacement; the basis of the review's reclassification of that entry and of its caution about auditing guarantee claims.","marker":"[Behboudian et al., 2020, 2022]"},{"why":"Bi-level shaping-weight optimisation that adjusts the signal against measured true reward; the representative T3/G4 method behind the observed empty T3–G1 cell.","marker":"[Hu et al., 2020]"},{"why":"Shows human evaluative feedback is policy-dependent and best treated as an advantage, grounding the claim that interaction-driven adaptation is not an MDP reward.","marker":"[MacGlashan et al., 2017]"},{"why":"Belief reward shaping whose prior influence decays with evidence; the G2 exemplar distinguishing exact recovery in the limit from exact preservation.","marker":"[Marom and Rosman, 2018]"},{"why":"Bayes-adaptive formulation in which potential-based pseudo-rewards are necessary and sufficient for Bayes-optimality in meta-RL and eventually approximately optimal in ordinary RL.","marker":"[Lidayan et al., 2025]"}],"fun_headline_variants":["Reward shaping safety hinges on experience, not performance","All provably safe dynamic shaping methods are experience-driven","New taxonomy maps where reward shaping guarantees survive","In deep RL, reward shaping optimality is conditional","Dynamic reward shaping unified under information-driven lens"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's central patterns, especially the empty cell joining performance-driven adaptation with exact preservation, depend on the candidate set being representative of the field, but that set was assembled by citation tracing from a theory seed rather than a systematic database search, so the empty cells could be artifacts of selection rather than properties of the field.","fun_headline_variants_meta":{"raw":{"variants":["Reward shaping safety hinges on experience, not performance","All provably safe dynamic shaping methods are experience-driven","New taxonomy maps where reward shaping guarantees survive","In deep RL, reward shaping optimality is conditional","Dynamic reward shaping unified under information-driven lens"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001014,"raw_usage":{"total_tokens":4351,"prompt_tokens":1082,"completion_tokens":3269,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":3197}},"tokens_in":698,"tokens_out":3269,"duration_ms":24998,"temperature":1.0,"reasoning_tokens":3197,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:19:15.075470+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic, query-logged database search for peer-reviewed additive reward-shaping methods whose signal is revised by an outer loop against measured task return and whose paper proves preservation of the exact optimal-policy set; finding one, or constructing a tabular MDP with a known optimal policy in which a performance-driven update rule provably satisfies the chronological potential pairing at every update, would refute the observed emptiness of the T3–G1 cell.","supporting_citations":[],"review_version":1}