{"id":"ecacec51-d853-4800-ad49-590820cc1f02","arxiv_id":"2607.07893","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"A discounted reach-cost Bellman operator is contractive, has a unique fixed point identical to the HJ reachability value function, and can be approximated by sample-based RL.","lead":"The paper derives a non-additive Bellman operator whose unique fixed point exactly equals a discounted reachability value function from Hamilton-Jacobi theory, and proves discounting makes the operator contractive. This gives a semantics-preserving bridge so reinforcement learning can approximate reachable sets without rewriting the safety meaning of the value function.","discovery_kind":"unification","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies the paper’s strongest claim and correctly notes that Assumptions 2 and 8 are used for regularity and viscosity theory. However, those assumptions are not load-bearing for the contraction itself; the contraction follows from the discount factor alone. Because the equivalence theorems rest on both the contraction (Banach) and the classical viscosity uniqueness results, and both pieces are standard under the stated hypotheses, no internal inconsistency or hidden gap threatens the central claim. The numerical experiments are consistent and the open FVI question is already acknowledged. Consequently the reader’s CONDITIONAL verdict with low correctness risk remains appropriate; no adjustment is warranted.","tokens_in":26975,"tokens_out":459,"duration_ms":4730,"concrete_test":"Independently re-derive the contraction inequality of Lemma 9 from Definition 5 using only the non-expansiveness of inf and the 1-Lipschitz property of min; confirm that the factor e^{-λσ} appears solely from the continuation branch and that no trajectory-Lipschitz estimate is invoked. If the derivation succeeds without Assumption 2, the contraction claim is secure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the non-additive operators T_{σ,λ} and T^∞_{σ,λ} are contractions whose unique fixed points coincide with the viscosity solutions of the corresponding HJVIs—holds under the paper’s stated hypotheses. The contraction proofs (Lemmas 9 and 12) rely only on the 1-Lipschitz property of min and the factor e^{-λσ}<1; they do not require Lipschitz continuity of f. Lipschitz continuity of f and boundedness of g are used for regularity of W (Lemmas 3–6) and for the viscosity characterization (Theorems 1–2), which are classical under those assumptions. The equivalence theorems (Theorems 7–8) therefore rest on standard, well-supported ingredients. The open FVI convergence question (Remark 8) is correctly flagged by the authors and does not undermine the operator-theoretic result.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper derives a non-additive Bellman operator for the semantics-preserving discounted reach-cost value function of Choi et al., proves that discounting renders the operator a contraction on a complete space of bounded functions, and shows that its unique fixed point coincides with the unique viscosity solution of the associated Hamilton–Jacobi variational inequality (finite- and infinite-horizon). Discrete-time consistent approximations and a fitted-value-iteration scheme are developed, and RL is interpreted as sample-based approximation of the same fixed-point equation. Numerical experiments on a double-integrator reach problem and a Dubins-car avoid problem show close agreement with HJ reference solutions and alignment of zero level sets.","tokens_in":27190,"tokens_out":1074,"duration_ms":26655,"significance":"If the results hold, the paper closes a genuine gap between HJ reachability and RL: prior discounted formulations either sacrificed exact reachability semantics for contraction (e.g., MDR) or preserved semantics on the HJ side without a matching Bellman fixed-point characterization. The operator-theoretic bridge (Definitions 5 and 7; Theorems 3–8) is clean, uses standard Banach and viscosity tools under classical Lipschitz/boundedness assumptions, and is complementary to concurrent travel-cost formulations. The explicit flagging of the open FVI convergence question (Remark 8) is a strength. The contribution is primarily theoretical; empirical support is limited to low-dimensional systems but is consistent with the claims that are actually proven.","major_comments":[{"comment":"Lemma 9 (and the parallel Lemma 12): the written contraction proof treats the rollout horizon as σ throughout and pulls out the factor e^{-λσ}, but Definition 5 uses h(τ)=min{τ,σ}. When τ<σ the continuation term lands on the fixed boundary Ψ(0,·)=g(·), so the difference is actually zero; Remark 4 correctly notes that the boundary condition is essential for uniform contraction near τ=0. The claim is true, but the proof as written does not case-split on h(τ). Please rewrite Lemma 9 (and the discrete analogue Lemma 14) with an explicit case analysis so that the argument is self-contained and matches Definition 5.","section":null},{"comment":"Abstract, Introduction, and Section VIII: the paper repeatedly frames the contribution as uniting HJ reachability with reinforcement learning and as enabling “scalable, data-driven computation … in high-dimensional systems.” What is actually constructed is fitted value iteration with exact rollouts of a known dynamics model (explicitly identified as ADP in Section VIII), and Remark 8 correctly leaves open whether the composite “regress-then-apply-operator” map converges to the discrete fixed point. The operator-theoretic core (Theorems 3–8) does not depend on this claim, but the abstract and title overstate what is proven. Soften the RL/scalability language to match the proven fixed-point equivalence and the acknowledged open FVI question, or supply partial error bounds / high-dimensional evidence.","section":null}],"minor_comments":[{"comment":"Table I is a useful running reference; consider adding a one-line pointer to it at the first appearance of W_Bell / ĉW / Ψ_θ so readers do not miss the tier structure.","section":null},{"comment":"Section IX-B: for the double integrator the infinite-horizon reachable set fills the domain, so zero-level-set comparison is impossible; this is explained, but a short finite-horizon double-integrator panel (or a reach-avoid variant) would make the zero-level-set claim more uniform across both examples.","section":null},{"comment":"Eq. (42) / Definition 5: the notation U[0,h(τ)] vs U[0,σ] is slightly inconsistent with the later discrete operators; unify the control-interval notation.","section":null},{"comment":"Proposition 3 writes V(x) rather than V(t,x); fix the argument list for consistency with Definition 1.","section":null},{"comment":"Algorithms 1–2 are clear; a brief note on how U_d is chosen (grid density, effect on the min) would help reproducibility of the numerical scheme in Section VII.","section":null},{"comment":"Related-work discussion of Hsu et al. and the concurrent travel-cost paper [27] is fair; a single sentence clarifying that the present operator is non-additive min-structured (not a running-cost integral) would further reduce possible confusion with additive discounted formulations.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The concurrent self-citations ([27], [28]) are complementary rather than load-bearing; no novelty-disclosure concern. The manuscript is a solid theory contribution for eess.SY / control journals; the main risk is overclaiming the RL bridge relative to the open FVI question. Minor revision that tightens Lemma 9 and softens the abstract should be sufficient."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that this paper finally supplies the missing operator-theoretic link. Choi et al. already gave a discounted reach-based value that keeps the original sign semantics and satisfies an HJVI. What was missing was a Bellman operator whose unique fixed point is exactly that same function. El-Hajj et al. derive it (finite- and infinite-horizon versions), prove it is a contraction under discounting despite the non-additive min structure, and show the fixed point coincides with the viscosity solution. That is the real contribution.\n\nThe math is standard Banach-plus-viscosity work and looks carefully done under the usual global Lipschitz and bounded-cost assumptions. The discrete consistency lemmas and the explicit FVI algorithms are clean. Experiments on the double integrator and Dubins car show pointwise agreement with a semi-Lagrangian HJ solver and, more importantly, aligned zero level sets, so the reachability semantics survive. They also correctly flag the open question of whether fitted value iteration with function approximation converges to the discrete fixed point.\n\nSoft spots are real but proportionate. The FVI convergence gap is left open (they say so). Experiments stay low-dimensional; there is no high-dim demonstration that would actually stress the “scalable” claim. No code is released. The free parameters (λ, σ, SIREN frequencies) matter in practice. None of this undercuts the central operator result.\n\nThis is for people working at the HJ–RL interface or on safety certificates in continuous control. The theory is solid enough that a serious referee should see it. I would bring it to reading group and would cite the operator and equivalence theorems if I am writing in this area. Send it to peer review.","headline":"Clean, usable completion of the HJ–RL bridge: a contractive non-additive Bellman operator whose fixed point exactly matches Choi’s semantics-preserving discounted reach value.","tokens_in":27794,"tokens_out":444,"would_cite":true,"duration_ms":12958,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A non-additive Bellman operator for discounted reach-cost has a unique fixed point that is exactly the Hamilton-Jacobi reachability value function.","keywords":["Hamilton-Jacobi reachability","non-additive Bellman operator","reinforcement learning","reach cost","safety-critical control","viscosity solution","discounted value function","fixed-point theorem"],"falsifier":"On a low-dimensional system whose true Hamilton-Jacobi value function can be computed accurately (for example the double-integrator or Dubins-car examples already used in the paper), run fitted value iteration with the proposed operator and check whether the learned zero level set fails to align with the reference zero level set, or whether the pointwise residual fails to go to zero as the discretization parameters vanish.","tokens_in":27890,"feed_emoji":"🎯","tokens_out":905,"duration_ms":8752,"temperature":0.7,"pith_summary":"Hamilton-Jacobi reachability gives formal safety and reachability certificates for continuous-time systems, but grid-based solvers hit the curse of dimensionality. Reinforcement learning scales better, yet standard RL is built for additive cumulative rewards, while reachability is a non-additive stopping problem. Earlier discounted formulations either changed the meaning of the reachable set or never produced a true Bellman fixed-point equation. This paper starts from a discounted reach-cost value function that keeps the original sign-based semantics and constructs a matching non-additive Bellman operator. Discounting makes the operator a contraction, so Banach's theorem yields a unique fixed point that coincides with the Hamilton-Jacobi viscosity solution. Reinforcement learning is then simply a sample-based method for solving that same fixed-point equation, which means learned value functions can still be read as rigorous reachable sets.","feed_headline":"Non-additive Bellman operator matches Hamilton-Jacobi reachability","feed_subtitle":"Discounting makes the operator contractive; its fixed point is exactly the safety value function.","key_machinery":"The reachability-preserving Bellman operator T_{σ,λ} (and its infinite-horizon counterpart). It replaces the usual sum of rewards by a min of a stopping branch and a discounted continuation branch; the discount factor e^{-λσ} makes the whole map a contraction whose unique fixed point is the true reachability value function.","core_discovery":"The discounted reach-cost value function that already appears in the Hamilton-Jacobi literature is exactly the unique fixed point of a non-additive Bellman operator. Discounting alone is enough to make that operator contractive on a complete space of bounded functions, so existence, uniqueness and convergence of value iteration follow, and the Hamilton-Jacobi and Bellman characterizations are identical.","pith_inferences":["The same construction should extend almost immediately to reach-avoid problems by swapping the min for a suitable max-min structure, yielding a single operator that encodes both safety and liveness.","Once the contraction property is established, residual-based post-hoc certification methods already developed for additive Bellman operators become available for these non-additive reachability operators.","The explicit separation of Bellman step σ from integration step Δt suggests a practical schedule that trades contraction speed against numerical consistency, which could be optimized automatically during learning."],"forward_implications":["Learned reachability value functions can be interpreted as rigorous safety certificates rather than heuristic scores.","Value iteration and fitted-value methods become legitimate numerical solvers for continuous-time reachable sets once the non-additive operator is used.","The same contraction argument applies to both finite-horizon and infinite-horizon stationary problems, giving a uniform theory.","High-dimensional systems that are intractable for grid-based Hamilton-Jacobi solvers become candidates for data-driven reachable-set computation while preserving semantics."],"fun_headline_variants":["Discounted reach-cost value is unique fixed point of non-additive Bellman","Non-additive Bellman fixed point equals Hamilton-Jacobi reachability value","Contractive non-additive Bellman unites HJ reachability and RL","Semantics-preserving Bellman operator matches HJ discounted reach-cost","RL samples the fixed point of a reachability-preserving Bellman operator"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The system dynamics must be globally Lipschitz continuous in the state, uniformly in the control, and the reach-cost function must be bounded; without those two conditions the contraction proof and the uniqueness of the fixed point no longer hold on the whole space.","fun_headline_variants_meta":{"raw":{"variants":["Discounted reach-cost value is unique fixed point of non-additive Bellman","Non-additive Bellman fixed point equals Hamilton-Jacobi reachability value","Contractive non-additive Bellman unites HJ reachability and RL","Semantics-preserving Bellman operator matches HJ discounted reach-cost","RL samples the fixed point of a reachability-preserving Bellman operator"]},"model":"grok-4.5","effort":"low","cost_usd":0.003588,"raw_usage":{"total_tokens":1240,"prompt_tokens":867,"num_sources_used":0,"completion_tokens":99,"cost_in_usd_ticks":35880000,"prompt_tokens_details":{"text_tokens":867,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":274,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":867,"tokens_out":99,"duration_ms":3338,"temperature":1.0,"reasoning_tokens":274,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T15:53:29.723849+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a low-dimensional system whose true Hamilton-Jacobi value function can be computed accurately (for example the double-integrator or Dubins-car examples already used in the paper), run fitted value iteration with the proposed operator and check whether the learned zero level set fails to align with the reference zero level set, or whether the pointwise residual fails to go to zero as the discretization parameters vanish.","supporting_citations":[],"review_version":1}