{"id":"a6d0723a-f682-4443-9648-653ae51ff0e0","arxiv_id":"2607.25241","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"LURE is a multiply robust, asymptotically normal estimator for off-policy evaluation when the true action is hidden and only a noisy surrogate is observed, using the next state as the identifying proxy.","lead":"Off-policy reinforcement-learning estimators assume the actions in historical logs are correct; this paper treats them as noisy proxies for hidden true actions and builds a multiply robust estimator, LURE, that uses the next state as an internal proxy to recover the value of a target policy. The result matters because real-world datasets—electronic health records, transaction logs—routinely misrecord the action actually taken.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's identification is outsourced to an unverified self-cited preprint; no proof that Assumption 3.1 alone identifies the misclassification matrix.","rationale":"The reader's weakest_assumption was Assumption 3.1(iii), the conditional-independence assumption. That is indeed a strong and important assumption, but it is a modeling assumption on which the paper's theorems are explicitly conditional, so it is not an internal gap in the argument. The more load-bearing issue for the central claim is that the identification step itself is not actually demonstrated in the paper: Theorem 1's proof delegates the key identifiability of the misclassification matrix to a self-cited, non-archived preprint and does not reproduce the argument. If that external result has hidden conditions or does not apply to this MDP setting, V(π) is not shown to be identifiable and the entire LURE construction loses its target. This is an omitted proof at the root of the logical chain, and it is explicitly flagged by the manuscript's own citation pattern. I therefore focus the stress-test on this step. The concrete test—an independent derivation of identifiability from Assumption 3.1—would settle the concern: if it succeeds, the paper only needs to include the proof; if it fails, the central claim is unsupported. Either way, the reader's CONDITIONAL verdict remains appropriate, so no change to the verdict is needed.","tokens_in":39584,"tokens_out":17135,"duration_ms":182009,"concrete_test":"Independently re-derive identification of Pr(Ã|A,S) from Assumption 3.1. Concretely, write the observed conditional density of (Ã,R,S') given S as a two-component mixture with Ã⊥R⊥S'|S,A, and show that the 2×2 matrices E[h_X(X)|S,A] for X=Ã,R,S' are identified up to the label switch resolved by Assumption 3.1(iv). If the derivation requires a completeness or support condition not stated in Assumption 3.1, Theorem 1 is not proven; if the derivation succeeds, the concern is resolved. Also check whether Zhou and Tchetgen Tchetgen (2024) Theorem 4 uses the same conditional independence with one binary surrogate and one continuous proxy, and whether it requires any additional assumption absent here.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The logical foundation of the paper is Theorem 1 (identification of V(π)). Its proof in Appendix A.2 does not contain an identification argument: it asserts that Zhou and Tchetgen Tchetgen (2024) proved Pr(Ã=a'|A=a,S=s) is identifiable under Assumption 3.1, then inverts the 2×2 misclassification matrix in Eq. (10). The cited paper is a self-cited arXiv preprint with a static hidden-treatment structure; the present paper never states which of its conditions map to Assumption 3.1, nor does it reproduce the argument. If that external result requires an additional proxy or completeness condition not included here, then V(π) may not be identifiable from O=(S,Ã,R,S') and every subsequent theorem is conditional on an unsupported premise. This is not a matter of whether Assumption 3.1 is realistic; it is an omitted proof step at the root of the central claim. The paper's own proof of Lemma 2 also explicitly relies on Zhou and Tchetgen Tchetgen (2024), but the key identification step is never shown.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies off-policy evaluation (OPE) in infinite-horizon discounted Markov decision processes when the true action A is unobserved and only a noisy surrogate \\tilde A is recorded. Under Assumption 3.1 — in particular the conditional independence R ⊥⊥ \\tilde A ⊥⊥ S' | (S,A) — the paper claims identification of the policy value V(π) using the next state as a proxy. It then derives an observed-data influence function (Theorem 2), constructs the cross-fitted LURE estimator, proves a second-order error bound and multiple robustness (Theorem 3, Corollary 1), and establishes asymptotic normality (Theorem 4). A novel EM-style algorithm is proposed to estimate the latent-action nuisance components. The method is evaluated in simulations (tabular, continuous-state, CartPole) and on a MIMIC-III sepsis analysis. The central claim is that LURE is the first OPE method that is valid when recorded actions are only noisy proxies for the true actions.","tokens_in":39928,"tokens_out":5985,"duration_ms":62304,"significance":"If the identification and asymptotic results are correct, the paper addresses a genuinely important gap: standard OPE methods can be severely biased when actions are misclassified, and the paper provides a principled alternative. The influence-function construction is nontrivial and the multiple-robustness decomposition in Theorem 3 is a strength: the remainder is explicitly bounded by products of nuisance-error terms, and the four consistency cases in Corollary 1 are clearly delineated. The paper also provides extensive simulations and a real-data application. However, the central identification theorem is not proved in the manuscript: it delegates the key step to an external self-cited preprint, and the practical nuisance estimator (Algorithm 2) is not given any convergence analysis. These are load-bearing gaps that prevent the results from being fully verified as stated.","major_comments":[{"comment":"The proof of Theorem 1 does not contain an identification argument. It states that 'Zhou and Tchetgen Tchetgen (2024) proved that Pr(˜A=a'|A=a,S=s) is identifiable' and then inverts the 2×2 misclassification matrix. No conditions from that external preprint are stated, no mapping to Assumption 3.1 is provided, and no argument is reproduced. If the cited result requires an additional proxy or a completeness condition not implied by Assumption 3.1, then V(π) may not be identifiable from O=(S,˜A,R,S'). Since Theorem 1 underpins every subsequent theorem, this is a load-bearing omission. The authors should either give a self-contained proof of identifiability of the misclassification matrix under Assumption 3.1, or state and prove a precise theorem from Zhou and Tchetgen Tchetgen (2024) and verify all its conditions.","section":"Appendix A.2 (Proof of Theorem 1)"},{"comment":"Lemma 2 is the core identity behind the influence function: E[g_a(R,˜A,S)|S,A] = E[g'_a(S',˜A,S)|S,A] = 1(A=a). The proof is explicitly 'based on the proof of Theorem 4' of the same self-cited preprint and does not show the calculation for the product contrast. While the identity is plausible and can be verified by direct algebra under Assumption 3.1, the manuscript does not actually prove it. Given that Lemma 2 is essential for Theorem 2 and for the cancellation of the linear drift in Theorem 3, the proof should be self-contained.","section":"Lemma 2 (Appendix A.1)"},{"comment":"Theorems 3 and 4 assume rate conditions on the nuisance estimators (α_min > 1/4, label-selection condition, etc.), but no theory is supplied for the proposed EM-style algorithm (Algorithm 2). It is not shown that the EM iterates converge to the true posterior η(O;a), nor that the resulting weighted regressions achieve the assumed rates. Since LURE is implemented with this algorithm, the asymptotic theory is conditional on an unverified premise about a central algorithmic component. The authors should either prove convergence/rates under explicit conditions, or clearly state the rate conditions as assumptions on the generic nuisance estimators and provide guidance on when Algorithm 2 satisfies them. At minimum, a consistency analysis of the EM procedure is needed.","section":"Section 5 / Algorithm 2"},{"comment":"The conditional independence R ⊥⊥ ˜A ⊥⊥ S' | (S,A) is strong and is a key premise for Lemma 2 and Theorem 1. The paper's motivating EHR example (delayed charting, documentation errors) naturally raises the possibility that recording error is correlated with unmeasured patient acuity or with the reward/transition beyond the true action. The paper offers no sensitivity analysis or discussion of how violations of this assumption affect LURE. Given that this assumption is load-bearing for identification, the authors should at least provide a sensitivity model (e.g., allowing bounded residual dependence) or a clear statement of the limits of the method. Without this, the practical scope of the central claim is not quantified.","section":"Assumption 3.1(iii)"}],"minor_comments":[{"comment":"Coverage is reported only for LURE because the baselines are biased. It would be helpful to also report bias and RMSE for all methods in a table, so readers can assess the magnitude of the bias reduction and the relative efficiency of LURE across misclassification rates.","section":"Section 6.1, Table 1 and Figures 3–5"},{"comment":"The choice of l(S') is not formalized in the theory. The simulations select the coordinate of S' with largest partial correlation with ˜A given S. This data-dependent selection is not accounted for in the influence-function derivation. If l is treated as fixed in the theory, state this explicitly; if it is selected from data, discuss the impact on inference.","section":"Section 4, proxy selection"},{"comment":"The matrix P_{˜A,A}(s) is defined with columns corresponding to A=0 and A=1. The inversion step is correct under Assumption 3.1(iv), but the notation could be clearer about the column ordering to avoid confusion.","section":"Appendix A.2, Equation (10)"},{"comment":"Theorem 3 defines L_n = max_a ||δθ~A(·,a)||^κ, but the proof of label-selection error uses the L2 norm and derives O_p(||δθ~A||^κ). The role of κ is not pinned down; in Theorem 4 the condition is stated as max_a ||δθ~A||^κ = o_p(n^{-1/2}) for some κ>0. Since this condition is effectively a rate condition on ||δθ~A||, it would be clearer to state it directly in terms of the L2 error rate and avoid the unspecified κ.","section":"Theorem 3 and Appendix A.4"},{"comment":"The discussion does not mention the conditional-independence assumption's limitations or the EM convergence issue. A short paragraph on these two points would help readers understand the scope of the claims.","section":"Section 8 / Discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially a solid contribution to OPE with misclassified actions, and the influence-function/remainder analysis is careful. However, the identification theorem is outsourced to a self-cited unpublished preprint, and the proposed EM estimator is not supported by convergence theory. Both are load-bearing and must be fixed before the claims are credible. The self-citation pattern (Theorem 1 and Lemma 2 both rely on the co-author's preprint) is a concern from an editorial standpoint; it is not itself grounds for rejection, but it increases the need for a fully self-contained proof. If the authors can supply the missing identification proof and either prove or transparently assume the EM rates, the paper could be suitable for publication. I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper is worth engaging with: it studies a real and underappreciated problem—off-policy evaluation when recorded actions are noisy surrogates—and the LURE estimator is a genuine extension of the hidden-treatment idea to infinite-horizon RL. The next-state variable does double duty as proxy and continuation, and the influence function has a product form that isn't in the static paper. The multiple-robustness theorem and the asymptotic expansion look carefully done; the product-error remainder is the right structure, and the simulations show a clean bias reduction across tabular, continuous, and Gym settings.\n\nThe soft spot is the one the stress-test flagged, and I think it's real. Appendix A.2 does not prove identification from Assumption 3.1. It says Zhou and Tchetgen Tchetgen (2024) proved the misclassification matrix is identifiable and then inverts it. No conditions from that preprint are restated, and no argument is given that Assumption 3.1 alone is sufficient. That's a load-bearing citation at the root of Theorem 1. It may well be true—the static result is plausible and Hu (2008) has related arguments—but the paper needs to show it, not cite it. If the external result needs an extra proxy or completeness condition, the whole edifice is conditional.\n\nEverything else is softer. The EM pipeline for nuisance estimation is a heuristic; no convergence proof and no rates, yet Theorem 4 assumes rates. The proxy-coordinate selection is data-driven but not accounted for in the inference. No code is released. Assumption 3.1(iii) is strong and there's no sensitivity analysis. Those are standard reviewer asks, not fatal.\n\nNet: the central idea is defensible, and the influence-function machinery looks right. The identification gap is the one thing that could sink it, but it's fixable. I'd send this to peer review, with instructions to require a self-contained identification proof and code. I'd also want the authors to state which conditions of the cited preprint are being imported.\n\nI'd cite this if I worked on hidden-action OPE, and I'd bring it to reading group.","headline":"Worth engaging with, but the core identification is outsourced to a self-cited preprint and needs to be brought into the paper before I trust the theorems.","tokens_in":40363,"tokens_out":2369,"would_cite":true,"duration_ms":24544,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under a conditional-independence condition, the value of a target policy in offline reinforcement learning is identifiable and estimable even when the recorded actions are only noisy proxies for hidden true actions.","keywords":["off-policy evaluation","hidden actions","latent action","multiply robust estimation","influence function","offline reinforcement learning","action misclassification","sequential decision making"],"falsifier":"Simulate an MDP in which the misclassification probability $Pr(\\tilde A \\neq A \\mid S, A)$ depends on an unobserved acuity variable that also shifts the reward, while keeping the other assumptions intact; if LURE's estimate is biased or its confidence intervals undercover, the central claim fails. Concretely, set $R = \\theta_R(S,A) + U$ and $logit Pr(\\tilde A \\neq A) = \\nu U$ with $U$ a latent state correlate, then check whether the estimator still centers on the true value as $\\nu$ grows.","tokens_in":39453,"feed_emoji":"🎯","tokens_out":4570,"duration_ms":48831,"temperature":0.7,"texified_at":"2026-08-05T21:45:47.262569+00:00","pith_summary":"The paper takes on offline reinforcement learning in the setting where the actions that actually drove the system were never recorded; only noisy surrogates are available. It claims that, under a conditional-independence condition, the value of a target policy is uniquely identified from the observed states, surrogate actions, rewards, and next states. The proposed estimator, LURE, replaces the missing action indicators with contrast weights built from the surrogate action, reward, and next state, and is multiply robust and asymptotically normal. If correct, this gives practitioners a way to evaluate policies from error-prone logs such as electronic health records without assuming the recorded action is the delivered action.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":6637,"prompt_tokens":694,"completion_tokens":5943,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":694,"completion_tokens_details":{"reasoning_tokens":5314}},"feed_headline":"Hidden actions no longer block offline RL value estimation","feed_subtitle":"A new estimator stays valid when recorded actions are only noisy proxies for true actions.","key_machinery":"The central object is the observed-data influence function $\\phi_\\pi(O)$ for the policy value. It replaces the missing action indicator $1(A=a)$ with product contrast weights: $g_a$ uses the surrogate action and reward, and $g'_a$ uses the surrogate action and a feature of the next state, each normalized by the difference between action-specific conditional means. These weights are constructed so that, conditional on the true state and action, their expectation is exactly $1(A=a)$; the next state plays the role of the extra proxy variable that makes this replacement possible. A weighted EM-style algorithm iterates between estimating latent-action posterior probabilities and re-estimating the nuisance functi","core_discovery":"The paper establishes that in an infinite-horizon discounted Markov decision process with binary hidden actions, the policy value $V(\\pi)$ is identified from the observed distribution of $(S, \\tilde A, R, S')$ under coverage, relevance, conditional independence, and a surrogate-informativeness assumption. The next state $S'$ serves as the crucial second proxy: together with the reward $R$ and the surrogate action $\\tilde A$, it permits construction of functions $g_a$ and $g'_a$ whose conditional expectations equal the unobserved indicators $1(A=a)$. The estimator LURE is the sample mean of the corresponding observed-data influence function, with nuisance functions estimated by a weighted EM-style procedure an","pith_inferences":["My inference: the conditional-independence assumption — that reward, surrogate action, and next state are independent given the true state and action — is the spot most likely to fail in practice, for instance when documentation delay is correlated with patient severity. A sensitivity analysis that perturbs this independence and reports how LURE's estimate drifts would be a natural companion to th","My inference: the same contrast-weight construction could be adapted to other error structures, such as a coarsened or aggregated version of a continuous action, or misclassification probabilities that depend on an auxiliary recorded covariate.","My inference: the label-alignment step relies on a separation condition; with limited data or weak separation, labels could switch across cross-fitting folds and inflate variance. Diagnostics for label-switching would be a practical safeguard for real deployments."],"forward_implications":["LURE gives asymptotically valid confidence intervals for the value of a target policy when only surrogate actions are recorded, enabling hypothesis tests and policy comparisons with quantified uncertainty.","The multiple-consistency structure means a practitioner can obtain consistent estimates from several different combinations of correctly specified nuisance components, such as the density ratio and surrogate-action model without a correct reward model.","The framework extends standard doubly robust off-policy evaluation to a measurement-error setting, replacing the assumption of perfectly recorded actions with a conditional-independence assumption about the proxies.","In clinical datasets where documented treatments may differ from delivered ones, using LURE can change which treatment policy is preferred relative to methods that treat the documented action as the delivered action, as the sepsis application illustrates."],"fun_headline_variants":["Offline RL value estimation survives hidden actions","LURE: robust offline RL with unobserved actions","Hidden actions no longer break offline RL","New estimator handles noisy action proxies in offline RL","Offline RL: learning from proxies when true actions are hidden"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is Assumption 3.1(iii): conditional on the true state and action, the recorded surrogate action, the reward, and the next state are mutually independent.","fun_headline_variants_meta":{"raw":{"variants":["Offline RL value estimation survives hidden actions","LURE: robust offline RL with unobserved actions","Hidden actions no longer break offline RL","New estimator handles noisy action proxies in offline RL","Offline RL: learning from proxies when true actions are hidden"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1080,"prompt_tokens":668,"completion_tokens":412,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":340}},"tokens_in":412,"tokens_out":412,"duration_ms":4109,"temperature":1.0,"reasoning_tokens":340,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:00:55.894417+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an MDP in which the misclassification probability $Pr(\\tilde A \\neq A \\mid S, A)$ depends on an unobserved acuity variable that also shifts the reward, while keeping the other assumptions intact; if LURE's estimate is biased or its confidence intervals undercover, the central claim fails. Concretely, set $R = \\theta_R(S,A) + U$ and $logit Pr(\\tilde A \\neq A) = \\nu U$ with $U$ a latent state correlate, then check whether the estimator still centers on the true value as $\\nu$ grows.","supporting_citations":[],"review_version":1}