{"id":"703c0e3a-c9d8-4c93-84f6-8391f32228b8","arxiv_id":"2608.13461","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A targeted-regularized, doubly robust estimator for causal effects on post-click conversion rates, with theoretical convergence rates and experiments showing gains over existing CVR causal estimators.","lead":"This paper proposes a new doubly robust estimator for the causal effect of a treatment on post-click conversion rate, built from semiparametric influence functions and a targeted regularization loss. It targets e-commerce and advertising settings where clicks and conversions form a two-stage chain and click-only analyses are biased.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The advertised double robustness does not cover a misspecified click model μ1: the Lemma 4.1 remainder includes an unmultiplied (μ̂1−μ1)^2 term, so the §4.2 estimator is inconsistent when only μ1 is wrong.","rationale":"I read the paper's central claim as proposing a doubly robust estimator for the CVR estimand ψ_a, with the advertised guarantee that consistency survives misspecification of one nuisance model. The most load-bearing condition is that the remainder in the von Mises expansion indeed vanishes when at least one nuisance is correct. The paper's own Lemma 4.1 shows this is not true for a misspecified click model µ1: the Rµ term contains (µ̂1−µ1)^2 with no vanishing cross-factor. This is an internal inconsistency, independent of any debate about continuous treatments or Dirac measures, and it directly refutes the abstract's 'one of the nuisance estimators is inconsistent' statement. The reader's rationale already flagged this exact issue, although the reader's named weakest assumption was instead the Dirac-delta/continuous-treatment concern; that Dirac concern is also real, but the µ1 remainder issue is sufficient by itself to reject the headline double-robustness claim. Because my concern reinforces the reader's REJECT verdict rather than moving it, I recommend UNCHANGED. A revision would need to either weaken the double-robustness claim to permit only π or µ2 misspecification, add conditions under which µ1 misspecification is harmless, or show that the practical targeted-regularization estimator has a different bias structure that restores robustness to µ1 errors.","tokens_in":18306,"tokens_out":6313,"duration_ms":66693,"concrete_test":"Analytic check: set π̂=π, µ̂2=µ2, and µ̂1≡c in the §4.2 estimator and in Lemma 4.1, then compute E[ψ̂dr_a] conditional on X. If the derived expression E[µ2(2c−µ1)/c²] is not equal to ψ_a = E[µ2/µ1] for c≠µ1, the double-robustness claim fails. Computational confirmation: generate data from the Appendix A.4 synthetic process with oracle π and µ2, replace µ̂1 by a constant c, and run the one-step estimator for n = 10^3, 10^4, and 10^5; if the estimation error does not shrink to zero as n grows, the misspecified-µ1 case is not covered by the claimed property.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the §4.2 estimator 'remains consistent even if one of the nuisance estimators is inconsistent,' with Theorem 4.2 giving root-n consistency when all three nuisances converge at o_p(n^{−1/4}). The load-bearing condition for double robustness is that the second-order remainder R2 in Lemma 4.1 vanishes when at least one nuisance model is correct. Inspecting R2, the first two integrals are products of (π−π̂) with (µ2−µ̂2) or (µ1−µ̂1), so they vanish if π̂ or the corresponding outcome nuisance is correct. However, the third term Rµ is bounded by C[(µ̂1−µ1)^2 + |µ̂1−µ1||µ̂2−µ2|]; because of the unmultiplied (µ̂1−µ1)^2 term, it does not vanish when only µ̂1 is misspecified. Concretely, set π̂=π, µ̂2=µ2, and µ̂1≡c≠µ1. Then the one-step estimator's expectation is E[µ2(X,a)(2c−µ1(X,a))/c²], which differs from ψ_a = E[µ2(X,a)/µ1(X,a)] by O(1). Thus the estimator is consistent when π alone or µ2 alone is wrong, but not when µ1 alone is wrong. The abstract and Section 4.2 thus state a stronger double-robustness property than Lemma 4.1 supports, directly undermining the paper's headline theoretical contribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a semiparametric doubly robust estimator for the causal effect on post-click conversion rate (CVR), targeting the estimand ψ_a = E[E(Y2(a)|X)/E(Y1(a)|X)] under a chain-structured outcome setting. The estimator is a one-step bias-corrected estimator built from three nuisance functions: the click probability μ1, the conversion probability μ2, and the treatment density π. The headline theoretical claims are that the estimator remains consistent even if one of the nuisance estimators is inconsistent, that it achieves root-n consistency when all nuisances converge at o_P(n^{-1/4}), and that a targeted-regularization framework inherits these properties with a slower but still fast rate. The paper also reports experiments on synthetic, semi-synthetic, and real-world data showing strong empirical performance of the proposed multi-task network.","tokens_in":1736,"tokens_out":3366,"duration_ms":139933,"significance":"If the theoretical claims were correct, the paper would make a useful and nontrivial contribution: a doubly robust estimator for a ratio estimand with chain-structured outcomes, with rigorous rates and a practical neural implementation. The empirical evaluation is extensive, and the targeted-regularization framework is a plausible route to stabilization. However, the central double-robustness claim is contradicted by the paper's own von Mises expansion: when only the click model μ1 is misspecified, the remainder does not vanish, so the estimator is not consistent in that case. In addition, the root-n theory for continuous treatment is not established because the influence function uses a Dirac measure without smoothing. These two issues undermine the paper's main claimed contribution.","major_comments":[{"comment":"The claimed double robustness property—consistency even if one of the nuisance estimators is inconsistent—is not supported by Lemma 4.1. The remainder R2 contains the term ∫ Rμ dP(x) with |Rμ| ≤ C[(μ̂1−μ1)^2 + |μ̂1−μ1||μ̂2−μ2|]. If π̂=π and μ̂2=μ2 but μ̂1 is misspecified (e.g., μ̂1≡c≠μ1), the first two integrals in R2 vanish due to the (π−π̂) factors, but the Rμ term does not vanish because it includes the unmultiplied (μ̂1−μ1)^2. In that setting the population limit of the one-step estimator is E[μ2(X,a)(2c−μ1(X,a))/c²], which differs from ψ_a=E[μ2(X,a)/μ1(X,a)] by an O(1) quantity. Thus the estimator is consistent only if μ1 is correctly specified and at least one of π or μ2 is correct; it is not consistent when only μ1 is wrong. This contradicts the abstract and §4.2 claims and propagates to Theorem 5.1 through the T2 term in its proof.","section":"§4.2, Lemma 4.1, Theorem 4.2"},{"comment":"The theoretical analysis treats the treatment A as continuous while the influence function contains the Dirac measure δ(A=a). For a continuous treatment, the one-step estimator as written is not a well-defined statistic because observations with A=a occur with probability zero, and the empirical-process term (P_n−P){φ_a(P)} would not have finite variance under a Dirac weight. The paper calls δ a 'notational device' but does not provide the smoothing, kernel, or discretization conditions under which a Dirac-based influence function supports root-n consistency. Standard continuous-treatment doubly robust estimators require kernel smoothing and attain slower rates (e.g., Kennedy et al., 2017). Theorem 4.2 therefore overstates the theoretical guarantee for exactly the continuous-treatment setting used in the experiments.","section":"§3 and Theorem 4.2"},{"comment":"The proof of Theorem 5.1 imports the bound ‖ε̂n−ε̌n‖=O_p(n^{-1/3}√log n) from the proof of Lemma 3 of (Nie et al., 2021) with the statement that it is independent of the target estimand, but the targeted-regularization loss R in this paper is a different functional involving μ̂1 in denominators and a cross-product of residuals. The conditions under which the spline approximation result transfers to this modified loss are not verified, so the claimed rate for the targeted-regularized estimator is not rigorously established.","section":"Appendix A.3, Theorem 5.1"}],"minor_comments":[{"comment":"The definitions of μ1 and μ2 have unbalanced parentheses: the displayed formulas are missing a closing parenthesis.","section":"§3"},{"comment":"The 'Ours' row appears to contain repeated entries in the CVR and CTR columns, making the reported AMSE values ambiguous.","section":"Table 2"},{"comment":"The norm ∥·∥ in conditions 2–4 is not specified; the authors should state whether it is the sup-norm or the L2(P) norm, since the proof of the remainder bound uses different norms.","section":"Theorem 4.2"},{"comment":"The sentence 'due to the limitations of publicly available datasets, our method may be underestimated' is vague; please specify which limitation causes the underestimation and how AUUC/QINI relate to the averaged estimand defined in §3.","section":"§6.3"},{"comment":"The paper uses δ(A=a) in the influence function but later mentions 'kernelized IF' for continuous treatment; this connection should be made explicit when the estimator is first introduced to avoid confusion about the role of the Dirac measure.","section":"§4.2"}],"recommendation":"reject","confidential_remarks":"The paper contains a practically oriented framework and a set of experimental results that may be of interest to the community. However, the central theoretical claim—double robustness with respect to any one of the three nuisance estimators—is false, and the failure case (misspecified μ1) is exactly the click model that is often hard to estimate in CVR applications. The continuous-treatment root-n claim also lacks the necessary smoothing assumptions. These are load-bearing errors that cannot be fixed by local editing; the estimator itself would need to be modified or the claims substantially downgraded. I would be open to reviewing a revised version that either constructs a genuinely triply robust estimator or clearly characterizes the actual consistency conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the advertised double-robustness property does not hold. Lemma 4.1's remainder has a term bounded by (μ̂1−μ1)^2 with no multiplicative π or μ2 error, so if π and μ2 are correct but μ1 is misspecified, R2 is O(1) and the one-step estimator has a nonzero bias. The abstract's \"even if one of the nuisance estimators is inconsistent\" is therefore wrong, and this is the paper's central theoretical claim—a load-bearing flaw, not a footnote.\n\nWhat's genuinely useful: the paper identifies a real gap. The ideal-loss literature on CVR prediction optimizes an unbiased loss but gives no guarantee about the final estimand. Framing CVR causal effect as a ratio estimand and building a targeted-regularization estimator on top of it is a reasonable and novel application. The empirical sections are thorough: synthetic, semi-synthetic News, CRITEO-UPLIFTv2, ablations, and a fair comparison with a naive IPS-based loss-debiasing alternative. The proposed method wins, and the ablation shows the targeted regularizer matters. That part is credible and competently done.\n\nSoft spots beyond the DR claim: Theorem 4.2 asserts root-n consistency for continuous treatment using a Dirac-delta influence function. For pointwise dose-response curves, that is not standard; you usually need kernel smoothing and you get slower rates (e.g., Kennedy et al. 2017). The paper calls δ a 'notational device' but never states the regularity conditions under which the Dirac-based IF is legitimate. That is a genuine gap. Also, the DR property, even if corrected, is weaker than advertised: consistent when π alone or μ2 alone is wrong, but not when μ1 alone is wrong. The authors could fix this by reframing the contribution as a targeted-regularization method with rate guarantees for the correctly specified μ1 case, and dropping the 'any single nuisance' claim.\n\nThe math is otherwise standard semiparametric one-step/TMLE machinery. The citations look right—DragonNet, VCNet, Kennedy's review, the ideal-loss papers. No citation issues jump out.\n\nWho is this for? Applied researchers in adtech/e-commerce who want a practical CVR dose-response estimator. They will get a well-tested method, but they should ignore the 'consistent even if μ1 is wrong' slogan.\n\nRecommendation: this deserves a serious referee—the problem is important and the empirical work is substantial—but I would not accept it until the theoretical claims are corrected. If the authors scale back to what the math actually supports, it becomes a solid applied paper. If they keep the DR claim, it's a reject.","headline":"The paper's advertised double robustness fails when the click model μ1 is misspecified; the empirical framework is solid but the central theory needs correction.","tokens_in":19168,"tokens_out":5356,"would_cite":false,"duration_ms":52653,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper constructs a doubly robust semiparametric estimator for causal effects on post-click conversion rate that stays consistent when one of its nuisance models fails.","keywords":["causal effect estimation","post-click conversion rate","doubly robust estimation","semiparametric theory","influence function","von Mises expansion","targeted regularization","continuous treatment"],"falsifier":"Simulate the synthetic continuous-treatment scenario from the appendix with a fixed treatment value $a$ and a known true $\\psi_a$; compute the one-step estimator with cross-fitting, replicate many times, and check whether the empirical RMSE decays at an $n^{-1/2}$ rate without any kernel smoothing or discretization of $\\delta(A=a)$. If the rate is slower or the variance diverges, the Dirac-based influence function is not valid for pointwise continuous-treatment inference.","tokens_in":18041,"feed_emoji":"📈","tokens_out":9727,"duration_ms":92055,"temperature":0.7,"pith_summary":"The paper's claim is that causal effects on post-click conversion rate (CVR) — the probability of conversion among users who would click under a treatment — can be estimated consistently from full observational data, not just from the clicked subset. Existing practice either trains on clicked samples, which biases any causal estimate, or debiases a prediction loss without ensuring the final estimate is unbiased. The authors derive the influence function of the estimand $\\psi_a = \\mathbb{E}[P(Y_2(a)=1 \\mid Y_1(a)=1, X)]$, construct a one-step doubly robust estimator from its von Mises expansion, and prove it is root-n consistent and asymptotically normal when the nuisance functions (click probability, conversion probability, treatment density) converge at $o(n^{-1/4})$ or faster, so the target converges faster than any nuisance estimate. They then turn the correction into a targeted-regularization loss inside a multi-task network, and report that this estimator outperforms standard causal baselines and a naive combination of loss debiasing with causal estimators on synthetic and real data.","feed_headline":"New estimator removes selection bias from CVR causal effects","feed_subtitle":"Semiparametric theory corrects click-selection bias and keeps root-n consistency for the target estimand.","key_machinery":"The load-bearing object is the influence function $\\phi_a$ and its von Mises expansion, which turns the estimation problem into a distributional Taylor expansion: the leading bias of the plug-in ratio $\\mu_2/\\mu_1$ is subtracted off by the correction term, and the leftover remainder is quadratic in nuisance errors. In the practical framework, that same correction is realized as the targeted-regularization term, which trains a low-dimensional spline parameter $\\epsilon(a)$ to make the influence-function correction nearly zero, replacing the hard one-step correction with a soft regularizer and avoiding cross-fitting. The theoretical engine is the remainder bound $R_2 = O(\\|\\pi-\\hat\\pi\\|(\\|\\mu_1-\\hat\\mu_1\\|+\\|\\mu_2-\\hat\\mu_2\\|)+\\|\\mu_1-\\hat\\mu_1\\|^2+\\cdots)$, which translates nuisance convergence rates into the convergence rate of the target estimate.","core_discovery":"The central discovery is a closed-form influence function for the CVR causal estimand: $\\phi_a(P)=\\frac{\\delta(A=a)}{\\pi \\mu_1^2}\\big[(Y_2-\\mu_2)\\mu_1 - (Y_1-\\mu_1)\\mu_2\\big] + \\frac{\\mu_2}{\\mu_1} - \\psi_a(P)$, where $\\pi$ is the treatment density and $\\mu_1, \\mu_2$ are the conditional click and conversion probabilities. From the corresponding von Mises expansion, the paper obtains a one-step estimator whose error is dominated by a second-order remainder consisting of products of nuisance errors, which yields double robustness: the estimator remains consistent if one of the nuisance models is wrong, and it converges at $\\sqrt{n}$ under $o(n^{-1/4})$ nuisance rates. Framing CVR this way, the paper argues, is what previous loss-debiasing CVR work lacks: unbiased loss does not imply an unbiased estimator, whereas the influence-function construction targets the estimand directly. The same remainder analysis is used to design the targeted-regularization loss for the practical multi-task framework.","pith_inferences":["If the Dirac-based gradient is given rigorous footing for continuous treatments, pointwise dose-response curves for chain outcomes would be estimable at root-n without kernel smoothing, a stronger guarantee than standard kernelized continuous-treatment results; a testable middle step is to replace $\\delta(A=a)$ with a kernel and compare the rates.","Because the remainder structure depends only on products of nuisance errors, the estimator should transfer to longer funnel chains (impression to click to purchase, or install to registration) and to ratio-type estimands generally, provided the intermediate event is observed for all units.","The real-data evaluation at the individual level (AUUC/QINI) sits outside the average-level theory; a natural extension is a pseudo-outcome regression for heterogeneous CVR effects, which the paper does not develop.","The comparison with inverse-propensity reweighting of the loss suggests that unbiasedness of a training loss is neither necessary nor sufficient for unbiasedness of an estimand; the same logic could be tested in other selection-biased prediction settings."],"forward_implications":["Platforms can estimate CVR causal effects from observational logs over the full population, so the click-selection bias that affects standard estimators is removed.","Neural networks and other flexible nonparametric models can be used for the nuisance functions without sacrificing the root-n convergence of the target effect estimate.","The final estimate stays consistent when the conversion model or the click model is misspecified, as long as the other nuisance parts are estimated consistently.","The multi-task design gives decision-makers joint estimates of the treatment's effect on clicks and on post-click conversion, so a policy that boosts clicks but hurts conversion can be detected.","Loss-debiased CVR prediction does not by itself yield valid causal estimates; the semiparametric estimator is the component that makes the final quantity trustworthy."],"supporting_citations":[{"why":"Supplies the influence-function and von Mises-expansion framework used to derive the estimator and its remainder.","marker":"Kennedy, 2024"},{"why":"Introduces targeted regularization for neural treatment-effect estimation, the template the paper extends to the CVR estimand.","marker":"Shi et al., 2019"},{"why":"Provides the continuous-treatment machinery: varying-coefficient networks, spline propensity estimation, and the spline approximation bound used in the proof of Theorem 5.1.","marker":"Nie et al., 2021"},{"why":"Originates targeted maximum likelihood estimation, whose perturbation-step idea the targeted-regularization loss adapts.","marker":"Van der Laan et al., 2011"},{"why":"Supplies the double/debiased machine-learning framework and the cross-fitting and nuisance-rate assumptions underlying root-n consistency.","marker":"Chernozhukov et al., 2018"},{"why":"Shows that deep neural network estimators can achieve the $o(n^{-1/4})$ nuisance convergence rates required by Theorem 4.2.","marker":"Farrell et al., 2021"},{"why":"Demonstrates that clicked-only CVR effect estimates mislead and provides the baseline the experiments must outperform.","marker":"Huang et al., 2024"}],"fun_headline_variants":["Double-robust CVR estimator kills selection bias","CVR causal effects without click-bias distortion","Influence-function estimator debiases CVR causally","Targeted regularization yields double-robust CVR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument hinges on treating the point-mass weight $\\delta(A=a)$ in the influence function as a legitimate mathematical object for a continuous treatment; if that pointwise influence function requires smoothing to be valid, the root-n consistency claim at a fixed treatment value does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Double-robust CVR estimator kills selection bias","CVR causal effects without click-bias distortion","Influence-function estimator debiases CVR causally","Targeted regularization yields double-robust CVR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000511,"raw_usage":{"total_tokens":2528,"prompt_tokens":1032,"completion_tokens":1496,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":1433}},"tokens_in":648,"tokens_out":1496,"duration_ms":10910,"temperature":1.0,"reasoning_tokens":1433,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:31:19.123798+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the synthetic continuous-treatment scenario from the appendix with a fixed treatment value $a$ and a known true $\\psi_a$; compute the one-step estimator with cross-fitting, replicate many times, and check whether the empirical RMSE decays at an $n^{-1/2}$ rate without any kernel smoothing or discretization of $\\delta(A=a)$. If the rate is slower or the variance diverges, the Dirac-based influence function is not valid for pointwise continuous-treatment inference.","supporting_citations":[],"review_version":1}