{"id":"8c7d2709-7ae7-4323-ab8d-b43692a26a7d","arxiv_id":"2608.00915","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"UpliftBench finds that on IHDP Qini shows no detectable alignment with effect accuracy (+0.07 rank correlation, CI includes zero) while AUUC is consistently more aligned, and that on Jobs ranking metrics fail at sign-threshold policy selection.","lead":"This paper introduces UpliftBench, a benchmark for uplift modeling, and reports two metric-mismatch findings: Qini can misrank models on continuous outcomes where AUUC does not, and ranking metrics alone cannot select good sign-threshold policies. The work quantifies the cost of choosing the wrong evaluation metric and offers a practical remedy, threshold calibration.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"F2's Jobs selection-regret CI excludes zero only for split dependence ρ̄≤0.12, yet observed cross-split risk correlations are +0.41; empirical F2 is not yet statistically supported.","rationale":"The paper's F1 is well-supported: it persists across tuning-off/AUUC-tuned arms, estimator exclusions, base-learner swaps, and the all-100 extension; the paired AUUC-over-Qini gap remains positive with CIs excluding zero even with DR/R removed (+0.28/+0.34 on all 100). The affine-invariance lemma and the counterexample are clean. So my concern is not with F1. It is with F2. The empirical magnitude is the load-bearing part of F2 because Proposition 5.3 alone (rank metrics are invariant to score shifts while sign policies are not) is a known non-identifiability statement; the '14–15% regret' number and the 'risk-selection beats random while ranking metrics do not' contrast are what make the objective-mismatch point concrete. That contrast is computed across ten overlapping re-splits of one sample, and the paper's own sensitivity shows the CI excludes zero only for between-split dependence ρ̄ ≤ 0.12. The observed raw correlation of per-model risk vectors across splits is +0.41, well above the threshold; although that raw correlation includes genuine model signal, the paper does not estimate ρ̄ of the differenced statistic, so the possibility that the true dependence exceeds 0.12 is not excluded. In that case the reported CIs understate uncertainty and the empirical F2 half may be a within-sample artifact. The four 'convergent' diagnostics all derive from the same overlapping observations and are not independent. I would not reject the paper: the disclosure is exemplary, the artifact is reproducible, and the structural argument is sound. But I would condition acceptance on either (a) a design-adjusted inference that estimates ρ̄ directly and shows the result survives, or (b) a disjoint Jobs evaluation. Otherwise the F2 finding should be reported purely as a descriptive illustration, not as a measured operational cost.","tokens_in":36099,"tokens_out":13195,"duration_ms":116132,"concrete_test":"Recompute the F2 cross-repeat regret 95% CI with a design-effect correction using an estimate of ρ̄ obtained from the released per-split prediction artifacts: compute the mean pairwise correlation of the per-split selection-evaluation gains across the ten splits, then rescale the cluster-bootstrap standard error by sqrt(1 + 9*ρ̄_hat). If the CI for the risk-selection gain includes zero at the estimated ρ̄_hat, the empirical F2 claim does not survive the paper's own dependence concern.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing concern is the empirical half of Finding 2 (Jobs, Section 5.2). The cluster bootstrap treats the ten released Jobs re-splits as exchangeable clusters, but these are re-partitions of one LaLonde sample, so the effective independent information may be far smaller than ten. The paper's own design-effect sensitivity (Appendix E) states that the 95% regret CI for direct policy-risk selection excludes zero only for average between-split dependence ρ̄ ≤ 0.12 (Qini/AUUC; 0.06 for uplift-at-k), and the per-model risk vectors across splits already correlate at +0.41 (Section 5.2). If the true dependence of the differenced selection-evaluation statistic is near that raw correlation, the claimed 'lower benchmark regret than random' vanishes. The four supporting diagnostics (rotation regret, sign pattern, negative rank correlations, selector contrast) all read the same overlapping observations, so their convergence is not independent confirmation. The structural boundary (Prop. 5.3) is unaffected but is a standard non-identifiability point; the empirical magnitude on Jobs is what makes F2 a finding. The paper is honest about this fragility and reads F2 as 'modest', but the headline claim 'direct policy-risk selection yields lower benchmark regret than random while Qini, AUUC, and uplift-at-k do not' is not robust to the paper's own observed dependence level.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces UpliftBench, an outer-test-isolated benchmark of 12 uplift estimators across seven dataset families, scored by six objectives including ranking metrics, effect accuracy, and policy objectives. It reports two main findings. F1: on the semi-synthetic continuous IHDP benchmark, the unnormalized Qini ranking shows no detectable alignment with true effect accuracy (mean rank correlation +0.07, 95% CI [-0.03, +0.16] over all 100 IHDP realizations), while AUUC and uplift-at-k remain informative; the paired AUUC-over-Qini gap is +0.49 [+0.40, +0.59] and survives estimator exclusions, base-learner swaps, and tuning-objective changes. F2: on the Jobs benchmark, ranking metrics are structurally insufficient for sign-threshold policies because they discard score level (Proposition 5.3), and a within-sample split-rotation analysis suggests direct policy-risk selection has lower benchmark regret than Qini/AUUC/uplift-at-k selection, with threshold calibration removing 81% of Qini-selection regret. Both findings are explicitly bounded: F1 does not replicate on ACIC or Revenue-Synthetic, and F2 vanishes under a budgeted-value objective. The paper releases versioned loaders, protocol code, result artifacts, and a leaderboard.","tokens_in":36376,"tokens_out":8439,"duration_ms":82477,"significance":"The paper is a careful, well-scoped empirical study of a question that matters for benchmark design: metric choice, not just estimator quality, can drive published conclusions. F1's core statistic is strong: the paired AUUC-over-Qini gap is positive across all 100 IHDP realizations with confidence intervals excluding zero, and the result is tested against a wide set of confounders, including a Qini-tuned candidate panel, estimator exclusions, and implementation variants. The reference objectives are external ground-truth quantities, so the comparison is not circular. F2's structural proposition is correct, and the threshold-calibration remedy is practically actionable. The empirical half of F2 is appropriately labeled in the body as a descriptive within-sample case study, and the paper honestly discloses the sensitivity of its confidence intervals to between-split dependence. The reproducibility package appears unusually complete, with versioned data loaders, fixed protocols, git-hash-stamped result schemas, and documented regeneration targets.","major_comments":[],"minor_comments":[{"comment":"The abstract states as an empirical result that direct policy-risk selection 'yields lower benchmark regret than random model selection while Qini, AUUC, and uplift-at-k do not', but the body's own design-effect sensitivity shows the regret interval excludes zero only for between-split dependence rho-bar <= 0.12, with raw cross-split risk correlations around +0.41. I recommend using the same descriptive-case-study register in the abstract and contribution list as in Section 5.2, so the headline matches the disclosed evidential strength.","section":"Abstract and Contribution (4)"},{"comment":"The family-level ordering of the composite R (IHDP 3.52, ACIC 1.07, Revenue-Synthetic 0.39) is described as a 'one-in-six chance event'; because this comparison was selected after the gradient was observed, the one-in-six framing overstates its evidential value. The text later notes the absence of pre-registration credit, but the earlier sentence should be rephrased to avoid a post-hoc probability claim.","section":"Appendix E.1"},{"comment":"The statement that the 10K and 100K Criteo orderings 'did not agree' is based on a single pair of non-nested stratified subsamples, and the paper correctly interprets this as a resolution limitation. The wording could nonetheless more sharply distinguish 'limited resolution' from 'evidence of misranking' in the opening sentences of the resolution paragraph.","section":"Appendix I"},{"comment":"The labels 'Student-t3' and 'Log-n.' render awkwardly and should be typeset with proper mathematical notation, e.g. Student-t_3 and log-normal, to match journal conventions.","section":"Figure 3 and Table 8"},{"comment":"The sentence 'F1 is an artifact of neither metric's selection' is supported by the tuning-off and AUUC-tuning arms, but the reader must wait until Appendix E.2 for the actual numbers. Consider moving the one-sentence summary of Table 5 into the main text, since the Qini-tuned panel is a natural initial concern.","section":"Section 5.1 and Table 6"}],"recommendation":"minor_revision","confidential_remarks":"The paper is unusually honest about its limitations, and the F2 empirical fragility is a scope-and-wording issue rather than a correctness problem. I support publication after the abstract is aligned with the body's descriptive framing of the Jobs result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful benchmark paper, and F1 is the real contribution. On IHDP's continuous outcomes, the Qini coefficient's model ranking shows no detectable alignment with effect accuracy (mean +0.07, 95% CI [-0.03, +0.16] across all 100 realizations), while AUUC and uplift-at-k track it. The paired AUUC-over-Qini gap survives tuning-off, AUUC-tuning, XGBoost base learner, DR/R exclusion, and causalml implementation checks. That is new, practically important, and carefully bounded: the paper shows the failure is not a tail effect and appears on only one of three continuous families, with the mechanism left open. The lemma (Qini is treated-count-weighted AUUC) is a clean structural observation.\n\nF2 is softer. Prop 5.3 (rank metrics are invariant to score shifts, sign-threshold policies are not) is elementary but correct. The empirical claim — direct policy-risk selection beats random on Jobs while Qini/AUUC/uplift-at-k do not — depends on ten re-splits of one LaLonde sample treated as exchangeable clusters. The paper's own design-effect sensitivity shows the 95% CI excludes zero only for between-split dependence ρ̄ ≤ 0.12, while raw cross-split risk correlations are +0.41. So that specific claim is not robust to plausible split dependence, and the stress-test note lands. The paper says exactly this in Section 5.2 and reads the result as modest, but the abstract states the flat version. The calibration result (81% regret reduction) is a useful, less fragile takeaway.\n\nWhat earns credit: the artifact. Versioned loaders, git-hashed result parquets, documented reproduction targets, control arms for the main threat (Qini-tuned panel), external reference objectives, honest disclosure of propensity-nesting imperfections and F2's within-sample nature. The citation pattern is fine; closest prior work is explicitly addressed.\n\nThe main soft spot is framing, not evidence: F2's headline overstates a within-sample descriptive result whose CI vanishes under the paper's own dependence levels. A serious revision should add a disjoint validation protocol (listed as future work) or explicitly demote that claim. This deserves peer review: the reproducibility package and F1 alone justify refereeing, and F2 can be fixed or restated.","headline":"F1 is a solid, carefully-controlled empirical finding on Qini's silent failure on continuous outcomes; F2's empirical half is fragile but honestly disclosed, and the artifact deserves a serious referee.","tokens_in":36883,"tokens_out":2930,"would_cite":true,"duration_ms":28801,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On the IHDP benchmark, the Qini ranking shows no detectable alignment with effect accuracy while AUUC and uplift-at-k stay informative.","keywords":["uplift modeling","heterogeneous treatment effects","causal machine learning","evaluation metrics","benchmark design","model selection","policy evaluation","Qini coefficient"],"falsifier":"Estimate the actual between-split dependence of the Jobs policy-risk estimates directly, for example by a random-effects variance decomposition, and re-run the cross-repeat selection rotation; if the estimated dependence exceeds the threshold at which the paper's interval still excludes zero, the selection-regret claim includes zero and its empirical half collapses to the structural argument. For F1, a controlled sweep of the paper's own R-composite predictor across new continuous-outcome families with known effects would settle whether any ex-ante quantity can predict when Qini decouples from effect accuracy.","tokens_in":2413,"feed_emoji":"📊","tokens_out":8766,"duration_ms":161768,"temperature":0.7,"pith_summary":"This paper builds UpliftBench, an outer-test-isolated benchmark of 12 uplift estimators across 25 instances from seven dataset families, to test whether published disagreements about the best uplift model come from the models or from the metric. It establishes that the disagreement is largely about metrics: on the continuous IHDP benchmark, the Qini coefficient's model ranking shows no detectable correlation with effect accuracy (mean rank correlation +0.07, 95% CI [-0.03, +0.16] across all 100 realizations), while AUUC and uplift-at-k stay informative. On the binary Jobs study, ranking metrics fail to select models for sign-threshold policies because they discard the score level, and calibrating the decision threshold removes 81% of the Qini-selection regret. Both findings are bounded: F1 appears on one of three continuous families, and F2 vanishes under a budgeted-value objective.","feed_headline":"Uplift benchmark fights are about the metric, not the models","feed_subtitle":"Qini's model ranking misses effect accuracy on IHDP; on Jobs, calibrated thresholds cut 81% of selection regret.","key_machinery":"The argument is carried by four objects. (1) The unnormalized Qini coefficient, defined by the count-corrected cumulative gain g(k)=sum y_i t_i - (sum y_i (1-t_i)) T_k/C_k integrated against a chord baseline; Lemma 5.1 proves Qini is the treated-count-weighted prefix-mean AUUC, g(k)=T_k u(k), and Lemma 5.2 separates the score into a depth-weighted contrast plus an interleaving term, localizing the IHDP divergence to treated-count weighting. (2) The rank-invariance boundary (Proposition 5.3): any metric that depends on the model only through the induced ranking is invariant to strictly increasing transformations, including shifts, whereas the sign-threshold policy pi(x)=1[tau(x)>=0] is not, so ranking metrics cannot identify the threshold-optimal model from rank information alone. (3) The multi-regime benchmark design: seven dataset families with different outcome types and reference objectives, scored under outer-test isolation with cluster bootstraps over benchmark realizations as the inferential unit. (4) A hand-checkable 12-unit counterexample where a single large control outcome makes Qini prefer a near-random model while AUUC and effect accuracy prefer the better model.","core_discovery":"The central claim is that the choice of evaluation metric, not the estimators themselves, explains much of the disagreement among published uplift benchmarks. Two distinct failures are identified. On IHDP (continuous outcomes, known effects), the unnormalized Qini's model ranking does not track effect accuracy: across all 100 realizations its mean rank correlation with effect accuracy is +0.07 [-0.03, +0.16], and as pairwise concordance it is a coin flip (0.52 [0.49, 0.56]), while AUUC shows a paired advantage of +0.49 [+0.40, +0.59]. Structurally, Qini is treated-count-weighted AUUC (Lemma 5.1), so the divergence localizes to treated-count weighting. On Jobs (binary, experimentally identified policy objective), every ranking metric correlates negatively with policy value, and metric selection is indistinguishable from random at 14-15% regret, while direct policy-risk selection beats random; this follows a rank-invariance boundary (Proposition 5.3), and calibrated thresholds remove 81% of the Qini-selection regret. The paper frames both as bounded, benchmark-discovered failure signatures, not universal laws.","pith_inferences":["A testable extension implied by the paper's Lemma-based decomposition: any cumulative-sum functional that weights by treated count, not just Qini, should show the same decoupling on IHDP, while mean-based and depth-weighted functionals should stay aligned; rescoring stored predictions under alternative weight functions would settle whether treated-count weighting is the active ingredient.","If the composite predictor R (effect-to-heterogeneity ratio inflated by treatment imbalance) survives a controlled sweep, F1 could become a diagnosable condition: a practitioner could compute it ex ante from observed data and know when to distrust Qini.","The broader lesson extends beyond uplift modeling: whenever the scoring metric and the deployment objective are invariant to different transformations of the predicted scores, benchmark rankings can decouple from decision quality.","A direct stress test of F2's scope would replace the ten re-splits with genuinely independent randomized experiments; if ranking-metric selection still fails to beat random selection there, the regret figure becomes a general property of rank-only selection for threshold policies."],"forward_implications":["On continuous outcomes, practitioners should not use the unnormalized Qini as the sole selection criterion: on the standard IHDP benchmark its ranking is statistically indistinguishable from random relative to effect accuracy.","For sign-threshold deployment rules, a ranking metric must be paired with a threshold calibrated on held-out selection data; doing so cut the Qini-selection regret on Jobs by 81%.","When the deployment rule is rank-based (budgeted top-k targeting), ranking metrics transmit selection signal and the F2 gap disappears, so the metric must be chosen to match the intended decision rule.","Benchmark winner claims are protocol-scoped: standings should be reported per regime and per objective rather than as a single leaderboard.","Library defaults matter: causalml computes Qini on continuous outcomes with no warning, and its shipped score behaves like the analyzed object, so the failure is silent for downstream users."],"supporting_citations":[{"why":"Supplies IHDP, the continuous semi-synthetic benchmark with known individual effects; F1's evidence base across its 100 simulation realizations.","marker":"[13]"},{"why":"Provides the IHDP realizations and the ten Jobs re-splits, and defines the reference objectives (effect accuracy and policy risk) that F1 and F2 are measured against.","marker":"[31]"},{"why":"Introduces the Qini coefficient construction over cumulative gains versus a chord baseline; the object F1 shows to be outcome-regime sensitive on IHDP.","marker":"[29]"},{"why":"The package whose shipped Qini and cumulative-gain AUUC are scored on identical stored predictions, confirming the canonical object behaves like the analyzed one (rank agreement +0.99) and localizing F1's discrepancy to treated-count weighting.","marker":"[3]"},{"why":"The randomized job-training experiment behind Jobs, whose experimental subset supplies the identified policy-value objective used in F2.","marker":"[22]"},{"why":"ACIC 2016, the independent continuous-outcome validation family that bounds F1: paired AUUC-over-Qini gap +0.11 with the confidence interval covering zero.","marker":"[9]"},{"why":"Closest prior work; the paper's F2 observation echoes its finding that targeting quality and effect estimation can come apart, reframed here as a benchmarked decomposition.","marker":"[37]"}],"fun_headline_variants":["Uplift metric, not model, drives benchmark disagreements","Qini's ranking misses effect accuracy; AUUC aligns better","Calibrated thresholds cut 81% of Qini-selection regret","Metric mismatches, not weak models, explain uplift fights"],"cache_read_input_tokens":39040,"weakest_assumption_plain":"The Jobs selection-regret claim (F2) assumes the ten released re-splits are approximately exchangeable clusters: the 95% regret interval excludes zero only for between-split dependence rho <= 0.12, while per-model risk vectors across splits already correlate at +0.41, so if the true split dependence is larger, the empirical half of F2 collapses to the structural rank-invariance argument.","fun_headline_variants_meta":{"raw":{"variants":["Uplift metric, not model, drives benchmark disagreements","Qini's ranking misses effect accuracy; AUUC aligns better","Calibrated thresholds cut 81% of Qini-selection regret","Metric mismatches, not weak models, explain uplift fights"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001115,"raw_usage":{"total_tokens":4733,"prompt_tokens":1125,"completion_tokens":3608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":741,"completion_tokens_details":{"reasoning_tokens":3537}},"tokens_in":741,"tokens_out":3608,"duration_ms":23111,"temperature":1.0,"reasoning_tokens":3537,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:14:47.013674+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate the actual between-split dependence of the Jobs policy-risk estimates directly, for example by a random-effects variance decomposition, and re-run the cross-repeat selection rotation; if the estimated dependence exceeds the threshold at which the paper's interval still excludes zero, the selection-regret claim includes zero and its empirical half collapses to the structural argument. For F1, a controlled sweep of the paper's own R-composite predictor across new continuous-outcome families with known effects would settle whether any ex-ante quantity can predict when Qini decouples from effect accuracy.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the IHDP realizations and the ten Jobs re-splits, and defines the reference objectives (effect accuracy and policy risk) that F1 and F2 are measured against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Qini coefficient construction over cumulative gains versus a chord baseline; the object F1 shows to be outcome-regime sensitive on IHDP."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The randomized job-training experiment behind Jobs, whose experimental subset supplies the identified policy-value objective used in F2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ACIC 2016, the independent continuous-outcome validation family that bounds F1: paired AUUC-over-Qini gap +0.11 with the confidence interval covering zero."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior work; the paper's F2 observation echoes its finding that targeting quality and effect estimation can come apart, reframed here as a benchmarked decomposition."}],"review_version":1}