{"id":"54d8925d-c975-4ad7-a53e-3daa8f293315","arxiv_id":"1908.02922","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Trimmed Match estimates the ratio of incremental revenue to incremental ad spend by symmetrically trimming poorly matched geo pairs, and is shown to be more efficient than existing estimators on simulated and real data.","lead":"This paper develops a robust statistical estimator, Trimmed Match, for measuring the incremental return on advertising spend in randomized geo experiments. It targets settings with few, highly variable geographic units and budget constraints that create interference between units.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Homogeneity assumption θ_g=θ* is the load-bearing link; the only check is circular, and the sensitivity analysis changes the target.","rationale":"The symmetry result in Proposition 2 is mathematically clean, and the trimmed mean argument is coherent conditional on Assumptions 0 and 1. The GitHub implementation and reproducible simulations provide real independent support. However, Assumption 1 is not a minor technicality: it is exactly what converts the bivariate causal ratio into a univariate location problem. With heterogeneous θ_g, the residual carries an ad-spend-weighted deviation term that is assignment-dependent, so the estimating equation's root does not consistently target θ*. The paper's Section 8 verification is explicitly circular, and the sensitivity analysis shifts the estimand to a virtual all-treated experiment, so the demonstrated robustness is not for the target quantity of the actual paired experiment. This is the same load-bearing concern the reader identified. I see no reason to move the verdict: the concern is addressable but unresolved, which is what CONDITIONAL expresses. The reader's emphasis on Assumption 1 and its circular test is correct, and the proposed concrete check would directly assess whether the data support the assumption.","tokens_in":18362,"tokens_out":7802,"duration_ms":79016,"concrete_test":"For each of the three real case studies, construct a valid test of the model's symmetry implication that avoids plugging in a single estimated θ*: compute the 90% confidence interval for θ* by inverting the Wilcoxon signed-rank test (valid under Assumptions 0 and 1), then evaluate the Wilcoxon signed-rank p-value of residuals ϵ_i(θ) on a fine grid of θ values inside that interval. If the maximum p-value over the grid is below 0.05 (or the interval is empty), the data are incompatible with the symmetric-residual model and Assumption 1 is not supported; if p-values remain large, the circularity concern is mitigated for these datasets.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 2 and the Trimmed Match construction require Assumption 1 (θ_g = θ* for all geos). If geo-level iROAS is heterogeneous, then ϵ_i(θ*) = (Z_{i1}−Z_{i2})A_i + (θ_{i1}−θ*)S_{i1}A_i − (θ_{i2}−θ*)S_{i2}A_i (with roles swapped when A_i = −1). The extra terms involve the realized, assignment-dependent ad spend S, so ϵ_i(θ*) is not generally symmetric about 0 and equation (5.2) does not identify the population θ* of (1.2); the root is an assignment-dependent weighted average. The paper's only empirical check is Section 8's Wilcoxon test on residuals at the estimated θ*, which footnote 4 concedes is inaccurate because θ* is estimated; this is circular. The Section 7.2 sensitivity analysis does not repair the gap: when Assumption 1 is violated, the 'true' θ* is redefined via a virtual experiment in which all geos are treated, not the actual paired experiment, so the claimed robustness is for a different estimand. Because trimming decisions are based on ϵ_i(θ), bias in the estimating equation can also distort which pairs are trimmed, compounding the error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a framework and estimator for the incremental return on ad spend (iROAS) in randomized paired geo experiments. It introduces Assumption 0, that an unobserved 'uninfluenced response' Z_g is invariant to treatment assignments even under budget-induced interference, and Assumption 1, that the unit-level iROAS is constant across geos. Under these assumptions, Proposition 2 shows that the residual ϵ_i(θ*) = Y_i - θ* X_i is symmetric about zero for each pair, reducing iROAS estimation to a univariate location problem. The proposed Trimmed Match estimator solves the trimmed mean equation ϵ̅_{nλ}(θ)=0, with a data-driven choice of trim rate based on confidence interval width. The paper reports simulations showing efficiency gains over empirical, sign-test, and Wilcoxon rank-based estimators, three real case studies, an O(n^2 log n) algorithm for computation, and an existence theorem for the estimator.","tokens_in":18601,"tokens_out":4311,"duration_ms":46265,"significance":"If the two assumptions hold, the paper makes a useful contribution: Proposition 2 is cleanly derived, the estimator is distribution-free and interpretable (it trims poorly matched pairs), and the simulation evidence supports efficiency gains in heavy-tailed settings. The connection to Rosenbaum's instrumental variable framework is insightful, and the availability of a Python implementation is a practical strength. However, the entire inferential edifice rests on Assumption 1, and the manuscript's verification of that assumption is circular, while the sensitivity analysis in Section 7.2 changes the estimand rather than addressing the identification failure. These issues materially affect whether the estimator can be claimed to target the θ* of equation (1.2) in realistic heterogeneous-geo settings.","major_comments":[{"comment":"The only empirical check of Assumption 1 is a Wilcoxon signed-rank test applied to residuals ϵ_i(θ̂_trim) computed at the estimated parameter. Footnote 4 concedes that these p-values may not be accurate because θ* is estimated, so the test is circular: the residuals are by construction centered at the estimator's root, and the test does not account for estimation uncertainty. This is not a valid verification of symmetry of the true residuals under Assumption 1. The authors should either use a split-sample or resampling procedure that accounts for parameter estimation, or present the Wilcoxon result only as an informal diagnostic, and explicitly state that Assumption 1 remains an untestable identifying assumption.","section":"Section 8, footnote 4"},{"comment":"The sensitivity analysis redefines the 'true' θ* via a virtual experiment in which all geos are assigned to treatment with a doubled total incremental budget. This is a different estimand from the θ* defined in equation (1.2) for the actual paired experiment, because under budget-constrained interference the average incremental response and spend in (1.2) are assignment-dependent. Consequently, the robustness claims in Section 7.2 and the abstract, that estimates remain reliable when Assumption 1 is violated, do not apply to the original target. The manuscript should either adopt the virtual-experiment θ* as the explicit parameter of interest throughout, or provide sensitivity results for a well-defined assignment-conditional parameter, with a formal statement of what the estimating equation identifies under heterogeneity.","section":"Section 7.2"},{"comment":"When Assumption 1 is violated, the residual ϵ_i(θ*) includes the additional terms (θ_{i1}-θ*)S_{i1}A_i and (θ_{i2}-θ*)S_{i2}A_i, with roles depending on A_i, which depend on realized ad spend and treatment assignment. These terms are not generally symmetric about zero, so Proposition 2 fails and the trimmed mean equation (5.2) does not identify the population θ* of (1.2); instead, the root is an assignment-dependent weighted average of heterogeneous θ_g. This is a load-bearing gap because the paper's main claim, that Trimmed Match robustly estimates the overall iROAS, requires Assumption 1. The authors should provide a formal analysis of the bias as a function of the degree of heterogeneity and the spend distribution, or explicitly restrict the target to a random-coefficient model in which Assumption 1 is replaced by a defined aggregation.","section":"Proposition 2 and equation (5.2)"}],"minor_comments":[{"comment":"The set of untrimmed indices I in equation (5.5) depends on the estimated θ̂, but the manuscript does not specify how ties in ϵ_i(θ) at the trimming boundaries are handled in the point estimator; a tie-breaking rule would make the estimator fully defined.","section":"Section 5.1"},{"comment":"The data-driven choice of trim rate λ by minimizing confidence interval width is presented as a contribution of independent interest, but no theoretical justification is given for this criterion or for the recommended α0=0.5; the discussion would benefit from at least a heuristic argument or a small asymptotic analysis.","section":"Section 6"},{"comment":"The proofs of Lemma 2 and Theorem 1 are omitted. Since the algorithm and the existence result are central to the computational and inferential claims, these proofs should be provided in the appendix or a supplementary file rather than referenced as 'straightforward' and 'omitted for conciseness'.","section":"Appendix A"},{"comment":"Lemma 1 is a direct rearrangement of the definition of θ_g in (1.1), so describing it as the basis of a 'novel statistical framework' is somewhat overstated; the novelty lies in the robust estimation strategy under Assumptions 0 and 1, not in the lemma itself.","section":"Section 1, Lemma 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the Annals of Applied Statistics and the methodological core is sound when Assumption 1 holds, but the current verification of that assumption is circular and the sensitivity analysis changes the estimand. These are fixable in revision, so I recommend major revision rather than rejection. The authors should be encouraged to state clearly whether the target is the assignment-conditional or virtual-experiment iROAS, and to provide a non-circular diagnostic or a formal bias analysis under heterogeneity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a real contribution, not a repackaging. What's new is the Trimmed Match estimator and the framing that turns paired geo experiments under budget-constraint interference into a symmetric residual problem. Proposition 2 is correct under Assumption 0 and Assumption 1, and the efficiency gains over the empirical, sign, and Wilcoxon estimators in heavy-tailed simulations are real. The associated Python library is public, and the three case studies give the method a plausible practical ring. The citation pattern is appropriate—Rosenbaum's IV reasoning, Vaver and Koehler, Kerman et al., and the recent geo-experiment literature are all there.\n\nThe soft spot is exactly where the reader's report puts it: Assumption 1, the constant unit-level iROAS, is doing the work, and the paper's check of it is weak. The Wilcoxon test in Section 8 is applied to residuals computed at the estimated theta, and footnote 4 concedes the p-values may not be accurate. That is circular in a mild but real sense. More importantly, the Section 7.2 sensitivity analysis does not repair the gap: when theta_g varies, the 'true' theta* is computed from a virtual all-treated experiment, not from the actual paired experiment. So the claimed robustness is for a different estimand. I don't think this sinks the paper—for practical purposes, the estimator likely still targets a useful treatment-effect ratio—but it should be stated honestly.\n\nTwo smaller items. The plug-in trim rate is a fitted parameter, and the paper acknowledges the resulting confidence intervals may undercover in finite samples; the simulations show mild undercoverage, so this is a disclosure issue more than a flaw. And Theorem 1's proof is omitted. I trust it, given the algorithm and the continuity argument sketched, but a referee should ask for the proof or a reference.\n\nOverall: the math that is shown is clean, the estimator is well motivated, and the weakness is concentrated in one unverified, but plausible, assumption. This deserves serious peer review. I would send it out, with a request to be upfront about what Assumption 1 buys and what happens when it fails.\n\nSend it to the reading group; worth an hour of discussion.","headline":"A clean, practical estimator for a real advertising measurement problem, with one load-bearing assumption that the paper verifies only weakly.","tokens_in":19128,"tokens_out":3566,"would_cite":true,"duration_ms":38293,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G35","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Randomized paired geo experiments turn iROAS estimation into a robust symmetry problem, and the Trimmed Match estimator solves it by discarding poorly matched pairs.","keywords":["incremental return on ad spend","geo experiments","causal inference","interference","trimmed mean","effect ratio","distribution-free estimation","heavy-tailed data"],"falsifier":"One direct test: simulate a randomized paired geo experiment with known heterogeneous unit-level iROAS values (alternating $\\theta_g = \\theta_0(1 \\pm \\delta)$ with $\\delta$ near 1, as the paper's own sensitivity analysis does), compute the residuals $\\epsilon_i(\\theta^*)$ at the true $\\theta^*$, and apply a Wilcoxon signed-rank symmetry test at the paper's sample sizes; the residual distribution will be detectably asymmetric, and the trimmed-mean estimate will shift with the realized assignment, showing that both the symmetry claim and the estimator's target depend on Assumption 1 holding.","tokens_in":18124,"feed_emoji":"📈","tokens_out":14552,"duration_ms":133375,"temperature":0.7,"pith_summary":"The paper is trying to establish that in a randomized paired geo experiment, the incremental return on ad spend (iROAS) can be estimated without modeling the heavy-tailed distribution of ad spend and response. Its key reduction is Proposition 2: under two assumptions, the residual difference $Y_i - \\theta^* X_i$ between the treated and control geo in each pair is independent across pairs and symmetric about zero. That turns the ratio-estimation problem into a one-dimensional location problem, and solving a trimmed-mean version of that location equation gives the Trimmed Match estimator, which discards the worst-matched pairs. The authors argue that this estimator is distribution-free and interpretable, and that it is often more efficient than the empirical, sign-test, and Wilcoxon estimators when geo sizes are highly heterogeneous; the claim is supported by simulations and three real experiments.","feed_headline":"Trimmed Match estimates ad iROAS without modeling heavy tails","feed_subtitle":"It trims poorly matched geo pairs and stays distribution-free under budget constraints and heavy-tailed spend.","key_machinery":"The machinery is the residual identity $\\epsilon_i(\\theta) = Y_i - \\theta X_i = (Z_{i1} - Z_{i2}) A_i$, where $A_i$ is the fair-coin sign deciding which geo in pair $i$ receives treatment. Because $Z_{i1}$ and $Z_{i2}$ are non-random under Assumption 0, multiplication by $A_i$ makes $\\epsilon_i(\\theta^*)$ symmetric about zero and independent across pairs; this licenses sign-test, Wilcoxon, and trimmed-mean estimating equations for $\\theta^*$. The trimmed-mean equation $\\bar{\\epsilon}_{n\\lambda}(\\theta) = 0$ then does double duty: it defines the estimator as a ratio of trimmed sums of $Y_i$ and $X_i$, and it identifies poorly matched pairs through large $|\\epsilon_i|$ values for removal. Lemma 2, stating that the ordering of the residuals changes only at $\\theta_{ij} = (y_j - y_i)/(x_j - x_i)$, is what makes the computation $O(n^2 \\log n)$.","core_discovery":"The central claim is Proposition 2: with randomized paired assignment, Assumption 0 (the uninfluenced response $Z_g = R_g - \\theta^* S_g$ is invariant to all treatment assignments, so budget-constraint interference enters only through observed spend) and Assumption 1 (unit-level iROAS is constant, $\\theta_g = \\theta^*$ for all geos) imply that the residuals $\\epsilon_i(\\theta^*) = Y_i - \\theta^* X_i$ are mutually independent and symmetrically distributed about zero. Consequently $\\theta^*$ is the root of the trimmed-mean equation $\\bar{\\epsilon}_{n\\lambda}(\\theta) = 0$, and the Trimmed Match estimator $\\hat{\\theta}^{(\\mathrm{trim})}_\\lambda$ is defined as the root that minimizes symmetric deviation; when it exists it equals the ratio of the sums of $Y_i$ and $X_i$ over the untrimmed pairs. The paper also supplies a data-driven trim rate chosen by minimizing confidence-interval width, proves existence when the trimmed sum of $X_i$ is nonzero, and provides an $O(n^2 \\log n)$ algorithm.","pith_inferences":["If the symmetry reduction is as general as Proposition 2 suggests, other robust location estimators with higher efficiency at heavy tails—Huber-type M-estimators or adaptively weighted trimmed means—could be ported to effect-ratio estimation and would likely beat Trimmed Match in the same simulation designs.","The paper's residual-based check of Assumption 1 is indirect, because it tests residuals built from the estimated $\\hat{\\theta}^{\\mathrm{(trim)}}_{\\hat{\\lambda}}$; a direct test requires a holdout estimate of $\\theta^*$ or a permutation distribution that accounts for the estimation step, and until then the constant-iROAS assumption is supported only indirectly.","When Assumption 1 fails, Trimmed Match does not collapse but targets an assignment-dependent weighted average of geo-level iROAS values; a natural extension would let $\\theta_g$ depend on geo covariates and use the same residual symmetry as a diagnostic instead of an assumption.","Because the trim rate is chosen by minimizing confidence-interval width, the method implicitly trades coverage for power; with small $n$, reported intervals should be treated as optimistic unless calibrated by simulation, a point the paper itself raises when discussing undercoverage."],"forward_implications":["Under the paper's two assumptions, iROAS inference becomes distribution-free: any symmetry-based estimating equation applied to $\\epsilon_i(\\theta)$ yields valid point and interval estimates without modeling the joint spend-response distribution.","Trimmed Match can be more efficient than the empirical ratio, the sign-test estimator, and the Wilcoxon estimator when geo sizes are heavy-tailed; the simulated log-normal and half-Cauchy scenarios show the largest gains.","The data-driven trim rate makes the method adaptive: in settings where trimming does not reduce variance, the method can select a trim rate of zero and reduce to the empirical estimator, as in case B of the real studies.","The same framework applies to other matched-pairs effect-ratio problems, such as incremental cost-effectiveness ratios, because only the symmetry of residuals is used.","An $O(n^2 \\log n)$ algorithm and a studentized trimmed-mean $t$ approximation for confidence intervals make the method practical for routine advertiser experiments."],"supporting_citations":[{"why":"Introduces randomized paired geo experiments and the matching recommendation; supplies the model-based geo regression benchmark the paper contrasts with.","marker":"Vaver and Koehler (2011)"},{"why":"Defines the effect-ratio estimator as the midpoint of values minimizing a sign- or rank-based statistic, the direct precursor of Trimmed Match.","marker":"Rosenbaum (1996)"},{"why":"Supplies the instrumental-variable identification of a ratio of causal effects that the residual equation generalizes.","marker":"Angrist, Imbens and Rubin (1996)"},{"why":"Provides the distribution-free randomization-inference perspective and the stronger-instrument idea behind trimming poorly matched pairs.","marker":"Rosenbaum (2002)"},{"why":"Gives the studentized trimmed mean and winsorized variance used to construct Trimmed Match confidence intervals.","marker":"Tukey and McLaughlin (1963)"},{"why":"Proposes selecting a trimmed mean's trim rate by minimizing estimated variance, the idea the paper extends to confidence-interval-width minimization.","marker":"Jaeckel (1971)"},{"why":"Proves consistency of Jaeckel's adaptive trimmed mean, backing the data-driven trim-rate approach.","marker":"Hall (1981)"},{"why":"Shows a smaller study with a stronger instrument can be more powerful and less sensitivity-prone, motivating residual-based trimming.","marker":"Small and Rosenbaum (2008)"},{"why":"Provides a model-based time-series iROAS estimator from geo experiments that Trimmed Match is designed to complement.","marker":"Kerman, Wang and Vaver (2017)"}],"fun_headline_variants":["Trimmed Match: robust ad iROAS from geo pair trimming","Distribution-free iROAS with trimmed geo pairs","Paired geo trim handles budget caps and heavy tails","Geo-pair trimming yields distribution-free causal iROAS","Robust causal iROAS from trimmed paired geo experiments"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is Assumption 1, that every geographic unit has the same true incremental return on ad spend; if unit-level returns differ, the residual symmetry that powers the whole estimator breaks and the method estimates an assignment-dependent weighted average rather than the population ratio, while the paper's own check of this premise is indirect because it tests residuals computed with the estimated value of that ratio.","fun_headline_variants_meta":{"raw":{"variants":["Trimmed Match: robust ad iROAS from geo pair trimming","Distribution-free iROAS with trimmed geo pairs","Paired geo trim handles budget caps and heavy tails","Geo-pair trimming yields distribution-free causal iROAS","Robust causal iROAS from trimmed paired geo experiments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001261,"raw_usage":{"total_tokens":5186,"prompt_tokens":992,"completion_tokens":4194,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":4114}},"tokens_in":608,"tokens_out":4194,"duration_ms":30667,"temperature":1.0,"reasoning_tokens":4114,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:29:34.693326+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One direct test: simulate a randomized paired geo experiment with known heterogeneous unit-level iROAS values (alternating $\\theta_g = \\theta_0(1 \\pm \\delta)$ with $\\delta$ near 1, as the paper's own sensitivity analysis does), compute the residuals $\\epsilon_i(\\theta^*)$ at the true $\\theta^*$, and apply a Wilcoxon signed-rank symmetry test at the paper's sample sizes; the residual distribution will be detectably asymmetric, and the trimmed-mean estimate will shift with the realized assignment, showing that both the symmetry claim and the estimator's target depend on Assumption 1 holding.","supporting_citations":[{"cited_title":"Koehler , Jim J","cited_arxiv_id":null,"evidence_quote":"Introduces randomized paired geo experiments and the matching recommendation; supplies the model-based geo regression benchmark the paper contrasts with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the effect-ratio estimator as the midpoint of values minimizing a sign- or rank-based statistic, the direct precursor of Trimmed Match."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the instrumental-variable identification of a ratio of causal effects that the residual equation generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the distribution-free randomization-inference perspective and the stronger-instrument idea behind trimming poorly matched pairs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the studentized trimmed mean and winsorized variance used to construct Trimmed Match confidence intervals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proposes selecting a trimmed mean's trim rate by minimizing estimated variance, the idea the paper extends to confidence-interval-width minimization."},{"cited_title":"( 1981 )","cited_arxiv_id":null,"evidence_quote":"Proves consistency of Jaeckel's adaptive trimmed mean, backing the data-driven trim-rate approach."},{"cited_title":"Rosenbaum , Paul R","cited_arxiv_id":null,"evidence_quote":"Shows a smaller study with a stronger instrument can be more powerful and less sensitivity-prone, motivating residual-based trimming."},{"cited_title":", Wang , Peng P","cited_arxiv_id":null,"evidence_quote":"Provides a model-based time-series iROAS estimator from geo experiments that Trimmed Match is designed to complement."}],"review_version":1}