{"id":"a148c586-e136-45a2-9e3d-84f291941882","arxiv_id":"2411.16906","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Under a binary instrumental variables model with monotone treatment response, the joint distribution of potential outcomes among compliers is identified, allowing researchers to profile latent persuasion types.","lead":"This paper develops a method to identify, in randomized encouragement designs, the share of compliers who are always-voters, never-voters, or mobilised by the treatment, and to profile their characteristics. It also proposes a sharp test of the identification assumptions and applies the methods to voter mobilization experiments.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's core identification is internally sound, but the policy-relevant profiling estimates in Theorem 3.2 are not covered by the paper's own sensitivity analysis; modest violations of monotone treatment response could materially change the Bridgeport Democrat-share and cost-per-Democrat…","rationale":"The reader's weakest-assumption diagnosis is correct: monotone treatment response is the assumption that converts marginal complier outcome distributions into the joint distribution, and Theorem 3.2 inherits this dependence. My stress-test agrees with that diagnosis but sharpens it. The paper includes a sensitivity analysis, which is real evidence of good practice, but Table 6 only varies the demobilised mass in the joint outcome distribution. The headline applied results—especially the Bridgeport claim that 81.3% of mobilised compliers are Democrats, that about 28 Democrats were mobilised, and that the cost per Democrat is $1,066—are profiling results from Theorem 3.2, and no sensitivity analysis is reported for them. A small demobilised share can reallocate probability mass across the three persuasion types without changing the observed marginals, so the profiling moments can shift even when Table 6 looks reassuring. The proposed concrete test would resolve whether this concern actually moves the applied conclusions. I am not suggesting the theory is wrong; the identification argument is internally consistent and Proposition 3.1 is a clean consequence of the stated assumptions. The issue is robustness of the paper's most policy-relevant numbers to a violation the authors themselves take seriously enough to analyse. Because the reader's conditional verdict already asks for empirical corrections and frames the application as conditional on strong assumptions, my concern does not change the verdict; it strengthens the rationale for the condition. I mark agreement as partial because the reader identified the same weakest assumption but did not flag that the sensitivity analysis excludes the profiling results that carry the paper's applied conclusions.","tokens_in":35727,"tokens_out":10702,"duration_ms":95014,"concrete_test":"Recompute the Bridgeport profiling results with demobilised compliers allowed: fix delta = P[Yi(1)=0,Yi(0)=1|C] at 0, 0.03, 0.05, 0.08, and 0.10, impose the observed complier marginals and the observed moments that identify the always/never/mobilised equations, and re-estimate P[Democrat=1|mobilised,C], the implied number of mobilised Democrats, and the cost per Democrat. If these quantities move by more than 50% over the delta grid used in Table 6, the Section 5 conclusions are not robust to plausible MTR violations and should be presented as conditional on the assumption, not as empirical findings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 3.1 is correct as a statement about what follows from Assumption 2.1: with binary outcomes, monotone treatment response makes 'always-voter' equal to {Yi(0)=1}, 'never-voter' equal to {Yi(1)=0}, and 'mobilised' equal to {Yi(1)=1,Yi(0)=0}, so the joint distribution among compliers collapses to Imbens-Rubin marginals. The load-bearing point is that this equivalence is the entire identification; if P[Yi(1)=0,Yi(0)=1|C] = delta > 0, the three-cell model in Proposition 3.1 is misspecified and Theorem 3.2's profiling estimands are no longer identified. The paper's sensitivity analysis (Table 6) correctly shows that the mobilised share rises from 13.9% to 23.9% in Bridgeport when delta = 0.10, but it is restricted to the joint outcome distribution; it does not trace how delta affects the empirically emphasised quantities P[Democrat=1|mobilised,C], the 28 mobilised Democrats, or the $1,066 cost per Democrat. The sharp test in Section 4.2 is a joint test of the IA-IV plus MTR restrictions and did not reject, but non-rejection cannot establish delta = 0, and the test's power against small demobilised shares is not analysed. Thus the central theoretical claim is secure under MTR, while the applied payoff claims have a quantitatively unexamined fragility.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies identification in a binary treatment/binary outcome Imbens-Angrist IV model augmented by the monotone treatment response assumption (Yi(1) ≥ Yi(0)). It shows that the joint distribution of potential outcomes among compliers is point identified: the always-voter, never-voter, and mobilised shares are the Imbens-Rubin marginals, and the mobilised share equals the LATE. It extends Abadie's kappa weighting to moments of (Yi(t), Ti, Xi) and, under monotone treatment response, to moments of (Ti, Xi) conditional on always-voter, never-voter, and mobilised compliers. It also proposes a sharp test of the identifying assumptions based on a system of linear inequalities, gives conditions under which the DellaVigna-Gentzkow approximated persuasion rate equals the local persuasion rate, provides a sensitivity analysis for the monotone treatment response assumption, and applies the methods to Green et al. (2003).","tokens_in":35992,"tokens_out":19813,"duration_ms":174384,"significance":"If correct, the main theoretical result is useful and clean: with binary outcomes, monotone treatment response collapses joint persuasion types into marginal potential-outcome events, so the joint distribution among compliers is identified from standard Imbens-Rubin/Abadie results. The extension of Abadie's kappa to treatment-inclusive moments is a genuine contribution, and the proposed sharp test is a real refutation test rather than a fitted verification. The proofs are complete and rely on known external results, with no hidden free parameters in the identification argument. The contribution is somewhat incremental given Jun and Lee (2023) and the acknowledged independent work of Comey et al. (2023), but the paper provides a useful unified treatment and a concrete application.","major_comments":[{"comment":"The application's headline cost-effectiveness figures are internally inconsistent. In Bridgeport, Table 4 gives P[Yi(1)=1,Yi(0)=0|C]=0.139 and Table 5 gives P[Democrat=1|Yi(1)=1,Yi(0)=0,C]=0.813. With the first-stage complier share 0.277 and n=1,806 (about 500 compliers), this implies about 56 compliers who are both mobilised and Democrats, or 3.1% of the sample. The text reports exactly this 3.1% and then, one sentence later, reports '1.6%, or around 28 people', apparently multiplying by the 0.5 assignment probability. If the intended estimand is the number of complier-mobilised Democrats actually assigned to treatment (28), that needs to be stated explicitly and the 3.1% number cannot be used as the share of voters mobilised by the experiment without the same qualifier. The cost computation is also not transparent: the stated components ($3,000 + $30 per outreach voter) do not obviously sum to $29,350, and the implied cost per Democrat is roughly $524 if the 56 figure is used instead of $1,066. Please correct the arithmetic and define the target estimand precisely.","section":"Section 5.2; Tables 4-5"},{"comment":"The sensitivity analysis in Table 6 varies the demobilised share P[Yi(1)=0,Yi(0)=1|C] and recomputes the three joint outcome probabilities among compliers, but it never traces the effect of this violation on the profiling estimands emphasised in the application: P[Democrat=1|mobilised,C], the number of mobilised Democrats, or the cost per Democrat. This is a substantive gap because Theorem 3.2's mobilised-cell formula is derived from the identity E[g(z,X)(Y(1)-Y(0))1{C}], and once demobilised individuals (Y(1)=0,Y(0)=1) are allowed, the observed numerator no longer equals E[g(z,X)1{Y(1)=1,Y(0)=0}1{C}]; the profiling ratios are therefore not identified. Table 6 shows the mobilised share rising from 13.9% to 23.9% in Bridgeport at δ=0.10, and the same δ could shift the Democrat share and cost figures by an amount the paper does not quantify. The introductory claim that the paper 'provides a simple sensitivity analysis for the monotone treatment response assumption' should be scoped to the joint outcome distribution, or the analysis should be extended to the Theorem 3.2 estimands.","section":"Section 4.3; Table 6; Theorem 3.2"},{"comment":"The sharp test is a joint test of IA-IV plus monotone treatment response, and the paper is careful to state that non-rejection does not verify the assumptions. Since the paper's own Table 6 entertains demobilised shares as large as 0.10, the test's power against such alternatives should be assessed or at least discussed; otherwise the Section 5.3 conclusion that the assumptions are 'not rejected' gives little assurance for the application. Please report the numerical test statistics and subsampling p-values, and ideally a small simulation showing which demobilised shares the test can detect with the Green et al. sample sizes.","section":"Section 4.2; Section 5.3"}],"minor_comments":[{"comment":"The text refers to 'Lemma 3.1' and 'the identification results in Lemma 3.1', but no lemma with that number is stated in the main text; renumber the result or add the lemma statement.","section":"Section 4.3; Appendix A.2"},{"comment":"Appendix A.13 refers to 'Theorem 3.3', which is not defined anywhere in the paper; correct the cross-reference.","section":"Appendix A.13"},{"comment":"Several labels collide: 'Assumption 2.1' is used again in Appendix B.1, 'Proposition 4.1' appears in both the main text and Appendix D, and 'Proposition 2.1' appears in Appendix B.2; renumber the appendix items.","section":"Appendix B; Appendix D"},{"comment":"The test statistic T_n is defined with a constraint Bp=1, but the matrix B is never introduced; define B as the row vector of ones or write the constraint as the sum of the components of p being one.","section":"Section 4.2"},{"comment":"Part (2) of Proposition 4.3 says 'satisfies the restrictions in P0' before P0 is defined in equation (4.2); state the definition of P0 before or inside the proposition.","section":"Proposition 4.3"},{"comment":"The definition of L_n(t) sums over all N_n = C(n,b) subsamples, which is computationally impossible for n=18,933; state that random subsamples are used in practice and specify their number.","section":"Section 4.2"},{"comment":"The sentence 'we estimate that 3.1% of mobilised voters are also compliers and Democrats' is ambiguous; it should say '3.1% of the Bridgeport sample are mobilised compliers who are Democrats' (or whatever is intended), and a confidence interval for this joint share should be reported given the wide CI for the Democrat share among mobilised compliers.","section":"Section 5.2"},{"comment":"The table note contains 'among compilers' in the last sentence; it should be 'among compliers'.","section":"Table 6 note"},{"comment":"Proposition 5.1 contains the typo 'rull rank'; it should be 'full rank'.","section":"Appendix E.1"}],"recommendation":"major_revision","confidential_remarks":"The theoretical identification results are sound and well proved, and the paper gives appropriate credit to Jun and Lee (2023) and Comey et al. (2023). My main concern is that the empirical application, which is prominently featured in the introduction and abstract, contains arithmetic and interpretive errors in the '28 mobilised Democrats' and '$1,066 per Democrat' figures, and the sensitivity analysis does not cover the profiling estimands that drive those applied conclusions. These issues can be fixed within the manuscript's scope, so I am recommending major revision rather than rejection. If the journal requires replication materials, the author should provide code and data for the cost calculation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a real but modest theoretical contribution. The key identification result amounts to combining Imbens-Rubin marginal outcome distributions with monotone treatment response to get the joint distribution among compliers, and then extending Abadie's kappa to profile the three persuasion types. That is worth having, and the paper is honest that the MTR assumption does the heavy lifting. The sharp test and sensitivity analysis are useful additions. The appendix proofs look correct, and the writing is clear about what is new relative to Jun-Lee and Comey et al.\n\nThe empirical application is where I have real problems. The Bridgeport cost calculation has a straightforward arithmetic inconsistency: the text says 1.6% of mobilised voters are Democrats, or about 28 people, and then uses that to get $1,066 per Democrat. But Table 5 reports a point estimate of 81.3% Democrat among mobilised compliers, and Table 4 says 13.9% of compliers are mobilised. With 1,806 observations and a 27.7% complier share, that gives roughly 0.813 × 0.139 × 0.277 × 1806 ≈ 56 people, not 28. The reported 1.6% and 28 people are off by about a factor of two. This needs to be fixed, and it changes the cost estimate materially.\n\nThe sensitivity analysis is also too narrow. It varies the size of the demobilised group and shows how the mobilised share changes in the joint distribution, but it does not trace the effect on the profiling estimands—the Democrat share among mobilised compliers, the number of mobilised Democrats, or the cost per Democrat. Since those are the headline applied results, the fragility under MTR violations is unexamined for exactly the quantities the empirical section advertises. The sharp test did not reject, but non-rejection cannot establish that the demobilised share is zero, and the paper doesn't analyse the power of the test.\n\nNone of this undermines the theoretical identification results. Proposition 3.1 is correct, and Theorem 3.2 is a clean extension of Abadie's kappa. The paper deserves a serious referee. My advice is to send it out, but the empirical section needs correction, and a replication package would be appropriate for an applied econometrics journal. I would cite the identification results after the empirical fixes; the policy numbers are not trustworthy as they stand.","headline":"The theoretical identification results are sound and worth publishing after the empirical section's arithmetic error and the too-narrow sensitivity analysis are fixed.","tokens_in":36557,"tokens_out":1940,"would_cite":true,"duration_ms":18573,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"In a binary instrumental-variable model with monotone treatment response, the joint distribution of potential outcomes among compliers is point identified, so the shares and covariate profiles of always-voters, never-voters, and mobilised…","keywords":["instrumental variables","monotone treatment response","persuasion","compliers","Abadie kappa","local persuasion rate","sharp test","get-out-the-vote"],"falsifier":"A direct check would be a crossover or panel design in which the same individual is observed under both treatment and control: if a non-negligible share have $Y_i(1)=0$ and $Y_i(0)=1$, the monotone-response assumption fails and the point-identification claim collapses. Equivalently, applying the paper's sharp linear-system test to data with a known demobilised subpopulation should reject at a rate above the nominal size.","tokens_in":35463,"feed_emoji":"🗳️","tokens_out":11425,"duration_ms":93004,"temperature":0.7,"pith_summary":"This paper asks what an encouragement experiment can reveal about who is actually persuaded, not just how many are moved on average. It shows that under the standard binary Imbens-Angrist instrumental-variable assumptions plus a monotone treatment response, the joint distribution of potential outcomes among compliers is point identified. Because that joint distribution is exactly what separates always-voters, never-voters, and mobilised voters, each latent type's share and its profile in pre-treatment covariates become estimable from the observed data. The paper also gives a sharp test of the identifying assumptions, a sensitivity analysis for the monotone-response assumption, and an application to get-out-the-vote experiments.","feed_headline":"Persuasion types become identifiable from a binary instrument","feed_subtitle":"A single monotone-response assumption lets a binary instrument identify who persuasion actually moves","key_machinery":"The carrying object is the pair of binary potential outcomes with monotone treatment response, $Y_i(1) \\geq Y_i(0)$ almost surely. This restriction removes the demobilised type, so each persuasion type among compliers corresponds to an event whose marginal probability the Imbens-Rubin framework already identifies. The profiling results ride on an extension of Abadie's kappa weighting: Theorem 3.1 shows that any moment of $(Y_i(t), T_i, X_i)$ among compliers is identified, and Theorem 3.2 conditions those moments on the joint potential-outcome type. The test and sensitivity analysis use the same linear structure: cell probabilities are written as linear combinations of unobserved type probabilities, so the assumptions hold exactly when some nonnegative vector $p$ satisfies $A_{obs} p = b$.","core_discovery":"The central discovery is that the joint distribution of potential outcomes among compliers, usually treated as unidentified in instrumental-variable settings, is point identified when the outcome is binary and treatment response is monotone. Proposition 3.1 gives explicit formulas: the share of always-voters among compliers equals $(E[Y_i(1-T_i)|Z_i=0] - E[Y_i(1-T_i)|Z_i=1])/(E[T_i|Z_i=1]-E[T_i|Z_i=0])$, with analogous expressions for never-voters and mobilised compliers. Identification works because monotone treatment response collapses joint types to marginal events: always-voters are those with $Y_i(0)=1$, never-voters are those with $Y_i(1)=0$, and mobilised compliers are the remaining cell. Theorem 3.2 then identifies the conditional expectation of any measurable $g(T_i, X_i)$ given each persuasion type among compliers. The paper also characterises the identifying assumptions sharply as the existence of a nonnegative solution to a linear system and applies the method to the Green et al. (2003) get-out-to-vote experiments.","pith_inferences":["The same identification logic would apply to other binary-outcome encouragement settings, such as charitable giving, advertising, or job-training take-up, whenever the treatment plausibly moves outcomes only in one direction.","Because the point-identification result hinges on ruling out demobilised voters, the method is most credible when the treatment lowers the cost of an action; for counter-attitudinal or backfiring treatments, researchers would need the partial-identification version of the same linear-system argument.","The sharp test could be extended to continuous covariates by partitioning the covariate space and using high-dimensional linear-inequality inference, making the specification check practical in observational studies.","The sensitivity analysis suggests a routine robustness practice: report the estimated joint distribution as a function of the allowed share of demobilised compliers, which directly shows how much of the mobilised-voter conclusion depends on the monotone-response assumption."],"forward_implications":["Researchers can estimate the share of compliers who are always-voters, never-voters, and mobilised voters, and can profile each group by covariates such as partisanship or prior turnout.","The approach extends Abadie's kappa weighting: any moment of the joint distribution of treatment and covariates is identifiable conditional on a persuasion type, not merely for compliers as a whole.","A sharp test reduces the identifying assumptions to checking whether a known linear system has a nonnegative solution, so the validity of the instrument and of monotone treatment response can be jointly tested.","The comparison of persuasion-rate estimands pins down when the commonly used approximated persuasion rate coincides with the local persuasion rate under one-sided non-compliance.","Applied to the Green et al. (2003) experiments, the method estimates that roughly 8% of compliers in the full sample and 14% in Bridgeport were mobilised, with prior-turnout profiles consistent with habit formation."],"supporting_citations":[{"why":"Supplies the kappa weighting that Theorem 3.1 extends to moments of the joint distribution of (Yi(t), Ti, Xi) among compliers.","marker":"Abadie (2003)"},{"why":"Gives the LATE identification result that motivates the IA IV framework and the Wald-ratio denominators used throughout.","marker":"Imbens and Angrist (1994)"},{"why":"Identifies the marginal distributions of potential outcomes among compliers, which Proposition 3.1 builds on for always-voter and never-voter shares.","marker":"Imbens and Rubin (1997)"},{"why":"Defines the binary IV model of persuasion and the local persuasion rate that this paper extends and tests.","marker":"Jun and Lee (2023)"},{"why":"Introduces the linear-programming representation of potential-outcome restrictions that underlies the sharp test and sensitivity analysis.","marker":"Balke and Pearl (1997)"},{"why":"Establishes that instrument validity and related assumptions are refutable but nonverifiable, which the paper's sharp characterization uses.","marker":"Kitagawa (2015)"},{"why":"Provides the subsampling test for linear inequalities with known coefficients that implements the sharp test.","marker":"Bai et al. (2022)"},{"why":"Supplies the get-out-the-vote field experiment data used in the empirical application.","marker":"Green et al. (2003)"},{"why":"Supports the habit-formation interpretation of the profiling results on prior turnout.","marker":"Gerber et al. (2003)"},{"why":"Summarises the approximated persuasion rate whose equivalence to the local persuasion rate Proposition 4.1 characterises.","marker":"DellaVigna and Gentzkow (2010)"}],"fun_headline_variants":["Binary IV plus monotonicity exposes persuasion types","Identify complier persuasion types with a binary IV","New method pinpoints who persuasion actually moves","Profiling always, never, and mobilised voters","From a binary instrument to full persuasion profiles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that no one is dissuaded by the treatment: an individual who would take the action without the treatment also takes it with the treatment, which is what collapses the unobserved joint persuasion types onto identifiable marginal outcome events.","fun_headline_variants_meta":{"raw":{"variants":["Binary IV plus monotonicity exposes persuasion types","Identify complier persuasion types with a binary IV","New method pinpoints who persuasion actually moves","Profiling always, never, and mobilised voters","From a binary instrument to full persuasion profiles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1393,"prompt_tokens":947,"completion_tokens":446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":376}},"tokens_in":563,"tokens_out":446,"duration_ms":4771,"temperature":1.0,"reasoning_tokens":376,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:45:53.004010+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be a crossover or panel design in which the same individual is observed under both treatment and control: if a non-negligible share have $Y_i(1)=0$ and $Y_i(0)=1$, the monotone-response assumption fails and the point-identification claim collapses. Equivalently, applying the paper's sharp linear-system test to data with a known demobilised subpopulation should reject at a rate above the nominal size.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the kappa weighting that Theorem 3.1 extends to moments of the joint distribution of (Yi(t), Ti, Xi) among compliers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the LATE identification result that motivates the IA IV framework and the Wald-ratio denominators used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the binary IV model of persuasion and the local persuasion rate that this paper extends and tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the linear-programming representation of potential-outcome restrictions that underlies the sharp test and sensitivity analysis."},{"cited_title":"Santos, and A","cited_arxiv_id":null,"evidence_quote":"Provides the subsampling test for linear inequalities with known coefficients that implements the sharp test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the get-out-the-vote field experiment data used in the empirical application."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the habit-formation interpretation of the profiling results on prior turnout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Summarises the approximated persuasion rate whose equivalence to the local persuasion rate Proposition 4.1 characterises."}],"review_version":1}