Pith. sign in

REVIEW 3 cited by

Individualized Policy Evaluation and Learning under Clustered Network Interference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.02467 v3 pith:V4ZMU5YM submitted 2023-11-04 stat.ME cs.LGecon.EM

classification stat.MEcs.LGecon.EM
keywords estimatorinterferenceevaluationeffectsefficientlearnedlearningperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although there is now a large literature on policy evaluation and learning, much of the prior work assumes that the treatment assignment of one unit does not affect the outcome of another unit. Unfortunately, ignoring interference can lead to biased policy evaluation and ineffective learned policies. For example, treating influential individuals who have many friends can generate positive spillover effects, thereby improving the overall performance of an individualized treatment rule (ITR). We consider the problem of evaluating and learning an optimal ITR under clustered network interference (also known as partial interference), where clusters of units are sampled from a population and units may influence one another within each cluster. Unlike previous methods that impose strong restrictions on spillover effects, such as anonymous interference, the proposed methodology only assumes a semiparametric structural model, where each unit's outcome is an additive function of individual treatments within the cluster. Under this model, we propose an estimator that can be used to evaluate the empirical performance of an ITR. We show that this estimator is substantially more efficient than the standard inverse probability weighting estimator, which does not impose any assumption about spillover effects. We derive the finite-sample regret bound for a learned ITR, showing that the use of our efficient evaluation estimator leads to the improved performance of learned policies. We consider both experimental and observational studies, and for the latter, we develop a doubly robust estimator that is semiparametrically efficient and yields an optimal regret bound. Finally, we conduct simulation and empirical studies to illustrate the advantages of the proposed methodology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Targeting in Dynamic Systems

    stat.ME 2025-06 conditional novelty 6.0 of 10

    Optimal targeting in Markovian systems reduces to CADE thresholding with state-specific shadow-cost thresholds, estimable via state-level value iteration.

  2. Online Experimental Design With Estimation-Regret Trade-off Under Network Interference

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A Pareto-optimal trade-off between regret and treatment-effect estimation is derived and achieved for bandits with network interference, by compressing the action space through exposure mapping.

  3. Direct Profit Estimation Using Uplift Modeling under Clustered Network Interference

    cs.LG 2025-09 conditional novelty 4.0 of 10

    An AddIPW-based learning objective with cluster-level outcome transformations produces uplift policies that outperform naive methods under strong clustered network interference.

Pith tools