Pith. sign in

REVIEW

Identification of Subgroups With Similar Benefits in Off-Policy Policy Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.14272 v1 pith:4NQORQAH submitted 2021-11-28 cs.LG cs.AIstat.ME

classification cs.LGcs.AIstat.ME
keywords policydecisionpredictionsbaselineaccuratebetterevaluationhtes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Off-policy policy evaluation methods for sequential decision making can be used to help identify if a proposed decision policy is better than a current baseline policy. However, a new decision policy may be better than a baseline policy for some individuals but not others. This has motivated a push towards personalization and accurate per-state estimates of heterogeneous treatment effects (HTEs). Given the limited data present in many important applications, individual predictions can come at a cost to accuracy and confidence in such predictions. We develop a method to balance the need for personalization with confident predictions by identifying subgroups where it is possible to confidently estimate the expected difference in a new decision policy relative to a baseline. We propose a novel loss function that accounts for uncertainty during the subgroup partitioning phase. In experiments, we show that our method can be used to form accurate predictions of HTEs where other methods struggle.

Discussion (0). Sign in to comment.

Pith tools