Pith. sign in

Paper Citation Record · LEDGER

Off-policy estimation with adaptively collected data: the power of online learning

As of 14 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2411.12786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12786 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:58.594829Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy60
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdd83350-4a3f-4855-a636-79ceb0361447 · outbound

This paper cites Effective evaluation using logged bandit feedback from multiple loggers.

Off-policy estimation with adaptively collected data: the power of online learning Effective evaluation using logged bandit feedback from multiple loggers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.256076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.384377Z digest=sha256:a61413288dfaa685102f08f2a0f4a6ce3f4eb226bd03c1f357af883aa9e42124

Observation 13ea1470-4f55-4404-8e50-78c360375fc7 · outbound

This paper cites Thompson sampling for contextu al bandits with linear payoffs.

Off-policy estimation with adaptively collected data: the power of online learning Thompson sampling for contextu al bandits with linear payoffs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.388449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.388449Z digest=sha256:b512d580ba95822fd7ca4cc6dd71e9e872cc7aad792c61d733a03259c2c789e9

Observation ce097e39-82ee-4281-b98d-a34f8d4f4eef · outbound

This paper cites Finite-sample optimal e stimation and inference on average treatment effects under unconfoundedness.

Off-policy estimation with adaptively collected data: the power of online learning Finite-sample optimal e stimation and inference on average treatment effects under unconfoundedness

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.240714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.391895Z digest=sha256:bf4afcdc45a4ac61c82c73a67d728941d1837e0dd1138fbc80486dd2c597ccb4

Observation 8bf30fc0-a75e-412b-bc2b-87bb306bf313 · outbound

This paper cites Counter factual reasoning and learning systems: The example of computational advertising.

Off-policy estimation with adaptively collected data: the power of online learning Counter factual reasoning and learning systems: The example of computational advertising

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.230750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.395022Z digest=sha256:f0cdb6400fb4082a62c1b29997468efac7e16e2f3d76df06dada4b9e43de4748

Observation fd40c365-92ad-4ac5-bbc8-19596f448401 · outbound

This paper cites Double/debiased/neyman machine learning of treatment eff ects.

Off-policy estimation with adaptively collected data: the power of online learning Double/debiased/neyman machine learning of treatment eff ects

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.221129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.397978Z digest=sha256:5680e11dff5880c52cd7372184b135d0164156f555cd143e65428f729ab977e3

Observation e32cfe94-d42a-41ad-9186-b6f943c93bf2 · outbound

This paper cites Double/debiased machine learning for tre atment and structural parameters, 2018.

Off-policy estimation with adaptively collected data: the power of online learning Double/debiased machine learning for tre atment and structural parameters, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.212056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.401013Z digest=sha256:4de379166fcd54fc10099cbf7e94b7bf31af035409fdfe378da0e5625a5a0de7

Observation 2e61cfd6-12a8-4a77-a4c6-12f39503cd48 · outbound

This paper cites Semiparametric e fficient inference in adaptive experiments.

Off-policy estimation with adaptively collected data: the power of online learning Semiparametric e fficient inference in adaptive experiments

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.202829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.404140Z digest=sha256:b58dd6d2d7da8e9ec384b1c00b4c942f6ef7314bae5c0aebc52e39802bcc97ab

Observation 491d4f36-c436-434a-a1b3-0c90b1b30727 · outbound

This paper cites Clip-ogd: An experimental design for adaptive neyman allocation in sequential experiments.

Off-policy estimation with adaptively collected data: the power of online learning Clip-ogd: An experimental design for adaptive neyman allocation in sequential experiments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.193720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.406886Z digest=sha256:2aca857cbbae49af438b7aba2160601a2c879c2f60eb89e33d68c8706d545c4d

Observation dfdd3bd8-5e89-4030-b644-7e841cfa0b74 · outbound

This paper cites Do ubly Robust Policy Evaluation and Optimization.

Off-policy estimation with adaptively collected data: the power of online learning Do ubly Robust Policy Evaluation and Optimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.184495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.409474Z digest=sha256:e86dcd04b8e07538dac75867b8b22361263ee1f3e8a19a29e724a5201db9d860

Observation dff7eb14-9d79-402b-ad90-bf38a1f04644 · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

Off-policy estimation with adaptively collected data: the power of online learning Doubly Robust Policy Evaluation and Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.412086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.412086Z digest=sha256:ff5660e72d45c1e7739966cb68995c1b82a21b07058f7d6098a06c62c775c680

Observation 7d9430af-0516-438e-bd0d-a37e3f968b60 · outbound

This paper cites Overlap in observational studies with high-dimensional covariates.

Off-policy estimation with adaptively collected data: the power of online learning Overlap in observational studies with high-dimensional covariates

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.175673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.415711Z digest=sha256:1776cfebf62da6b11357bc944d56c49caf2c113eec40523fff74e44d973e982a

Observation 2b0c36a5-6efe-4731-b99e-726bed0c03d4 · outbound

This paper cites More robust doubly robust off- policy evaluation.

Off-policy estimation with adaptively collected data: the power of online learning More robust doubly robust off- policy evaluation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.166818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.419407Z digest=sha256:a3fb3270543882df46128887311f6e9dcdecad95e9669c5ed88844d93adec3dd

Observation a62be84f-d5b7-47bf-ae67-fe25208a1530 · outbound

This paper cites Off-policy evalua- tion with deficient support using side information.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy evalua- tion with deficient support using side information

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.157620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.422621Z digest=sha256:e5e7a3b74f5bd18efbbbb7d338c210db54b09ed09b782a4ad76a1f120577b9af

Observation 42f631f7-b5c5-4b08-ba6e-5e56dd33b700 · outbound

This paper cites On choosing and bounding pr obability metrics.

Off-policy estimation with adaptively collected data: the power of online learning On choosing and bounding pr obability metrics

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.148297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.425943Z digest=sha256:76abd51b9b7af1fb257574d21d4488eda752f4a87648cc867429f41f81fd21e1

Observation 4efef9cc-4459-402c-82b4-b956e067a7e3 · outbound

This paper cites Some limit theorems for empirical pro cesses.

Off-policy estimation with adaptively collected data: the power of online learning Some limit theorems for empirical pro cesses

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.138932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.429094Z digest=sha256:bf996083bb2519da47ea249d0892e6d610724df38a9efad11b2bd9ca7c30ce58

Observation 41617ce4-8cec-4cd5-a09a-1744a5b61858 · outbound

This paper cites Confidence Intervals for Policy Evaluation in Adaptive Experiments.

Off-policy estimation with adaptively collected data: the power of online learning Confidence Intervals for Policy Evaluation in Adaptive Experiments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.432460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.432460Z digest=sha256:a324119f53268e3765c3fbd7afd5278221fa242b8269d5b6741eb3c6f7489cd8

Observation 0766f0a5-fa30-4f30-a69c-f9b62ed2d9d7 · outbound

This paper cites Confidence inter- vals for policy evaluation in adaptive experiments.

Off-policy estimation with adaptively collected data: the power of online learning Confidence inter- vals for policy evaluation in adaptive experiments

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.129076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.436031Z digest=sha256:b76b534724f46a47eac75a1a4daf82a16e3bfa4495edcfac6917205fc6b62803

Observation 692f9d41-76df-4b6a-a04a-de2cafda0a08 · outbound

This paper cites Introduction to online convex optimization.

Off-policy estimation with adaptively collected data: the power of online learning Introduction to online convex optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.119387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.439220Z digest=sha256:0ede09dbe5bfb1e99328c72fcd4236ef35661c063d80be8b0ce6c3a907bedf5f

Observation 51e107c3-03f6-491f-bbdb-60360cedec4b · outbound

This paper cites Weighted average importance sampling and def ensive mixture distributions.

Off-policy estimation with adaptively collected data: the power of online learning Weighted average importance sampling and def ensive mixture distributions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.109820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.442438Z digest=sha256:eeca2d69ce3d85a3001df03f534f813c7ec5a31c65bbbb1a698bbb46057d3dac

Observation fddd75ce-9f02-4eaa-b5f8-1d1e2d9534fe · outbound

This paper cites Efficient estim ation of average treatment effects using the estimated propensity score.

Off-policy estimation with adaptively collected data: the power of online learning Efficient estim ation of average treatment effects using the estimated propensity score

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.100505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.445622Z digest=sha256:6962cdc8dd797354fcad7949e6a3ca2616195d017efd31ec97b2ce812be16940

Observation ff2cfef6-8ffe-4454-9e4f-cef40093e66d · outbound

This paper cites A generalization of sa mpling without replacement from a finite universe.

Off-policy estimation with adaptively collected data: the power of online learning A generalization of sa mpling without replacement from a finite universe

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.091187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.448907Z digest=sha256:176a8caee2aae46e881d0101e4cd3cf8c3fc34ace87cf73e7472a6770185c07a

Observation 765d4f6a-cc66-4f1b-98c3-af5a1bfbf90c · outbound

This paper cites Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon.

Off-policy estimation with adaptively collected data: the power of online learning Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.080484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.452186Z digest=sha256:500b5f6a2d178a831962d0b3a94016fa76b4b9d37ec242af6371aa3a932eca66

Observation 42d8ad55-67c2-44d3-b7e5-37b9018e7402 · outbound

This paper cites Nonparametric estimation of average treatme nt effects under exogeneity: A review.

Off-policy estimation with adaptively collected data: the power of online learning Nonparametric estimation of average treatme nt effects under exogeneity: A review

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.071509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.455371Z digest=sha256:58f1e891d40805885e7bf4cbe656c55bad64ee74ebeb3cb6a500106da9f99dff

Observation ab7db0ae-314d-4100-bf24-ca9e988eb639 · outbound

This paper cites Causal inference in statistics, social, and biomedical sci ences.

Off-policy estimation with adaptively collected data: the power of online learning Causal inference in statistics, social, and biomedical sci ences

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.062808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.458429Z digest=sha256:c1c35a2bda6696ef177e7a0c08b092a577fdab8612167bf57c0fa3e1640ee144

Observation 54c95186-535e-4c35-b68f-a97bb9c66e74 · outbound

This paper cites Truncated importance sampling.

Off-policy estimation with adaptively collected data: the power of online learning Truncated importance sampling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.053695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.461628Z digest=sha256:6a4536abe8a347ea82e347cb893b6af816fd5fac4e6ceb3529b6737a867fdc7b

Observation 33ecec2d-451b-42fb-8475-e8cb28bee0c2 · outbound

This paper cites Policy learning "without" overlap: Pessimism and generalized empirical Bernstein's inequality.

Off-policy estimation with adaptively collected data: the power of online learning Policy learning "without" overlap: Pessimism and generalized empirical Bernstein's inequality

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.464904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.464904Z digest=sha256:c4d7a9b4858564c9683bc46dcfd98ef81ba1f2b194ee36cd3ed2e0a74944e4fa

Observation 73586c0e-5b7c-464f-9f2b-1fd543b59673 · outbound

This paper cites Optimal off- policy evaluation from multiple logging policies.

Off-policy estimation with adaptively collected data: the power of online learning Optimal off- policy evaluation from multiple logging policies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.044609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.468517Z digest=sha256:02e8737c3c42a6c24e532def730b12a0d720f479f5a1be9153e33e265117b307

Observation 1415e787-1b6c-4ec2-8a14-6cd007dc1bf8 · outbound

This paper cites Policy evaluation and optimization w ith continuous treatments.

Off-policy estimation with adaptively collected data: the power of online learning Policy evaluation and optimization w ith continuous treatments

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.034761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.471851Z digest=sha256:8027da3160911ed72c3c6f16f362c9d3eb6a35c208986ddb03fc732c20ed83a3

Observation 7fe2337c-4657-423b-b733-7b91b5f3a839 · outbound

This paper cites an unresolved cited work.

Off-policy estimation with adaptively collected data: the power of online learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:59.024823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.474877Z digest=sha256:33db0d5c3028e38b6f3eac2b3fecd0efed6379f1c45cb8108b6a393481471bfd

Observation 2bd7d3f3-db7f-4c00-af86-588d5af2e177 · outbound

This paper cites Off-po licy confidence sequences.

Off-policy estimation with adaptively collected data: the power of online learning Off-po licy confidence sequences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.015260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.477767Z digest=sha256:d5bf7f12ff62a080c8a86c09f64af0d783a9eb230a07041d89fd52cdd8717598

Observation ef8513de-cbe9-4a65-8952-3db835d22697 · outbound

This paper cites Efficient Adaptive Experimental Design for Average Treatment Effect Estimation.

Off-policy estimation with adaptively collected data: the power of online learning Efficient Adaptive Experimental Design for Average Treatment Effect Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.481508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.481508Z digest=sha256:fdfcf972cf25d37ad56e6a493eb1f41a15875e24891e3d70f83e8541739cc82e

Observation 70e4033f-2abc-4256-af5e-2f3bbe37c276 · outbound

This paper cites Irregular identification, support conditions, and inverse weight estima- tion.

Off-policy estimation with adaptively collected data: the power of online learning Irregular identification, support conditions, and inverse weight estima- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.005109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.484430Z digest=sha256:1af4a2af884dafbd72767ec2d9b8cfd1081a2bd1a195f86da7a7ebe33338683c

Observation c9c49b9a-9ba4-470a-b4ba-53cfe7aa30da · outbound

This paper cites Asymptotically efficient ada ptive allocation rules.

Off-policy estimation with adaptively collected data: the power of online learning Asymptotically efficient ada ptive allocation rules

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.995176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.487162Z digest=sha256:b932653d628533304ea771bf7ac6e15bebb3f6693afe5280c4a61cd6739717e7

Observation 8340bea5-6a62-4d7d-9302-6f9027d95a2e · outbound

This paper cites Bandit algorithms.

Off-policy estimation with adaptively collected data: the power of online learning Bandit algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.489678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.489678Z digest=sha256:8235829ddf5727422e45282b6496ed391de7f42a4255341dc15c91a0f66dcd73

Observation 71145d63-232b-40ae-af3c-f59f7086d9c7 · outbound

This paper cites Local metric learning for off-policy evaluation in contextua l bandits with continuous actions.

Off-policy estimation with adaptively collected data: the power of online learning Local metric learning for off-policy evaluation in contextua l bandits with continuous actions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.979962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.492191Z digest=sha256:a908772d23c89e1d2ac6fec818db549b007f5d10b5872f6fe04c1e3f10e93bee

Observation 59b946f9-4f6c-4513-956b-4e5b40cb731d · outbound

This paper cites Distribution-free assessment of population overlap in observational studies.

Off-policy estimation with adaptively collected data: the power of online learning Distribution-free assessment of population overlap in observational studies

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.970619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.494876Z digest=sha256:cf7d28cb26497a41f9d018ba56214a93dfebf34bd064e893e57916825c5f518d

Observation 74c1930e-a5a0-49b4-a907-d89f17998b77 · outbound

This paper cites Sharp high-probability sample complexities for policy evaluation with linear function approximation.

Off-policy estimation with adaptively collected data: the power of online learning Sharp high-probability sample complexities for policy evaluation with linear function approximation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.961990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.497374Z digest=sha256:85e860d924ff4e9f4d96157c7f9eceee7b6aedeb6f94a7271e3ad474199aaa0f

Observation 514c7f5d-4fe5-41ad-b7f4-1a26385c3c6b · outbound

This paper cites Toward minimaxoff-policy value estimation.

Off-policy estimation with adaptively collected data: the power of online learning Toward minimaxoff-policy value estimation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.952463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.500400Z digest=sha256:7415b96f58f9fe0cac493a17c5e1aefc9465fade83a43c8fab6f30a70695be39

Observation 702998d8-e69f-46f9-a818-9ed76ae8d96e · outbound

This paper cites Statistical analysis with missing data , volume 793.

Off-policy estimation with adaptively collected data: the power of online learning Statistical analysis with missing data , volume 793

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.942657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.503586Z digest=sha256:164fd538b158cdc7df32348b2e362f5f09f3f37dbe31417097b09cba73ceb248

Observation 4f599360-83c2-46fe-a12f-0d273eef2325 · outbound

This paper cites Statistical infer ence for the mean outcome under a possibly non-unique optimal treatment strategy.

Off-policy estimation with adaptively collected data: the power of online learning Statistical infer ence for the mean outcome under a possibly non-unique optimal treatment strategy

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.933079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.506578Z digest=sha256:c26bc1f04083507c96d5b585174973b604ef6a17d05a57b7ef55b45f47cbcbcb

Observation 880a205a-4902-443e-ad93-6d7f137db380 · outbound

This paper cites Min imax off-policy evaluation for multi-armed bandits.

Off-policy estimation with adaptively collected data: the power of online learning Min imax off-policy evaluation for multi-armed bandits

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.923814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.509538Z digest=sha256:c4018821d49a9d6266ae2ec7baeba624a63e3ec02bf74f71c1d1c32d07a188ff

Observation bceb8c21-f1ac-4b72-a0dd-d67635c8d435 · outbound

This paper cites Off-policy estimation of linear functionals: Non-asymptotic theory for semi-parametric efficiency.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy estimation of linear functionals: Non-asymptotic theory for semi-parametric efficiency

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:43:58.640071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.512501Z digest=sha256:61ee6c84abb6a682a01c6e393c5b2d9d6ea43714c7e68a8bececaaab9b713549

Observation 54415f53-83e2-4302-bbeb-06c16943141e · outbound

This paper cites Efficient counter factual learning from bandit feedback.

Off-policy estimation with adaptively collected data: the power of online learning Efficient counter factual learning from bandit feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.914524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.516147Z digest=sha256:425c1e9aa0a372a08aeb3c2be5d240a821deb5a3ff234ec610ffd1aada517a67

Observation c620aecc-93d6-472a-8f1d-1ebf3afd61b4 · outbound

This paper cites Offline policy evaluation in large action spaces via outcome-oriented action group ing.

Off-policy estimation with adaptively collected data: the power of online learning Offline policy evaluation in large action spaces via outcome-oriented action group ing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.905104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.519430Z digest=sha256:864385944eb67e5969a4375d210f44559423b8834f3fd17a1f3f9f22f05b5400

Observation 70e47d55-6214-43f6-80ad-4964406a5e51 · outbound

This paper cites Online non-parametric regression.

Off-policy estimation with adaptively collected data: the power of online learning Online non-parametric regression

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.896509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.522574Z digest=sha256:181a40e4e57fcb1ee36bfb719616366b372e04cd76f41b46f1d1cf3d13e465b6

Observation 6a50acc1-fa07-4676-9d78-ace03cdeeb35 · outbound

This paper cites Seque ntial complexities and uniform mar- tingale laws of large numbers.

Off-policy estimation with adaptively collected data: the power of online learning Seque ntial complexities and uniform mar- tingale laws of large numbers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.888627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.525570Z digest=sha256:7cae632ab7018159bd196bf9c36c426502ac69a61f26baddb786012e7dc11a6f

Observation 0ad9acdd-50b3-4190-98a3-c0138b3c4caf · outbound

This paper cites Relax and ra ndomize: From value to algorithms.

Off-policy estimation with adaptively collected data: the power of online learning Relax and ra ndomize: From value to algorithms

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.879912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.528627Z digest=sha256:4c8f57c4d5f527f17879f98f3dd5662f033f741a56c44bf87ddb48d46a1e964f

Observation 61554459-13d2-46f6-85d9-e83c8c8689db · outbound

This paper cites Comment: Performance of double-robust estimators when” inverse probability” weights ar e highly variable.

Off-policy estimation with adaptively collected data: the power of online learning Comment: Performance of double-robust estimators when” inverse probability” weights ar e highly variable

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.870744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.531540Z digest=sha256:2dfab129adc7cafb3324100e3395304a76be90cea0a3d863705c2427fd02a002

Observation 17d2bccb-68fd-429f-bd91-b582d4ae5e5d · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.

Off-policy estimation with adaptively collected data: the power of online learning Semiparametric efficiency in multivariate regression models with missing data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.860460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.534555Z digest=sha256:67a172b1bd4b337be02993133705ac82dcce32931f3e75059f915b0c36e16c7b

Observation 2ef21956-0e7a-464d-b7e4-54e89463da98 · outbound

This paper cites Estimatio n of regression coefficients when some regressors are not always observed.

Off-policy estimation with adaptively collected data: the power of online learning Estimatio n of regression coefficients when some regressors are not always observed

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.851091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.537719Z digest=sha256:2e31e077cf5b34da4d7a049946f56776e15650c5b563824d981ec9946ce7c7c1

Observation d4c43c02-d4e0-4815-9509-efa4d20b4331 · outbound

This paper cites Analysis o f semiparametric regression models for repeated outcomes in the presence of missing data.

Off-policy estimation with adaptively collected data: the power of online learning Analysis o f semiparametric regression models for repeated outcomes in the presence of missing data

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.841575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.540801Z digest=sha256:a730a3cda34d919261cf84160b9c480314b9f54966dd994f7c28e5ae501417f5

Observation df24683d-c045-4d2b-b50f-09943e68d932 · outbound

This paper cites A tutorial on thompson sampling.

Off-policy estimation with adaptively collected data: the power of online learning A tutorial on thompson sampling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.831606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.543919Z digest=sha256:b27d650e1b2bda6e982744f2aa61b56c3859d7e293c0634ba1fb25022de6db0c

Observation 7dcc1f49-cef1-4b27-a663-9803d30b5ea5 · outbound

This paper cites Off-Policy Evaluation for Large Action Spaces via Embeddings.

Off-policy estimation with adaptively collected data: the power of online learning Off-Policy Evaluation for Large Action Spaces via Embeddings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.546929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.546929Z digest=sha256:ab2da2650b18814da4b334ede42c2551de4cb3ebf57ab85a2a4017795e7f5bb8

Observation de50d675-da0e-4cb4-8ef2-25962971d72f · outbound

This paper cites Off-policy ev aluation for large action spaces via conjunct effect modeling.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy ev aluation for large action spaces via conjunct effect modeling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.822558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.550231Z digest=sha256:5617787c8ddaec1288114608f411d496a0bd8d211d584c3d1af146d7e8cf9798

Observation 2b63066d-06c7-41c8-b289-5469d9008d7f · outbound

This paper cites Lear ning from logged implicit exploration data.

Off-policy estimation with adaptively collected data: the power of online learning Lear ning from logged implicit exploration data

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.814022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.553443Z digest=sha256:79a1023deb955fe31378cd292d92ebd92aa147b9ebcfc68831fdfbab3170ae4d

Observation b6bdbb72-811a-4e54-9836-42f2ed0d9109 · outbound

This paper cites Doubly robust off-policy evaluation with shrinkage.

Off-policy estimation with adaptively collected data: the power of online learning Doubly robust off-policy evaluation with shrinkage

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.805656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.556638Z digest=sha256:fb7f38bbd61b04d3e4bc8d3d80d81d136ed8f21b86cf1b7a007ed83727baa0d9

Observation 8b319e47-8d66-4cf5-ad49-7037c7300635 · outbound

This paper cites Cab: Continuous adaptive blend- ing for policy evaluation and learning.

Off-policy estimation with adaptively collected data: the power of online learning Cab: Continuous adaptive blend- ing for policy evaluation and learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.797671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.559742Z digest=sha256:e134a98e3338b9982e6f7df6b6be804fbd2ff9b43b1ec61ec4d09a3b9d99c3e1

Observation ecde691c-e9cc-4d95-b1c5-d1544122943d · outbound

This paper cites The self-normalized estimator for counterfactual learning.

Off-policy estimation with adaptively collected data: the power of online learning The self-normalized estimator for counterfactual learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.789061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.562987Z digest=sha256:b22885eb44012a7ea225ab6791285ddbf7a81a598624c24d48a5226b14035b25

Observation e38b07e5-19f9-49d3-a04b-42ea05405a31 · outbound

This paper cites Data-efficient off-policy policy eva luation for reinforcement learn- ing.

Off-policy estimation with adaptively collected data: the power of online learning Data-efficient off-policy policy eva luation for reinforcement learn- ing

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.780101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.566553Z digest=sha256:ca7a5baf73eaab8fbc011fe7bbf00bba831a117ebe447bfbefde35d6f00559d9

Observation c73b29ef-8313-497e-9d84-de6f5950c774 · outbound

This paper cites On the likelihood that one unknown probability ex ceeds another in view of the evidence of two samples.

Off-policy estimation with adaptively collected data: the power of online learning On the likelihood that one unknown probability ex ceeds another in view of the evidence of two samples

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.771008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.569592Z digest=sha256:c7ae6280b8e9412563ad4bf99d0732e4ff7aff57ec8c73342d2066adb3315235

Observation d4634878-2d7b-4c84-a5b6-a07f3bddd642 · outbound

This paper cites The construction and analysis of adaptive group sequential designs.

Off-policy estimation with adaptively collected data: the power of online learning The construction and analysis of adaptive group sequential designs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.761483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.572740Z digest=sha256:29c51e58e74a4e52ff2dfc9ad27b362009751a9ab85a0b299c1211bdca5d110a

Observation 55ff0e85-6000-421d-8fef-6a9d3b9cd250 · outbound

This paper cites High-dimensional statistics: A non-asymptotic viewpoint , volume 48.

Off-policy estimation with adaptively collected data: the power of online learning High-dimensional statistics: A non-asymptotic viewpoint , volume 48

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.752267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.575839Z digest=sha256:9eae1e705db85e43a22b2c3c19cda77d82756f4ffbd95e1e2bee1600890a683a

Observation 0525e20e-def8-4f05-90d4-6c53dfd5e888 · outbound

This paper cites Oracle-e fficient pessimism: Offline policy optimization in contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Oracle-e fficient pessimism: Offline policy optimization in contextual bandits

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.743269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.578928Z digest=sha256:1925f29a0251f45713de0d3d80d462d09dce8c0b409dab2c38f3ff832c967d4b

Observation 0ac61766-8332-43d3-b8f1-f7864532a8c4 · outbound

This paper cites Optimal and ad aptive off-policy evaluation in contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Optimal and ad aptive off-policy evaluation in contextual bandits

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.733433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.581648Z digest=sha256:1a736006ee75e42628e7938bbee52db240f262d5a064a124cc58d598e9f8064e

Observation 649e6e36-317d-4fea-b7b2-eea3bb886478 · outbound

This paper cites Anytime-valid off-policy inference for contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Anytime-valid off-policy inference for contextual bandits

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.724253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.584182Z digest=sha256:862d89e7f683a8d506e900a7f211c30a149fb7d741421925687199947d53d4b2

Observation 88d8437e-3d16-4f29-bd16-c7139087686d · outbound

This paper cites Asymptotic inference of causal effects with o bservational studies trimmed by the estimated propensity scores.

Off-policy estimation with adaptively collected data: the power of online learning Asymptotic inference of causal effects with o bservational studies trimmed by the estimated propensity scores

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.715417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.586736Z digest=sha256:ad1a280162a96f0cf12aebc4e51aa0f19d43bc293520476893fbb5710de51f96

Observation 101db0d1-b20d-4d34-b628-83ec133cee91 · outbound

This paper cites Off-policy evaluation via adaptive weighting with data from contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy evaluation via adaptive weighting with data from contextual bandits

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.706252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.589198Z digest=sha256:63972189be04e8c51ee60d5b16ee8721ce1356a9ee1389fa08f9319f0cde899c

Observation 7a83b08d-89d3-434a-b576-448588353f15 · outbound

This paper cites Policy learning with adaptively collected data.

Off-policy estimation with adaptively collected data: the power of online learning Policy learning with adaptively collected data

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.696569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.591909Z digest=sha256:e906c5e4def2deaac3b37dcae546997879f466300fd8b1205222dc381a9379c2

Observation 755424cb-481b-4825-a410-50b2442d6e48 · outbound

This paper cites Inference for batched bandits.

Off-policy estimation with adaptively collected data: the power of online learning Inference for batched bandits

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.686245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:58.594829Z digest=sha256:7ea3332e5f39fe03310b9436c47578a6189ffb7e86b21471526545f5959f9fa4

Pith citing papers

No inbound Pith citation observations are available.