Pith. sign in

Paper Citation Record · LEDGER

Off-policy estimation with adaptively collected data: the power of online learning

As of 15 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2411.12786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12786 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:58.594829Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy60
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdd83350-4a3f-4855-a636-79ceb0361447 · outbound

This paper cites Effective evaluation using logged bandit feedback from multiple loggers.

Off-policy estimation with adaptively collected data: the power of online learning Effective evaluation using logged bandit feedback from multiple loggers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.256076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.384377Z digest=sha256:e859435edb30df521d6392c50ebc92cd727f9c89d58707ce9e75f7e8191d22fe

Observation 13ea1470-4f55-4404-8e50-78c360375fc7 · outbound

This paper cites Thompson sampling for contextu al bandits with linear payoffs.

Off-policy estimation with adaptively collected data: the power of online learning Thompson sampling for contextu al bandits with linear payoffs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.388449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.388449Z digest=sha256:4fa6300450040897e414fbe4b0d62f0d9d39b7160ffcd8103bfba7fec8e825be

Observation ce097e39-82ee-4281-b98d-a34f8d4f4eef · outbound

This paper cites Finite-sample optimal e stimation and inference on average treatment effects under unconfoundedness.

Off-policy estimation with adaptively collected data: the power of online learning Finite-sample optimal e stimation and inference on average treatment effects under unconfoundedness

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.240714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.391895Z digest=sha256:1e2cca53456db57e0ccc10e1a2103ea5d44a946b9c5cc88534c7e3d3a5bc71ae

Observation 8bf30fc0-a75e-412b-bc2b-87bb306bf313 · outbound

This paper cites Counter factual reasoning and learning systems: The example of computational advertising.

Off-policy estimation with adaptively collected data: the power of online learning Counter factual reasoning and learning systems: The example of computational advertising

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.230750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.395022Z digest=sha256:c39b368701f4cefa303aa281cecad65bd98e8c79ecd64cbad5996b206cf91753

Observation fd40c365-92ad-4ac5-bbc8-19596f448401 · outbound

This paper cites Double/debiased/neyman machine learning of treatment eff ects.

Off-policy estimation with adaptively collected data: the power of online learning Double/debiased/neyman machine learning of treatment eff ects

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.221129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.397978Z digest=sha256:fe941d81b0fe4c6d3a4843ca9e4c765ea8e6645587016fb60e481084eed1887b

Observation e32cfe94-d42a-41ad-9186-b6f943c93bf2 · outbound

This paper cites Double/debiased machine learning for tre atment and structural parameters, 2018.

Off-policy estimation with adaptively collected data: the power of online learning Double/debiased machine learning for tre atment and structural parameters, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.212056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.401013Z digest=sha256:f5e93c3d82ed45b50ad966e5475d14608d757f69254b41be4d0afeb92058066f

Observation 2e61cfd6-12a8-4a77-a4c6-12f39503cd48 · outbound

This paper cites Semiparametric e fficient inference in adaptive experiments.

Off-policy estimation with adaptively collected data: the power of online learning Semiparametric e fficient inference in adaptive experiments

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.202829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.404140Z digest=sha256:f91f381ba9de80e627b5ca7cb83419ea91b27bb8e9b0672ed92548eb5ade5034

Observation 491d4f36-c436-434a-a1b3-0c90b1b30727 · outbound

This paper cites Clip-ogd: An experimental design for adaptive neyman allocation in sequential experiments.

Off-policy estimation with adaptively collected data: the power of online learning Clip-ogd: An experimental design for adaptive neyman allocation in sequential experiments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.193720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.406886Z digest=sha256:cc34a8b173885352e21a3dddff0813ad143abab660a94c313725062f6a455711

Observation dfdd3bd8-5e89-4030-b644-7e841cfa0b74 · outbound

This paper cites Do ubly Robust Policy Evaluation and Optimization.

Off-policy estimation with adaptively collected data: the power of online learning Do ubly Robust Policy Evaluation and Optimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.184495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.409474Z digest=sha256:37ea5f20317821a7dac83344162a61f71427d0d61b32ad9d7a5303d48d371bc2

Observation dff7eb14-9d79-402b-ad90-bf38a1f04644 · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

Off-policy estimation with adaptively collected data: the power of online learning Doubly Robust Policy Evaluation and Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.412086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.412086Z digest=sha256:bc80596e548b671ed6f0288b0f98cab01649b1288f3b8445b877fd4c8ecff9a2

Observation 7d9430af-0516-438e-bd0d-a37e3f968b60 · outbound

This paper cites Overlap in observational studies with high-dimensional covariates.

Off-policy estimation with adaptively collected data: the power of online learning Overlap in observational studies with high-dimensional covariates

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.175673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.415711Z digest=sha256:303626bc7d52986cd74f1924a0956825287feb86183427a2674c362367195011

Observation 2b0c36a5-6efe-4731-b99e-726bed0c03d4 · outbound

This paper cites More robust doubly robust off- policy evaluation.

Off-policy estimation with adaptively collected data: the power of online learning More robust doubly robust off- policy evaluation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.166818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.419407Z digest=sha256:d0d69203d39dba1506dd75828c24bbff116b76aa376c6c8a4f0b3bc6089c2458

Observation a62be84f-d5b7-47bf-ae67-fe25208a1530 · outbound

This paper cites Off-policy evalua- tion with deficient support using side information.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy evalua- tion with deficient support using side information

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.157620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.422621Z digest=sha256:a860e1f998e531dc5b516b3b352831c1c734460aa4558c108af6af62a5712457

Observation 42f631f7-b5c5-4b08-ba6e-5e56dd33b700 · outbound

This paper cites On choosing and bounding pr obability metrics.

Off-policy estimation with adaptively collected data: the power of online learning On choosing and bounding pr obability metrics

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.148297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.425943Z digest=sha256:caa4fe04a4809ac9f5500356a864d3c1cf1ca80ad89ba830f4b6236abdc23303

Observation 4efef9cc-4459-402c-82b4-b956e067a7e3 · outbound

This paper cites Some limit theorems for empirical pro cesses.

Off-policy estimation with adaptively collected data: the power of online learning Some limit theorems for empirical pro cesses

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.138932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.429094Z digest=sha256:2ca837a5fc1a9f9e82eda6329a26dea6612146f843fd0a858f22725bd93b5c51

Observation 41617ce4-8cec-4cd5-a09a-1744a5b61858 · outbound

This paper cites Confidence Intervals for Policy Evaluation in Adaptive Experiments.

Off-policy estimation with adaptively collected data: the power of online learning Confidence Intervals for Policy Evaluation in Adaptive Experiments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.432460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.432460Z digest=sha256:6ea31b1eabfb0f481131822bb7977a3d5b4c1cf7a6ef7173d0afba819a2a41f7

Observation 0766f0a5-fa30-4f30-a69c-f9b62ed2d9d7 · outbound

This paper cites Confidence inter- vals for policy evaluation in adaptive experiments.

Off-policy estimation with adaptively collected data: the power of online learning Confidence inter- vals for policy evaluation in adaptive experiments

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.129076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.436031Z digest=sha256:4364e9ea416e7224272bfebd4f98e197261127d3733853734600fec4100d06e4

Observation 692f9d41-76df-4b6a-a04a-de2cafda0a08 · outbound

This paper cites Introduction to online convex optimization.

Off-policy estimation with adaptively collected data: the power of online learning Introduction to online convex optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.119387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.439220Z digest=sha256:7a70144dafebe98ab2d82e1084a7fc8b8fe3d5ff92add93d21c3a6898649462f

Observation 51e107c3-03f6-491f-bbdb-60360cedec4b · outbound

This paper cites Weighted average importance sampling and def ensive mixture distributions.

Off-policy estimation with adaptively collected data: the power of online learning Weighted average importance sampling and def ensive mixture distributions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.109820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.442438Z digest=sha256:4e33d46c8551b1d037e61541eae25445ace2fe54f3ce35463b978a6f9ea50b64

Observation fddd75ce-9f02-4eaa-b5f8-1d1e2d9534fe · outbound

This paper cites Efficient estim ation of average treatment effects using the estimated propensity score.

Off-policy estimation with adaptively collected data: the power of online learning Efficient estim ation of average treatment effects using the estimated propensity score

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.100505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.445622Z digest=sha256:e8fbbe0aaca63e80190ad8b188822f3f046a9605e54f7c8c2739b8b09815be9a

Observation ff2cfef6-8ffe-4454-9e4f-cef40093e66d · outbound

This paper cites A generalization of sa mpling without replacement from a finite universe.

Off-policy estimation with adaptively collected data: the power of online learning A generalization of sa mpling without replacement from a finite universe

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.091187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.448907Z digest=sha256:1645e98844f763398ea436a68d6da0be8262e028ad1ffc114cec98ad414a9b51

Observation 765d4f6a-cc66-4f1b-98c3-af5a1bfbf90c · outbound

This paper cites Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon.

Off-policy estimation with adaptively collected data: the power of online learning Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.080484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.452186Z digest=sha256:e7953c38d1498de8d377136fe38763ae90e5d47f7fe22c8ddfb710547ec5ad93

Observation 42d8ad55-67c2-44d3-b7e5-37b9018e7402 · outbound

This paper cites Nonparametric estimation of average treatme nt effects under exogeneity: A review.

Off-policy estimation with adaptively collected data: the power of online learning Nonparametric estimation of average treatme nt effects under exogeneity: A review

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.071509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.455371Z digest=sha256:1b4fe3eb50fcf6f6fdd2b06c8a34619b777141714adeb7ce2868a07c5dd607b4

Observation ab7db0ae-314d-4100-bf24-ca9e988eb639 · outbound

This paper cites Causal inference in statistics, social, and biomedical sci ences.

Off-policy estimation with adaptively collected data: the power of online learning Causal inference in statistics, social, and biomedical sci ences

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.062808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.458429Z digest=sha256:726f120228611c720d47218c244426200f1c6d6c67d989d3878598e8a5d70b1a

Observation 54c95186-535e-4c35-b68f-a97bb9c66e74 · outbound

This paper cites Truncated importance sampling.

Off-policy estimation with adaptively collected data: the power of online learning Truncated importance sampling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.053695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.461628Z digest=sha256:5339aff4b448a1e3ab9aeb88559b8e4544684fa45920664b05a58d5f36442f9e

Observation 33ecec2d-451b-42fb-8475-e8cb28bee0c2 · outbound

This paper cites Policy learning "without" overlap: Pessimism and generalized empirical Bernstein's inequality.

Off-policy estimation with adaptively collected data: the power of online learning Policy learning "without" overlap: Pessimism and generalized empirical Bernstein's inequality

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.464904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.464904Z digest=sha256:a707be3e7081e48e53e828a57abe5ca0eccc9e8399d02208cf070045dbfd9599

Observation 73586c0e-5b7c-464f-9f2b-1fd543b59673 · outbound

This paper cites Optimal off- policy evaluation from multiple logging policies.

Off-policy estimation with adaptively collected data: the power of online learning Optimal off- policy evaluation from multiple logging policies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.044609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.468517Z digest=sha256:c51093d35322c27531d948eaa3bf2c54535e429f924b21644dc07f9af17916e3

Observation 1415e787-1b6c-4ec2-8a14-6cd007dc1bf8 · outbound

This paper cites Policy evaluation and optimization w ith continuous treatments.

Off-policy estimation with adaptively collected data: the power of online learning Policy evaluation and optimization w ith continuous treatments

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.034761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.471851Z digest=sha256:9181ed3a3fce76928bf17f949124db5577676aa3c603f13bc9e095d682a7d59f

Observation 7fe2337c-4657-423b-b733-7b91b5f3a839 · outbound

This paper cites an unresolved cited work.

Off-policy estimation with adaptively collected data: the power of online learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:59.024823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.474877Z digest=sha256:8ec1b0f2748413df2e6ff67d2eba9dfed89ee9c38a1d0f890c5ead4ed3f8b183

Observation 2bd7d3f3-db7f-4c00-af86-588d5af2e177 · outbound

This paper cites Off-po licy confidence sequences.

Off-policy estimation with adaptively collected data: the power of online learning Off-po licy confidence sequences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.015260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.477767Z digest=sha256:cd26f914cec9bc5de261d19e0a859abdbdb3842944eb84e950bd9c5569c7b390

Observation ef8513de-cbe9-4a65-8952-3db835d22697 · outbound

This paper cites Efficient Adaptive Experimental Design for Average Treatment Effect Estimation.

Off-policy estimation with adaptively collected data: the power of online learning Efficient Adaptive Experimental Design for Average Treatment Effect Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.481508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.481508Z digest=sha256:6d43080f777033ee4105707780f58fde0aa20104ac6dfdfc0ea3a2013916b6d5

Observation 70e4033f-2abc-4256-af5e-2f3bbe37c276 · outbound

This paper cites Irregular identification, support conditions, and inverse weight estima- tion.

Off-policy estimation with adaptively collected data: the power of online learning Irregular identification, support conditions, and inverse weight estima- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:59.005109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.484430Z digest=sha256:9a45d4a7ef076f7bd7e01548dfbb2d92aa67a83d2bc3538f6b1e0db4cc17fd63

Observation c9c49b9a-9ba4-470a-b4ba-53cfe7aa30da · outbound

This paper cites Asymptotically efficient ada ptive allocation rules.

Off-policy estimation with adaptively collected data: the power of online learning Asymptotically efficient ada ptive allocation rules

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.995176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.487162Z digest=sha256:1fec90a0a3f087b31558207dffbaa4d8de7bed5b75a564c94c5e29f02d9a0ca2

Observation 8340bea5-6a62-4d7d-9302-6f9027d95a2e · outbound

This paper cites Bandit algorithms.

Off-policy estimation with adaptively collected data: the power of online learning Bandit algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.489678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.489678Z digest=sha256:b2e66c1b11aca6c1baf10cd39c8fafa091d0e74b098e015e0d19283ca73db6a9

Observation 71145d63-232b-40ae-af3c-f59f7086d9c7 · outbound

This paper cites Local metric learning for off-policy evaluation in contextua l bandits with continuous actions.

Off-policy estimation with adaptively collected data: the power of online learning Local metric learning for off-policy evaluation in contextua l bandits with continuous actions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.979962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.492191Z digest=sha256:bb49b252361d64cd1aec56b84c1fe8dc6ad5163b7a5a130f5ae1a74198be51e7

Observation 59b946f9-4f6c-4513-956b-4e5b40cb731d · outbound

This paper cites Distribution-free assessment of population overlap in observational studies.

Off-policy estimation with adaptively collected data: the power of online learning Distribution-free assessment of population overlap in observational studies

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.970619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.494876Z digest=sha256:d98b503ce4dc72f32f41aa6d0122ae01565c3b13b1e04d290ab8d418f169088b

Observation 74c1930e-a5a0-49b4-a907-d89f17998b77 · outbound

This paper cites Sharp high-probability sample complexities for policy evaluation with linear function approximation.

Off-policy estimation with adaptively collected data: the power of online learning Sharp high-probability sample complexities for policy evaluation with linear function approximation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.961990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.497374Z digest=sha256:c82422edb3afd7409cdc58caadd17423cbfa288a04c8ea339b254c67c78e2737

Observation 514c7f5d-4fe5-41ad-b7f4-1a26385c3c6b · outbound

This paper cites Toward minimaxoff-policy value estimation.

Off-policy estimation with adaptively collected data: the power of online learning Toward minimaxoff-policy value estimation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.952463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.500400Z digest=sha256:5290df5306ad793f8294fc928fcaf3138ddd352cc7240bf51f9655a8160a2a00

Observation 702998d8-e69f-46f9-a818-9ed76ae8d96e · outbound

This paper cites Statistical analysis with missing data , volume 793.

Off-policy estimation with adaptively collected data: the power of online learning Statistical analysis with missing data , volume 793

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.942657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.503586Z digest=sha256:ead19033f149139ff8cb9682ee7c4a8d7bc80ab5caf8eb9a80efe7c36eacde3e

Observation 4f599360-83c2-46fe-a12f-0d273eef2325 · outbound

This paper cites Statistical infer ence for the mean outcome under a possibly non-unique optimal treatment strategy.

Off-policy estimation with adaptively collected data: the power of online learning Statistical infer ence for the mean outcome under a possibly non-unique optimal treatment strategy

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.933079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.506578Z digest=sha256:4c7cd74a5d8d6f0d25f64cc4e32a50177825a541000f8f878031ea7d2cbcbb21

Observation 880a205a-4902-443e-ad93-6d7f137db380 · outbound

This paper cites Min imax off-policy evaluation for multi-armed bandits.

Off-policy estimation with adaptively collected data: the power of online learning Min imax off-policy evaluation for multi-armed bandits

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.923814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.509538Z digest=sha256:0d936006ec3673eccf732ec88085c748eb6d79e9cb113ad615fe2d18aa708b14

Observation bceb8c21-f1ac-4b72-a0dd-d67635c8d435 · outbound

This paper cites Off-policy estimation of linear functionals: Non-asymptotic theory for semi-parametric efficiency.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy estimation of linear functionals: Non-asymptotic theory for semi-parametric efficiency

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:43:58.640071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.512501Z digest=sha256:a2e550a2da3d572af77dfb585f6c0e51f5dfc77f776e422c3bf0ef4772411297

Observation 54415f53-83e2-4302-bbeb-06c16943141e · outbound

This paper cites Efficient counter factual learning from bandit feedback.

Off-policy estimation with adaptively collected data: the power of online learning Efficient counter factual learning from bandit feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.914524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.516147Z digest=sha256:c04278bfe0a55ea9826dd16c1278b5da9081e896424a88d40c59264814b43775

Observation c620aecc-93d6-472a-8f1d-1ebf3afd61b4 · outbound

This paper cites Offline policy evaluation in large action spaces via outcome-oriented action group ing.

Off-policy estimation with adaptively collected data: the power of online learning Offline policy evaluation in large action spaces via outcome-oriented action group ing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.905104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.519430Z digest=sha256:413fa8b4750e8f0af8420a47d9ac00b1e0ee5e15e4c547b5ec29e5cbda10a9bb

Observation 70e47d55-6214-43f6-80ad-4964406a5e51 · outbound

This paper cites Online non-parametric regression.

Off-policy estimation with adaptively collected data: the power of online learning Online non-parametric regression

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.896509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.522574Z digest=sha256:77fd038662db6e3d2aa95ca78edd05dac9aba52022a17bc00462fdd8b3180f68

Observation 6a50acc1-fa07-4676-9d78-ace03cdeeb35 · outbound

This paper cites Seque ntial complexities and uniform mar- tingale laws of large numbers.

Off-policy estimation with adaptively collected data: the power of online learning Seque ntial complexities and uniform mar- tingale laws of large numbers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.888627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.525570Z digest=sha256:f65c1c3373667f51090ea10ed91e2f6f8cc34e072332989174ba4f50df52b857

Observation 0ad9acdd-50b3-4190-98a3-c0138b3c4caf · outbound

This paper cites Relax and ra ndomize: From value to algorithms.

Off-policy estimation with adaptively collected data: the power of online learning Relax and ra ndomize: From value to algorithms

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.879912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.528627Z digest=sha256:b579162ffd20e1b77083bb980874f2b1815eee48b7f1fc7cb7a139d56a53464b

Observation 61554459-13d2-46f6-85d9-e83c8c8689db · outbound

This paper cites Comment: Performance of double-robust estimators when” inverse probability” weights ar e highly variable.

Off-policy estimation with adaptively collected data: the power of online learning Comment: Performance of double-robust estimators when” inverse probability” weights ar e highly variable

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.870744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.531540Z digest=sha256:47104ea3d0ef2fd51257140aa187e660f5b6869c1137ee8291befcd867b8d7ee

Observation 17d2bccb-68fd-429f-bd91-b582d4ae5e5d · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.

Off-policy estimation with adaptively collected data: the power of online learning Semiparametric efficiency in multivariate regression models with missing data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.860460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.534555Z digest=sha256:439aba7359602744f5a7b041f29749803303add3565d1546afb5eaf99c00f57d

Observation 2ef21956-0e7a-464d-b7e4-54e89463da98 · outbound

This paper cites Estimatio n of regression coefficients when some regressors are not always observed.

Off-policy estimation with adaptively collected data: the power of online learning Estimatio n of regression coefficients when some regressors are not always observed

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.851091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.537719Z digest=sha256:1d7b391ddc753f3d07b3da914d457af8ed3b4f07bf25d9c130588e1dfa4f23f7

Observation d4c43c02-d4e0-4815-9509-efa4d20b4331 · outbound

This paper cites Analysis o f semiparametric regression models for repeated outcomes in the presence of missing data.

Off-policy estimation with adaptively collected data: the power of online learning Analysis o f semiparametric regression models for repeated outcomes in the presence of missing data

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.841575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.540801Z digest=sha256:646ce448af9e32a0e20602d3790d5d6ee7df963069ae913da6149a8ed8c22228

Observation df24683d-c045-4d2b-b50f-09943e68d932 · outbound

This paper cites A tutorial on thompson sampling.

Off-policy estimation with adaptively collected data: the power of online learning A tutorial on thompson sampling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.831606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.543919Z digest=sha256:93fcba54c40ca490a47af9e6a999ac45aaa19fcf277898f15f23f0fd6fb33af2

Observation 7dcc1f49-cef1-4b27-a663-9803d30b5ea5 · outbound

This paper cites Off-Policy Evaluation for Large Action Spaces via Embeddings.

Off-policy estimation with adaptively collected data: the power of online learning Off-Policy Evaluation for Large Action Spaces via Embeddings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:58.546929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:58.546929Z digest=sha256:95fd4e8d616dd5f5ec38bafb1d9722cc686a1388bd76833805770a719cad412b

Observation de50d675-da0e-4cb4-8ef2-25962971d72f · outbound

This paper cites Off-policy ev aluation for large action spaces via conjunct effect modeling.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy ev aluation for large action spaces via conjunct effect modeling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.822558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.550231Z digest=sha256:9ce86119fea5a82e7ab8c65c1438ba051ee39b005c13b5ca5215829e30db54c4

Observation 2b63066d-06c7-41c8-b289-5469d9008d7f · outbound

This paper cites Lear ning from logged implicit exploration data.

Off-policy estimation with adaptively collected data: the power of online learning Lear ning from logged implicit exploration data

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.814022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.553443Z digest=sha256:ec74873e1c355eb1a4a78e6ded6811c0f836ed664358cf35495154303825e765

Observation b6bdbb72-811a-4e54-9836-42f2ed0d9109 · outbound

This paper cites Doubly robust off-policy evaluation with shrinkage.

Off-policy estimation with adaptively collected data: the power of online learning Doubly robust off-policy evaluation with shrinkage

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.805656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.556638Z digest=sha256:c31bd4d040e25331fc96369fe7a4e210c4eece69ab08a7411e7b53f9dad01fa1

Observation 8b319e47-8d66-4cf5-ad49-7037c7300635 · outbound

This paper cites Cab: Continuous adaptive blend- ing for policy evaluation and learning.

Off-policy estimation with adaptively collected data: the power of online learning Cab: Continuous adaptive blend- ing for policy evaluation and learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.797671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.559742Z digest=sha256:214af68a8c1604aeb523077c273c00de0fb02dbc3256428da15d0105bc5c69bd

Observation ecde691c-e9cc-4d95-b1c5-d1544122943d · outbound

This paper cites The self-normalized estimator for counterfactual learning.

Off-policy estimation with adaptively collected data: the power of online learning The self-normalized estimator for counterfactual learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.789061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.562987Z digest=sha256:12dddda8b262ca32ecae10b8657e284d39679b4b4a6812b0fbc397fdba7c984a

Observation e38b07e5-19f9-49d3-a04b-42ea05405a31 · outbound

This paper cites Data-efficient off-policy policy eva luation for reinforcement learn- ing.

Off-policy estimation with adaptively collected data: the power of online learning Data-efficient off-policy policy eva luation for reinforcement learn- ing

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.780101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.566553Z digest=sha256:d39d327e65b9c44c13ef22cefd9c7310c3e221b11a039b159cd9615a49cb31d1

Observation c73b29ef-8313-497e-9d84-de6f5950c774 · outbound

This paper cites On the likelihood that one unknown probability ex ceeds another in view of the evidence of two samples.

Off-policy estimation with adaptively collected data: the power of online learning On the likelihood that one unknown probability ex ceeds another in view of the evidence of two samples

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.771008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.569592Z digest=sha256:aaddf0a685ce65d09d2241223932f54a9492901e5d1d50b252d9743df523f2cc

Observation d4634878-2d7b-4c84-a5b6-a07f3bddd642 · outbound

This paper cites The construction and analysis of adaptive group sequential designs.

Off-policy estimation with adaptively collected data: the power of online learning The construction and analysis of adaptive group sequential designs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.761483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.572740Z digest=sha256:8fa61c3d8e5815b60a2fa8bfe2c1f925b5fbfd4b856dae75130252b3b4aff60f

Observation 55ff0e85-6000-421d-8fef-6a9d3b9cd250 · outbound

This paper cites High-dimensional statistics: A non-asymptotic viewpoint , volume 48.

Off-policy estimation with adaptively collected data: the power of online learning High-dimensional statistics: A non-asymptotic viewpoint , volume 48

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.752267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.575839Z digest=sha256:9bd7278469b7179e5b359ef8eed28570b4daf3a4550ae62823731344f9c8be28

Observation 0525e20e-def8-4f05-90d4-6c53dfd5e888 · outbound

This paper cites Oracle-e fficient pessimism: Offline policy optimization in contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Oracle-e fficient pessimism: Offline policy optimization in contextual bandits

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.743269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.578928Z digest=sha256:64df220fc5256597ea42e1839944a3f0fe8d62619d7ed1761bb2ab6835ebfb4f

Observation 0ac61766-8332-43d3-b8f1-f7864532a8c4 · outbound

This paper cites Optimal and ad aptive off-policy evaluation in contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Optimal and ad aptive off-policy evaluation in contextual bandits

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.733433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.581648Z digest=sha256:3b6f14065c469a94ac34cb872c33f84c7505c81d917053e00f78f533279884c8

Observation 649e6e36-317d-4fea-b7b2-eea3bb886478 · outbound

This paper cites Anytime-valid off-policy inference for contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Anytime-valid off-policy inference for contextual bandits

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.724253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.584182Z digest=sha256:05642c1adba9004878f63000dc25389633a798266b9458b2503fb4ceb9b05115

Observation 88d8437e-3d16-4f29-bd16-c7139087686d · outbound

This paper cites Asymptotic inference of causal effects with o bservational studies trimmed by the estimated propensity scores.

Off-policy estimation with adaptively collected data: the power of online learning Asymptotic inference of causal effects with o bservational studies trimmed by the estimated propensity scores

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.715417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.586736Z digest=sha256:077629c42dcbfe935e16a02f7c2c785f34a3d550562080fe805120140eed7563

Observation 101db0d1-b20d-4d34-b628-83ec133cee91 · outbound

This paper cites Off-policy evaluation via adaptive weighting with data from contextual bandits.

Off-policy estimation with adaptively collected data: the power of online learning Off-policy evaluation via adaptive weighting with data from contextual bandits

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.706252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.589198Z digest=sha256:0981ecb886608521c62de0df772ecf6b96f7c2dabb68215badaa2ffafdbf4333

Observation 7a83b08d-89d3-434a-b576-448588353f15 · outbound

This paper cites Policy learning with adaptively collected data.

Off-policy estimation with adaptively collected data: the power of online learning Policy learning with adaptively collected data

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.696569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.591909Z digest=sha256:dc1e1f4f005292e9e5e70e525bd36e8640f555394309afccf93ce85b647b6103

Observation 755424cb-481b-4825-a410-50b2442d6e48 · outbound

This paper cites Inference for batched bandits.

Off-policy estimation with adaptively collected data: the power of online learning Inference for batched bandits

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:58.686245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:43:58.594829Z digest=sha256:037384be54f3e49e116e3667ef6294303281239ab30f3551b8e4bcd8dede2ebd

Pith citing papers

No inbound Pith citation observations are available.