Pith. sign in

Paper Citation Record · LEDGER

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.22012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22012 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:09:47.169916Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 633db9ce-0a57-4857-8b48-a2fca70f788c · outbound

This paper cites In contrast, the evaluation policyπis defined as π(a|x) = (1−ϵ)·I{a= argmax a′∈A q(x, a′, e;λ)}+ ϵ |A| ,(17) whereϵ∈[0,1]controls the quality ofπand we setϵ= 0.2as default.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits In contrast, the evaluation policyπis defined as π(a|x) = (1−ϵ)·I{a= argmax a′∈A q(x, a′, e;λ)}+ ϵ |A| ,(17) whereϵ∈[0,1]controls the quality ofπand we setϵ= 0.2as default

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:47.151022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:47.151022Z digest=sha256:11d485202ed2dcfa15a8f9165a148f8ea406cfab0526280508fb1e8f37f63f6d

Observation 122ae1d4-f65b-478f-9974-721b9f200f32 · outbound

This paper cites Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.509741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.509741Z digest=sha256:94f92df2db2c8ccda2c1fa98777b3246ee5c17f32d8ea82b5a1da00bf48dd8a1

Observation 3a37fb38-9220-4166-9679-d475bf072c8c · outbound

This paper cites an unresolved cited work.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:47.169916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:47.169916Z digest=sha256:e034459e3600b6a3b1863d2659b192126db43dcbbde6ae248ff6abae3fde9dbb

Observation f43878d4-3aa2-42f8-9c17-ec9da84d492d · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.724676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.724676Z digest=sha256:e6cde7d059dc31ce0f7c14966313e785f0c1084b23c180ff2d6bb8a33d5dd527

Observation 52ce0654-8abf-4bea-b0fc-55e4c07eed20 · outbound

This paper cites Our main motivation is to solve the prevalent problem of (completely) deterministic logging and new actions, issues that pessimistic techniques do not aim to address.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Our main motivation is to solve the prevalent problem of (completely) deterministic logging and new actions, issues that pessimistic techniques do not aim to address

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.979697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.979697Z digest=sha256:aa03738d1fc397ef8e0459255d50d3d97460e8088b7bd957684775fb390defcb

Observation a34e0857-7be8-45e8-9531-ed3fb22fc4a2 · outbound

This paper cites an unresolved cited work.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:47.049372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:47.049372Z digest=sha256:e63e036f5ddf7dd93a7597706a23604d18c8a0c79fd11e0137df897cbc453ee6

Observation d661228a-fee9-4fc6-a8a7-ae6578ee5f24 · outbound

This paper cites Cross-Validated Off-Policy Evaluation.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Cross-Validated Off-Policy Evaluation

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:45.859271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:45.859271Z digest=sha256:db6e0d1ea91f51d0276ba626ad6231ab6c66c32ac68119be27978ff7c583f961

Observation 6e3bcd95-e09c-47cf-b468-a4bbf3e91efa · outbound

This paper cites Offline policy evaluation in large action spaces via outcome-oriented action grouping.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Offline policy evaluation in large action spaces via outcome-oriented action grouping

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.375691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.375691Z digest=sha256:d9ee2cf8b2d7b62f26fbca77219132d4dcb2183105db8e8194fa1366c473eec5

Observation e0e8c9ce-0b0e-4625-8118-0f809c0deccf · outbound

This paper cites Counterfactual risk minimization: Learning from logged bandit feedback.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Counterfactual risk minimization: Learning from logged bandit feedback

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.599454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.599454Z digest=sha256:b7a4ecefbb75ed123b679be40d4d79d1e1fc6173e612b238acb7ba4c9170261a

Observation 272011fd-ccd5-4849-ac41-3f290ace972f · outbound

This paper cites $\beta$-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits $\beta$-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.794867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.794867Z digest=sha256:2ff6b42c2ef6f362b7f332d0468f3f2364898e49cef8d348d3b9719c8cd61c8e

Observation 431e59de-26fa-4e25-9b3c-4a01f2d271f4 · outbound

This paper cites To- wards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits To- wards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.167737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.167737Z digest=sha256:3ea07766fb303ee04f080c7832ffca275e7c22dded808836f5fa323759f2ecb1

Observation 2d06e4dd-1ee2-42cd-9766-037f380ac698 · outbound

This paper cites Off-policy evaluation for large action spaces via policy convolution.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Off-policy evaluation for large action spaces via policy convolution

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.441229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.441229Z digest=sha256:f091449c75087b05b27d0febd105f1354f8adfb7387c2790e58b50e76df4a248

Observation 6d493504-e007-463b-8e6f-30455ed2b5ba · outbound

This paper cites Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.306510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.306510Z digest=sha256:7d6995b9e0561d2fe4ba61399ee69e05d6658e49b3679160c05bc6cc4880f58a

Observation b52975e1-4161-43af-924c-ea6aeb460c81 · outbound

This paper cites Automated Off-Policy Estimator Selection via Supervised Learning.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Automated Off-Policy Estimator Selection via Supervised Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:45.952279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:45.952279Z digest=sha256:b52d76eb2c719ebad87a1f6461fc75a81c9683d3757c7a081c552eb22e6cf3a8

Observation a21cf437-399f-4666-bfc1-5514840d3af8 · outbound

This paper cites Towards assessing and benchmarking risk-return tradeoff of off-policy evaluation.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Towards assessing and benchmarking risk-return tradeoff of off-policy evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.095096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.095096Z digest=sha256:135b6fe171f7214f6c7076e0d7362917ad68bb6938b433062cd9f0a323d224df

Observation 6d7f93dc-faa5-4c3c-a981-84838885bb72 · outbound

This paper cites an unresolved cited work.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.860596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.860596Z digest=sha256:48f8e6117bae6d7c2c204c27a04ac82b3576350cbb9cb2f6c87576587316016d

Pith citing papers

No inbound Pith citation observations are available.