Pith. sign in

Paper Citation Record · LEDGER

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.22012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22012 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:09:47.169916Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 633db9ce-0a57-4857-8b48-a2fca70f788c · outbound

This paper cites In contrast, the evaluation policyπis defined as π(a|x) = (1−ϵ)·I{a= argmax a′∈A q(x, a′, e;λ)}+ ϵ |A| ,(17) whereϵ∈[0,1]controls the quality ofπand we setϵ= 0.2as default.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits In contrast, the evaluation policyπis defined as π(a|x) = (1−ϵ)·I{a= argmax a′∈A q(x, a′, e;λ)}+ ϵ |A| ,(17) whereϵ∈[0,1]controls the quality ofπand we setϵ= 0.2as default

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:47.151022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:47.151022Z digest=sha256:3fb88eff9b7ad437969cd73fa0d74fb0e9950e5ec4d09930df3865a54c69a4ba

Observation 122ae1d4-f65b-478f-9974-721b9f200f32 · outbound

This paper cites Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.509741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.509741Z digest=sha256:f743d4ad8960fc8e94c4a9f5639093bdd5578f4d1982dac2cd1ba71c1d0932f8

Observation 3a37fb38-9220-4166-9679-d475bf072c8c · outbound

This paper cites an unresolved cited work.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:47.169916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:47.169916Z digest=sha256:d27a19195840d2191f7a06e4110bd317ccba20c98109b56c752b6eb6f5b553e9

Observation f43878d4-3aa2-42f8-9c17-ec9da84d492d · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.724676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.724676Z digest=sha256:b0e0a218112b5a83fce4741d3003277d0e289a3a529f50879ff9c0ee70f442e4

Observation 52ce0654-8abf-4bea-b0fc-55e4c07eed20 · outbound

This paper cites Our main motivation is to solve the prevalent problem of (completely) deterministic logging and new actions, issues that pessimistic techniques do not aim to address.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Our main motivation is to solve the prevalent problem of (completely) deterministic logging and new actions, issues that pessimistic techniques do not aim to address

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.979697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.979697Z digest=sha256:64fd68ffb3c1a1fae3684c18750cc2911bb42cd328ccec1f42a9494c1adeca5e

Observation a34e0857-7be8-45e8-9531-ed3fb22fc4a2 · outbound

This paper cites an unresolved cited work.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:47.049372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:47.049372Z digest=sha256:feee6f17bf8dabdae7b57b8b116148fd139601813d36a1c6686a17ae6b1da1e8

Observation d661228a-fee9-4fc6-a8a7-ae6578ee5f24 · outbound

This paper cites Cross-Validated Off-Policy Evaluation.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Cross-Validated Off-Policy Evaluation

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:45.859271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:45.859271Z digest=sha256:623a6aabbd9cbe5bc652907370817378ea8734407523f7ab4c23d4534e90caa8

Observation 6e3bcd95-e09c-47cf-b468-a4bbf3e91efa · outbound

This paper cites Offline policy evaluation in large action spaces via outcome-oriented action grouping.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Offline policy evaluation in large action spaces via outcome-oriented action grouping

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.375691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.375691Z digest=sha256:fbab946926b61ddb83f17b1dfd59aa834243a377db6c50ef15bf9c2ff8768e4b

Observation e0e8c9ce-0b0e-4625-8118-0f809c0deccf · outbound

This paper cites Counterfactual risk minimization: Learning from logged bandit feedback.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Counterfactual risk minimization: Learning from logged bandit feedback

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.599454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.599454Z digest=sha256:03601facc71bdaef9605a1bce5ab72a22f12f47287d1ff4e4e638256d1074cfb

Observation 272011fd-ccd5-4849-ac41-3f290ace972f · outbound

This paper cites $\beta$-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits $\beta$-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.794867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.794867Z digest=sha256:2faa52fc5ef26890cd200d0162fca822fb5991a77899cb158eb06ca9c94721aa

Observation 431e59de-26fa-4e25-9b3c-4a01f2d271f4 · outbound

This paper cites To- wards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits To- wards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.167737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.167737Z digest=sha256:01147245bac95af79f83e5bff6f4e5f4bcedb04ba3ff67bc2f5e656d618b50a3

Observation 2d06e4dd-1ee2-42cd-9766-037f380ac698 · outbound

This paper cites Off-policy evaluation for large action spaces via policy convolution.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Off-policy evaluation for large action spaces via policy convolution

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.441229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.441229Z digest=sha256:d44ab18f159e3b803fcd75219c9a32f4108379624e850177da467c74ac137029

Observation 6d493504-e007-463b-8e6f-30455ed2b5ba · outbound

This paper cites Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.306510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.306510Z digest=sha256:47647ab2764b2803b6594b77b332496be1f6646621f7acb581dd96fad9114768

Observation b52975e1-4161-43af-924c-ea6aeb460c81 · outbound

This paper cites Automated Off-Policy Estimator Selection via Supervised Learning.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Automated Off-Policy Estimator Selection via Supervised Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:45.952279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:45.952279Z digest=sha256:aed192d72fa580c0cbc2515da4aeb33d70c24e8404178a34c094971d0fdc7a3c

Observation a21cf437-399f-4666-bfc1-5514840d3af8 · outbound

This paper cites Towards assessing and benchmarking risk-return tradeoff of off-policy evaluation.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Towards assessing and benchmarking risk-return tradeoff of off-policy evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.095096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.095096Z digest=sha256:b9ccc4dc4f456b14e98dbe442bf53b08b13fd4f8614f8d364441c740dae9f807

Observation 6d7f93dc-faa5-4c3c-a981-84838885bb72 · outbound

This paper cites an unresolved cited work.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.860596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.860596Z digest=sha256:14c3a2767b1988142767e513cbc3ffc88dd910e23cf5ba0520ad4a07b24d5aeb

Pith citing papers

No inbound Pith citation observations are available.