Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:09:47.169916Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.22012.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:09:47.169916Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 633db9ce-0a57-4857-8b48-a2fca70f788c · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits In contrast, the evaluation policyπis defined as π(a|x) = (1−ϵ)·I{a= argmax a′∈A q(x, a′, e;λ)}+ ϵ |A| ,(17) whereϵ∈[0,1]controls the quality ofπand we setϵ= 0.2as default
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122ae1d4-f65b-478f-9974-721b9f200f32 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a37fb38-9220-4166-9679-d475bf072c8c · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43878d4-3aa2-42f8-9c17-ec9da84d492d · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ce0654-8abf-4bea-b0fc-55e4c07eed20 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Our main motivation is to solve the prevalent problem of (completely) deterministic logging and new actions, issues that pessimistic techniques do not aim to address
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34e0857-7be8-45e8-9531-ed3fb22fc4a2 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d661228a-fee9-4fc6-a8a7-ae6578ee5f24 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Cross-Validated Off-Policy Evaluation
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3bcd95-e09c-47cf-b468-a4bbf3e91efa · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Offline policy evaluation in large action spaces via outcome-oriented action grouping
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0e8c9ce-0b0e-4625-8118-0f809c0deccf · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Counterfactual risk minimization: Learning from logged bandit feedback
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272011fd-ccd5-4849-ac41-3f290ace972f · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits $\beta$-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 431e59de-26fa-4e25-9b3c-4a01f2d271f4 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits To- wards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d06e4dd-1ee2-42cd-9766-037f380ac698 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Off-policy evaluation for large action spaces via policy convolution
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d493504-e007-463b-8e6f-30455ed2b5ba · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b52975e1-4161-43af-924c-ea6aeb460c81 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Automated Off-Policy Estimator Selection via Supervised Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21cf437-399f-4666-bfc1-5514840d3af8 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Towards assessing and benchmarking risk-return tradeoff of off-policy evaluation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7f93dc-faa5-4c3c-a981-84838885bb72 · outbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.