Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:09.116234Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2502.02516.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:09.116234Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:54:00.193317Z
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f81defeb-c1ed-4883-805f-fa8c0d194bce · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Define now the policy π(u|s) = P (u|s, π(s))
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 18a61155-b45b-43a1-b567-3bf2bd923696 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 095938ef-068c-4a9e-a1e7-1d1be90e62f5 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 38ebd29f-8653-48fa-bda1-0ce07e72259f · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9273a68d-b490-4b9a-985f-f1ef5d939d49 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5cf64e25-2ccc-4458-a23b-fa870086c120 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation sufficient
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9dbe6623-bd68-49ff-8fda-fbac3738d90c · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation This choice encourages to select under-sampled actions for β >0, while for β = 0 we obtain a uniform forcing policy πf,t(a|s) = 1/A
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c40abb8-cd42-4a39-ac6e-c57ced12dd03 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Then, such solution induces an ergodic (irreducible and aperiodic) chain by Assumption 3.1 and Assumption 5.1
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a96dabc-1749-4a6c-b267-a3e29c11f82d · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9846c2a1-5e65-4537-9661-21aa528875c4 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd763a3e-bbe3-4d47-a287-9b0accf62d59 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation On the other hand, in the single-policy scenario we use a default target policy policy πdef that is different for each environment
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73fd5463-4a96-43fc-8001-b8559bbc02da · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation For the reward free case we use Rcanon to perform evaluation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32fadcb6-8f4e-4849-9e20-e03dda2e8ab1 · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b03f0c00-833f-4b14-9a34-9ed35b420fcb · outbound
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Clustered KL-barycenter design for policy evaluation
Reference 418
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 661319d6-b19e-4270-8ce8-866f141e0195 · inbound
Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.