Pith. sign in

Paper Citation Record · LEDGER

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

As of 12 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2502.02516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02516 v3

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:57:09.116234Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:54:00.193317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f81defeb-c1ed-4883-805f-fa8c0d194bce · outbound

This paper cites Define now the policy π(u|s) = P (u|s, π(s)).

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Define now the policy π(u|s) = P (u|s, π(s))

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.396885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.029865Z digest=sha256:707e47d5fa51e05a8c2707c02a0410754ab71c73819070cb49dd717c72681bd5

Observation 18a61155-b45b-43a1-b567-3bf2bd923696 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.366261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.041873Z digest=sha256:6f1dd1da23d711ab65f5ce6724a4b81a1919fbe5e1519751e907d658f324ecf1

Observation 095938ef-068c-4a9e-a1e7-1d1be90e62f5 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.381993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.035422Z digest=sha256:2c36a9cf4371eb8ebd435cfec55a271ae6a3870c78425c1d788e88f0d1751707

Observation 38ebd29f-8653-48fa-bda1-0ce07e72259f · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.269649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.091686Z digest=sha256:2e5142cc05681000bd87183a207c7e1a5209aeef7ae42c4e07246e679a69fe98

Observation 9273a68d-b490-4b9a-985f-f1ef5d939d49 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.350101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.048682Z digest=sha256:17a977728fc69f2b2b8c1dde96106aa2c4024a82d9631f453a58cc6a4d884d0d

Observation 5cf64e25-2ccc-4458-a23b-fa870086c120 · outbound

This paper cites sufficient.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation sufficient

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.335207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.054503Z digest=sha256:f4f08b49cd1764ff7720a96d4094988ea96c908a2aef9b06871ba3801c56211a

Observation 9dbe6623-bd68-49ff-8fda-fbac3738d90c · outbound

This paper cites This choice encourages to select under-sampled actions for β >0, while for β = 0 we obtain a uniform forcing policy πf,t(a|s) = 1/A.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation This choice encourages to select under-sampled actions for β >0, while for β = 0 we obtain a uniform forcing policy πf,t(a|s) = 1/A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.319452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.064185Z digest=sha256:90736936c0842c1106df631688aa9f05006562590c827d63d49d859255ca323b

Observation 4c40abb8-cd42-4a39-ac6e-c57ced12dd03 · outbound

This paper cites Then, such solution induces an ergodic (irreducible and aperiodic) chain by Assumption 3.1 and Assumption 5.1.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Then, such solution induces an ergodic (irreducible and aperiodic) chain by Assumption 3.1 and Assumption 5.1

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.303737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.075723Z digest=sha256:219981d20e6e691ecfcb1638a44d244645925723ace05e0692a8baacf3545921

Observation 9a96dabc-1749-4a6c-b267-a3e29c11f82d · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.287755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.084632Z digest=sha256:039d178e2b34b83ed56eb2ebc58e2c5b1b5e219132e4561f38d0ec2941f23f5b

Observation 9846c2a1-5e65-4537-9661-21aa528875c4 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.253084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.098369Z digest=sha256:9a2f8e4f820ce885cf2d2cab842f902e9242f53bb1fa55d1a2feb7031e80a2e9

Observation dd763a3e-bbe3-4d47-a287-9b0accf62d59 · outbound

This paper cites On the other hand, in the single-policy scenario we use a default target policy policy πdef that is different for each environment.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation On the other hand, in the single-policy scenario we use a default target policy policy πdef that is different for each environment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.236354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.103954Z digest=sha256:b8d92914b88473afbc2ed67b9a657be1a516e0825fab86288c59df46e9b23608

Observation 73fd5463-4a96-43fc-8001-b8559bbc02da · outbound

This paper cites For the reward free case we use Rcanon to perform evaluation.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation For the reward free case we use Rcanon to perform evaluation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:57:09.216143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.110031Z digest=sha256:ae2d7a0bb02caa01622b5a0bea98a9197640e27ed59715c04f1fcd12ab00bfc5

Observation 32fadcb6-8f4e-4849-9e20-e03dda2e8ab1 · outbound

This paper cites an unresolved cited work.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:57:09.196812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.116234Z digest=sha256:da6ac9dd3cfda2b876a17111274477df06bd24d0a8738cb59f3d882c6b67c3c6

Observation b03f0c00-833f-4b14-9a34-9ed35b420fcb · outbound

This paper cites Clustered KL-barycenter design for policy evaluation.

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation Clustered KL-barycenter design for policy evaluation

Reference 418

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:57:09.177406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-09T11:57:09.021310Z digest=sha256:ef966874207409fedc944f49fd285ccb8ef948da3d44b8bd48b2bb85c83aeaf1

Pith citing papers

Observation 661319d6-b19e-4270-8ce8-866f141e0195 · inbound

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning cites this paper.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.193317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.193317Z digest=sha256:8f681d88f8796e7216476613c7f86ea88c263f38b7a6807d5aca9d6b4788dbad