Pith. sign in

Paper Citation Record · LEDGER

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

As of 5 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 4 inbound Pith citation observations for arXiv:2602.08499.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08499 v2

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:23:08.113394Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.275053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T23:04:01.364747Z

Reference resolution

5 of 5 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49629e8f-5cea-4869-863b-130081a16299 · outbound

This paper cites an unresolved cited work.

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:23:07.732084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:23:07.732084Z digest=sha256:06b7d9e18e58e459848e5331633bd1167627bee16ec92b65120cfa8d4bb06ccc

Observation 584d726f-0bda-445f-a56b-29751b511513 · outbound

This paper cites an unresolved cited work.

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:23:07.866263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:23:07.866263Z digest=sha256:559813c5ca9a19d1b7b26f98282d7dbfdac8956147045347f15c63f93886aadc

Observation 053d8066-c87f-4a15-a22d-8e998a294c37 · outbound

This paper cites an unresolved cited work.

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:23:07.981528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:23:07.981528Z digest=sha256:2e2d164005063aa96f8e3e7ec403505bd44522409ef537755e520dcce273c902

Observation 8929999b-17ca-42a7-9dea-fea12315fa84 · outbound

This paper cites The largestnis 1533.

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards The largestnis 1533

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:23:08.113394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:23:08.113394Z digest=sha256:15c1de85cc0fb36e958b5d520465014abaaff4edec66f572e558643e26006665

Observation 4e194c70-e8a8-4203-8667-81903fef3168 · outbound

This paper cites 11 Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards A.

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards 11 Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards A

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:23:07.634282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:23:07.634282Z digest=sha256:4b1c69a5ec4e5eb3c41115c0139ce3fbe952745f74fac5065b93534a27c55282

Pith citing papers

Observation e3588103-6d95-4d47-8f70-046e88853374 · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:11.606338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:86c321bbfdb2939e36f4c7eb53b2459da4e8a66179bc3fcb61df981737741aac

Observation 427e1e51-9686-4af8-893e-8b5616a06d2d · inbound

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards cites this paper.

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.366182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:58:46.313028Z digest=sha256:ee3cb1b2555bb606eccff4cb46d423a66ceb8282eecbbce569bf8953ddea4c81

Observation 55a80e4a-cd6d-4993-bf4f-5eb67519a838 · inbound

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications cites this paper.

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Reference 247

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:43:15.723332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:36:23.776293Z digest=sha256:e063849ea4425b18dee51533847ee88e5e8e77309c1c15976b7e5ba6e8180cdd

Observation 1760ee1a-2a76-4fac-b3b8-d602362fe35f · inbound

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning cites this paper.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.275053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.275053Z digest=sha256:21e86f383f8c62f487cf1329864d0a5553d95054636ab4a2e3d81e93eda2d239