Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:23:08.113394Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 4 inbound Pith citation observations for arXiv:2602.08499.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:23:08.113394Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.275053Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-29T23:04:01.364747Z
5 of 5 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 49629e8f-5cea-4869-863b-130081a16299 · outbound
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584d726f-0bda-445f-a56b-29751b511513 · outbound
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053d8066-c87f-4a15-a22d-8e998a294c37 · outbound
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8929999b-17ca-42a7-9dea-fea12315fa84 · outbound
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards The largestnis 1533
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e194c70-e8a8-4203-8667-81903fef3168 · outbound
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards 11 Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards A
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3588103-6d95-4d47-8f70-046e88853374 · inbound
Policy Improvement Reinforcement Learning Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 427e1e51-9686-4af8-893e-8b5616a06d2d · inbound
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55a80e4a-cd6d-4993-bf4f-5eb67519a838 · inbound
World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Reference 247
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1760ee1a-2a76-4fac-b3b8-d602362fe35f · inbound
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.