Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2306.02231.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:50.418201Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:53:16.384194Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 03c377c5-e3ab-436d-8210-6ccd562be404 · inbound
Thompson Sampling in Online RLHF with General Function Approximation Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d541982-153e-4321-8f10-3826e4d6b35f · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2ee59d-f3c4-464d-8a00-1d3a097ef619 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd3c7038-4d99-4149-9d54-430804dd56a6 · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50061053-d40d-49bc-bc12-ee3ff89a5d00 · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf3a4c5a-57d2-4e9e-ad79-db6e2d1f44ec · inbound
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0eb34602-8b01-4ae2-aec5-b8a4ddf6ede3 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 164
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 182af3d4-ac06-4c18-b103-917c2cb86351 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 165
Source-reported events for the cited work
Unavailable: canonical work link unavailable.