Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T13:40:39.598985Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2606.10968.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T13:40:39.598985Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T08:49:37.615924Z
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation da798e84-75d3-4c64-ab8b-74fa70eaf887 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 749ea2ef-3962-42c1-a1be-13b71d402f2b · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International conference on machine learning , pages=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7641eaf3-7e1e-49e6-81d9-7b7b932eb043 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Proceedings of the nineteenth international conference on machine learning , pages=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4e07276-ef90-4037-86e7-8af3931f4821 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International conference on machine learning , pages=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9611386-c5dd-4155-af39-025814280182 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c75c9dc-d30b-48e6-bfa1-6ba694efd28b · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Riedmiller , title=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b8f5a20-3205-4698-a8af-7db4de249624 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Francis Song and Abbas Abdolmaleki and Jost Tobias Springenberg and Aidan Clark and Hubert Soyer and Jack W
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c612ccec-95a9-4578-8cc1-07d7baa69d94 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b893bb6-924a-434c-80a2-6da0f348ec9d · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 2019 , cdate=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bceba189-06f8-41bf-890a-d888b595aa6e · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International Conference on Learning Representations , year=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6b9655-0f05-42d3-a70e-9efa17eea6f1 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning International Conference on Learning Representations , year=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f58bc07-fe6a-4a4c-873b-61022426dbad · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning The annals of mathematical statistics , volume=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dd911a5-1551-47b9-87ae-1e79f6199b1e · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1422971e-36f9-4ce6-9813-6b991d84c1d9 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Advances in neural information processing systems , volume=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e2c108-c115-4b57-abb2-1ce1f1a51311 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204eb0be-40c8-4979-8f94-28b87591052b · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721df89d-fe93-443b-aaf9-e63f58b5f088 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 2026 , url=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ab2e50-8c7c-4c37-ad8c-ec3a83a62571 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b5f9014-2376-492d-b3e7-ae5b5c56e784 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6702c7ff-0e30-4f28-8e6a-45d6766bf672 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6759555a-8bdd-4595-bba1-86fb34d9f78c · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Soft Adaptive Policy Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e36b273-1c89-4306-8a8a-fcabc4df73ce · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning arXiv preprint arXiv:2601.22718 , year=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22ee3450-61f3-4f6a-b286-a32905d15bfe · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515e78cc-88b2-4214-ad1b-581355e8abc7 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Second Conference on Language Modeling , year=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad0c724c-5297-4c94-99a9-550e54b433a9 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning 2025 , url=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe439ed9-0450-4c35-bcf8-9aaa0cfb9b18 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44674da9-7e50-4e80-9487-9e9220f8ee14 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d6f86a76-b645-4da9-a1ec-829abc32a93d · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning CoRR , volume=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61ff2c1-7276-4f66-b0d7-96adf902a2fb · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Rethinking the Trust Region in LLM Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 87dc1642-f0b4-4789-a449-6f0e14205f74 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cdaeb9c-f790-426c-9086-d08b0766c84d · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Fipo: Eliciting deep reasoning with future-kl influenced policy optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0dbbe045-60e1-40ef-8f5e-d26700439a33 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2587ab3f-3995-4846-914c-118d62cd7f08 · outbound
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning The Fourteenth International Conference on Learning Representations , year=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb8cef3-6fd1-42ae-b224-2feaf23564c4 · inbound
Predictive Divergence Masks for LLM RL Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.