Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:41:31.820379Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2509.02737.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:41:31.820379Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7d35bf1a-e57d-4d8f-8a1b-5dc505c3d3a9 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Contrastive policy gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8e58389-6fe2-4ffb-9a72-4a59c7d60f94 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7b9e5a-49bb-47da-be6b-bf81ae7030e3 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient All experimental results are average over20 random seeds and we choose the best from three learning rates
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bedbb2b9-5aee-49cc-b0bc-574766382c09 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Inducing Neural Collapse to a Fixed Hierarchy-Aware Frame for Reducing Mistake Severity
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45e855d4-ff40-48cc-9e16-7eb3a5b7474d · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Continuous control with deep reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb96999-d049-4458-a9c1-d8ad2bd83976 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Neural Collapse with Cross-Entropy Loss
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7667d8f6-d6e8-4567-af9b-59a8740588ff · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Asynchronous methods for deep reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a05b1657-feca-4682-b4ec-62170104d518 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Instruction Tuning with GPT-4
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53350eeb-dbf1-454d-a348-85532088681d · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Linguistic Collapse: Neural Collapse in (Large) Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 420bd0d2-310f-46e7-b139-d779ab608147 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Huaqing Xiong, Tengyu Xu, Lin Zhao, Yingbin Liang, and Wei Zhang
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3c2ad1-eed3-4b28-a8e7-0e04cf3f2c4d · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Convergence and iteration complexity of policy gradient method for infinite-horizon reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 831869e0-15e0-4b83-a0d7-865c67f2d115 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02e43c6a-ac4c-4afe-82f7-5a13f2499e98 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient On the Global Convergence of Natural Actor-Critic with Two-layer Neural Network Parametrization
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9966f96a-fc25-44d9-8b33-5bf80b3e4f46 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Learning classifiers on positive and unlabeled data with policy gradient
Reference 1999
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05fc2478-8a1a-49cc-8652-54b738a9b238 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Global Optimality Guarantees For Policy Gradient Methods
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e0d5b91-927e-4028-9c8c-657d004fab70 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc20b73-a249-4daf-a3be-f04b21089ac2 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient TD Convergence: An Optimization Perspective
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c029449-dc7f-41c3-9c3b-709a55ba1910 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Understanding of a convolutional neural network
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f8e7448-b3dd-43dd-86ff-fb971f93b4d7 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient Neural collapse with unconstrained features
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b845a436-ed3a-4a49-ba40-bab28aabbdeb · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient OpenAI Gym
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdeee453-98d9-47a6-a565-64e2b28942a6 · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient A General Language Assistant as a Laboratory for Alignment
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed67a22-9d53-41f1-95b1-e161316c5c3a · outbound
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient doi: 10.18653/v1/2024.emnlp-main.1190
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.