Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:56:41.084268Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2507.20150.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:56:41.084268Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:19:34.840954Z
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a7b6546-5f47-4b35-ac84-a2438a216ce7 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e01e78-4e71-473c-94a6-2deb1c3f5060 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef49e7f3-9dab-4f2f-8752-72992aa824ae · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Defense Against Reward Poisoning Attacks in Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08201b88-4bc2-466f-99b0-7405c37ae0b0 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc187ff-f529-490d-ac32-1ed0bab4e3bd · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ae26018-5913-4345-bf2f-319ccf727136 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00fe4cc6-6b70-4895-b35d-e92c2d86bd2c · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803f4fed-0a88-4288-a3d0-ba591b2e511d · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d513ce0-0a9e-4640-9569-6636b90a197a · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 15b88e43-d525-4a12-8c13-5251498e696f · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models spurious reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 78913d69-3cbb-4bef-9023-1c2d92327ffc · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 1963
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70c39aa-a11a-4a9d-816c-329a2cd21bb1 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond
Reference 2005
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50b2653-10c6-49bf-975f-081d70d6f876 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Proximal Policy Optimization Algorithms
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb82b6c-d6ae-483f-bd21-7cc8eca0f49b · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6183f9fc-4be1-4bdf-953a-7cd4438515d7 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd17ebcf-4b39-4eac-af62-47d1c12e5df4 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c93f222-a987-42fb-b239-849d117f537f · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c63d43f-c85a-4e12-882d-bfe71776b201 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3826fd40-049f-4b4b-bdde-f91ab48284b9 · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc062070-efd6-41d5-ba8a-75fbb2cef70e · outbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Qwen2.5-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86ffbcb-f091-4df2-8a86-c7aa1ba77b3b · inbound
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.