Pith. sign in

Paper Citation Record · LEDGER

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2507.20150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20150 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:56:41.084268Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:19:34.840954Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a7b6546-5f47-4b35-ac84-a2438a216ce7 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:40.998568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:40.998568Z digest=sha256:6d0662e21e11b480b3274710c6a73a3b57a15c4d9707bdb73dfce47fbe81e29e

Observation b1e01e78-4e71-473c-94a6-2deb1c3f5060 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.008415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.008415Z digest=sha256:5d558cb2ff9b6100310675305bb87c75e6f130e7aeecf96eb8deeacd661a4c3b

Observation ef49e7f3-9dab-4f2f-8752-72992aa824ae · outbound

This paper cites Defense Against Reward Poisoning Attacks in Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Defense Against Reward Poisoning Attacks in Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.018945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.018945Z digest=sha256:ce5729a2cb93f212ed236f516228f41910edc0bdacb8473f6f9308262e2ca253

Observation 08201b88-4bc2-466f-99b0-7405c37ae0b0 · outbound

This paper cites Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.028809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.028809Z digest=sha256:5ff35589a865e66583c040cd52eb57bbd4d9ddb802ce31cc1d8fdfd77aa21519

Observation 0fc187ff-f529-490d-ac32-1ed0bab4e3bd · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.033080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.033080Z digest=sha256:c9d0f4c3448abcb41df706d0c49d321d90a91eec85a981f44da717a29905e9bb

Observation 0ae26018-5913-4345-bf2f-319ccf727136 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.062832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.062832Z digest=sha256:249260aea0f12a38d1b79e731fc36516c850dfba67cec1b091381740a2aea46c

Observation 00fe4cc6-6b70-4895-b35d-e92c2d86bd2c · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.067681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.067681Z digest=sha256:6612dbc46a4f78780d9630d4ba19755a7beda1130a0631d0a2599d29175a8b42

Observation 803f4fed-0a88-4288-a3d0-ba591b2e511d · outbound

This paper cites Qwen3 Technical Report.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.075895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.075895Z digest=sha256:d99fcc0f7d3e6db43e5fca69b1774a3701c4b391399f15108c7f814dfa94e179

Observation 5d513ce0-0a9e-4640-9569-6636b90a197a · outbound

This paper cites an unresolved cited work.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:56:41.491657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:56:41.079766Z digest=sha256:f1120041376bdc925a6e56215f9bff0e591a8638a063ad50540d77a2730d577c

Observation 15b88e43-d525-4a12-8c13-5251498e696f · outbound

This paper cites spurious reasoning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models spurious reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:56:41.477462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:56:41.084268Z digest=sha256:4428fa0b48512dbca5245787b51e720a93016e9c865a9a92c9df6f8167b5cb69

Observation 78913d69-3cbb-4bef-9023-1c2d92327ffc · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.023330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.023330Z digest=sha256:6e082d4a6198e7f0f65c816ceb086eda8b168e582aa9fcbadf1bd1d36c1b3e86

Observation d70c39aa-a11a-4a9d-816c-329a2cd21bb1 · outbound

This paper cites A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.050554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.050554Z digest=sha256:788ebe0f322670e670faec8e3810989431a78072f6728c45ad35fe632ee73706

Observation f50b2653-10c6-49bf-975f-081d70d6f876 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.054541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.054541Z digest=sha256:86d4aa64626b00319ccfc3de636fcb4de3a7fdc1a4259303a4e52a7dfa66aecf

Observation 4eb82b6c-d6ae-483f-bd21-7cc8eca0f49b · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.058636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.058636Z digest=sha256:fc1d358f034d20adc67b16c3248406737ff588dd1b9e8bb2ac46463506ec46a3

Observation 6183f9fc-4be1-4bdf-953a-7cd4438515d7 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.041784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.041784Z digest=sha256:a64466199d09bba47867cadd406b7abacb3e7aed3fdebae7f0d35d5ea541f518

Observation fd17ebcf-4b39-4eac-af62-47d1c12e5df4 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.071930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.071930Z digest=sha256:325de19d6055f38a25d316c8e7d303e82881913ceb307e31a069e1c1281f9a93

Observation 4c93f222-a987-42fb-b239-849d117f537f · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.014266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.014266Z digest=sha256:b6703229d507162577c0f888ae47d2d5fa4f410f86cb0249fefff0adbbb8637f

Observation 0c63d43f-c85a-4e12-882d-bfe71776b201 · outbound

This paper cites MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.046033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.046033Z digest=sha256:1363d0dc041151252f43cf60568338b34ecf14e2b8ce8fee72e13581d87a869a

Observation 3826fd40-049f-4b4b-bdde-f91ab48284b9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.037293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.037293Z digest=sha256:97ebc8a44b3c11b47cb15279441d4e818c0fb23e9202188b116df8ffd08242de

Observation fc062070-efd6-41d5-ba8a-75fbb2cef70e · outbound

This paper cites Qwen2.5-VL Technical Report.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.003583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.003583Z digest=sha256:7c490db87055db9b84f8cc1c1806b708f42ebb2d73fcb341c106e1eb8111e71b

Pith citing papers

Observation b86ffbcb-f091-4df2-8a86-c7aa1ba77b3b · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:34.840954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:34.840954Z digest=sha256:6dea1c640dd4a2ee1e2d399599b8010b810d39431c823e686407929a662a06fa