Pith. sign in

Paper Citation Record · LEDGER

Stabilizing Policy Optimization via Logits Convexity

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2603.00963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.00963 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:53:09.484847Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T20:58:19.104482Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:56:01.761633Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 404904fe-7036-4f2d-a49d-842460e67567 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Stabilizing Policy Optimization via Logits Convexity MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.595577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.595577Z digest=sha256:b102375259b68d7fd3ae8d0ce2d9f87d4041fa55742ecdd13b267b6a76955890

Observation 73a0700a-0d2e-4a57-a14a-421d727feebb · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Stabilizing Policy Optimization via Logits Convexity The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.706557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.706557Z digest=sha256:daff96fea984e5fc5e995593fef431d632bbec1e1b7eee831c201948648d4f8a

Observation 59dbce43-6432-4300-b4d8-c551396d6bbe · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Stabilizing Policy Optimization via Logits Convexity MiniLLM: On-Policy Distillation of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.813705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.813705Z digest=sha256:4178dbbacb1fa5ca2a075e81be52cd75fcae6a95fa4cf7d923d0c23c52e5df87

Observation 1ffb8ba8-62cd-47d8-aa2a-61471182c523 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Stabilizing Policy Optimization via Logits Convexity High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.307029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.307029Z digest=sha256:641561cff42a71e30c265c20f6508458f22ea4e3cc5d6ebe3a0312986fb3c8bd

Observation c822d43c-a000-4fea-8898-e4eba2f046e8 · outbound

This paper cites Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs.

Stabilizing Policy Optimization via Logits Convexity Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.502936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.502936Z digest=sha256:f7056a2b1ad3e8adb2594b453dedde652604a8d89173970e623977ec0819a083

Observation 9b95b6dd-9cba-47c4-ba27-802b4ecd1a1f · outbound

This paper cites Qwen3 Technical Report.

Stabilizing Policy Optimization via Logits Convexity Qwen3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.621723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.621723Z digest=sha256:76fc2d20ab6013c999f760fb8a7272073024ceb8b628fef492b05a9a5fcab6b4

Observation 7951a39a-8e29-4105-9e88-6df39b8ef26c · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Stabilizing Policy Optimization via Logits Convexity What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.739032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.739032Z digest=sha256:8aeaf000dc71c3cab01d375d78fd6f0dc6d8b67610dd60b4148d2ddba9c6633f

Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.856440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.856440Z digest=sha256:b328a103a63fc05697c5598fadf18801bb4fb02521c13880900a8811811ba9af

Observation 0b4f0e5c-2109-440b-8fed-60bde5b5d0b8 · outbound

This paper cites Group Sequence Policy Optimization.

Stabilizing Policy Optimization via Logits Convexity Group Sequence Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.971429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.971429Z digest=sha256:947c6955ff22fe8d980f75e8ccbfa9f91ee8cb2b8ab92b48b8c047fdeb59bd96

Observation 221083bd-c429-4d9a-9d64-47c01145b324 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Stabilizing Policy Optimization via Logits Convexity DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.115910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.115910Z digest=sha256:c6bff0039aa9ec9264c1eb3909f8bf34a5d6911cfc7b199128b2b38b97a17958

Observation 2391cd77-dfb3-4a1b-886f-c5b2b48ee6cd · outbound

This paper cites CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning.

Stabilizing Policy Optimization via Logits Convexity CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.204030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.204030Z digest=sha256:8d491ed7eed415bbc60f8513229c6496dd1a2ce3b6efaba43066e9e9b72536fb

Observation 15919a99-9074-41fb-b1e1-c97b3525a8ad · outbound

This paper cites To address this, they propose a pretraining procedure for the value model, and decouple the λ in GAE for the policy and value model computations.

Stabilizing Policy Optimization via Logits Convexity To address this, they propose a pretraining procedure for the value model, and decouple the λ in GAE for the policy and value model computations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.278564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.278564Z digest=sha256:8e6a820ba5a967cc32351623f88eeaedf488d9ddf30e948ce0a546b6a52d4fbf

Observation 972b9aed-1cc3-410a-bc62-018c17593ba8 · outbound

This paper cites Similarly, Shao et al.

Stabilizing Policy Optimization via Logits Convexity Similarly, Shao et al

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.354384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.354384Z digest=sha256:49cf10be075d06f0e92fd3f6a78168e5887766f726870adcd14eee0b535bb648

Observation 04fcd6b7-ccd5-4a9d-ac89-fe4776e09aef · outbound

This paper cites Building upon the same idea, DCPO (Yang et al., 2025b) addresses the limitation in DAPO, where the same clip range is set for different positions.

Stabilizing Policy Optimization via Logits Convexity Building upon the same idea, DCPO (Yang et al., 2025b) addresses the limitation in DAPO, where the same clip range is set for different positions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.484847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.484847Z digest=sha256:ae3e8fe1a481d7f0710c64a0448694c8210b9d3438481d9ed9f3a94c4af49b55

Observation 906355fb-d127-40b4-8a78-ba2ec430365a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stabilizing Policy Optimization via Logits Convexity DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.385115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.385115Z digest=sha256:876eea921530e7a2c4b7d8f18b3b75f29df9459ece2ae692151e246dc47dfdcb

Observation 43196cf2-9d1d-4c7e-8452-a3ed04166d7f · outbound

This paper cites Crafting papers on machine learning.

Stabilizing Policy Optimization via Logits Convexity Crafting papers on machine learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.113911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.113911Z digest=sha256:804c5ae81b7236f8ba45d1257796195dcdef3652626821cb517d02a7fd895004

Observation 8a908f5f-d786-4004-a0c4-5edab5a2b6e7 · outbound

This paper cites Generalist Reward Models: Found Inside Large Language Models.

Stabilizing Policy Optimization via Logits Convexity Generalist Reward Models: Found Inside Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.222944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.222944Z digest=sha256:2d484c1f8ab1c369a1382812ca5633260064d4e8d88ae9949105925e2cb569f0

Observation faadda65-7e26-489b-a2a4-0e551a95f1db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Stabilizing Policy Optimization via Logits Convexity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.923480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.923480Z digest=sha256:c01e303489aa5293141075aa9d69f7d7fe648528b53625d98035e4fae800c661

Observation 69356ac8-eeec-4bed-9ebf-82ed49ed2f7a · outbound

This paper cites TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback.

Stabilizing Policy Optimization via Logits Convexity TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.503407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.503407Z digest=sha256:98026d84ccda64160d0954ac4aa763e3bc67f61831c222c58c71988d99514039

Observation 64088392-ed46-4f74-a279-93297f8bb234 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Stabilizing Policy Optimization via Logits Convexity REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.005719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.005719Z digest=sha256:8c93cd21e62b392e318b33f9b3f93eab260b843080ee1194cbf0f75aa4c97b07

Pith citing papers

Observation b0a4bd84-7b6f-40d8-85b9-1a1e00f6396a · inbound

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization cites this paper.

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization Stabilizing Policy Optimization via Logits Convexity

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:05.363647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:45:15.962965Z digest=sha256:96282887a4f4982e3453187bc67bb4b390c7d1682514de8eec3dc41a532578d1

Observation 22b2d736-9ff7-4af2-aba3-6e4e5900bae4 · inbound

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization cites this paper.

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization Stabilizing Policy Optimization via Logits Convexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T20:58:19.104482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:58:19.104482Z digest=sha256:a472e64776ebfd2d5ce957a921a1e37450f1445bd6d313bf91bfb4bab6eade78