Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T19:53:09.484847Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2603.00963.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T19:53:09.484847Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T20:58:19.104482Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T09:56:01.761633Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 404904fe-7036-4f2d-a49d-842460e67567 · outbound
Stabilizing Policy Optimization via Logits Convexity MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a0700a-0d2e-4a57-a14a-421d727feebb · outbound
Stabilizing Policy Optimization via Logits Convexity The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59dbce43-6432-4300-b4d8-c551396d6bbe · outbound
Stabilizing Policy Optimization via Logits Convexity MiniLLM: On-Policy Distillation of Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffb8ba8-62cd-47d8-aa2a-61471182c523 · outbound
Stabilizing Policy Optimization via Logits Convexity High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c822d43c-a000-4fea-8898-e4eba2f046e8 · outbound
Stabilizing Policy Optimization via Logits Convexity Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b95b6dd-9cba-47c4-ba27-802b4ecd1a1f · outbound
Stabilizing Policy Optimization via Logits Convexity Qwen3 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7951a39a-8e29-4105-9e88-6df39b8ef26c · outbound
Stabilizing Policy Optimization via Logits Convexity What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · outbound
Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4f0e5c-2109-440b-8fed-60bde5b5d0b8 · outbound
Stabilizing Policy Optimization via Logits Convexity Group Sequence Policy Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221083bd-c429-4d9a-9d64-47c01145b324 · outbound
Stabilizing Policy Optimization via Logits Convexity DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2391cd77-dfb3-4a1b-886f-c5b2b48ee6cd · outbound
Stabilizing Policy Optimization via Logits Convexity CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15919a99-9074-41fb-b1e1-c97b3525a8ad · outbound
Stabilizing Policy Optimization via Logits Convexity To address this, they propose a pretraining procedure for the value model, and decouple the λ in GAE for the policy and value model computations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972b9aed-1cc3-410a-bc62-018c17593ba8 · outbound
Stabilizing Policy Optimization via Logits Convexity Similarly, Shao et al
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04fcd6b7-ccd5-4a9d-ac89-fe4776e09aef · outbound
Stabilizing Policy Optimization via Logits Convexity Building upon the same idea, DCPO (Yang et al., 2025b) addresses the limitation in DAPO, where the same clip range is set for different positions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906355fb-d127-40b4-8a78-ba2ec430365a · outbound
Stabilizing Policy Optimization via Logits Convexity DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43196cf2-9d1d-4c7e-8452-a3ed04166d7f · outbound
Stabilizing Policy Optimization via Logits Convexity Crafting papers on machine learning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a908f5f-d786-4004-a0c4-5edab5a2b6e7 · outbound
Stabilizing Policy Optimization via Logits Convexity Generalist Reward Models: Found Inside Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faadda65-7e26-489b-a2a4-0e551a95f1db · outbound
Stabilizing Policy Optimization via Logits Convexity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69356ac8-eeec-4bed-9ebf-82ed49ed2f7a · outbound
Stabilizing Policy Optimization via Logits Convexity TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64088392-ed46-4f74-a279-93297f8bb234 · outbound
Stabilizing Policy Optimization via Logits Convexity REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a4bd84-7b6f-40d8-85b9-1a1e00f6396a · inbound
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization Stabilizing Policy Optimization via Logits Convexity
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22b2d736-9ff7-4af2-aba3-6e4e5900bae4 · inbound
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization Stabilizing Policy Optimization via Logits Convexity
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.