Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:14:53.084480Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.20543.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:14:53.084480Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d4a1d0e2-fa81-418a-a580-9e9cd53aefda · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab93d74a-9181-4856-9c4d-24c6a3300412 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Evaluating Large Language Models Trained on Code
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bacf7d46-6f06-4ecf-8be7-69c1a5f7eb4d · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4263efc-e780-4a88-91e7-e55119155782 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f381000e-02f5-4e38-8cfe-d9ce89303776 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Soft actor- critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4adbbb55-7d08-476c-bbd5-bc638c57dcb5 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c438dce3-e397-4c15-a545-0243ad685036 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d2689a0-509a-4c1f-98de-03ca1e3d9ba6 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Reinforcement learning with KL penalties is better viewed as Bayesian inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b5b5ab-3351-4b8e-8e77-36f226585886 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Conservative Q- learning for offline reinforcement learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21403e57-5303-457e-920f-67a325d5f2b3 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa5c229-72ee-42f7-9314-b035d470078f · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a284f46-fa53-43e6-a885-84592237f648 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Let's Verify Step by Step
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df64507-abac-4e22-a073-fc1f704e22fb · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Understanding R1-Zero-Like Training: A Critical Perspective
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc83e24-01f4-4a0b-802b-6a0595fbe205 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Manning, and Chelsea Finn
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86fb70e2-15d2-4099-ac17-d04c00ab3f38 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Vanishing gradients in reinforcement finetuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1678a4-98ae-45fc-babc-2415442f5739 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea90215b-1e69-4660-a087-aa37df089232 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47a8065-aa39-41cf-b2de-f2fa2ed0dc2d · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion HybridFlow: A Flexible and Efficient RLHF Framework
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968ff127-977e-4155-87ca-a1faf74458a4 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Qwen2.5 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e01598-9ac0-470e-b5e6-3681d2c157b6 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b8fd97-5e54-4efc-9280-b1c8ed3769a4 · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Understanding diversity collapse in RLVR via the lens of over- training.arXiv:2606.15455, 2026
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25c59c6-96e1-4bca-b08a-cfe6eda755dc · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa86c831-0274-4832-9eba-19bc8f4e8c0e · outbound
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.