Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2501.11651.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:08:21.609464Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T22:25:39.474855Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 81251f2c-dffb-488b-8f22-598a5490d6d0 · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 269
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3abbd6aa-3f04-4587-8587-14516d4153ce · inbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecd65069-43f6-4e93-b30d-fac8feff78d3 · inbound
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c480f1c5-9c68-4f25-9b8e-77c4d3e525d8 · inbound
Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac01cefd-0396-434f-892a-51412916c5c1 · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16c6a00-d1fe-4f2f-a93f-369528e211ca · inbound
Reinforced Language Models for Sequential Decision Making T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a32a9fd8-f044-4e31-8089-d72f7c94858b · inbound
Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12d51e4-c978-4d5f-8b2b-c7e22818bcdb · inbound
Self-Forcing++: Towards Minute-Scale High-Quality Video Generation T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d89408a0-8077-47b2-95f1-fd431cf9722b · inbound
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 824f44cf-2477-4107-bf83-d736628ddf0d · inbound
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fbac5259-3219-4704-9dae-3d7634003ab4 · inbound
Beyond the Sampled Token: Preserving Candidate Support in RLVR T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caaa32b5-751e-44d1-b34d-a7d5220f2f0f · inbound
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f8d439-e2b5-45f7-a112-384252096432 · inbound
Targeted Exploration via Unified Entropy Control for Reinforcement Learning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f39bcd0d-440c-402f-a67a-3420cc2047bf · inbound
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92b1c155-540f-4788-9ef3-23d07d1f64ca · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0e72bd3-cb95-465d-8f94-3ec23d0f8b01 · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d79e701-e8c6-4cc7-ba42-410b598735f4 · inbound
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 212e8307-e81a-479c-a1aa-106cd0950245 · inbound
Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.