Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T10:52:32.933138Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 4 inbound Pith citation observations for arXiv:2506.05760.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T10:52:32.933138Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T06:31:27.811744Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T19:50:10.420475Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd4c933f-4c63-4c16-99e1-4f53b3374b70 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1f2d634-6b4f-45b6-b7b1-fe67c5da8546 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 34c7f713-f302-4dc3-9a69-3b21c0818fbc · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7f916c0f-343d-4939-b8c4-b798ac8b7d93 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Language Models can Self-Lengthen to Generate Long Texts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5038f393-df94-4407-8ad9-d8331a54f79a · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f61d086a-d96b-4c33-a3fe-7d28d77515fd · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8afab734-64e6-4d04-85a6-82fb60bf5675 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning In our experiment, we use the proximal policy optimization (PPO) (Schul- man et al., 2017) algorithm with generalized advan- tage estimation (GAE) as the advantage estimator
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca231d3c-cb09-4e4e-8dfe-3ae3945d7818 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning We utilize a rollout strategy based on the vLLM engine with a tensor model parallel size of 2
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa636e58-32c7-4b3a-90e0-a3295bdda004 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5355a77-c4e0-4cc3-b671-2d1b42020328 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 48b9fa43-a8b2-4c83-9554-1703df330bb8 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7451109f-68ab-4e35-8388-d83d2366c6e8 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ee47962-5671-4932-8109-b7b727c7f555 · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7172521-7a6a-4f25-81bf-f0573e744d6d · outbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Analysis
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62f63ad2-eab9-444a-8fd3-cb25742f770b · inbound
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93602bd8-b9c2-4e1a-b81a-e912928bcd5d · inbound
Self-Evolving Deep Research via Joint Generation and Evaluation Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 385c9300-c51a-4d61-98c9-7011035a35d6 · inbound
IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 007fffc9-e60a-4768-bbc6-20577928739f · inbound
OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.