Pith. sign in

Paper Citation Record · LEDGER

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 4 inbound Pith citation observations for arXiv:2506.05760.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05760 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T10:52:32.933138Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:31:27.811744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T19:50:10.420475Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd4c933f-4c63-4c16-99e1-4f53b3374b70 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:53:02.879690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:0e3f78b4e25b3f47575389970169fad5462d54f86f7982d01f5e728f2bf8095b

Observation b1f2d634-6b4f-45b6-b7b1-fe67c5da8546 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T10:53:02.884676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:bcb51dc34427b52a324612bb49649fd9bbe3edd919acdeb830047edfdb08121a

Observation 34c7f713-f302-4dc3-9a69-3b21c0818fbc · outbound

This paper cites GPT-4 Technical Report.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning GPT-4 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:53:02.879025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:749d496304ef11a5012d7fe5992b52d94c8d56084fd4ef18f93f53c85ff316de

Observation 7f916c0f-343d-4939-b8c4-b798ac8b7d93 · outbound

This paper cites Language Models can Self-Lengthen to Generate Long Texts.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Language Models can Self-Lengthen to Generate Long Texts

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T10:53:02.877270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:efaccc956f1a87c1d2a57796e0d0249c55802208bb52e97d57a963d58494e8df

Observation 5038f393-df94-4407-8ad9-d8331a54f79a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T10:53:02.882074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:724d1fb3d4e883785ef5892c67699520bc8b594502cc6107762d47d822a090b8

Observation f61d086a-d96b-4c33-a3fe-7d28d77515fd · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:53:02.871689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:5de8d72fa8ead3524e455af8a7ecd65859a87435aee966a277cdfef02270b18e

Observation 8afab734-64e6-4d04-85a6-82fb60bf5675 · outbound

This paper cites In our experiment, we use the proximal policy optimization (PPO) (Schul- man et al., 2017) algorithm with generalized advan- tage estimation (GAE) as the advantage estimator.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning In our experiment, we use the proximal policy optimization (PPO) (Schul- man et al., 2017) algorithm with generalized advan- tage estimation (GAE) as the advantage estimator

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T10:53:02.943562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:782cdc26086ed381ce614cea38d18f2d78117650ea8bd95c65a57efb3de38e65

Observation ca231d3c-cb09-4e4e-8dfe-3ae3945d7818 · outbound

This paper cites We utilize a rollout strategy based on the vLLM engine with a tensor model parallel size of 2.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning We utilize a rollout strategy based on the vLLM engine with a tensor model parallel size of 2

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T10:53:02.933572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:987bf00d2fd59415c6eadf25c01a4f3e28b64be1ff063c08e338f987a9e5c485

Observation fa636e58-32c7-4b3a-90e0-a3295bdda004 · outbound

This paper cites an unresolved cited work.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-19T10:53:02.938983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:2e30b2c2166f5872b4c5bb29e8bb0a8f4cc6fc6e02fde43819c0b862d452ee0a

Observation f5355a77-c4e0-4cc3-b671-2d1b42020328 · outbound

This paper cites an unresolved cited work.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-19T10:53:02.941874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:957191358f92cb4096a1a6077a4c92087f6441acd3895e6ba10f915a0edea696

Observation 48b9fa43-a8b2-4c83-9554-1703df330bb8 · outbound

This paper cites an unresolved cited work.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-19T10:53:02.940410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:8d106e9529ed385d64a325b6b29d42bba29ceb7cf841be132463e848ff8f97e7

Observation 7451109f-68ab-4e35-8388-d83d2366c6e8 · outbound

This paper cites an unresolved cited work.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-19T10:53:02.945256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:8bfb611fa16413bacbc8a552fa60a94f0f283020f4f36b59a1cd609b07c2cd6a

Observation 2ee47962-5671-4932-8109-b7b727c7f555 · outbound

This paper cites an unresolved cited work.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-19T10:53:02.937373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:34b5078158f2635c179b273ad65d509277b2601ea9cf34b1fa4dac5af83fd2bc

Observation c7172521-7a6a-4f25-81bf-f0573e744d6d · outbound

This paper cites Analysis.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning Analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T10:53:02.935668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:61a737d931204d659ba6f2e90a83f9064983e846df5cb57532addc4e93f9f7a4

Pith citing papers

Observation 62f63ad2-eab9-444a-8fd3-cb25742f770b · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:16:19.365835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:f9efb6b87ea891e29b6326dc40a5f6d2f2ee67aea12525d16fa5e75d056e85aa

Observation 93602bd8-b9c2-4e1a-b81a-e912928bcd5d · inbound

Self-Evolving Deep Research via Joint Generation and Evaluation cites this paper.

Self-Evolving Deep Research via Joint Generation and Evaluation Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.579000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T06:31:27.811744Z digest=sha256:882a8e4549cce72ead3483c2b65f50141be24998536fcc68b3ba585065b8633e

Observation 385c9300-c51a-4d61-98c9-7011035a35d6 · inbound

IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking cites this paper.

IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.871621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T16:25:57.418334Z digest=sha256:ba0f024389f121f8ecfc45e54d046892e1afefc1197c21978161cb39f8d8e010

Observation 007fffc9-e60a-4768-bbc6-20577928739f · inbound

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning cites this paper.

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:50:10.422142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:00:37.797967Z digest=sha256:226876293567b9f801f109fad59bbf488424248c9c1109f33adc605790e8f73f