Pith. sign in

Paper Citation Record · LEDGER

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2601.15141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.15141 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:01:51.020857Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:34:52.336080Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.265122Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4ab3d3b-027d-4553-895e-a50d9a704d3c · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.349112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.349112Z digest=sha256:0b1010b4adb0f483557dca2f460edc20d8fe3a76a84f0e3d15c8ee20a860217a

Observation e714dde6-3275-43e7-887e-7dc0ec4f00d8 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.654520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.654520Z digest=sha256:1131fe5a46a1a696cae35a6f82e16deb9af4ead8640a76862bed1c7b717a538e

Observation bbaf49e3-f50a-4ba9-8bc2-a838fd1b49c6 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.820681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.820681Z digest=sha256:c858b2e403333bd4b18a9a7a05f68523ae090bf7a83d44f751d4fbbea3ffd655

Observation 180c1cac-d97e-4ecd-8ddb-c91f4bfcbdc5 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.960580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.960580Z digest=sha256:5afdaab49d5b4faec7bdc6eb5fc7ab713aed983c5b334558c2a051b81698be09

Observation ed3f8c17-543b-47e7-81eb-eed34f2ded00 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Skywork Open Reasoner 1 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.068086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.068086Z digest=sha256:ab1d9e302d6415f3f7c8281f67a154b5314053d7e8e016019cb736e3576a85ef

Observation b5601a98-0cf0-4a9b-ab4b-e18507318bbb · outbound

This paper cites Qwen2.5-Coder Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Qwen2.5-Coder Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.175019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.175019Z digest=sha256:63de641273633180c162c2776f28f8d864bbd035bd03bd9e669ecc2b2c663cef

Observation 11165d18-2286-4f7f-bb6d-bb09a955bfec · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.238425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.238425Z digest=sha256:a98adff055c88d7686efb9e27ead05358716ef97f3aa69d5509efe692b5ce4e4

Observation 142c089e-79e0-4e3d-89ff-a7f711dcfe86 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.282180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.282180Z digest=sha256:89424508af26dea956521a82d54cc5c522d5013bbc3d4440456873eae224b2d6

Observation 32cf656d-bd86-4703-97a2-04814d70c89e · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.333636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.333636Z digest=sha256:78ea1e4dfba82c43e91d98198a95c7015141eb4db5b4b9d4907c6f59caf55dc4

Observation 4a7dc707-789b-4e8c-920a-913bc6f4425d · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.390340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.390340Z digest=sha256:3d09872be4bbe5bfd594fc67206b987f693730c5c297c37d55737833b161402f

Observation a2ca8d26-6970-4f63-b54e-ed5efa763351 · outbound

This paper cites DeepSeek-V3 Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.434593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.434593Z digest=sha256:9fd319fb1fff7e71a2cf2bb0b94dbe0a4dae0675e83d475b6a2fab30c58106cf

Observation 434df131-b215-494a-a712-39f06697e8c7 · outbound

This paper cites TALM: Tool Augmented Language Models.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning TALM: Tool Augmented Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.506280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.506280Z digest=sha256:85082db2edd51ead0e548ef6cbceff185614a38ae644b18fa14b2cdb3cbe64a7

Observation ab254e2c-0c71-4078-b29c-ad48dd7389c4 · outbound

This paper cites Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.563297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.563297Z digest=sha256:f0a87444e81c4d1ed1708e9bed02ab855a77ab152daa293629d4a51cfeb88d8d

Observation 7ab0f697-fc9d-4e97-8423-7290699715d7 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.606749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.606749Z digest=sha256:0c3b0bd1a8ad66f0408db2a3df77f92880f66c84cd84cb7910f19a01bc4b5564

Observation de3eaf9c-8587-4984-a51f-5b8bc5bbad01 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.682153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.682153Z digest=sha256:64f5586362584ae14d3487e6599fc2f39f8820efc1094ff4cac461e5ee872e76

Observation 98bcfdf8-07fa-49b6-8372-7ea0a030e119 · outbound

This paper cites pocoo.org/2025/10/17/code/.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning pocoo.org/2025/10/17/code/

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.821208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.821208Z digest=sha256:2bc4977ca088aef260c6f56db9a302a9a79bd596466c397a589ff74d26fde408

Observation d2bb13be-c000-42e5-a325-75a7376d07c8 · outbound

This paper cites rStar2-Agent: Agentic Reasoning Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.933722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.933722Z digest=sha256:16f3aabb9c7b1473f1f93867e6f8f55e0477ed82c708ac85192c1ff224b8a3d2

Observation 5928fef3-ab27-4ab0-9044-f5aa5a6d68eb · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.047393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.047393Z digest=sha256:54853e8024f41941c18bc98391ab14690b6652bf86204ea81f62b4ee7f11ec9a

Observation 7323dc5c-2272-404c-9b85-c0a26531bf90 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.181248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.181248Z digest=sha256:53d700bdc3fadf87798f59f4c0986c5f74fcea1696faab505c31e5c0bf755daf

Observation 7d71d9eb-b6e0-40a5-ae43-97b93d7b3733 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.527848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.527848Z digest=sha256:4966c95dcb287e6268ee87505138558e9c5671e3207b349d7674126c9d1da054

Observation 3be6609d-e3ea-4139-b1a0-b2be409ba534 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.680369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.680369Z digest=sha256:c787724fc8193dd2c9fb1d047bbc74926bc7a4919ebb07dd8cf7e1525b5eec68

Observation e7050150-7c8a-43fd-814b-2a4d581f2c0d · outbound

This paper cites True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.797041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.797041Z digest=sha256:274112f70f62301a508f88f0502407615f6d4fb7a375434152d05eb332e9200b

Observation bd56bd04-6138-4ab8-bd2f-20d03ae4cd79 · outbound

This paper cites DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.824574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.824574Z digest=sha256:150a6330cf10a82c964d0fc8d5b9513e74e8f61fa13aa9495e1b994307610b78

Observation 8d8949bb-f3fb-444a-a76b-3f016ae956c1 · outbound

This paper cites Qwen3 Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.828166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.828166Z digest=sha256:df77f74a9cbc5dbfe47f5e02b7f276ed9ca243e4609bf3281f27c98bfcd79b0c

Observation e2a7845b-c40d-4353-9094-ae815c6317e5 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.876731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.876731Z digest=sha256:1b9bd95982e5b5e538aecb1a939db4bf818befb8aafab7b545145e4ef3b39086

Observation 39931dbd-5fe6-461a-8fa0-ca30c8a38a7b · outbound

This paper cites purified.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning purified

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:51.020857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:51.020857Z digest=sha256:375619b95729d2728d1096d4dcb8e4565c0c1077d12be246047013611df060b6

Observation b564bf72-8e55-4866-ab6c-5958a9e452f4 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.366780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.366780Z digest=sha256:01bf7e28473e9281cff2be9178eb8e092569522078d548852071330eb4c0abe2

Observation 9246615c-e139-4311-addc-cc4d0d27038d · outbound

This paper cites Agentic entropy-balanced policy optimization.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Agentic entropy-balanced policy optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.539543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.539543Z digest=sha256:5c236f3051018139a6884c8bbdf633cf95c2aea43b3939f01b9520199539077b

Observation 588c30f6-4fc9-4a2e-9190-e5a041ed1f75 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.015328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.015328Z digest=sha256:0ac005ac4a80b7a2400674a15018e19de6c916ce41baefc6b8683773b32bc5ea

Observation 3d2ac0bf-b775-47b6-b766-fa0590755ef2 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.241440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.241440Z digest=sha256:0db3000bceb943bf5310674c47c8d88b4681b2b14ae2ed837e5c34f947b01c4f

Observation 022a5203-10bc-4fa8-970e-a6c44630f96e · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.458376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.458376Z digest=sha256:5ac396aadb1b7db8c8b478a30ed78740f93eeaefee638fe172f5d89966c22383

Pith citing papers

Observation 5b4f6607-78ba-464c-acf5-87d8e6f19b8d · inbound

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents cites this paper.

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:02.718893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T06:10:26.185447Z digest=sha256:dcca4522b5a774de918f2412f50ebad34b1606c286db5d5680881e85bbe231f0

Observation a0d08ac2-af13-4bb9-9029-08f8a770c63f · inbound

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents cites this paper.

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:02.718893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:34:52.336080Z digest=sha256:8de8b4e9e563d64292b49aa205ad0395287c88d77d61535623c37bb7eec686be