Pith. sign in

Paper Citation Record · LEDGER

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 2 inbound Pith citation observations for arXiv:2510.24636.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.24636 v3

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:43:26.153787Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T15:12:55.978703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T03:26:29.328767Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4017fda2-d60e-4662-ab92-ced0d4c72468 · outbound

This paper cites LitSearch: A Retrieval Benchmark for Scientific Literature Search.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning LitSearch: A Retrieval Benchmark for Scientific Literature Search

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.166612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.166612Z digest=sha256:7e2a98e97bb35f0730c10006378ee6f0de4928040325a42ab57735cfde88bf4e

Observation a669659d-eb68-4d3f-b5f3-51fdfb80fb1c · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.842667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.842667Z digest=sha256:032b758930dbb230e56251581324178cca375dcfacfbe8f9f18c876e49fe5198

Observation 0d203271-3140-4f3f-baaa-ff3b3defe4d4 · outbound

This paper cites A Survey on LLM-as-a-Judge.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning A Survey on LLM-as-a-Judge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.975882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.975882Z digest=sha256:e1bffe384dd510cd4067735eaea079eea2fbc242d3bf811b52a8f8997f34eca9

Observation c37af3d7-af75-4336-af67-24164bd603c7 · outbound

This paper cites Reward Reasoning Model.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Reward Reasoning Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.151892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.151892Z digest=sha256:2ca9d7ab859ef66c1703960bec3716d4829993f3e0a2a671366274701d61b2f3

Observation 4536026d-9fbc-432f-ab63-f788f4e274dd · outbound

This paper cites GPT-4o System Card.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.321802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.321802Z digest=sha256:7ce0c839206e4c62cb8054b7592d861c1d083ef294186f138a23fd5718e7352c

Observation 9c9888e2-16e0-4a62-a3a7-75b701758423 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.471106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.471106Z digest=sha256:7c6775ba693e70a8e28625c5a9c352129f198046dafa08e896d67e1ca59edab6

Observation bb48b874-e55b-4f2c-8572-c3f6a9ed6cb4 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.665067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.665067Z digest=sha256:98c07febff4ac63b4d1b01815634987ad3bf7d75431f90de79ee445d4502881c

Observation 2ec46496-2ed5-4882-87c3-f8db1312b954 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.808467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.808467Z digest=sha256:10e57844a9b9f3e86c62ceb865c00a5bbf6329580333d261b20998e01ce35b51

Observation 76c6d061-bfe2-4bd6-bed9-7a813641a7a9 · outbound

This paper cites DeepSeek-V3 Technical Report.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.944019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.944019Z digest=sha256:e842bc714496c5d80eda1dc6ad4d3419be5efe35ef67da10ca50e04d2cc6a46b

Observation 0825a5ca-74e1-48e8-97be-1945e1419bf0 · outbound

This paper cites ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.107788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.107788Z digest=sha256:a0226329d232c09bdbe11d333793dd8e8871dd0ac8dc6575201995129e455787

Observation df821748-0ef0-4ed6-a211-7d8dc675b0d9 · outbound

This paper cites Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.470188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.470188Z digest=sha256:a3a8f4192999a3048b63fa74ecbeda5f54d2af04469d9201de2b2b552d7f5299

Observation 620c362c-7cd2-42bc-ab82-327944e133cc · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.595789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.595789Z digest=sha256:93a898ac95d9c954d091c547e49941f2b8c6f65157e4f78e16939751726c258d

Observation 5494b6d3-cc63-4ce5-aa65-7cff8bc3b62f · outbound

This paper cites DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.717056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.717056Z digest=sha256:7083b1d81bb900d0ceb3e5313cf144013ac50d73c7608e012a4a4443a77bd5eb

Observation 9154d8a7-6529-4d91-a2b9-cbb063127e87 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.904052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.904052Z digest=sha256:2ec174da5caf53751315ece0fc750fa3600c6fb69a87767b6a7fe1a1786a3c66

Observation d1803f90-09a4-43e3-9cd1-ab8b8054d950 · outbound

This paper cites <search>.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning <search>

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:26.153787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:26.153787Z digest=sha256:873bd3538ce0fa97f308a98c180f806a835f044d0b9974ccd9ee88652635ef7b

Observation 0e72a6e3-9c4c-430a-b864-1cc7972b1e72 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.323549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.323549Z digest=sha256:7c23a167e83ed67c08e8dbead8cc6126b7692788c4cc5e7ab9dbb381ce24e77a

Observation fe62aaab-6e99-4cff-a2d1-4f62d6b61ad2 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.625509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.625509Z digest=sha256:be73cecdc531ee212e6fe5d062775cae8fe9d7798bd3f16de0dbd10f159c9df5

Observation 277f7038-62bf-4de2-9117-c86f581cb9d5 · outbound

This paper cites Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025a.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025a

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.483080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.483080Z digest=sha256:ca0c12ff54602b3d59553331893854156855b83ac356b412c988081b601ab74e

Observation 6210a4bb-f318-477e-b054-a9acfed1291d · outbound

This paper cites TravelAgent: An AI Assistant for Personalized Travel Planning.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning TravelAgent: An AI Assistant for Personalized Travel Planning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.344859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.344859Z digest=sha256:152d9ca9a492aecb4dfc47b8f1e4e5e4a8a89fd240867946952b68dcab6555f3

Pith citing papers

Observation 2fa5c8cd-2e1d-4666-a6ab-e196f6207750 · inbound

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation cites this paper.

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:26:29.330841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:03:54.185019Z digest=sha256:05ddc8a67284ac1ce1d6d1bd42ab1ec6ac1de604fde114aafd4017b06e30990f

Observation f31f01ec-4405-47ac-808b-e0a1c6ec45da · inbound

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents cites this paper.

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-04T15:12:55.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:12:55.978703Z digest=sha256:f798fa92c5ccb57a29a152e121ced2932d2e896eb6f2935581075b608cf2497c