Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:43:26.153787Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 2 inbound Pith citation observations for arXiv:2510.24636.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:43:26.153787Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:12:55.978703Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T03:26:29.328767Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4017fda2-d60e-4662-ab92-ced0d4c72468 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning LitSearch: A Retrieval Benchmark for Scientific Literature Search
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a669659d-eb68-4d3f-b5f3-51fdfb80fb1c · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d203271-3140-4f3f-baaa-ff3b3defe4d4 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning A Survey on LLM-as-a-Judge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c37af3d7-af75-4336-af67-24164bd603c7 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Reward Reasoning Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4536026d-9fbc-432f-ab63-f788f4e274dd · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning GPT-4o System Card
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9888e2-16e0-4a62-a3a7-75b701758423 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb48b874-e55b-4f2c-8572-c3f6a9ed6cb4 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec46496-2ed5-4882-87c3-f8db1312b954 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76c6d061-bfe2-4bd6-bed9-7a813641a7a9 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepSeek-V3 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0825a5ca-74e1-48e8-97be-1945e1419bf0 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df821748-0ef0-4ed6-a211-7d8dc675b0d9 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 620c362c-7cd2-42bc-ab82-327944e133cc · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5494b6d3-cc63-4ce5-aa65-7cff8bc3b62f · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9154d8a7-6529-4d91-a2b9-cbb063127e87 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning TTRL: Test-Time Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1803f90-09a4-43e3-9cd1-ab8b8054d950 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning <search>
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e72a6e3-9c4c-430a-b864-1cc7972b1e72 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe62aaab-6e99-4cff-a2d1-4f62d6b61ad2 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277f7038-62bf-4de2-9117-c86f581cb9d5 · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025a
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6210a4bb-f318-477e-b054-a9acfed1291d · outbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning TravelAgent: An AI Assistant for Personalized Travel Planning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa5c8cd-2e1d-4666-a6ab-e196f6207750 · inbound
AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f31f01ec-4405-47ac-808b-e0a1c6ec45da · inbound
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
Reference 174
Source-reported events for the cited work
Unavailable: canonical work link unavailable.