Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T18:54:12.287974Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2605.26958.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T18:54:12.287974Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:45:20.213019Z
A source-named dated measurement, never combined with another source.
Source: cited_works
9 of 9 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ea3269b1-723f-4083-9629-43f89a2235b6 · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b42d9c75-6cb1-46af-a9df-d7ae210f34f9 · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ad1ebb8-5f75-45c0-a3ce-b4aaa5d68dd8 · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Understanding R1-Zero-Like Training: A Critical Perspective
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 678d383f-070b-4a2a-9e8c-e8fcfcd89991 · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Proximal Policy Optimization Algorithms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 17665943-1fa0-4311-a2bb-a08751e4ff4c · outbound
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b5ebe72-64ad-49f6-9bd7-a05aeb16b743 · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4de462b2-b73d-4c42-9feb-49b5030b0633 · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8caf034-7943-4c82-afe0-0a5e22074f2d · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Each tool cycle must follow: <call_tool name="...">...</call_tool> <tool_output>...</tool_output> <think>...</think>
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 758c813c-19dc-48a8-92ef-cbc25efd514d · outbound
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation google_search
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e80624-d368-4dce-b65b-1becc619e022 · inbound
SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.