Pith. sign in

Paper Citation Record · LEDGER

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation

As of 5 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2605.26958.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26958 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T18:54:12.287974Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:45:20.213019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ea3269b1-723f-4083-9629-43f89a2235b6 · outbound

This paper cites DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:03:51.676587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:b38cb50fccab17f4b06dc40ea1ce8d90ac3e9a380fbdc5c65009f7141dc7fa50

Observation b42d9c75-6cb1-46af-a9df-d7ae210f34f9 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:03:51.679831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:de846f7438f95544dbdbfcd4742ea49e8812ec219558c526ababdd5e0e6711d3

Observation 5ad1ebb8-5f75-45c0-a3ce-b4aaa5d68dd8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Understanding R1-Zero-Like Training: A Critical Perspective

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:03:51.672444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:4acd07c028148fed241bd0fd475e2088369937bfc1162f0c8bc12824897ab3c0

Observation 678d383f-070b-4a2a-9e8c-e8fcfcd89991 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Proximal Policy Optimization Algorithms

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:03:51.676298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:f947478872088de975d79099f6891cc2511d79a7fd0dd85d1b8072c5fc068e26

Observation 17665943-1fa0-4311-a2bb-a08751e4ff4c · outbound

This paper cites query":.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation query":

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T18:54:12.287974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:66c796a622f33fdbf3c196f70b2a06e2bd4822c1294f379ccf6d231f0b2adf4a

Observation 9b5ebe72-64ad-49f6-9bd7-a05aeb16b743 · outbound

This paper cites an unresolved cited work.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T18:54:12.287974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:051b3687715acaaba2061d41d3d131ac89bbc43fa41b68fcd17e50b1abcf7e8d

Observation 4de462b2-b73d-4c42-9feb-49b5030b0633 · outbound

This paper cites an unresolved cited work.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T18:54:12.287974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:705b53877d8278030e65d9c4f875d9cae69c67139c39d866df3e163c500df929

Observation f8caf034-7943-4c82-afe0-0a5e22074f2d · outbound

This paper cites Each tool cycle must follow: <call_tool name="...">...</call_tool> <tool_output>...</tool_output> <think>...</think>.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation Each tool cycle must follow: <call_tool name="...">...</call_tool> <tool_output>...</tool_output> <think>...</think>

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T18:54:12.287974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:1b165bb86c94f803301ef574c059e3d639f7a24d6e5c585c6933eab3cbc9749e

Observation 758c813c-19dc-48a8-92ef-cbc25efd514d · outbound

This paper cites google_search.

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation google_search

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T18:54:12.287974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:54:12.287974Z digest=sha256:44dac8f9bd99aa14bcd5547c0f587ff7c390f1fe9ab21ffa63c4c64f58500c72

Pith citing papers

Observation 29e80624-d368-4dce-b65b-1becc619e022 · inbound

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task cites this paper.

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T03:45:20.213019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:45:20.213019Z digest=sha256:09c9c5a01a70e88b371ede9e824e6aee0bcff47d52330d18e7fb006f1a45f85f