Pith. sign in

Paper Citation Record · LEDGER

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.04242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04242 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:42:48.585661Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 178ee030-9241-4f65-80b5-5bf31c43f516 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:5809ef5d0881837a2cb0430fcc1a6153b5504ef35eab85b9935c60c7249b2328

Observation 93976136-347c-4e67-a07b-36fa8d26a94d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:2de0b7eb1d02e096fcea3ab102031ae9524107fe5814565199f92fe118904619

Observation 0a866a45-338b-463d-8713-e5a0ec7c328e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:6de64c46d52f3bf0582acf4be5081f41184303720bc42f51bf2847bafe3f5818

Observation dabedd51-f93d-4cff-9dcc-ab4ffe9f28ab · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:cbb873f964b8715e15779fca5edb5bda9e513aa47afe65e470d849b2c019735d

Observation edf4455c-ab29-4126-ad7f-f65d9057ac9b · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:d6f7f1681ac8034ac563b1ba928f9d9296fc7e89e3c3ded15f461a4d59100230

Observation 64e0cb62-eda7-4d44-9cf7-0e300816ebda · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:2d3187db0019c4200ad161d27f1d5d46b660846af2b14a799d0455935191af0e

Observation 749e3263-caa9-4f95-ad41-54336055499e · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning AgentBench: Evaluating LLMs as Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:0536bb3a0acda8237754cf67481df3da05e0ad1f9878cc793b799ac1294d64f1

Observation 8e4a3955-d5ed-4530-8ff1-9ed8a0726123 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:3451f5ef75f73e13df0b7d61682be0c96545eab2144f5bc86ed60ca4ca882f8a

Observation e87dcbcb-7037-4b0b-b2fb-8c8199957eee · outbound

This paper cites Proximal Policy Optimization Algorithms.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:1a96321032a6a147e30bca64e3371eec6784f898690c0746fa53169ebec9abc0

Observation 441ad118-9f5f-4ae6-a75d-57c47984b60c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:f2d698549c17094e6496b0afa80d00985482e6f2b581ea52f0af37ca709d14bc

Observation 81905e31-a397-4083-8ada-1d9aa26f3c55 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:16ef3db821153130cc0ab1b7584b9341f6cf2e2b853ce0b3e33dd57c136b9d3a

Observation 8a6fa4ea-990c-445c-9f08-767973d964ec · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:195b4b246d00b23e61676544c01521feff7bc5e8f98aaa43d1184f067f5ff21b

Observation 9787f238-e3d7-4b05-8b14-a3bbd516f723 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T20:42:48.585661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:42:48.585661Z digest=sha256:02e6b6e4aa27840a8cec238a217acf7dca237f476f401f907841a985348a99aa

Pith citing papers

No inbound Pith citation observations are available.