Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:42:48.585661Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.04242.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:42:48.585661Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 178ee030-9241-4f65-80b5-5bf31c43f516 · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93976136-347c-4e67-a07b-36fa8d26a94d · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a866a45-338b-463d-8713-e5a0ec7c328e · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dabedd51-f93d-4cff-9dcc-ab4ffe9f28ab · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf4455c-ab29-4126-ad7f-f65d9057ac9b · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64e0cb62-eda7-4d44-9cf7-0e300816ebda · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 749e3263-caa9-4f95-ad41-54336055499e · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning AgentBench: Evaluating LLMs as Agents
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e4a3955-d5ed-4530-8ff1-9ed8a0726123 · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87dcbcb-7037-4b0b-b2fb-8c8199957eee · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 441ad118-9f5f-4ae6-a75d-57c47984b60c · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81905e31-a397-4083-8ada-1d9aa26f3c55 · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a6fa4ea-990c-445c-9f08-767973d964ec · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9787f238-e3d7-4b05-8b14-a3bbd516f723 · outbound
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.