Pith. sign in

Paper Citation Record · LEDGER

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2508.13023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13023 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:54:32.886708Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:45:00.969348Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c77f7630-e7d6-44fb-b63a-597701da933d · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:05:31.574588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:95271f31ae1905ac4fdca10a56358e9d0d0d308264943d2f87fdbe943275acd3

Observation 7199d10d-8201-4edd-bc17-a5419835fcc6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:48.905173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:4e76d28926e26fefa742d0d31e5881b8f3ece6514223a113239c5021583ce54b

Observation 965b4760-9fc6-4212-9403-9f045c9e2cf2 · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:28:52.857138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:28:06.253200Z digest=sha256:c3bd6e6ab3abbc5e98ae6442a3a0acb2e5b54165133076a99bd014e55765238b

Observation 352b23a7-b5db-4e71-9ddd-108c87290dcf · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:00.970934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:37:18.563704Z digest=sha256:cfee15e3e25ed6c1e1d44617d64cdab3e101f20c22328e5d6bdd13342ddf6b54

Observation a7c748f7-a37a-483c-9dcc-c030f20fa417 · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.886708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.886708Z digest=sha256:7a52985d6aa371cd31ccacf845590edaf44c9abcfd09c036ae61545970dd47c6