Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2502.18548.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:40:05.242510Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T20:27:21.677423Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation af526fd5-5d42-4b8f-adaf-ee33582649f1 · inbound
LaViPlan : Language-Guided Visual Path Planning with RLVR What is the Alignment Objective of GRPO?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceacf06a-fc2f-4d2f-a80c-e3815ef21d0e · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges What is the Alignment Objective of GRPO?
Reference 295
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9453a80a-4f89-45ac-adb3-85b33cfdf8d6 · inbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning What is the Alignment Objective of GRPO?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe3e218-3c61-4a80-9874-7cd98d65b5ff · inbound
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance What is the Alignment Objective of GRPO?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad4b882-8e19-4fbc-a89a-e9cc12206ddc · inbound
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment What is the Alignment Objective of GRPO?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 782d47d7-3fa1-432a-8f54-e29f6e876ae9 · inbound
Small Agent Group is the Future of Digital Health What is the Alignment Objective of GRPO?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d04fe06-516c-4049-9ca4-eb37bc0c4bb0 · inbound
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization What is the Alignment Objective of GRPO?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03924e2d-e892-473e-a75f-a02ff046bda8 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex What is the Alignment Objective of GRPO?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cec13215-b508-4b05-bf47-65fbd128998a · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex What is the Alignment Objective of GRPO?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8eeebef-6f78-470d-87f7-7ed8b233b72e · inbound
Predictable GRPO: A Closed-Form Model of Training Dynamics What is the Alignment Objective of GRPO?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ab400f7-ce37-4488-b9f4-c6e13e59df5b · inbound
Predictable GRPO: A Closed-Form Model of Training Dynamics What is the Alignment Objective of GRPO?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.