Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2510.00977.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:01:06.623978Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T17:58:47.896297Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 964beff2-08cf-4683-916e-3735a1a8ba7c · inbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models It Takes Two: Your GRPO Is Secretly DPO
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation da0e14bb-d075-4a25-bd41-f43183387ad6 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex It Takes Two: Your GRPO Is Secretly DPO
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d8860af1-1bc7-47e7-81df-981c76b9dd47 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex It Takes Two: Your GRPO Is Secretly DPO
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation aac73811-644f-48b5-93ff-571c0e86ac05 · inbound
Interactive Critique-Revision Training for Reliable Structured LLM Generation It Takes Two: Your GRPO Is Secretly DPO
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bb338d90-c674-4863-ad86-9e49946b8157 · inbound
How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR It Takes Two: Your GRPO Is Secretly DPO
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9068efe5-88a0-4fcb-b288-3f4441d23828 · inbound
Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs It Takes Two: Your GRPO Is Secretly DPO
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 42e87114-b9c6-46d0-be4a-b409ff485a96 · inbound
Rethinking Groups in Critic-Free RLVR It Takes Two: Your GRPO Is Secretly DPO
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 467aa4b5-4c6a-4ae4-927f-8a362a877226 · inbound
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity It Takes Two: Your GRPO Is Secretly DPO
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8cdab405-7cf5-4d94-be43-5d3540f04d9a · inbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models It Takes Two: Your GRPO Is Secretly DPO
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.