Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2602.05547.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T20:14:46.147637Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T18:30:01.596419Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 344a2ab8-f8f0-43e7-9408-7abc9bd20e44 · inbound
Target Policy Optimization Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab6fb01a-aee1-4930-a62d-f254786773b2 · inbound
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cbc1a8d-a12b-4309-af7d-e50be85a6bf9 · inbound
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78273ebe-f991-4a52-ae1b-550e622d8d86 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0dfe8bba-30b5-4330-953a-8bdde47a9ed5 · inbound
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d463cd78-c6e6-4721-ba5e-6c78df6acba7 · inbound
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bf43f96-847a-42d8-80b3-2a1b7f22338b · inbound
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.