Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:00.896630Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2506.23667.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:00.896630Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-15T21:51:10.972744Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T21:51:40.816217Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bef74cd5-63f2-434e-a12e-c43090396fe9 · outbound
L0: Reinforcement Learning to Become General Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8608f17b-f4a3-47d9-a94b-ba8cc8b87810 · outbound
L0: Reinforcement Learning to Become General Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff586b0-bb82-4514-b5a2-ad612c94b987 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 899f4bf4-0f05-44c5-bd4e-30cd78ef346f · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72fa6a89-92e5-439a-b933-439aa24f0a6a · outbound
L0: Reinforcement Learning to Become General Agents RAGent: Retrieval-based Access Control Policy Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7acf283a-68bb-4849-8f3f-ee404e39634e · outbound
L0: Reinforcement Learning to Become General Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7a83f0-e58a-4379-86a5-f54943bd6301 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0371d9fe-77f0-4a83-8b55-39be7cf64e4a · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f9a0550-db7d-461b-b08a-5dabcfa900b1 · outbound
L0: Reinforcement Learning to Become General Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8063a002-7b35-4e2b-bd16-40aca3581f7f · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10349242-37f1-4f2e-b007-f0b20d7ab42a · outbound
L0: Reinforcement Learning to Become General Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6394dd20-57cb-42e8-b990-3c0703ba3894 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b553d063-6f8b-41cf-a758-fb45b1b0bfd6 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7fbddeb-c4cf-4e61-afac-b28e34b7a3a4 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e091268f-41eb-4008-bc86-6a39117b640e · outbound
L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a587171-c7c8-4a4d-b714-114a72425a2e · outbound
L0: Reinforcement Learning to Become General Agents Qwen3 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f948a6-9df1-4632-a191-5819d4b31832 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9caf3835-4d4d-403e-ba6c-4f0e747b91a4 · outbound
L0: Reinforcement Learning to Become General Agents Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f6b2a6f-fb70-41df-906c-053903ef2397 · inbound
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation L0: Reinforcement Learning to Become General Agents
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1309c335-5c4e-4801-9ba6-78bcd1cdb231 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning L0: Reinforcement Learning to Become General Agents
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.