Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:43:39.668059Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.04713.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:43:39.668059Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 70ec9769-f60b-48ca-a6d3-49fda3f234b9 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Chain-of-thought prompting elicits reasoning in large language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0ffeb1-de11-4ff5-8bf3-b5482cd8a25f · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4201ee10-0574-48d4-a117-6a8ac4175376 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e925755-31ef-486a-909a-c0cd2daee820 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents BloombergGPT: A Large Language Model for Finance
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605a49f4-67c4-4e74-b45d-e8dfd122697a · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef1abde-2a08-4c32-a5a2-acce9e039fa7 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1905b2-6269-4463-af20-d0b1baabb352 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0784c5f-b2df-45c8-b930-432776fbb764 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Understanding R1-Zero-Like Training: A Critical Perspective
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7116b1-df75-4351-8918-b3fe713ebf6a · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56506d97-9f4d-4208-b9cf-dd0c17b7a2ed · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c434fbc-c40c-4bbe-92b9-e9566d9fb4e0 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac00b5aa-cc46-41ed-bd85-c273d00a1419 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee0e01e-e82a-47cb-9367-2b1a951be17b · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd16f28-233e-4c80-b702-71843a10e079 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6334ba09-e76b-4280-834e-37610ffbec87 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Qwen2.5 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10736b2-4857-4832-b1a0-d2152486afcb · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Let’s verify step by step
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77fa213-c602-48da-bd09-6f39ec1b5b8b · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2457a0-b480-46a9-9026-77784f6a0602 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efeabeec-43e6-4bfe-b139-2dab33c180c1 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db4b711-2172-4457-8992-12fff20a36e2 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7314226-47b9-4a3f-8f6b-d4863991494b · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Group Sequence Policy Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9eddf99-272c-4c12-adcb-85ed3ed4f4ca · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Learning to Reason under Off-Policy Guidance
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879668fc-aa9e-43c0-b887-e99bf21414bd · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RePO: Replay-Enhanced Policy Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 894710be-2893-4b06-bcb2-48495fcf0ba9 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af044139-4124-4971-8e38-abb3f30c6dae · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc658aa5-335e-456c-bbc8-569f751eb622 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ToRL: Scaling Tool-Integrated RL
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbe32b4-f7f7-4a81-ac5a-285c08b48206 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f0c4dd-5b3b-4455-be53-92fa6e44ec70 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e2f95c-bb53-4699-83d5-c095d79d9888 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Reinforcing multi-turn reasoning in llm agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be19805-6591-421b-a979-1c82e3d5867a · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agentic Reinforced Policy Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d857ae08-4bbd-4f34-887b-8ddda9891f45 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 777721e3-c21d-47d6-9062-6ca4cca5e100 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Learn the ropes, then trust the wins: Self-imitation with progressive exploration for agentic reinforcement learning.arXiv preprint arXiv:2509.22601, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56129498-2c5f-46f7-964c-f26bd9083420 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Self-imitation learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183b1700-f357-449a-a6ee-def7c82df27d · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agent Learning via Early Experience
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1fbede7-0b23-4089-96b2-ef3e8aae2deb · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.arXiv preprint arXiv:2510.14967, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4235b4-367b-4b80-bcde-edbb89703962 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cafa9543-7684-49c9-b7f4-a20beb77d3ed · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Process Reinforcement through Implicit Rewards
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caed79d0-a3c6-45d1-9366-664bf124d333 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Process vs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75ab180-2145-46ec-b701-018afdc389b7 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReAct: Synergizing reasoning and acting in language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce433ff1-3deb-421c-a259-1799afc7b68c · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents MIT press Cambridge, 1998
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38321f94-a052-4498-a382-c60bb5f72119 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Generalized proximal policy optimization with sample reuse.Advances in Neural Information Processing Systems, 34: 11909–11919, 2021
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929ef970-5392-4ea9-ac2b-1416cc842043 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aac4249-21e2-42b9-a01a-447c14971c0c · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc32df4-1653-4950-899b-3a9ce3dc25de · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Soft Adaptive Policy Optimization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d2b845-f56d-43d0-8e9b-a4288db39207 · outbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Concrete Problems in AI Safety
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.