Pith. sign in

Paper Citation Record · LEDGER

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.04713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04713 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T14:43:39.668059Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70ec9769-f60b-48ca-a6d3-49fda3f234b9 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Chain-of-thought prompting elicits reasoning in large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:1bea4b8782a69abb458183076da1c007edb9114d130d99e759f9b03ebeb7f9c7

Observation 6f0ffeb1-de11-4ff5-8bf3-b5482cd8a25f · outbound

This paper cites GPT-4 Technical Report.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:29671c82c59ae0b58839878bfcc4a01301b2b31768e92519acfa33a521ac7e6a

Observation 4201ee10-0574-48d4-a117-6a8ac4175376 · outbound

This paper cites Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b29c56a990ac96eb29bf3a26cf49def194f976ba84d5068cdaf7ffa5f0483b4e

Observation 5e925755-31ef-486a-909a-c0cd2daee820 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents BloombergGPT: A Large Language Model for Finance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:3e9af432ec067044bd54e505d0559e9fec5ca4b4dbfb7a45f1dac33d948869fd

Observation 605a49f4-67c4-4e74-b45d-e8dfd122697a · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b2ba992d2e60ef8aa3a53d3a8121c4f60575e7802deab08f330dc19057dff580

Observation 0ef1abde-2a08-4c32-a5a2-acce9e039fa7 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:397a4175bfe18c2c5d5ab48a5d8f1413dfdac19854621f5551c3b39b806d6816

Observation fe1905b2-6269-4463-af20-d0b1baabb352 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:8098b3c352a73097d96599e9bdbd391dee97f26ab9e9b3be26459f0e108a21e2

Observation d0784c5f-b2df-45c8-b930-432776fbb764 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Understanding R1-Zero-Like Training: A Critical Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:8ec41e21f3044e544bb3b31273adb954dda7c3e7b08c539085aa51e7b0bc8ef2

Observation 0d7116b1-df75-4351-8918-b3fe713ebf6a · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:5b668343004cb3401313b790b9223b396da9e899cc16c4cd63a6285dfc62505f

Observation 56506d97-9f4d-4208-b9cf-dd0c17b7a2ed · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:8898bbbf2332e52ca71bff6553cf8a9b0c0866a9e2c7fabc34ead59f22650870

Observation 4c434fbc-c40c-4bbe-92b9-e9566d9fb4e0 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:06e36542f4faf84bc1d814b3a38a50a4e1a9d8fe0697b4b4293e3a3841fb01ec

Observation ac00b5aa-cc46-41ed-bd85-c273d00a1419 · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b9ffefec13f036f6f2029469ba560ee6f9dd66539f3796c81b2979209528b49e

Observation 2ee0e01e-e82a-47cb-9367-2b1a951be17b · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:f2f32ab81c3d3474ddf618acf3896c4c6203d09a5449ce706e5db24e834f0987

Observation bfd16f28-233e-4c80-b702-71843a10e079 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:1c8231a38d7a5ab255a7a1bbdfc375be8908fdcefe7b91821f51e16129595556

Observation 6334ba09-e76b-4280-834e-37610ffbec87 · outbound

This paper cites Qwen2.5 Technical Report.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:ea9235620f4446540ee2ce0820f94c89a98af2b23834ff4cba6e08ee0d6b83dd

Observation c10736b2-4857-4832-b1a0-d2152486afcb · outbound

This paper cites Let’s verify step by step.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Let’s verify step by step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:d7509e64697b0c811641918c0526bdcdfc0bdb9138f1717610ebb35418edbb1e

Observation c77fa213-c602-48da-bd09-6f39ec1b5b8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:3e9daaaae4fcb1d98dd159a3b139f12c3f923cc4b001510899748c0a5ab87fc4

Observation ea2457a0-b480-46a9-9026-77784f6a0602 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:999c62a1ed51893fa45875819dd5bb323dfb405e22b8f6c6e8727f14a87e5304

Observation efeabeec-43e6-4bfe-b139-2dab33c180c1 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b69934eaef76e3b15fb87c4388e09169716e1a94001bdef2716264fa4f6e6b1f

Observation 5db4b711-2172-4457-8992-12fff20a36e2 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:3b445a3b528449da9db1c16b276ee119854717d97fd73575a9a806d1bc6e77b8

Observation d7314226-47b9-4a3f-8f6b-d4863991494b · outbound

This paper cites Group Sequence Policy Optimization.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Group Sequence Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b74f378927ab288e4975f5640352f82b467bc942d84b6af554e5c11b3ea561fe

Observation b9eddf99-272c-4c12-adcb-85ed3ed4f4ca · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Learning to Reason under Off-Policy Guidance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:83ab4f797148c3e088daa196f2b636e2e8259834b161fd3497071e7357630b0f

Observation 879668fc-aa9e-43c0-b887-e99bf21414bd · outbound

This paper cites RePO: Replay-Enhanced Policy Optimization.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RePO: Replay-Enhanced Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:a2e271f24397bbe8ce639726d6d68fe671d90b07fb664afc35c9d27ecf7c3bc6

Observation 894710be-2893-4b06-bcb2-48495fcf0ba9 · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:1834715b5705650aec0fd6830a78b5dff14f42e2b6ab64c2bc43c5f6cb67ea7a

Observation af044139-4124-4971-8e38-abb3f30c6dae · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:850cd9d667b5f67933b782ec4d15487173facfb0c336b39e005a1cbddd3ed01e

Observation bc658aa5-335e-456c-bbc8-569f751eb622 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:bea7c1f679683bf1479d11a8a079667863d378932a63290e6a5b608b261503d5

Observation abbe32b4-f7f7-4a81-ac5a-285c08b48206 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:4fdce2bcab46553a829470822dbcf6f9b447ddb77f3c42070a87382f6989c236

Observation a2f0c4dd-5b3b-4455-be53-92fa6e44ec70 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:78f95a9d026a0f7730f9f70b9d08e6fdb3d7b572d1071cb06d32dc1b05717d6f

Observation 12e2f95c-bb53-4699-83d5-c095d79d9888 · outbound

This paper cites Reinforcing multi-turn reasoning in llm agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Reinforcing multi-turn reasoning in llm agents via turn-level reward design.arXiv preprint arXiv:2505.11821, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:9ba26b8fc6e35205764a0e8b295daf150ae39bba9efcef479ae5289a062d63ba

Observation 6be19805-6591-421b-a979-1c82e3d5867a · outbound

This paper cites Agentic Reinforced Policy Optimization.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agentic Reinforced Policy Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:bd1e4f869551e94c5d6e00e8037c581b6a7541ba74ae196b4b32e25e38c8a9e0

Observation d857ae08-4bbd-4f34-887b-8ddda9891f45 · outbound

This paper cites Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:fcb18b407dec99b66e97979bf636308115a4df612deedc3572237ad76825326f

Observation 777721e3-c21d-47d6-9062-6ca4cca5e100 · outbound

This paper cites Learn the ropes, then trust the wins: Self-imitation with progressive exploration for agentic reinforcement learning.arXiv preprint arXiv:2509.22601, 2025.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Learn the ropes, then trust the wins: Self-imitation with progressive exploration for agentic reinforcement learning.arXiv preprint arXiv:2509.22601, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:135380f3a506b903f8d28d209cbe8be2b9b23e3db9989a273542dbc5010280c8

Observation 56129498-2c5f-46f7-964c-f26bd9083420 · outbound

This paper cites Self-imitation learning.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Self-imitation learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:f474363dbba4e8c866ed168428085aa9989437f91ad4a22b2686936a48a1035c

Observation 183b1700-f357-449a-a6ee-def7c82df27d · outbound

This paper cites Agent Learning via Early Experience.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Agent Learning via Early Experience

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:2d9bec50ec3b56b88b89b58c2bda378bea3194aa8ec6b97acc8601034bdbda11

Observation b1fbede7-0b23-4089-96b2-ef3e8aae2deb · outbound

This paper cites Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.arXiv preprint arXiv:2510.14967, 2025.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.arXiv preprint arXiv:2510.14967, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:6c84318fb623dca1aa6e17e742ba4c06cc0bd625b9d5186d2282db20f29acb84

Observation 0a4235b4-367b-4b80-bcde-edbb89703962 · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:6ac792e5547d54f75f19fc81bb0a3c52b9db3f4a1bd267f95b7768dda16264ae

Observation cafa9543-7684-49c9-b7f4-a20beb77d3ed · outbound

This paper cites Process Reinforcement through Implicit Rewards.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Process Reinforcement through Implicit Rewards

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:979de7368aaa5dcc37fb15af64fc2b4d7e8f56e10acdc3699e90636d8609fa08

Observation caed79d0-a3c6-45d1-9366-664bf124d333 · outbound

This paper cites Process vs.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Process vs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:7e5968cce6a2361e18c0953f825b223c7281901c3f9377f956e8b210e2ee0171

Observation c75ab180-2145-46ec-b701-018afdc389b7 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReAct: Synergizing reasoning and acting in language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:39897c51d6d88a98c02cbae953af7a5d1b31fcff18fe9b1d97202dc67272897a

Observation ce433ff1-3deb-421c-a259-1799afc7b68c · outbound

This paper cites MIT press Cambridge, 1998.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents MIT press Cambridge, 1998

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:39fa478e56c41c81a2077ee4efe6e0a7f43da0b810999e522c407de8fa4bdf05

Observation 38321f94-a052-4498-a382-c60bb5f72119 · outbound

This paper cites Generalized proximal policy optimization with sample reuse.Advances in Neural Information Processing Systems, 34: 11909–11919, 2021.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Generalized proximal policy optimization with sample reuse.Advances in Neural Information Processing Systems, 34: 11909–11919, 2021

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:20bc06be218681f1fc7149c0017a7342cd8f0b8e037438a52d7884e7baf6e2b5

Observation 929ef970-5392-4ea9-ac2b-1416cc842043 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:c801923727d9c471a8833cb7a8f4522677176b34ec9c9d7281ff88d738c7aefa

Observation 6aac4249-21e2-42b9-a01a-447c14971c0c · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:6ff8a4caea21c95e3923ea838d7c7e877a26e2b2089423c1b688cb8208d18045

Observation 1bc32df4-1653-4950-899b-3a9ce3dc25de · outbound

This paper cites Soft Adaptive Policy Optimization.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Soft Adaptive Policy Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:0e889ac72e6cbd3b8e6936913b0c40ffc8b63f99f2b56431705093673a6aa22c

Observation f7d2b845-f56d-43d0-8e9b-a4288db39207 · outbound

This paper cites Concrete Problems in AI Safety.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Concrete Problems in AI Safety

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b9dbb672d60746ca82d15d1d5966682c391dd08c5e5863d4863dc7effd521b0a

Pith citing papers

No inbound Pith citation observations are available.