Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning for Long-Horizon Interactive LLM Agents

As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 51 inbound Pith citation observations for arXiv:2502.01600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01600 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:56:00.329027Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:51.026135Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation be444522-3b84-4f00-a6d2-27129c6c827f · outbound

This paper cites write newline.

Reinforcement Learning for Long-Horizon Interactive LLM Agents write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.175690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.175690Z digest=sha256:3ed319a66da890e1899a6793241a106907f9f8a4f119c9371cf118ada608983d

Observation 9e9d7b82-4b1c-40d7-969d-86de32af08b9 · outbound

This paper cites Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLMs.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLMs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.929673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.180044Z digest=sha256:40048c1f0f978cb257f10e4b51383a4b2c0d9b64567b1f99ad389fb2bcde981b

Observation 9357bc52-4d7e-4185-964c-58cfaf9368c7 · outbound

This paper cites Thinking fast and slow with deep learning and tree search.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Thinking fast and slow with deep learning and tree search

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.920816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.183508Z digest=sha256:34a6baa3a3be43e9d739af5880501325db5af2fbe434b333d80b8510a1e2bf54

Observation 3c600c5b-3e54-43f7-96d3-3ae9a1749549 · outbound

This paper cites DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.187032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.187032Z digest=sha256:1dbf30f5d936aea4f0bbc43204cf2fdfa3840730304e774f60a5e4a32e0b5f14

Observation 122fa5cb-c3db-4f03-89ea-9df8ea0b834b · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Grounding large language models in interactive environments with online reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.912383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.190655Z digest=sha256:f57afabfc0a81b685bb68d268e528e4c0314d3efeb5139267b4b72457a1cfea1

Observation e48f99d2-9cf3-4d39-a050-74ea028dfc96 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents FireAct: Toward Language Agent Fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.193783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.193783Z digest=sha256:2c5cc8ab53cd59daf24c34e60bf6bc3f9241b416e8ba4ec2a0f9fe0b59b1dbda

Observation 97a8fec7-3e07-4384-8931-71b1345616b9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.197270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.197270Z digest=sha256:58cf14ee03d073fa972b7ff093aed635911186bdd782649f93b467bd329a8840

Observation b5796287-07ce-4448-a87b-b07710f8b6f5 · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcement Learning for Long-Horizon Interactive LLM Agents The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.200570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.200570Z digest=sha256:f71f89e1de01354c4091efa85dbf923caf6d8f32b4a614e5537577f2bc35c4f9

Observation d841b603-7417-495e-86db-f2062e30d551 · outbound

This paper cites D., Oosterhuis, H., de Rijke, M., and Shukla, S.

Reinforcement Learning for Long-Horizon Interactive LLM Agents D., Oosterhuis, H., de Rijke, M., and Shukla, S

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.203577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.203577Z digest=sha256:7eac4ad7212ce649416df7a420ee3987827653941f6df14f77a225cd1ce34118

Observation 5844cb03-48df-4753-95b5-fc08356f6552 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.206362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.206362Z digest=sha256:b802c634425e8ab320888345fa8456a86dae2c811f8d3d9b9e8f7b11d82c1bca

Observation f3e6733f-79b7-4d4f-8ce6-dd4c2300cce0 · outbound

This paper cites J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Reinforcement Learning for Long-Horizon Interactive LLM Agents J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.903345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.211006Z digest=sha256:677e0e7cf80c55a7dcb2c0529c3009fa26002be7ae9a2c6cd1118a5e4e7cad77

Observation 20058a19-7f85-4784-97f4-44a1377fe38d · outbound

This paper cites P., Littman, M.

Reinforcement Learning for Long-Horizon Interactive LLM Agents P., Littman, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.894633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.214221Z digest=sha256:2dec936025fe5e00e0937841a99c37cec6fa93856f5109b578196f111fe71964

Observation 70bc154c-7af5-47f7-8fe4-2d74ec3217ac · outbound

This paper cites and Langford, J.

Reinforcement Learning for Long-Horizon Interactive LLM Agents and Langford, J

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.217235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.217235Z digest=sha256:9186dbe8f7874c428127e83cdb99e6d2d9760f68040414adcb03f5ecae0c943e

Observation d8060d98-83f4-4402-b46b-a9de15137134 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Reinforcement Learning for Long-Horizon Interactive LLM Agents VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.220169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.220169Z digest=sha256:2116542a8d9622e7bdf1f42d1c8aced53ef7c9c42c60cdabe15f2ae89204d7e3

Observation e1852615-2ed0-4e33-b45a-294be7c124c1 · outbound

This paper cites Language models can solve computer tasks.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Language models can solve computer tasks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.880119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.223411Z digest=sha256:c885a66f0ddc6d86e1f01a8a1570c36854c8bb57b4bf33315cf7043ec7e93c49

Observation c7027d52-7992-4140-a450-a787d6c6bd78 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! In ICLR 2019 Workshops, 2019.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Buy 4 reinforce samples, get a baseline for free! In ICLR 2019 Workshops, 2019

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.871327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.226220Z digest=sha256:b0d67313e97bc6f4d715719930e153c0c9a9f9e9d67e0b928ec71596b907a3ae

Observation f2c3648e-c89e-4fb3-8f93-2b4fcb7a5209 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Reinforcement Learning for Long-Horizon Interactive LLM Agents H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.862493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.229213Z digest=sha256:23750deaae0d7d5628a7e612eab91b7461da8a37d4593f9d805a98e4dd83778f

Observation 9b6f215a-b079-49ea-ac85-70693b68404e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.232188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.232188Z digest=sha256:eca87855d8fc3a7a62e862d0aa87bef837ec16a4e9666f3adbc1e409c399b4c7

Observation d7f30d4d-343c-439f-8a42-8433b614696b · outbound

This paper cites AgentInstruct: Toward Generative Teaching with Agentic Flows.

Reinforcement Learning for Long-Horizon Interactive LLM Agents AgentInstruct: Toward Generative Teaching with Agentic Flows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.235295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.235295Z digest=sha256:82322dcb4dda6cbbbb3c7fed63e28450cc38048b1b43e2f342b2135cff7c8223

Observation d447578b-626f-48f9-ae56-3057b45e943e · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Reinforcement Learning for Long-Horizon Interactive LLM Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.238491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.238491Z digest=sha256:0d11f1eb94118f94b1341f9380eaed78e0bc488b9358fc9c70f3bfd9264a1085

Observation 1c93a203-4b43-4420-b096-0ee13fc603c2 · outbound

This paper cites D., and Barzilay, R.

Reinforcement Learning for Long-Horizon Interactive LLM Agents D., and Barzilay, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.853634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.241690Z digest=sha256:73e7d4538b9ca18e5713500ecf6f89850fbccd62a7485baec0574f33249f16be

Observation 52c2be5b-19a6-4376-83ea-4ef330ac51b6 · outbound

This paper cites Introducing OpenAI o1, 2024.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Introducing OpenAI o1, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.844786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.244583Z digest=sha256:c38d845d1fe40ea3df835f626fed46ad09ebf1f3d0595a6b78128507504fb354

Observation ea2a00a5-9396-4569-91ee-7b72038ef238 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.247681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.247681Z digest=sha256:b45ca7cc65487ca0176732fdd969acf5d258fe496b58922ad5360b81ef3793eb

Observation 832d4fbe-3772-4836-ab60-ac71b6e5e84d · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.250658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.250658Z digest=sha256:b25c19d5dc5e48cbce36709cc2a444c9960f3dabde2aaf68e321c936266b5241

Observation 3f99d275-8b4b-4e73-9702-44873ce220c8 · outbound

This paper cites ToolLLM : Facilitating large language models to master 16000+ real-world APIs.

Reinforcement Learning for Long-Horizon Interactive LLM Agents ToolLLM : Facilitating large language models to master 16000+ real-world APIs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.831141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.253897Z digest=sha256:42eb2eb1e492e0b49ca73c098f779022b265c045aa154f5c1b53c2edec4d298f

Observation ed9810bc-ea2f-4e26-a9c2-836040c6721a · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Toolformer: Language models can teach themselves to use tools

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.823270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.256781Z digest=sha256:3854066124276112359b87f6e32d7aa54ffea52097c2abbdccb6776813d17f8d

Observation 25e64932-c570-4820-adec-0a67c5139792 · outbound

This paper cites I., and Abbeel, P.

Reinforcement Learning for Long-Horizon Interactive LLM Agents I., and Abbeel, P

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.814874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.260358Z digest=sha256:dcdd7b0f6714bd837f7154047616c2e24aee07faf8b775d0088602cc8f60e41b

Observation 1a64858a-7356-4e9a-84af-064cb4b44654 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.263354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.263354Z digest=sha256:ea4117aa71909efe48fcd310045f2c3c9223304d091d5647aa2c509833ca63bc

Observation a20db280-cd63-4995-9cba-3a91d9e32b3a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning for Long-Horizon Interactive LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.266419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.266419Z digest=sha256:08e332efb50ec5c2df72589e1fa4ae4c5caca2a6e9cdf3db8131881996409afa

Observation da300191-1603-4e48-ae49-8c0440154f38 · outbound

This paper cites Direct multi-turn preference optimization for language agents.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Direct multi-turn preference optimization for language agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.806541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.269423Z digest=sha256:8a4057372b327f2e7eb3790078c459097990849735177c12d3606067f7f11405

Observation 264b9455-85ab-447e-9360-3ccafe36cfbe · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Reflexion: Language agents with verbal reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.798035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.272361Z digest=sha256:3d453089cf5c5a8c0d3d8c3d88d20f0dc83ec86d3701530ec7a52d937dd15cde

Observation 8fc187c8-522a-4390-81c4-0effd84dfc5c · outbound

This paper cites D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P.

Reinforcement Learning for Long-Horizon Interactive LLM Agents D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.789415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.275289Z digest=sha256:46c0b526ab8da80686154bff0bd8bb321a3b2bf0b9e1ca91b00a686d42172608

Observation a14132ce-3559-4151-93b0-a07490148853 · outbound

This paper cites M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P.

Reinforcement Learning for Long-Horizon Interactive LLM Agents M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.780972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.278327Z digest=sha256:714b7d864f57418c7921139481f005d5a8f7de132911dced6736f17e8326e5de

Observation ad18a1a7-50cc-4f3f-8baa-34fb59e4d741 · outbound

This paper cites torchtune: PyTorch's finetuning library, April 2024.

Reinforcement Learning for Long-Horizon Interactive LLM Agents torchtune: PyTorch's finetuning library, April 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.770623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.281144Z digest=sha256:51fcdf3f09246fcd0ab4d0db5f06fd81abb6991d791530a9387e7cc11783cfd1

Observation 7b3adcb7-ac58-4d3b-b959-4641a548575d · outbound

This paper cites A pp W orld: A controllable world of apps and people for benchmarking interactive coding agents.

Reinforcement Learning for Long-Horizon Interactive LLM Agents A pp W orld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.762825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.284213Z digest=sha256:6f93798d0c0941186e565c6d764e3efe0044f17d68c2359a4c4d0c1479def18d

Observation df69eb9d-2b8d-4aba-9585-39ff767a9699 · outbound

This paper cites Executable code actions elicit better LLM agents.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Executable code actions elicit better LLM agents

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.755002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.287166Z digest=sha256:24014980529a5a3ce2d82f2e3fd5f078a877d85c71f4df5857c8c8a432893641

Observation 225abf2b-3099-434c-9383-93941f9946cc · outbound

This paper cites DD-PPO : L earning near-perfect PointGoal navigators from 2.5 billion frames.

Reinforcement Learning for Long-Horizon Interactive LLM Agents DD-PPO : L earning near-perfect PointGoal navigators from 2.5 billion frames

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.746486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.289910Z digest=sha256:b8adaa026fb699e4fab7e3e44698c65fbffc8d190b2fa5710b0e78d4724c1090

Observation d620ba99-cd16-46b3-8e9a-23932052479a · outbound

This paper cites Cut Your Losses in Large-Vocabulary Language Models.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Cut Your Losses in Large-Vocabulary Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.293146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.293146Z digest=sha256:8c1299c15be847670b854b300786ff4a16afd655c62ec02446711f6c04e09506

Observation 21c700fd-8839-4879-b517-dacc14d99b3a · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:56:00.738369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.296441Z digest=sha256:8bfff3991f24960e5006365b45a63bd98c3f3ff94b367e51a50f29fb4356a0ad

Observation 7a73299b-08cb-421b-97a5-8b925761da1c · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.299288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.299288Z digest=sha256:a38d1ff3b701074a86a18270a04025ba7bb6239a3a7293712ff4d9b9dd7bbef9

Observation 16832fe2-692d-44bf-9b16-297241d0c33e · outbound

This paper cites Intercode: standardizing and benchmarking interactive coding with execution feedback.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Intercode: standardizing and benchmarking interactive coding with execution feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.729514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.302540Z digest=sha256:43e94caa59ff1976b00a67b4044221f1deb2821bc0ee4265840158bc68094714

Observation 03e338e7-1acd-47d8-a0c4-9cbd0eb20305 · outbound

This paper cites Keep CALM and explore: Language models for action generation in text-based games.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Keep CALM and explore: Language models for action generation in text-based games

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.720721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.305354Z digest=sha256:7112593dd23e089319fbe2913c64a379897b88cc2736df69926f360ce71108e7

Observation fbf0f105-5b4d-4ae1-b854-8867a3877da5 · outbound

This paper cites WebShop : Towards scalable real-world web interaction with grounded language agents.

Reinforcement Learning for Long-Horizon Interactive LLM Agents WebShop : Towards scalable real-world web interaction with grounded language agents

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.711390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.308532Z digest=sha256:9a10d7054369eec28b36ffcb5bc7d2e87d652b098361e730962d5d3750dd980f

Observation 1f641c1d-725b-40a4-9d96-1d1ae04fc5c5 · outbound

This paper cites R., and Cao, Y.

Reinforcement Learning for Long-Horizon Interactive LLM Agents R., and Cao, Y

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.701404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.311496Z digest=sha256:92804448b5349975bcafd29be102da9f5706f715d372f51e8c26f709d04489b3

Observation 46657d16-3287-408d-9117-e9057026d4fe · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.314373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.314373Z digest=sha256:175ab8413c8aa5d896a0574715cc390e5459bb7668459e9d785345d94b3baab5

Observation db68ef43-d623-4056-a443-933467be9fc4 · outbound

This paper cites ST ar: Bootstrapping reasoning with reasoning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents ST ar: Bootstrapping reasoning with reasoning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.692776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.317548Z digest=sha256:4a7966a5f355d3fb6776bddbaef4c47be9cd85f51c5d6cf87306162db90ee26f

Observation e4826752-2c93-4639-8fb9-f9c367cf720b · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.684076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.320483Z digest=sha256:bd28b428b49cff07febaea04035be66cd083cc34dd345cb1c78a95ea4ce5d8f9

Observation 9745ed1c-470d-4890-bf8a-09c68c3c99f1 · outbound

This paper cites PyTorch FSDP : Experiences on scaling fully sharded data parallel.

Reinforcement Learning for Long-Horizon Interactive LLM Agents PyTorch FSDP : Experiences on scaling fully sharded data parallel

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.675175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.323357Z digest=sha256:87148247f4dd1186118b1f5b31ab166be32f20a7c9e1770b8e4df05406a31389

Observation 9429a772-2349-4b75-822e-4b617f6fdd67 · outbound

This paper cites and Zanette, A.

Reinforcement Learning for Long-Horizon Interactive LLM Agents and Zanette, A

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:56:00.664792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:56:00.326131Z digest=sha256:686839c48d4621d05e76f09e5024fceed4d5d6912dbda5ad3251bc3f94b91768

Observation 5208e95e-bee9-4698-a6de-66cfa3ba12d6 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reinforcement Learning for Long-Horizon Interactive LLM Agents Fine-Tuning Language Models from Human Preferences

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.329027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.329027Z digest=sha256:9c9c10b1c2df3165b4da953c631bfc4a7549e1ae32e03d22a496b7af60c6abcc

Pith citing papers

Observation e4bb1187-2562-46e3-9dc3-335f6fe88efa · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:09.374887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:1e9cffe16df439cb4e216db838fb8f50a1c725b8a95d18760a9e5f704a0307e0

Observation 8e668ac4-ae4d-4216-b8b9-cc1bdf5cac79 · inbound

Group-in-Group Policy Optimization for LLM Agent Training cites this paper.

Group-in-Group Policy Optimization for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:15:08.493759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T09:15:08.193357Z digest=sha256:02e0c4c555dc767824d638c36f1225833a404d423e307a7f3e7dacd848b447a1

Observation 0d667442-84a5-4408-9b81-d21b07ebe5a5 · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.633951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:eff19afa6847113ea71ac766fe3ffa1cba489ef418d8986b537faefff1a1e7bb

Observation e8e64e95-ca1d-49b7-bed5-f699951ef985 · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.130495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.130495Z digest=sha256:bbb4bfca268bd2fbbcbeb1eeb0ccdbc6425d0cb22c9f0edf7a8e070bb5f2e4be

Observation 0b5bd41e-3c99-44ec-a815-beaad8d3931f · inbound

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training cites this paper.

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:51:23.567751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T12:48:32.123998Z digest=sha256:4db848023f2ecdf5170458227d03814f95b935b24a69acf96ff817646d5baecc

Observation 925a5ce1-d2d7-4178-bfe5-2fb732e95e7c · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.232880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.232880Z digest=sha256:bb66ac9d600f644ec8c553a7f660d05948fd843f85fa1c6bded982e3eec1f29c

Observation ad7c6cd3-d20e-4af2-b1dd-39bbcd235b9c · inbound

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search cites this paper.

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
verified exact
orphan_title_repair, observed 2026-05-13T17:16:39.680321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T17:16:09.927267Z digest=sha256:b7cc985e389fddd84127d1c7ea4ea89e42298639c09fe002ae76afe1a9d2a461

Observation 512c6297-bb55-4f2d-bdfc-82ef2af826de · inbound

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories cites this paper.

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:36.606515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:29:37.642356Z digest=sha256:81d01da0b83fdafc754cbbfe777f12af09f3c9195b17f77d5b88346ff697835a

Observation ccb1273e-df61-4d0b-9635-b572b7a75b4d · inbound

PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent cites this paper.

PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:06:09.596177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:17:59.156121Z digest=sha256:c05ed0c03ee6982b0940f1483fed99c328582ce4ad2f145e6167a557e929b809

Observation 7feaa6e3-7208-4040-9399-24d247f613d7 · inbound

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents cites this paper.

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:01.132467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T15:28:07.981488Z digest=sha256:aced485668d28f3151a8f620e3ea3a7da6bb6378889139508421d837b9f61849

Observation 3725c6d8-76ba-44c5-95b1-bf18b55c5d5c · inbound

A Survey on LLM-based Conversational User Simulation cites this paper.

A Survey on LLM-based Conversational User Simulation Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:11.955831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T03:21:09.118243Z digest=sha256:b882463ef79a84628a372d1e01ffce6c356ca7e798fd27a6fcb7873c150ee1e2

Observation e020dca7-2282-48be-b8af-941aec689ae8 · inbound

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards cites this paper.

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:16:27.727101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T11:35:24.834237Z digest=sha256:6c52e97d2d4a4c0c004a6190ccbf57ac3d61ad301ff52ba4c6c5a7b70697e347

Observation 210dde5e-a747-4990-bdaa-1517e3848351 · inbound

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards cites this paper.

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.399865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T03:13:17.539820Z digest=sha256:fe3752405b50e417ab093d9a541869586dc5d36b69c2108167bc41884bd8dcb2

Observation 3271b7ac-b713-4635-852e-de953a3b9536 · inbound

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards cites this paper.

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:50.901886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T01:45:31.691494Z digest=sha256:0ff009a093a5cca6da96ed0d37f7f77be1b278370d0351bad2380473e434b10a

Observation 25088d7d-5bec-46f7-bc20-57dbb390f5d3 · inbound

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards cites this paper.

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:42.017929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T17:20:07.747499Z digest=sha256:e5b773bf652d2097d6800ffb6aaaefcdbcca70dd538fff1ba4074c55c6f625b3

Observation ec1bfbce-90ef-4b2a-b5d2-a13ec96ed22e · inbound

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning cites this paper.

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:17.815362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-09T19:56:27.170349Z digest=sha256:40bf9bfb4d11806d0543f6b2bf218b7f048d4cd0cf6c72e7b7b5f7018647d38e

Observation 7f592c01-f6c7-43d2-a313-fd7547883156 · inbound

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning cites this paper.

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:55:59.108341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T00:58:25.685484Z digest=sha256:d75a02832844f229f65e322cd98fc84990334423c3eaa728d2c6acb2c70c9183

Observation 150cb53b-4bd5-449b-950c-265fbaa860c0 · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.599054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:14520c469f50e03a79487d850ed3411b457519ec90ed90cc2d91d4cbe45947e3

Observation 0c6d15d9-6bb8-41dc-a7cf-8dc0bfac18cc · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:08.634344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:60127e068bc4b2ca8011a81eb6915f8333826c3041aae17ef8ceaee7e045fe0a

Observation 8c63fbe8-01a5-46d0-b1b2-d810a5660313 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:05:55.145588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:906af3f2ffa793edae266c0805ca616691b8dbc44ae9d605317e1ea0c94b4cf1

Observation fc4b1324-6e69-4bb9-b185-37d0cbaf4b7b · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:17:28.298280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:2764196f4c2259a128df94f34eb666871f6ad5eb3d228b0298b198f59395b455

Observation 5002871f-5011-490c-8051-83c5be015592 · inbound

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents cites this paper.

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.236958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T01:10:01.008358Z digest=sha256:03773a9c5a5aabeb84fe35768684262d388411ae4b3933f5cde3f32bbd2782a0

Observation 70cf5383-fd29-4172-9093-e0c0ee7f2ffb · inbound

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment cites this paper.

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:55:51.066392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T01:54:49.131461Z digest=sha256:7b9707b94e18cf8b8ffc76e2841012a255f090a08411472bebe6506175e52964

Observation f4304d3e-0096-4a0d-9b56-6c8036d8642a · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.478495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:26a37cdaf2bacbb67c4d9f797d9f32c8e679fae7823ccec82df10fa9db4a51fc

Observation 9f7caa68-52f5-452f-99dc-cf86f58209e2 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.157089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.157089Z digest=sha256:6da84c742e152951b9e21f8c6f88ce2f54ad9190a9c14334acd8170e6fb733d8

Observation 8d01e1ee-1956-4538-94b6-9d60d3da26f7 · inbound

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy cites this paper.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:48:28.541838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T01:46:24.724553Z digest=sha256:0d2804e008fd48c2ffd44b9450f896e3c2fb090f504d2b1436a0904fdbb348af

Observation 225e3096-042b-4d89-9c8e-8faa4995535c · inbound

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code cites this paper.

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:08:05.103512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T05:06:46.360174Z digest=sha256:1ec0ea046cd708fbc2029a75973e552d8c2c5f265896ab4cbd7f404667f84f94

Observation fe5798ea-0ed8-40ee-bd67-1950eebb2d83 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.569721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:7e66b35e23e40fffd910d37b7020aecfdaaec8ef3c49c60068670cace033bbc2

Observation 32850909-3bfb-461d-8a2e-bb147a226f0c · inbound

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving cites this paper.

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:01.558000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-29T22:44:39.024780Z digest=sha256:ff96cbd57294e026d7ae1f7df5cd247c989eadc24f6f0d5082ec8e9d287a4cef

Observation 2e23d26e-160b-4cee-8734-d8aaf6893079 · inbound

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training cites this paper.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.914071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:98e89042567725e460e2fff035fe4363dcd9f5d4b9b156e9abebd0a38273e5a7

Observation eabaf675-ae68-45a3-9704-648abfd971f8 · inbound

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning cites this paper.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.163648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:add18fa8e8597b76ce782b5adb8c532df24a58129f65bfe92e5816b5ff4764be

Observation f2f634f9-c348-4780-b013-cf731755b35d · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.938852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:c520875ca33853dc4288202f35bb646fa739dc281babcf5db404fb510199f2b0

Observation 32262e4f-cb6c-4ad6-8590-788e32e41381 · inbound

Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework cites this paper.

Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.488885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T18:21:57.096578Z digest=sha256:b3059bade2480f059b3026a188c1070a75910d7ebba7f2f17cb5e27ce1de58ed

Observation ca03c586-95e5-467b-a080-52bc8ff08e3b · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:58.371419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:2d1e199f94c2a3963da0c9745738a0b785dafacb6204e4b40e5e67315b504478

Observation 5972265b-c31e-475e-bb50-4ecf7eb7d38e · inbound

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning cites this paper.

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.778295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T08:44:23.085858Z digest=sha256:bc49c245f4560073b342cce4e94b5e5954523f3113b0aad3182a509af09de4f9

Observation 151b103e-0bf4-4ebc-a0a6-affd7a86fcfa · inbound

Diagnosing Task Insensitivity in Language Agents cites this paper.

Diagnosing Task Insensitivity in Language Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:51.584015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T04:58:11.929896Z digest=sha256:55125b989f259fbfa84403a2b20a1cea3c7e7b0603062236a9b752058a22a78a

Observation b17426f0-44c3-46f9-a691-930410cb760e · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.045805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:d22df44c95765aab1f9d2aaa0992488a6a632d0667492722be448421d94ee13c

Observation c5c2b98d-1af9-4e9c-b0eb-147293a4d3f2 · inbound

Rank-Then-Act: Reward-Free Control from Frame-Order Progress cites this paper.

Rank-Then-Act: Reward-Free Control from Frame-Order Progress Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:45.215284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T17:42:44.309030Z digest=sha256:06d3a77c4648d9f3a04dd97340b63411f70efe1aa474378005c0e5167b54cc3d

Observation 07643e0c-ac7f-4345-8fa3-6f137fc0a8db · inbound

WorldSample: Closed-loop Real-robot RL with World Modelling cites this paper.

WorldSample: Closed-loop Real-robot RL with World Modelling Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.459374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T10:57:40.128651Z digest=sha256:8928e5bc321979f642ce9b6685350c9c81b9fee8dcb6e1de8052bca87e5dd985

Observation 1060e7b5-7a1f-44ad-ae9b-224b1b3a6971 · inbound

Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making cites this paper.

Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T05:22:49.815699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:22:49.815699Z digest=sha256:ebc3965f1c5819426dd594a6ae337f98e899c4ee52167e93767e19304db29e45

Observation b644b619-26d8-48ea-91df-0d76deedb835 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a84ff4b18e54cd3b2c94485f333e6b3473de0cbf129ab1ade9ec4a83bfe401d3

Observation 5fa8e51b-6044-4349-84d4-f8d597be9f87 · inbound

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL cites this paper.

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:29:10.250529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:29:10.250529Z digest=sha256:0e69b10cc3486b1c6086a73112794519f38c9098e82d5d906c01abab7e3ef642

Observation 74360757-7012-4b6a-9243-770addcc35bc · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.629933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.629933Z digest=sha256:627ade84928dbb3459ce242df7c49624a26596284d4d65859578c050a81fc81a

Observation 9adefe45-cc4b-402a-a50d-2d780189b19a · inbound

Agentic Reinforcement Learning with Self-Distilled Reward Shaping cites this paper.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.724043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.724043Z digest=sha256:f06745a91325bc3e6c66169aa9dbab78e42850a779d54a6a5918ab7a88918270

Observation 52bf7d1e-7a78-4cbc-9950-ef6e8145ea06 · inbound

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks cites this paper.

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:49.695690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:49.695690Z digest=sha256:dea3467d36e2761ce660f8f730e9365335e3a19f4e99aa6074cda828a1e26fdc

Observation f02e847f-6de2-4d70-9e96-b18755c4a78b · inbound

ADIAS: Automated Design of Interactive Agentic Systems cites this paper.

ADIAS: Automated Design of Interactive Agentic Systems Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:27.019408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:27.019408Z digest=sha256:43504bf743f8588cdeecea9da86c1fdfd1758fa4db6abb7b5611ead993213be2

Observation a1b4e244-6b66-4eee-8d81-a7f6395fbe2e · inbound

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents cites this paper.

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T19:29:14.787945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:29:14.787945Z digest=sha256:98c9cec6d23e53e3c13be834ae9e09def2249fd0025d80de40e22a1a79934a4a

Observation 604af2c4-39c3-4a80-8fb9-b61d552f1bdd · inbound

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training cites this paper.

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:00:52.751101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:00:52.751101Z digest=sha256:9939e8853dbed09aeb5eda375eb3810ec2783e12394152857df364cdffe28245

Observation e12b3ee5-81bf-4fa0-b109-8e6a4395e104 · inbound

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents cites this paper.

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:51.026135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:50:51.026135Z digest=sha256:1dce4e96ae5fcd51d2043b58f763b20c533263638908ae1030c6c02368a6cdb6

Observation d90394ae-8bf4-421a-81a6-97acf69a6280 · inbound

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning cites this paper.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.881986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.881986Z digest=sha256:a4a8c45ee5bfce796230ba92115193c745edb24f4481b100a5b14d64c74986fa

Observation 7903429c-8806-47f5-980f-186ef96cb48c · inbound

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents cites this paper.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.186850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.186850Z digest=sha256:886b667c2cb8a1852e7475d798762c15e913856762c75753c7d2d7cd7d638ac7