Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:56:00.329027Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 51 inbound Pith citation observations for arXiv:2502.01600.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:56:00.329027Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:51.026135Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
50 of 50 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation be444522-3b84-4f00-a6d2-27129c6c827f · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9d7b82-4b1c-40d7-969d-86de32af08b9 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9357bc52-4d7e-4185-964c-58cfaf9368c7 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Thinking fast and slow with deep learning and tree search
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3c600c5b-3e54-43f7-96d3-3ae9a1749549 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122fa5cb-c3db-4f03-89ea-9df8ea0b834b · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Grounding large language models in interactive environments with online reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e48f99d2-9cf3-4d39-a050-74ea028dfc96 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents FireAct: Toward Language Agent Fine-tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a8fec7-3e07-4384-8931-71b1345616b9 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5796287-07ce-4448-a87b-b07710f8b6f5 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d841b603-7417-495e-86db-f2062e30d551 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents D., Oosterhuis, H., de Rijke, M., and Shukla, S
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5844cb03-48df-4753-95b5-fc08356f6552 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Teaching Large Language Models to Reason with Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e6733f-79b7-4d4f-8ce6-dd4c2300cce0 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 20058a19-7f85-4784-97f4-44a1377fe38d · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents P., Littman, M
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 70bc154c-7af5-47f7-8fe4-2d74ec3217ac · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents and Langford, J
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8060d98-83f4-4402-b46b-a9de15137134 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1852615-2ed0-4e33-b45a-294be7c124c1 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Language models can solve computer tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c7027d52-7992-4140-a450-a787d6c6bd78 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Buy 4 reinforce samples, get a baseline for free! In ICLR 2019 Workshops, 2019
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2c3648e-c89e-4fb3-8f93-2b4fcb7a5209 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents H., Gonzalez, J., Zhang, H., and Stoica, I
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9b6f215a-b079-49ea-ac85-70693b68404e · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f30d4d-343c-439f-8a42-8433b614696b · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents AgentInstruct: Toward Generative Teaching with Agentic Flows
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d447578b-626f-48f9-ae56-3057b45e943e · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents WebGPT: Browser-assisted question-answering with human feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c93a203-4b43-4420-b096-0ee13fc603c2 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents D., and Barzilay, R
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 52c2be5b-19a6-4376-83ea-4ef330ac51b6 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Introducing OpenAI o1, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ea2a00a5-9396-4569-91ee-7b72038ef238 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832d4fbe-3772-4836-ab60-ac71b6e5e84d · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f99d275-8b4b-4e73-9702-44873ce220c8 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents ToolLLM : Facilitating large language models to master 16000+ real-world APIs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ed9810bc-ea2f-4e26-a9c2-836040c6721a · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Toolformer: Language models can teach themselves to use tools
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 25e64932-c570-4820-adec-0a67c5139792 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents I., and Abbeel, P
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1a64858a-7356-4e9a-84af-064cb4b44654 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20db280-cd63-4995-9cba-3a91d9e32b3a · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da300191-1603-4e48-ae49-8c0440154f38 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Direct multi-turn preference optimization for language agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 264b9455-85ab-447e-9360-3ccafe36cfbe · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Reflexion: Language agents with verbal reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8fc187c8-522a-4390-81c4-0effd84dfc5c · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a14132ce-3559-4151-93b0-a07490148853 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad18a1a7-50cc-4f3f-8baa-34fb59e4d741 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents torchtune: PyTorch's finetuning library, April 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b3adcb7-ac58-4d3b-b959-4641a548575d · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents A pp W orld: A controllable world of apps and people for benchmarking interactive coding agents
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df69eb9d-2b8d-4aba-9585-39ff767a9699 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Executable code actions elicit better LLM agents
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 225abf2b-3099-434c-9383-93941f9946cc · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents DD-PPO : L earning near-perfect PointGoal navigators from 2.5 billion frames
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d620ba99-cd16-46b3-8e9a-23932052479a · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Cut Your Losses in Large-Vocabulary Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c700fd-8839-4879-b517-dacc14d99b3a · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7a73299b-08cb-421b-97a5-8b925761da1c · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Qwen2.5 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16832fe2-692d-44bf-9b16-297241d0c33e · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Intercode: standardizing and benchmarking interactive coding with execution feedback
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 03e338e7-1acd-47d8-a0c4-9cbd0eb20305 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Keep CALM and explore: Language models for action generation in text-based games
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fbf0f105-5b4d-4ae1-b854-8867a3877da5 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents WebShop : Towards scalable real-world web interaction with grounded language agents
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1f641c1d-725b-40a4-9d96-1d1ae04fc5c5 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents R., and Cao, Y
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 46657d16-3287-408d-9117-e9057026d4fe · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db68ef43-d623-4056-a443-933467be9fc4 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents ST ar: Bootstrapping reasoning with reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4826752-2c93-4639-8fb9-f9c367cf720b · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9745ed1c-470d-4890-bf8a-09c68c3c99f1 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents PyTorch FSDP : Experiences on scaling fully sharded data parallel
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9429a772-2349-4b75-822e-4b617f6fdd67 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents and Zanette, A
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5208e95e-bee9-4698-a6de-66cfa3ba12d6 · outbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents Fine-Tuning Language Models from Human Preferences
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bb1187-2562-46e3-9dc3-335f6fe88efa · inbound
A Survey of Scaling in Large Language Model Reasoning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8e668ac4-ae4d-4216-b8b9-cc1bdf5cac79 · inbound
Group-in-Group Policy Optimization for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0d667442-84a5-4408-9b81-d21b07ebe5a5 · inbound
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e8e64e95-ca1d-49b7-bed5-f699951ef985 · inbound
Kinetics: Rethinking Test-Time Scaling Laws Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b5bd41e-3c99-44ec-a815-beaad8d3931f · inbound
World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 925a5ce1-d2d7-4178-bfe5-2fb732e95e7c · inbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7c6cd3-d20e-4af2-b1dd-39bbcd235b9c · inbound
OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 512c6297-bb55-4f2d-bdfc-82ef2af826de · inbound
ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ccb1273e-df61-4d0b-9635-b572b7a75b4d · inbound
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7feaa6e3-7208-4040-9399-24d247f613d7 · inbound
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3725c6d8-76ba-44c5-95b1-bf18b55c5d5c · inbound
A Survey on LLM-based Conversational User Simulation Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e020dca7-2282-48be-b8af-941aec689ae8 · inbound
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 210dde5e-a747-4990-bdaa-1517e3848351 · inbound
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3271b7ac-b713-4635-852e-de953a3b9536 · inbound
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 25088d7d-5bec-46f7-bc20-57dbb390f5d3 · inbound
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ec1bfbce-90ef-4b2a-b5d2-a13ec96ed22e · inbound
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7f592c01-f6c7-43d2-a313-fd7547883156 · inbound
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 150cb53b-4bd5-449b-950c-265fbaa860c0 · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0c6d15d9-6bb8-41dc-a7cf-8dc0bfac18cc · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8c63fbe8-01a5-46d0-b1b2-d810a5660313 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fc4b1324-6e69-4bb9-b185-37d0cbaf4b7b · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5002871f-5011-490c-8051-83c5be015592 · inbound
Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 70cf5383-fd29-4172-9093-e0c0ee7f2ffb · inbound
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4304d3e-0096-4a0d-9b56-6c8036d8642a · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9f7caa68-52f5-452f-99dc-cf86f58209e2 · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d01e1ee-1956-4538-94b6-9d60d3da26f7 · inbound
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 225e3096-042b-4d89-9c8e-8faa4995535c · inbound
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe5798ea-0ed8-40ee-bd67-1950eebb2d83 · inbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 32850909-3bfb-461d-8a2e-bb147a226f0c · inbound
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2e23d26e-160b-4cee-8734-d8aaf6893079 · inbound
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eabaf675-ae68-45a3-9704-648abfd971f8 · inbound
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2f634f9-c348-4780-b013-cf731755b35d · inbound
When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 32262e4f-cb6c-4ad6-8590-788e32e41381 · inbound
Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ca03c586-95e5-467b-a080-52bc8ff08e3b · inbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5972265b-c31e-475e-bb50-4ecf7eb7d38e · inbound
Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 151b103e-0bf4-4ebc-a0a6-affd7a86fcfa · inbound
Diagnosing Task Insensitivity in Language Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b17426f0-44c3-46f9-a691-930410cb760e · inbound
SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c5c2b98d-1af9-4e9c-b0eb-147293a4d3f2 · inbound
Rank-Then-Act: Reward-Free Control from Frame-Order Progress Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07643e0c-ac7f-4345-8fa3-6f137fc0a8db · inbound
WorldSample: Closed-loop Real-robot RL with World Modelling Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1060e7b5-7a1f-44ad-ae9b-224b1b3a6971 · inbound
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b644b619-26d8-48ea-91df-0d76deedb835 · inbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa8e51b-6044-4349-84d4-f8d597be9f87 · inbound
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74360757-7012-4b6a-9243-770addcc35bc · inbound
TCPO: Turn-Level Credit Policy Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9adefe45-cc4b-402a-a50d-2d780189b19a · inbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bf7d1e-7a78-4cbc-9950-ef6e8145ea06 · inbound
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02e847f-6de2-4d70-9e96-b18755c4a78b · inbound
ADIAS: Automated Design of Interactive Agentic Systems Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b4e244-6b66-4eee-8d81-a7f6395fbe2e · inbound
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604af2c4-39c3-4a80-8fb9-b61d552f1bdd · inbound
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12b3ee5-81bf-4fa0-b109-8e6a4395e104 · inbound
CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d90394ae-8bf4-421a-81a6-97acf69a6280 · inbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7903429c-8806-47f5-980f-186ef96cb48c · inbound
Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.