Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:55:38.734605Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 6 inbound Pith citation observations for arXiv:2509.06283.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:55:38.734605Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:48:42.025933Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-05T15:11:10.856386Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 556dd4b0-eeac-47b2-b240-88876f1a0a2b · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents gpt-oss-120b & gpt-oss-20b Model Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2304a2d2-ab78-47ad-9972-0abd7ef86a24 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c4b1357-78e4-46a9-9a29-58a20e0c6eaa · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Open Deep Search: Democratizing Search with Open-source Reasoning Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2ca88e-e8ab-4202-9231-a75aaafd9225 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents A survey on rag with llms.Procedia computer science, 246:3781–3790, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdec6b1e-b6f1-49c5-82ff-c03aa979db81 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Mindsearch: Mimicking human minds elicits deep ai searcher.arXiv preprint arXiv:2407.20183, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cdfc913-c152-4621-9cd7-ffd49e7a17a2 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Benchmarking Deep Search over Heterogeneous Enterprise Data
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb1f8fa2-b117-40a7-9948-593371d358ff · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Agentic Reinforced Policy Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c532bec-d396-4af8-bca5-184e04e40f98 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aecedd8-ebcb-427e-81ae-6d6372bca32f · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Try deep research and our new experimental model in gemini, your ai assistant
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9fcbe86e-28a9-4ce7-ad97-b1ade57395c9 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc4a0e3-83bf-42e1-995e-9ea455e02077 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Search-Time Data Contamination
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f0b15b-15eb-460c-bf54-2dae793b1fdc · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dacb85f-9cdc-4029-913b-85458931071d · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Characterizing Deep Research: A Benchmark and Formal Definition
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4610cf-c874-4f1b-bdf1-02dddb23c70d · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11cd8f17-d9da-4a14-93b0-28399a1e1dfe · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7dc80c9-2381-4982-a1c3-f58a47d0d2be · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Open deep research github
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b0957b8b-f8e2-456b-8c82-272e135b53da · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents WebSailor: Navigating Super-human Reasoning for Web Agent
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdc812e-5f22-4df7-b1b8-81b57d114aa0 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae1fb89-5328-4d81-bb50-0f1726832d97 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ec85fa-eb31-4f0e-a535-a8228a55399a · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44d5100-5f37-478c-8a4d-474f38a0c806 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents ToRL: Scaling Tool-Integrated RL
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18f1d8da-4204-4d18-a5a1-75c1911ed184 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Openmanus: An open-source framework for building general ai agents, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201cd6c1-3ec3-4890-b5e3-f4830eae074a · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96554e1-13b7-4d1d-a4c2-1b88cbf1d2e1 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Gaia: a benchmark for general ai assistants
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08a0fcf2-eee7-420a-a8e9-f4518f11d5e9 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Miromind open deep research
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d193feae-a9c3-4020-b05e-9089673fefd0 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Kimi-researcher: End-to-end rl training for emerging agentic capabilities
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8c76193a-9380-4d52-96be-cc508fc38800 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents SFR-RAG: Towards Contextually Faithful LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a6288bf-87fe-4713-b27b-e4e87c40e12d · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Gpt-5 system card
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ab292b9c-057c-4cfb-93d9-485bc0db0d26 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Openai o3 and o4-mini system card
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b7c8aca-cdcf-424d-a383-096e779aeb92 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Deep research system card
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 39c9a6dd-caea-4804-8c19-dbb3e0fca7c2 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Training powerful llm agents with end-to-end reinforcement learning, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6672358b-388b-4097-809c-69c1f972f419 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Introducing perplexity deep research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2ce42c27-37fa-4e47-b93d-f7bf0d47e88d · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Humanity's Last Exam
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d27b3b-770e-44f8-901f-9e23d6022008 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 573eb4d8-919a-42b3-b9a7-f4a316fbadf8 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Infogent: An Agent-Based Framework for Web Information Aggregation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cbc5e7d0-1a9f-4078-8ac9-0e73c1fea282 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation accfd666-2cbb-45af-b544-ae9267e41775 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d799053a-5fc0-4899-81d5-7c5721de4bfb · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a742768-3e2c-4bb6-847c-f3905eb78c54 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ad2de3-9f12-4099-a269-a768eac06bed · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Kimi K2: Open Agentic Intelligence
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aeb0479-c9c6-410d-a767-867b9125d94d · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67ed1eb-141d-4ab0-9e51-34aca1d61fd8 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.arXiv preprint arXiv:2505.16421, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff5359e-33df-468a-828f-c5c2c23a5954 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8(3):229–256, 1992
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa69e0b8-42bc-40d0-9a86-17b090bdf7d7 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d373ae-bc12-4996-aec3-9f18b32aa864 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Simpletir: End-to- end reinforcement learning for multi-turn tool-integrated reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2cce8f09-d5e8-4da4-991b-772f91bade9b · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Qwen2.5 Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 920c0cf4-e4c3-4c43-aad7-994578339158 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Qwen3 Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 068c808d-451e-43f2-bf59-8bc47270f5b2 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Cohen, Ruslan Salakhutdinov, and Christopher D
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5810b7ed-efcf-4f3d-893e-9a391375bf96 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6bd401b-f6eb-4c37-9660-28cbd0042548 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0faf9576-456a-43bb-9b57-b8aeb3a084db · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa87bd3-b355-42fd-9290-6d99c0869844 · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Group Sequence Policy Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab21033a-882b-4c3c-8ad4-d2e676b3465b · outbound
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 43fac9e8-f59f-4309-a598-e273cd7bff31 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
Reference 289
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d7e92b7d-3697-4d93-9b3e-5474dbeedd6e · inbound
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64481b9-d429-4109-a892-fac6de8f8057 · inbound
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b01e0d67-d5da-4a6a-8b09-5631a4e29744 · inbound
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2989f097-56e3-408e-8468-03fa761b66fa · inbound
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a8d41468-a08f-4be4-98fc-b5866c6e64c1 · inbound
Divide and Cooperate: Role-Decomposed Multi-Agent LLM Training with Cross-Agent Learning Signals SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.