Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:55.346259Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 17 inbound Pith citation observations for arXiv:2505.15117.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:55.346259Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:22:18.164811Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T00:55:12.126083Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a9c57f54-6a3f-411f-9cd5-c481acc54943 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b28661e-4667-436f-ad52-9e167fd09ce4 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbacb48-e9af-47ef-884b-eb32fcb592d5 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Self-rag: Learn- ing to retrieve, generate, and critique through self-reflection
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c45f28c-547c-4ea1-a66b-c9457ded3b34 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f6f5d45-770b-499d-a63b-ccc844135ad7 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents ELI5: Long Form Question Answering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b898ac7-96ce-4443-a0e5-0a274b00fb5d · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56345dbc-a3ed-4a1c-bf4e-68a59f2c5470 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9547765-daaf-4d6a-b000-24d7e114b491 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585497cc-a9c9-4f4b-81bb-fa58ff551145 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7630dd5-73ac-4352-be50-f252bdc34117 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Training Compute-Optimal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6534b44-fd37-4592-977e-7376e1c91ff2 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Long-context llms meet rag: Overcoming challenges for long inputs in rag
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 717a137d-4795-4d92-aa45-2f3bd6b12ce5 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents LLM Alignment as Retriever Optimization: An Information Retrieval Perspective
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b058ef94-0010-4640-97d2-ae7ecdca5340 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3821500-bdf1-4fca-a0d4-f43e0ba4cc4c · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Reinforcement learning: A survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 19e1688c-cfa7-418e-a7fc-348c0c3c9b8c · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Scaling Laws for Neural Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d37d64-fccf-4d4d-aba6-85a748414895 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Dense passage retrieval for open-domain question answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca007945-7e15-4249-a24c-414673a7b03d · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents A survey of reinforcement learning from human feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c94b57a-382e-430c-a0c6-0d0c5fc7a1a3 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Natural questions: a benchmark for question answering research
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f0a0b5-8016-4d30-81f7-4fe71c83a2e4 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f416758-8579-4885-887e-dad10bb3f9d6 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents RewardBench: Evaluating Reward Models for Language Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f145531-f0a2-41ec-880f-42eba51ab072 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Large language models in finance: A survey
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c74cf3b1-65ef-40d7-b028-718a802ccbb9 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Rec-r1: Bridging generative large language mod- els and user-centric recommendation systems via reinforcement learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f92054-4f7e-4285-b9cf-38cd12aa6c34 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Ra-dit: Retrieval-augmented dual instruction tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fca8fc0a-f564-42ca-9bd9-479b52f91105 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341d1ffc-de10-4f4f-8fa8-20f9a82917f2 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Fin-r1: A large language model for financial reasoning through reinforcement learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f61c753-65e0-45a0-ab56-5b0dee9f94a6 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Chatqa: Surpassing gpt-4 on conversational qa and rag
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e7884401-3cf6-45d4-ad39-d78d65af419e · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 595bedd6-b420-4b9a-827b-64e5da074ae9 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 195f0ec1-8ef3-4d27-9893-765da6383166 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents A study of generative large language model for medical research and healthcare
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65ef205c-a837-428d-8b96-ef6feb42b081 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Measuring and Narrowing the Compositionality Gap in Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d464510-268c-4b36-b755-855e0b957174 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Direct preference optimization: Your language model is secretly a reward model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06269762-9b6c-417f-aa33-40161e185fee · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Gpqa: A graduate-level google-proof q&a benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 900386d5-fcba-4598-9226-f38f09ab511a · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents The probabilistic relevance framework: Bm25 and beyond
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40aa8960-f10e-434e-bb78-93857c0be9d5 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Toolformer: Language models can teach themselves to use tools
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b9a6faf-50e9-4218-a32b-4af6ddb29e7d · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9023e9e1-0bed-4620-812b-fabe429d2868 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a6294c-588d-4f46-8ec0-7470b7666aad · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c40fc6e-ebcd-4414-b549-c8fb80e980cf · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents ASQA: Factoid Questions Meet Long-Form Answers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd1d6d0-871e-4444-896f-883123a0a9ee · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4333ad5d-3a13-4dc5-b02c-a338220a6faf · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9903980d-f83a-470c-b809-9297b5fa6ff9 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f818db25-cce8-4a26-b6db-dc0742739ac5 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13cd128c-5533-4061-9103-9dd198580236 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1454a45e-770d-4911-bd2a-429324bdc637 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68b8865-606f-4ebd-82f2-24226fc8b8a3 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Chain-of-thought prompting elicits reasoning in large language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29f35dad-8602-4e0c-91fc-4b05780a2a6c · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Measuring short-form factuality in large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f124a3f5-834f-435a-9c0e-b74acc1ed4fd · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Simple statistical gradient-following algorithms for connectionist reinforce- ment learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bfd1f691-7fa0-4e1b-8cd1-5f61af21abd6 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be230af-cc11-4b04-9d36-a1791b7f2d74 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92849fa8-cbaf-4c54-a339-5130b46e8ab6 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Qwen2.5 Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2d6c4b-4939-4028-9f89-0c0439cf2f94 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6b9cce-e7d1-4489-ae11-6655f4e0edc3 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents React: Synergizing reasoning and acting in language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 33aab401-0b34-4dc2-9521-4f7dd023735e · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa196788-e30a-4150-803f-794700ddbea6 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a4234d-b062-4944-b633-d5c847f3be3c · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3747c8b-8768-4e14-af6c-d6ce749ab8c3 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Benchmarking large language models for news summarization
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a765309a-41f1-4b5f-a97f-40bee11c104a · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents A Survey of Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46073e24-a2eb-43b6-9e9a-90fc8e3369fd · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb5f631-5897-445b-b8eb-09fd812f0456 · outbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Delicatessen
Reference 1953
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 38b539da-c53c-43a2-8365-fcfcebcd6817 · inbound
A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87aa6c1-0808-4f41-9a4d-e7235a8a94d3 · inbound
MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11891e05-7f2a-46ad-a05b-74e367e31854 · inbound
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e06a672-a89a-4390-a2df-320dc727fe33 · inbound
SSRL: Self-Search Reinforcement Learning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1aa506-a3b4-4efe-be0a-52109793cfd1 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 244
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0fcab628-5b5b-4eea-af6c-c1a9c9c9a03d · inbound
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6118e055-e4d6-4a0d-a75a-e3427ebcc2c3 · inbound
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785a30ee-2b24-48a5-af96-a455692b9848 · inbound
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85761933-03a2-4669-961f-021236f6c4fc · inbound
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f189d6d9-658d-4f1c-98a6-3b1e506dc583 · inbound
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb12523-4345-4d84-a68f-d2958bf61d59 · inbound
$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 600bc524-c3a1-494f-99ec-69a625a52d40 · inbound
$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8193e3d-2d96-468a-b0f0-524b0649de2b · inbound
Verbal-R3: Verbal Reranker as the Missing Bridge between Retrieval and Reasoning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e03c2656-8aee-4852-8bbf-20f7a3d434a2 · inbound
LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8fcae104-3086-451d-888d-490e03e36875 · inbound
DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe7c2143-8f14-4272-87fe-6dcc857aabba · inbound
TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f450e9f-b910-4ae1-be33-eae85bac52ea · inbound
TCPO: Turn-Level Credit Policy Optimization An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.