Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:12:12.793725Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2508.09303.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:12:12.793725Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:00:11.440671Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T22:32:44.029593Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7f254660-1248-4545-8767-f6d6303d4f1e · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f23d39-ef57-499b-bc83-f6391a3f148f · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60debba3-24aa-4666-9f1f-d21fd3254f03 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning A review of factors influencing user satisfaction in information retrieval
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cf214fc-bd03-4f87-b810-6759d0682650 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926ad8e5-0928-40fc-98fd-ff6f831f0d3b · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e51088c-9156-452a-8866-ac64a9521b25 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab708f3-dc0f-46a0-8c7d-86458547d365 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Retrieval augmented language model pre-training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 26e37146-a3b6-4175-ba2c-26680d9e7982 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7e6ef23-f227-443f-aa1e-de78ad7f9904 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning ORPO: monolithic preference optimization without reference model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7a300d47-82ef-406e-89b7-5e75e66634e9 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Survey of hallucination in natural language generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation afe7e04b-abae-4a2b-96fd-ebc3a778285c · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6e171a5-39d2-41d1-aa9d-37d2404192fc · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa8d116-1426-4df6-b0c9-70faad697aa7 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Weld, and Luke Zettlemoyer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38f07e05-e737-479d-8b23-f5f8e54beeb0 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Dense passage retrieval for open-domain question answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85cf4b4f-2a8f-42ac-ac19-09a0614378ef · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning A survey of reinforcement learning from human feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 499862f2-3f98-46d6-9e1f-15841e55680a · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3733b631-9114-4e2c-bd79-634a9bc85621 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76aff35a-0537-44c0-b9bd-34f2948e9609 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Large language models in finance: A survey
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40799862-6eab-41dd-9159-cb04c7eeef93 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db71740e-3dcd-4e2e-9a94-e4126188259d · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd37ca0f-46cc-4136-87e3-e5f3dc0f416d · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2c768b6-3469-4726-a7d1-178707bc66ec · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c510312e-8ad1-4642-9804-b8d81dcab088 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Iterative reasoning preference optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88087109-df13-4f68-9096-5b191e5b10e7 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Smith, and Mike Lewis
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f20ea70-41b3-4e1d-ad52-84df737196e9 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fe1ab6b-48d0-449b-8029-bcf0917ca917 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Toolformer: Language models can teach themselves to use tools
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81d2ab38-8b65-413e-8d7a-fe17069685e5 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2adda67e-b61d-4bfa-a97b-07329c450e82 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64186f76-6f6c-4130-a97c-02440d495a05 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3fca3a1-431d-40fe-a749-5676a5df1ae8 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Sutton and Andrew G
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d2d5a50-a4e4-4e37-bc26-1ccd91cc38aa · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Multihop- RAG : Benchmarking retrieval-augmented generation for multi-hop queries
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b75f438-f758-478e-99dc-a32f4f18f262 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a37d315-32e8-4778-82d9-7287715f7736 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Musique: Multihop questions via single-hop question composition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3f928e5-172d-4496-8e0e-f6a3ea73ec76 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d89a6b42-e2b5-4355-8559-36edbc6133c3 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdbb16b-0648-4682-840d-c6757cefeb33 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c3f128-cd6b-4ca1-b174-dfc382795218 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84dd0aec-2dac-43c6-b372-30f4b55a0925 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Reasoning or memorization? unreliable results of reinforcement learning due to data contamination
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc6a072b-c9c2-460c-8d00-ee5ef2d0a98b · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b525432-1de0-4413-ac36-c9212bd25ad7 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Qwen2.5 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64153c56-4968-433a-b4d7-a68af4fd5931 · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Cohen, Ruslan Salakhutdinov, and Christopher D
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 549c28d9-061a-40bb-a086-9bd47816549a · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Narasimhan, and Yuan Cao
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a035ada7-ce11-4bbb-8836-820e9b7b309f · outbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d5ae84-a72a-457b-a78b-f5679185acc6 · inbound
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f3c35a-f61f-4efb-a4ee-568e19fcdf03 · inbound
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d40a41f-8c7f-4ca1-bec9-e8442c3461bc · inbound
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e0200f79-ff96-474c-947e-e6565d15f8f8 · inbound
LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ea68c04-6587-46e0-964f-045f83cba423 · inbound
HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5213207-8b55-42ad-bdcb-e74528680afa · inbound
HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1b3fcab-7c07-4700-a910-7f540d833870 · inbound
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.