Pith. sign in

Paper Citation Record · LEDGER

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2508.09303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09303 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:12:12.793725Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:00:11.440671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T22:32:44.029593Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f254660-1248-4545-8767-f6d6303d4f1e · outbound

This paper cites GPT-4 Technical Report.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.149968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.149968Z digest=sha256:6381f0491b33b315777634fff2062c64d5a816f6dd3ca050f328c23055b50d14

Observation 61f23d39-ef57-499b-bc83-f6391a3f148f · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:17.666111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.265720Z digest=sha256:e9a27b500baa512d2e5a16863273e1d879aacd308de5f62f8aff5760843e62dd

Observation 60debba3-24aa-4666-9f1f-d21fd3254f03 · outbound

This paper cites A review of factors influencing user satisfaction in information retrieval.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning A review of factors influencing user satisfaction in information retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:17.330163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.330610Z digest=sha256:c6540d2022027afb90d0f09d79a74e25c029a16f1512b882397f1c3f4a7bf5fd

Observation 9cf214fc-bd03-4f87-b810-6759d0682650 · outbound

This paper cites The Llama 3 Herd of Models.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.409056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.409056Z digest=sha256:1d42d58d8a7560e987ccdf619803ad2acbc255d676746f313f759016f49f28f1

Observation 926ad8e5-0928-40fc-98fd-ff6f831f0d3b · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.477639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.477639Z digest=sha256:eab00e6d03aaa7bb03714b0e080c2f96d721e47e8bfd73cad6734927029cbfec

Observation 2e51088c-9156-452a-8866-ac64a9521b25 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.565175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.565175Z digest=sha256:9ddbecb1a21fd30ba02688a2354d04d60a7e7da28393c04eb2f75e443fc3a06a

Observation 6ab708f3-dc0f-46a0-8c7d-86458547d365 · outbound

This paper cites Retrieval augmented language model pre-training.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Retrieval augmented language model pre-training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:17.102872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.645837Z digest=sha256:8648d09f96cb0775cfac206e27e6d6e3298f9425d64cdbd1c8a545e3cfbeedf4

Observation 26e37146-a3b6-4175-ba2c-26680d9e7982 · outbound

This paper cites Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:16.861437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.710959Z digest=sha256:b70fd3c785a595bbf912ecab7b0f6078be4eea16414e04b39aa362dd7d8865be

Observation c7e6ef23-f227-443f-aa1e-de78ad7f9904 · outbound

This paper cites ORPO: monolithic preference optimization without reference model.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning ORPO: monolithic preference optimization without reference model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:16.593997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.741836Z digest=sha256:a23d839a53ab4c4cc36e2678bf6eb82a67a1f033fb5a2051fd3b7ea8347f8110

Observation 7a300d47-82ef-406e-89b7-5e75e66634e9 · outbound

This paper cites Survey of hallucination in natural language generation.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Survey of hallucination in natural language generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:16.354547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.773407Z digest=sha256:a805e7d5203173dd0c70c0c5131e54ef4ac3b4b796b6c1c5b378887eebfdeaa2

Observation afe7e04b-abae-4a2b-96fd-ebc3a778285c · outbound

This paper cites an unresolved cited work.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:12:16.098728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.826331Z digest=sha256:933318daf7c8c94c09561442ba11506a0c1a8436d9ecf8fa53879fd996dfdf25

Observation d6e171a5-39d2-41d1-aa9d-37d2404192fc · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:10.890698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:10.890698Z digest=sha256:018103dbc6109b1f2323ca220b74b9f241deb5ab24dafbde0b51f8275676d183

Observation 2aa8d116-1426-4df6-b0c9-70faad697aa7 · outbound

This paper cites Weld, and Luke Zettlemoyer.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Weld, and Luke Zettlemoyer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.848643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:10.958980Z digest=sha256:75adc4854781e90b61ce95d42fc6bf31bb5f4ec046e708ee815fc9449fbf247c

Observation 38f07e05-e737-479d-8b23-f5f8e54beeb0 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Dense passage retrieval for open-domain question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.722055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.008903Z digest=sha256:13e338ad646e99ac80cee6de727b37e719d1bd2584e3e00fcca88c76543f0280

Observation 85cf4b4f-2a8f-42ac-ac19-09a0614378ef · outbound

This paper cites A survey of reinforcement learning from human feedback.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning A survey of reinforcement learning from human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.064527Z digest=sha256:412a97876dd2233a97283acbaf375aa3f73444d2ed5f13afe9349d32e49d8615

Observation 499862f2-3f98-46d6-9e1f-15841e55680a · outbound

This paper cites Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.594794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.160381Z digest=sha256:7344d6955318c82c079d58558936026381687fff113d0da5b7f4a5d268e43e94

Observation 3733b631-9114-4e2c-bd79-634a9bc85621 · outbound

This paper cites Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.477053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.216923Z digest=sha256:65339b5bfe51b3442980f643d5626999e8377977d5ee8bd724f53bf005d0c289

Observation 76aff35a-0537-44c0-b9bd-34f2948e9609 · outbound

This paper cites Large language models in finance: A survey.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Large language models in finance: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.303679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.274785Z digest=sha256:db9f72fad63af519c4f92a88e14d05f2ff33de1b80c54151b673eec248121727

Observation 40799862-6eab-41dd-9159-cb04c7eeef93 · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:15.155608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.328199Z digest=sha256:6b3aaa498ae86f264e72d7738963e3170a9479fe6ecdb9f9e407e4f3c52e1d11

Observation db71740e-3dcd-4e2e-9a94-e4126188259d · outbound

This paper cites O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.380262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.380262Z digest=sha256:68cfb92230b11d624991d6970ea37baa163c6c190cdc1ab19914ee18b662051a

Observation dd37ca0f-46cc-4136-87e3-e5f3dc0f416d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.987611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.469102Z digest=sha256:2cee75515154ca9316a526ff34d4d843de9e01a32a9a758915bc14f291d1f004

Observation d2c768b6-3469-4726-a7d1-178707bc66ec · outbound

This paper cites an unresolved cited work.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:12:14.854532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.531717Z digest=sha256:5933597b89ffac3b2018edbe4b2af0fe7c9f4e15ffbf9e137e32a19a9af58afc

Observation c510312e-8ad1-4642-9804-b8d81dcab088 · outbound

This paper cites Iterative reasoning preference optimization.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Iterative reasoning preference optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.620173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.614989Z digest=sha256:530a2bccadc86be5cc2ba667f1b62d86dfff80fa0acf7baf5e7aa45b353c2ce8

Observation 88087109-df13-4f68-9096-5b191e5b10e7 · outbound

This paper cites Smith, and Mike Lewis.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Smith, and Mike Lewis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.444756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.660534Z digest=sha256:77a192c5d90c7c1ad595983e6346d6fd9aec39757e507b3e7695e3d04c3a2dc4

Observation 0f20ea70-41b3-4e1d-ad52-84df737196e9 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.294887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.710877Z digest=sha256:044f28f37efc8b067d5322f523bb7d034fab2be335a1cb87e0ab693d61e1b118

Observation 9fe1ab6b-48d0-449b-8029-bcf0917ca917 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Toolformer: Language models can teach themselves to use tools

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:14.082981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.771234Z digest=sha256:c37fda7f4b5e474b3dba14fdcfe8341e05fc600e4d85ad50cfff4c67f803fccd

Observation 81d2ab38-8b65-413e-8d7a-fe17069685e5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.811800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.811800Z digest=sha256:fc790f49fa1dc9dc401b7f359870741aefbea38a7038cf1474d8f56e18300698

Observation 2adda67e-b61d-4bfa-a97b-07329c450e82 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.863783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.863783Z digest=sha256:88c55161bb7ad2773e52429371ad7aea0787f81897fd26bfeaf6dae130dd02eb

Observation 64186f76-6f6c-4130-a97c-02440d495a05 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.902466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.902466Z digest=sha256:299a68d59f1c7d1ec30127a5753188af3dc96b1b39a50d3d3066809021fbb85b

Observation f3fca3a1-431d-40fe-a749-5676a5df1ae8 · outbound

This paper cites Sutton and Andrew G.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Sutton and Andrew G

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.926922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:11.938138Z digest=sha256:3ce1a583b8e1f52dd5e5e42b74fcdf99d30bed5b429e8334f1ee4a6f809b09af

Observation 6d2d5a50-a4e4-4e37-bc26-1ccd91cc38aa · outbound

This paper cites Multihop- RAG : Benchmarking retrieval-augmented generation for multi-hop queries.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Multihop- RAG : Benchmarking retrieval-augmented generation for multi-hop queries

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.800825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.022501Z digest=sha256:d066d990f45b7826fcaa68efa64cd843fe6164e81bb4824bd21b551387979b0f

Observation 0b75f438-f758-478e-99dc-a32f4f18f262 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.111745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.111745Z digest=sha256:f2423dbfcc6df2cd196fe7e387c1e65ba40a2480c34c4a41a838a39e75083946

Observation 2a37d315-32e8-4778-82d9-7287715f7736 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Musique: Multihop questions via single-hop question composition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.171420Z digest=sha256:7d1f914dc8733e2c4ef6684f0763e43ab0ba49a34074b9a123420e55a99c0175

Observation c3f928e5-172d-4496-8e0e-f6a3ea73ec76 · outbound

This paper cites Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.509777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.220676Z digest=sha256:be99ef688c1b14c60f7e3ba20479568eef7cca2de3c450dc59d48503e674c89b

Observation d89a6b42-e2b5-4355-8559-36edbc6133c3 · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.308929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.308929Z digest=sha256:a4acb93814df45e847993eeefc6d852c14c2eb9869483385fc45d96ec3284216

Observation 8cdbb16b-0648-4682-840d-c6757cefeb33 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.389009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.389009Z digest=sha256:6e440a4c411a3277c06f70be4e8b8382b16f93bde29b6f4b293095615219dec4

Observation 82c3f128-cd6b-4ca1-b174-dfc382795218 · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.451731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.451731Z digest=sha256:3366ce5bd706510f1bcaf1bdbc3f5bd7e24d4ca960d00fdce9eddae2946b5a02

Observation 84dd0aec-2dac-43c6-b372-30f4b55a0925 · outbound

This paper cites Reasoning or memorization? unreliable results of reinforcement learning due to data contamination.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Reasoning or memorization? unreliable results of reinforcement learning due to data contamination

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.500175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.500175Z digest=sha256:e535d6eaea5b0e75a20a8343063493810d8005a751d8cfaf1292c36114aca483

Observation dc6a072b-c9c2-460c-8d00-ee5ef2d0a98b · outbound

This paper cites MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.567646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.567646Z digest=sha256:ca7bd1f4887c7dc2512ce636df4d14573b3234790dd38ab3001925ce20dc9999

Observation 8b525432-1de0-4413-ac36-c9212bd25ad7 · outbound

This paper cites Qwen2.5 Technical Report.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.658091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.658091Z digest=sha256:f443ae46d258cc6776ee027f8f5f1e430e0e07ab6ef6dc64a8b25b521c543095

Observation 64153c56-4968-433a-b4d7-a68af4fd5931 · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.389477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.715033Z digest=sha256:2ec4f9b51c8ba7684d659a03afbdb3cf4b9896b8914307e408e06011fa250705

Observation 549c28d9-061a-40bb-a086-9bd47816549a · outbound

This paper cites Narasimhan, and Yuan Cao.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Narasimhan, and Yuan Cao

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:12:13.239715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T21:12:12.760593Z digest=sha256:2a0b3649a5694404ecf3a7d0801866c83ef0e1afdf0c50c3c1fa06c49f1bde19

Observation a035ada7-ce11-4bbb-8836-820e9b7b309f · outbound

This paper cites R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.793725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.793725Z digest=sha256:40e0ce1bbfc0cbc1bf15c8f5ead0cc73df43362e94aafba256c20920073e23cb

Pith citing papers

Observation 66d5ae84-a72a-457b-a78b-f5679185acc6 · inbound

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning cites this paper.

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:11.440671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:00:11.440671Z digest=sha256:dde85f63167c2a22bc3a24d73a9a588fb2368975cee8cdb52cf939826e41c788

Observation 74f3c35a-f61f-4efb-a4ee-568e19fcdf03 · inbound

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs cites this paper.

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:18.137585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:06:20.058342Z digest=sha256:2843c0d05f1981158b57cf420eb33f5e960055de11a93c1fecdc5bb69fce8b12

Observation 5d40a41f-8c7f-4ca1-bec9-e8442c3461bc · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.974701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:e633eeefa9ae8a4315c50a017cdad6613354185a3b1c5d0015c605817c54cb84

Observation e0200f79-ff96-474c-947e-e6565d15f8f8 · inbound

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG cites this paper.

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:12.210309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T10:27:00.257353Z digest=sha256:b3d2190772c7431944959ead20a189ffa7601aa1ed1bf08ab3735be08162c27f

Observation 4ea68c04-6587-46e0-964f-045f83cba423 · inbound

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents cites this paper.

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.310629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:28:36.266167Z digest=sha256:eb465e2bb0b89d79ab9fb06a09e1e1a64f1b17e39a7b6b6607d66e1f666a08cf

Observation b5213207-8b55-42ad-bdcb-e74528680afa · inbound

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents cites this paper.

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:19.032968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:18:01.006274Z digest=sha256:0b1a09103d195925e9b2f80ea2504d092f7c13f41321cc2f87ba66c57d6ab1b4

Observation c1b3fcab-7c07-4700-a910-7f540d833870 · inbound

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward cites this paper.

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.031157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T22:30:00.735630Z digest=sha256:b42e1220d16af0767fe894dd06cb3f29990c1423dc00b04918f7f861accdd625