Pith. sign in

Paper Citation Record · LEDGER

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

As of 23 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 90 inbound Pith citation observations for arXiv:2011.01060.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2011.01060 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T07:39:58.366423Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 90 of 90 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:18:44.136454Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:17:50.841893Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b36a629a-3b0f-4c30-bbd9-ae4fa763a8e7 · outbound

This paper cites In Proceedings of the 1993 ACM SIGMOD International Conference on Management of Data, SIGMOD ’93, page 207–216, New York, NY , USA.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of the 1993 ACM SIGMOD International Conference on Management of Data, SIGMOD ’93, page 207–216, New York, NY , USA

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.446622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:4ef2f6a53d29a9825521ea6f99785fc3d9fd4c6a45cddd89486d6a7f02c11551

Observation 7432c0d9-f441-4f71-bf49-e932fe90b9f0 · outbound

This paper cites In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , pages 1533–1544, Seattle, Washington, USA, October.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , pages 1533–1544, Seattle, Washington, USA, October

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.453648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:77cbba36180e78242091443304f3cc1052fd868ff2f03a3a19d3b8b2b772544b

Observation 99b0b215-057c-451f-94e2-d0cca7acc3e1 · outbound

This paper cites Large-scale Simple Question Answering with Memory Networks.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps Large-scale Simple Question Answering with Memory Networks

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T07:39:58.399296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:bb707d7618e9b65ae952cbfb0938a018d17dd58fde89e5ce365f3115f4d58e74

Observation 9df89687-67e3-4778-8798-3decd40b9cc9 · outbound

This paper cites an unresolved cited work.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:39:58.457160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:51519273d59a1b8665a0befda2100d5bb47f5c2e71be2bdf4a735a26e04193d7

Observation 8b5b711c-46c3-41b4-946f-a7fd4c879c23 · outbound

This paper cites HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.388251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:5f2ca0d67ffc11bf122b36de28b442bd5e7db924e1ba7cc2b4dda31f87279c04

Observation 552106ad-af2d-481e-8349-27bae1bc1adb · outbound

This paper cites an unresolved cited work.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:39:58.410989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:e17368f04b0577265a832940b0deb4c0ed7ae8c6f741764bc8c71c04766ef44c

Observation 7a87f1b8-3a3e-4485-86e7-8d24c8a0541a · outbound

This paper cites In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers , pages 2956–2965, Osaka, Japan, December.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers , pages 2956–2965, Osaka, Japan, December

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.416128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:e13dd8cb798beb387ef83692f6024e2284a17e325fd3249e9c7e020448776e3d

Observation c076e7a4-fb88-4667-a4fe-3aeaeca42fdf · outbound

This paper cites In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2021–2031, Copenhagen, Denmark, September.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2021–2031, Copenhagen, Denmark, September

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.421682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:8ad0d4fb273820e7c874d7574dbbf71508375786d072bfeaf5528d51317d25eb

Observation 2329a960-8cb7-4f8a-8101-57f7deb6abee · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T07:39:58.406542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:dd8189d4cc508009182cbefff97a0685fdcc0dba5f8c5036d85d2437291765af

Observation 4e3d2b22-2778-40bc-961b-99321182a240 · outbound

This paper cites In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas, November.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas, November

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.426466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:43bac3a4a96397dee361ababfcf6ce634b1e392cb37e5320efcd55f0e355a9ce

Observation 45f0188f-60ca-4dbc-8da8-c4d64668043e · outbound

This paper cites In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 1088–1098, Cambridge, MA, October.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 1088–1098, Cambridge, MA, October

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.431050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:f03ba38ad07f92a27dc259fa757f9ae7f3b08b17f0cde0e9f2807ae857ebbac9

Observation 70991d22-2845-4549-9c7c-e4999d799c4c · outbound

This paper cites Association for Computational Linguistics.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps Association for Computational Linguistics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.434892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:9a43956e820ea1f69caae060c595b23c5260565e9fe3be28e4622a7aaaefdafb

Observation 4da38f82-70cc-4859-9f23-aa44603f3a31 · outbound

This paper cites an unresolved cited work.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:39:58.438862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:178fe522c2441713b2acebdfc9ccac18f3e4a6bb849243d5a74e3c0d509b4a22

Observation 6bcf5f79-439d-490a-9bf2-a5433923ad3c · outbound

This paper cites In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages 2369–2380, Brussels, Belgium, October-November.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages 2369–2380, Brussels, Belgium, October-November

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:39:58.442595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:cc8c9e813e3d9fc4d44494af5f85b10bfa5d452358d62db7e7a2762fefe3d199

Observation ad1860ed-25b2-4403-89d6-0d2c59d374c8 · outbound

This paper cites an unresolved cited work.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:39:58.450122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:e04f8d181f857fd9708fee07f8667fca241c2cfebbc33e5cffa2b173e673f0b7

Pith citing papers

Observation ae5eec23-4273-41eb-9d42-1e5ef63d67eb · inbound

Retrieval-Augmented Generation for Large Language Models: A Survey cites this paper.

Retrieval-Augmented Generation for Large Language Models: A Survey Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 119

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:13:57.263558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T05:10:25.171044Z digest=sha256:01fe1e18f6507d4c8da0ce8a65ec2891bd2bbdd91d1cd14471e099c77ab395f2

Observation b575b6ef-4675-4025-bea6-114e200c6a41 · inbound

Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction cites this paper.

Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:27.462350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:12:27.462350Z digest=sha256:4cd72ddc1c8ee5057b60741db38e636dd9bcb3b09198dadecd153f996349a3ca

Observation 9cd2e587-5741-423b-8238-216849f4d0be · inbound

Zero-Indexing Internet Search Augmented Generation for Large Language Models cites this paper.

Zero-Indexing Internet Search Augmented Generation for Large Language Models Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:12:54.206055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:12:54.206055Z digest=sha256:1503438d0454a65660c8689c3c0482cb16d546d1fbcf1992cf9b6a05be430cd9

Observation 85512ca1-b096-4e75-a4e7-4dcf9e2c5833 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.682753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.682753Z digest=sha256:1ec5c70e7f600bec7cbfdee646611d0ad0a3acbf6545e550d17becfb75c6ca65

Observation 7c2523c9-da3f-473d-b5f0-af0c55d4f7bb · inbound

Review-Then-Refine: A Dynamic Framework for Multi-Hop Question Answering with Temporal Adaptability cites this paper.

Review-Then-Refine: A Dynamic Framework for Multi-Hop Question Answering with Temporal Adaptability Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:41:21.222983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:41:21.222983Z digest=sha256:2304fa9b4a47473722a425276c75915a61ae90298b35ada559694c725d70a0b2

Observation 04c5b463-0714-4264-a42a-5631638fdc3a · inbound

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge cites this paper.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.664341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.664341Z digest=sha256:0f55d51366863c09b644343d919b2774c7022befc95dcbca69924c750ae7aa3f

Observation 68a679ee-d8cb-40d6-b0c2-64c22fdd4820 · inbound

PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation cites this paper.

PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:08.533733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:08.533733Z digest=sha256:ce32bb0c138a9138def2a792e34997d81e4473fb537b7c26852f5dd562590655

Observation e52b249a-740e-4685-9fd1-f4255a219cc3 · inbound

Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads cites this paper.

Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:41:05.112914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:41:05.112914Z digest=sha256:05cd4ce65cb8c9db687db64b3f5b6ee22173d1fb98a7ec939e29b47f06931589

Observation 90c8e3d9-7f28-4b28-ace2-fee368a1be7c · inbound

Parametric Retrieval Augmented Generation cites this paper.

Parametric Retrieval Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T13:55:03.963226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:55:03.963226Z digest=sha256:84518a9c60c298fe6a0d79251107b42e16b5fbbfcbe4a756f2a0de88cef713d4

Observation 819165ce-9185-4f63-90b9-ad869203cfcd · inbound

CryptoX : Compositional Reasoning Evaluation of Large Language Models cites this paper.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.724372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.724372Z digest=sha256:df23665e1e7ae86083fc08ff3b362260f6fcb0b9a75dcb7c11193b21f6f3d230

Observation b6874ff6-3ed3-4ca9-a573-0dea88361b3d · inbound

Synergizing RAG and Reasoning: A Systematic Review cites this paper.

Synergizing RAG and Reasoning: A Systematic Review Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:18:44.136454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:18:44.136454Z digest=sha256:bc27f7e95a0b128941f5a05d48e78544da25b6e10fd48970fd097c8052783fa9

Observation d83ef06d-d7e7-487b-a580-ca0e2ed29e3f · inbound

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family cites this paper.

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:25:14.848855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:25:14.848855Z digest=sha256:a81153d6806558d6aee08dd5eec0c4ffbf5deb6962e88faa162e331d0619941d

Observation dd0020fb-070e-4471-8847-de50cade18db · inbound

ZeroSearch: Incentivize the Search Capability of LLMs without Searching cites this paper.

ZeroSearch: Incentivize the Search Capability of LLMs without Searching Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T17:44:13.310155Z digest=sha256:4f73a19f04dd351e6d0f3b3f56026685754f9f4d78cce4f72efa940bf6904aac

Observation d28a5430-c1ac-4c5e-8573-a16d1fc5e446 · inbound

ZeroSearch: Incentivize the Search Capability of LLMs without Searching cites this paper.

ZeroSearch: Incentivize the Search Capability of LLMs without Searching Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:06:45.837584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T16:05:04.715678Z digest=sha256:4aa0b7d2597bec0c007c1f1fc21d07273146857bfb4deabdd71849e93f4046ea

Observation 3d86b65a-ddb7-49e1-992b-ec0509af1c5d · inbound

Group-in-Group Policy Optimization for LLM Agent Training cites this paper.

Group-in-Group Policy Optimization for LLM Agent Training Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T09:15:08.193357Z digest=sha256:b0fa09a8f373716e77e9f4ce53e8c5f88e02e07c8fc947c2d6c4879e9bef298e

Observation 9e8994a4-5f6c-4f8e-b65c-761b0d3809f3 · inbound

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective cites this paper.

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:18.991105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:28:18.991105Z digest=sha256:ded9be3d5bc76fd15b439c590b2611f7930cf3fc9c8b13e45b48484f7d2071b5

Observation 9a671e0c-9e19-4301-8e17-06234cd5b2c1 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.224356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.224356Z digest=sha256:64919b856b45daeb4d32919eefe444a68761eac29d1645457a58dea9a3879e2f

Observation e5d7070d-1e11-4d88-a629-25d2cce2a0fa · inbound

Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation cites this paper.

Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.825922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.825922Z digest=sha256:0818f754db1a842a13fe88d7a1fd4aaeb5727e02f15a62000d67c27ba36ae330

Observation b9d38361-37fb-49af-8833-aa69ee60a63a · inbound

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation cites this paper.

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:15.233813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:15.233813Z digest=sha256:8078f188aa8c0f9bfc8ad19a2d3e9485f7494b92d2385995fbb312134658002b

Observation 215e9b38-a1b0-4bc7-9897-9dd826d031a4 · inbound

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering cites this paper.

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:47.840206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:47.840206Z digest=sha256:1f66e8089488e09bc1cbdda12bf9c401bf8c7258f0429ae47df1eedb7a63765c

Observation 8fd37d16-3671-4b6d-ac6b-9315399543ae · inbound

GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis cites this paper.

GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:53.778386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:30:53.778386Z digest=sha256:a8dc1e4c31705f221273f371844872e51a77f0ff8fcaf8492acd1793bef622f2

Observation 80f108ca-f080-49c1-9fb5-deb496e9de7a · inbound

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation cites this paper.

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:52:19.986230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T13:50:30.090068Z digest=sha256:0e80c8a683a225ea80257bb5be04b684324aecb636d0c3088d73dcf0e32f9f4d

Observation 355a31a2-51be-473f-8173-2ca586c76456 · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.299421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:56.299421Z digest=sha256:553918dd95541616c60f52ef9da1df9bbaaf5d39eb153cb5f2ec91ae1a90f20e

Observation 34e9275a-2a5e-4599-a125-4aa3579187e3 · inbound

WebDancer: Towards Autonomous Information Seeking Agency cites this paper.

WebDancer: Towards Autonomous Information Seeking Agency Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:04.990968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:04.990968Z digest=sha256:65adcc20dfac4a5c4868a108f193069da2349b3a1f7bd43db57e5177f3f77e58

Observation b6918372-803c-41a9-981e-114b841b3793 · inbound

Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning cites this paper.

Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:24:16.677572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:24:16.677572Z digest=sha256:36e4a4b56854a22996bca0607917faf4b7d7fa5d23e9bb0fcde9a528cd1ab3c0

Observation 9a6544ff-3691-471b-98e4-9a9173de2af9 · inbound

ComposeRAG: A Modular and Composable RAG for Corpus-Grounded Multi-Hop Question Answering cites this paper.

ComposeRAG: A Modular and Composable RAG for Corpus-Grounded Multi-Hop Question Answering Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:50.772066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:50.772066Z digest=sha256:f6f0a0634cb233782e34efc743b60dc71b3641f2dd0d4c190d78e1b8d0ffbb08

Observation da03794f-a6c1-433d-a9b3-e7bf9a975659 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:52:16.266090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:63b0e7635a830f653f466cf0122953c611addddd11d271441b0fe2a8aba06234

Observation 16979356-ba25-461f-81cb-6f5e52cc047e · inbound

Reinforcement Fine-Tuning for Reasoning towards Multi-Step Multi-Source Search in Large Language Models cites this paper.

Reinforcement Fine-Tuning for Reasoning towards Multi-Step Multi-Source Search in Large Language Models Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:04.556573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:04.556573Z digest=sha256:a1475896a43ee5cb100df6573a1a057ccd3342042ddc23c33ec13c071709a5a8

Observation c57f8f76-13b3-4a30-8ae8-34c19beb42a6 · inbound

Maximally-Informative Retrieval for State Space Model Generation cites this paper.

Maximally-Informative Retrieval for State Space Model Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:45.256425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:45.256425Z digest=sha256:2823c010cd9ba35f0d61d099f01064395fe0fd5134c6c87fcc0196819cbe132f

Observation b61b70dd-269d-46bb-8ae5-58921b341781 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.655749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.655749Z digest=sha256:4d3348960c9a4f03d83eec8d1f0867910c1f2f9264101e3f1839d1660ce26cbb

Observation 741dd78a-e701-4532-96fa-7b09cd18633c · inbound

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning cites this paper.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.044949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.044949Z digest=sha256:7b652ac41c9b037e8ce18756d57992fcf952f6d236f93ff33eb75f2645599547

Observation 7efa4012-cfc2-4164-a482-b47b391bd626 · inbound

HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving cites this paper.

HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:31.529044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:31.529044Z digest=sha256:1094d213397037ceb98ad8578c43c043dc5d0f6458f076ed7e2e95f9b7788761

Observation 2535f612-6327-41e0-aad9-08445c68fd72 · inbound

RAVine: Reality-Aligned Evaluation for Agentic Search cites this paper.

RAVine: Reality-Aligned Evaluation for Agentic Search Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:29.947160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:29.947160Z digest=sha256:2b85b4344de3dec5fa8d63797fff459673719466c6267f23e70b8405bfe71a39

Observation 7aa075bc-2c36-413a-8e52-3d98c125f4a6 · inbound

From Ranking to Selection: A Simple but Efficient Dynamic Passage Selector for Retrieval Augmented Generation cites this paper.

From Ranking to Selection: A Simple but Efficient Dynamic Passage Selector for Retrieval Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:54.268174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:01:54.268174Z digest=sha256:157c9c8cf389d7abb6837e113d75442b239b10d4f41f148c5002c99b4fe3d287

Observation e79197a4-5345-4fe4-8716-3f031433b673 · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.215321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.215321Z digest=sha256:f3c7fed9086ea1b1839b9bda2e0a7bc8d0089a99c373235e08b600e0ba5ba61b

Observation 98002a85-c5cf-4f65-ab05-e125803d4274 · inbound

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward cites this paper.

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:21:48.746644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:21:48.746644Z digest=sha256:3c467acc69799b12efb17b02bc6127ba02245301516981d25d89595e8210d836

Observation b85af062-d1d0-499d-b050-0154e419afa7 · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:32.528248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:32.528248Z digest=sha256:4292d284cf3a5d70629cadbd8a9cf5125aeebc4af26b86766adfdd93cabc2d4a

Observation 8a2b7885-43c1-4ae6-8006-a6aaf04d1a96 · inbound

Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information cites this paper.

Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:23:27.961276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:23:27.961276Z digest=sha256:a536acf47c845dd3f57560aced55a62619a5a489a0f2eb18a3eba919dfb904ae

Observation 94d2e65e-c47d-4cd9-a324-766b641c7266 · inbound

Retrieval Feedback Memory Enhancement Large Model Retrieval Generation Method cites this paper.

Retrieval Feedback Memory Enhancement Large Model Retrieval Generation Method Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:50:55.639018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:50:55.639018Z digest=sha256:f056f0fe752b786e842323e3b21ec8336acbd9e1454afdb2a68c61e3d8128571

Observation cb4010de-50f8-43f7-be0b-e937c1cc50c6 · inbound

Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning cites this paper.

Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:57.366013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:29:57.366013Z digest=sha256:1f1a57b835cc4048dc135c2abd40dfc5b453a25197cc942889860571051b94f1

Observation df384e68-eb04-4108-971d-c394512442a0 · inbound

Open Data Synthesis For Deep Research cites this paper.

Open Data Synthesis For Deep Research Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:45:11.820398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:45:11.820398Z digest=sha256:e624fad8b105e500c7b12e1423a9a30cea0caf4301e0646d44404bde46d9ab92

Observation f2508cad-2f92-408d-98d6-4eee797ec5a1 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.070949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.070949Z digest=sha256:d540d4a84a6a58ba239b95427c161c063237eef420ef6b365b5be1d0bc387595

Observation 88d67d04-10e9-4b9e-91fb-5cbfd1af2c05 · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:28.604509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:28.604509Z digest=sha256:bb7e61bd7065da2cff5dde31e3f88be014adb1a96f85ccc2a85e3b5eafb4e529

Observation 59ce93f3-c8e8-4245-86b5-a06d6cbc2b90 · inbound

BRoverbs -- Measuring how much LLMs understand Portuguese proverbs cites this paper.

BRoverbs -- Measuring how much LLMs understand Portuguese proverbs Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:58:53.792689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:58:53.792689Z digest=sha256:be582f9b7b0dc653e8fffc766b5544b06a3015e3015789e8453b078cbe39579b

Observation 8a1376b0-98e0-453f-9a7b-6f45d307f65a · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.112469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.112469Z digest=sha256:9e73e8ce999f04eb856943a3dd242b7d94fb8b203e1fbc65c1d92785babaf68c

Observation 2f9bd869-8670-4d95-b032-6ec5bd7a3aa5 · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.365586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.365586Z digest=sha256:42e2bb8d629d192a4d4aa350d787b6719f0ca323fe03a6c9100dc637f765e5fc

Observation 43168c99-88ac-4736-9b1b-7433bd3e196b · inbound

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs cites this paper.

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:11:18.210123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T11:06:20.058342Z digest=sha256:2f33fad67ff9082358badcbdee7fa55ff0d9f2eb65b0255d7f09282b2237daf8

Observation b003c741-9e5f-46c9-88f4-8ef24f77c6b4 · inbound

Question-Adaptive Graph Learning for Multi-hop Retrieval Augmented Generation cites this paper.

Question-Adaptive Graph Learning for Multi-hop Retrieval Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T07:11:48.249136Z digest=sha256:9eab728a986cec6dd4d26f35037cec1609026b5a744dbca0c057287c6579825f

Observation a98880f0-e466-4bcf-984d-e1e9c9050839 · inbound

EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle cites this paper.

EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T06:19:44.360734Z digest=sha256:c077dcb163a5e3d758d1f7fb34db613a7b03332561c79aef4e289b70c2b07d2c

Observation 2cabc30a-8a71-4219-a5ee-ae8465f18a02 · inbound

EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle cites this paper.

EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:50:36.439332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:50:06.642976Z digest=sha256:4f4d6ad5916eaff0ae61ec8be3e65336ac478985d093e01833399c72d9a88161

Observation 9e879b36-c0bc-45f1-9f7b-7eccd5082b8e · inbound

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping cites this paper.

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T03:10:57.839146Z digest=sha256:310b8101a94e7a28431e17da64913d20c489cb36bdf23d561d78fd0b06a0135a

Observation 461e1e52-7266-457c-a79f-6bfbbb36bcac · inbound

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning cites this paper.

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T00:57:25.902674Z digest=sha256:d01d7e8abc204eb5bff03871db00e6a0eba4f73fc8ade9684f83fb4f057fe1b0

Observation 0e924b27-a3d9-4b1d-961a-a6bb325df8f7 · inbound

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling cites this paper.

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:44:18.744201Z digest=sha256:ec52c50a0147b194eeb1e7c91be2db1099c69805a2a1ebda1e19a5677717faa1

Observation 9be4fa95-cb86-4366-bc53-2f5f809e3e78 · inbound

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning cites this paper.

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T21:38:56.042227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:38:56.042227Z digest=sha256:f78e35341eea966f40db6879109b9ac586ecddda98f7f66b9d5ee7de21d24bac

Observation bee7e3bc-b35a-4479-97aa-965fc489d581 · inbound

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services cites this paper.

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:21.132323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:21.132323Z digest=sha256:4f33f4c960469799026c5f3f484b8909e4d3e010dd558669a663d526f1168349

Observation c98c48de-7d5b-41cb-a0df-86cef427f83a · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.105728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:38.105728Z digest=sha256:5d8dc6701879ac83aded23352e0c48be8c301066ae455f6c6bf2d77e8920e2f8

Observation f322f208-5048-4895-9dba-2754ff805845 · inbound

Leveraging Spreading Activation for Improved Document Retrieval in Knowledge-Graph-Based RAG Systems cites this paper.

Leveraging Spreading Activation for Improved Document Retrieval in Knowledge-Graph-Based RAG Systems Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T15:45:17.483759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:45:17.483759Z digest=sha256:3868f712aac1d4a1f318cb3cd09ba3db1314fd7a38cf7fde340c24521e51c86a

Observation 7e9f187f-9e0d-4882-8e7a-cff50d3c057c · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:0cac080d881f1adf8b534202dac4fa81dc75377455abcc2da8d5d78955fd2a99

Observation 44943469-5037-495a-9788-3dac570e3991 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:23.846039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:23.846039Z digest=sha256:54f9088273084493c1a2ca09c8848c32249066a7bb0836cce7c78cf0ef2594d6

Observation 40674f8f-a68f-4f4f-8132-81250b5f2863 · inbound

Adaptive Information Control for Search-Augmented LLM Reasoning cites this paper.

Adaptive Information Control for Search-Augmented LLM Reasoning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:39:27.537439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:39:27.537439Z digest=sha256:2fdb4640ae8fb38c6412af40efcdedf5d7f2622129489611bf998c7463f339e9

Observation ccecdd0d-125a-4631-a6ce-5cd4740c5371 · inbound

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent cites this paper.

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:38.547858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:46:38.547858Z digest=sha256:51b2a602d00788d93ea2f0f615c4aff36a62f30f900af387d56d5e402b439171

Observation 956d9de2-e2f6-47ee-bd88-f3450e643a6b · inbound

MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens cites this paper.

MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T16:00:03.578685Z digest=sha256:48337f18fc3bf9d9046366ebc51d740fb16a202debef8fd6d795a3c23cc33345

Observation 8fc202cb-85db-4836-96c9-7cbe4707b7b6 · inbound

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning cites this paper.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:f3a85ac4fb79ec2b2a17b7804f203fa821941cfea316f77394a45fdb0214e55c

Observation 2b25c39b-d3e0-4bbb-9739-9083f855146b · inbound

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search cites this paper.

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 8

Resolution
verified exact
orphan_title_repair, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T17:16:09.927267Z digest=sha256:b8602b6c1950fcd988ef7dc07860ea7f8d08e83b351ccf24afec27daeeea07f0

Observation 3eacbc76-875d-4df7-ac24-333baa2b351a · inbound

Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems cites this paper.

Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T22:35:18.954951Z digest=sha256:046ff7180f70befee0d23ce3318de2316446d358f306b47cfeed08dc54ec5e0f

Observation a446c81c-b306-4b3f-9e82-29d563975943 · inbound

Transforming External Knowledge into Triplets for Enhanced Retrieval in RAG of LLMs cites this paper.

Transforming External Knowledge into Triplets for Enhanced Retrieval in RAG of LLMs Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:42:59.063869Z digest=sha256:3e17d0cc66406f1fa845ba0d852d2ce7bd4399891c509df1c1d6ced32148060e

Observation 0c1ea5b9-cf19-424e-86ff-6ea1ba61678e · inbound

HeadRank: Decoding-Free Passage Reranking via Preference-Aligned Attention Heads cites this paper.

HeadRank: Decoding-Free Passage Reranking via Preference-Aligned Attention Heads Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T06:36:22.111810Z digest=sha256:5022e14210b0480160004210dab674763c588debbaafe7fb1d46a03bcf292c15

Observation b437cd33-2d0b-4b5e-b03a-e836aeba4b31 · inbound

MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search cites this paper.

MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T06:15:11.432788Z digest=sha256:746f30760bd9300ad2e7561490459126d0d9845b61006c4bc903ee49b49c7299

Observation 46c4525e-fb01-4b87-9bc9-ff19219d2ba7 · inbound

EHRAG: Bridging Semantic Gaps in Lightweight GraphRAG via Hybrid Hypergraph Construction and Retrieval cites this paper.

EHRAG: Bridging Semantic Gaps in Lightweight GraphRAG via Hybrid Hypergraph Construction and Retrieval Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T05:43:04.813867Z digest=sha256:a50e7cc119f8607236fe86cfc0dcc28efb5bc84101e39fcc72f1ca78f6a5f7e1

Observation 5a18daf1-b189-4fb4-9b8c-256aee05d83b · inbound

AtomicRAG: Atom-Entity Graphs for Retrieval-Augmented Generation cites this paper.

AtomicRAG: Atom-Entity Graphs for Retrieval-Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T03:43:36.163220Z digest=sha256:f0de9e9e87d04dee164bfef46a4be9d08902f1cbf6d4b38054f8b86982692612

Observation 71898f25-1322-4658-9064-4eab9dabec81 · inbound

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference cites this paper.

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:39:58.458269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:37:52.545943Z digest=sha256:c897dc762565d6d3f9c563cc4ef9efbc307161123a788578146624b6ba11596f

Observation 34be19f9-2c42-48e1-bc6d-cf50ad4c0d76 · inbound

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering cites this paper.

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:13:44.963412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T17:03:46.019948Z digest=sha256:218ae44372fa8f19ac3075de6163094215313f9584d79700fa42775677cec448

Observation 840b8620-2233-4ff9-b976-b6626e99a153 · inbound

MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation cites this paper.

MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:36:08.106342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T22:24:25.409345Z digest=sha256:5d30777f339c652d61d37f6011b5360a9ba17228578696c439d681af8e68d138

Observation 922a8a2a-aefe-4b32-a77a-6bb10bdbfe9a · inbound

MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation cites this paper.

MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:42:37.996504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T18:20:43.092561Z digest=sha256:d93c2a410380cfd472a1a16ff26f0009215cec932299b5896741ac3be59e64ce

Observation e3366a89-d66a-4fe9-8fb0-8e132299b4a6 · inbound

Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking cites this paper.

Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.837176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:38:21.242007Z digest=sha256:0111c715a5ac7564ce7ade272752db5667ca0a74d347517dc8a443c60a4eae07

Observation 76007fe2-3744-45f9-827d-38802c53fbe5 · inbound

Policy and World Modeling Co-Training for Language Agents cites this paper.

Policy and World Modeling Co-Training for Language Agents Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:06:17.178935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T15:40:35.531232Z digest=sha256:db02f9206870d5721a00793614a786be62f1cf5fde964c87a741327a3f39ccc4

Observation 155c6d3f-b9e7-43e3-a831-2cd7182b4323 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:57.465624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:9720683ae2842126fcc6471b3963d7ad87f1146e52fd962234e5ba7b6baa621a

Observation 42bded46-782b-4d79-9915-8adf1c24620f · inbound

Selection Integrity for LLM Graph Memory: An Accumulability Criterion for Information-Flow-Blind Retrieval cites this paper.

Selection Integrity for LLM Graph Memory: An Accumulability Criterion for Information-Flow-Blind Retrieval Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:58:06.265956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:13:58.485088Z digest=sha256:c741016a341df7d26b452f0ca826a1b78fcf164ba264dac52de917abb903cb6a

Observation 1fb04bc6-912e-4048-9266-453a5115a3d8 · inbound

Agents-K1: Towards Agent-native Knowledge Orchestration cites this paper.

Agents-K1: Towards Agent-native Knowledge Orchestration Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:33.948394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T06:32:47.387422Z digest=sha256:3647d20875ed1ed55ed662ff1418427876850c943ffc4eca29057fda8f34402d

Observation c6885ddd-dfcd-44ff-895b-b47c24e904d8 · inbound

Agents-K1: Towards Agent-native Knowledge Orchestration cites this paper.

Agents-K1: Towards Agent-native Knowledge Orchestration Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:24:38.430155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T11:16:50.072447Z digest=sha256:8eade69665911a48e5edc08529bab9661d4790c997470e1cfedf804815826c92

Observation f1300452-95d2-46e2-a638-26489191cced · inbound

Agents-K1: Towards Agent-native Knowledge Orchestration cites this paper.

Agents-K1: Towards Agent-native Knowledge Orchestration Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T11:40:44.791538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:40:44.791538Z digest=sha256:3697b8d9dd697ffb19de588ad5441187e994cb551232ffb0aadd0186dda7082c

Observation 0d3a9853-1de3-4856-88e9-53f07e45e6cb · inbound

FlowRAG: Synergizing Explicit Reasoning via Frequency-Aware Multi-Granularity Graph Flow cites this paper.

FlowRAG: Synergizing Explicit Reasoning via Frequency-Aware Multi-Granularity Graph Flow Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T21:28:58.887483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T00:32:36.540334Z digest=sha256:863f0daa05ac6697d8fdd4bda460b9bb071047a1e101b5059b57fa459c6ce59c

Observation 21d41ae4-e7c1-4b7a-90ad-7ded4e39e8af · inbound

R$^2$-Searcher: Calibrating Retrieval and Reasoning Boundaries for Agentic Search cites this paper.

R$^2$-Searcher: Calibrating Retrieval and Reasoning Boundaries for Agentic Search Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:34:05.227281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T00:33:59.778294Z digest=sha256:90a35f4a2c81c5f9c199346b3e9eb5d13e1052db2c63b646b1255eb9644c10cd

Observation ede3a1b8-1754-4d9b-aea1-d9301445b575 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:073ee478d45133d52f0c27c26644c82ce048c00deb84fb60c84a8f2729122992

Observation 472c4520-63b8-4094-9577-a4ce9154dbf7 · inbound

Retrieving a Set, Not Independent Passages: Set-Level Compatibility Learning for Efficient Set Exploration cites this paper.

Retrieving a Set, Not Independent Passages: Set-Level Compatibility Learning for Efficient Set Exploration Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:17:50.867183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-11T03:17:20.012125Z digest=sha256:2e057b76513d74927cf1e9040922c9dc0f6fb841276d7edc8f57c13d64c65e09

Observation 04829529-6ee2-487f-bc6f-849892b168f8 · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:0f87299bd5de14311f2299e280a6df11a54736ca411bbacd5c739af22de6329e

Observation 435da1a2-d03b-4139-a24d-6d6fb00a3ecf · inbound

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning cites this paper.

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:13.855193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:13.855193Z digest=sha256:f2a8a38247aa39de4d46d517329cb0d3cfde3ebffa6c419d993ddd628d52e8c0

Observation a74b787b-e60c-4ba7-88e0-0eee59648c1b · inbound

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training cites this paper.

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:49.539339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:49.539339Z digest=sha256:5290cb7295a86c7a9934d2da33b517855d0ad4420a61e1e9c4ab0f2f5388e979

Observation 0ab5ee6d-5006-4113-bf0c-2aafd94fb84c · inbound

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents cites this paper.

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T15:27:03.040495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:27:03.040495Z digest=sha256:4e78cbe81a8312ef6b213a68db99a098a6c8c7923e4e89a39b708d0f2419ac4e

Observation bf81d290-3523-4c20-b077-1c4ee61a10b0 · inbound

Interpreting Language Model Hidden States at Scale cites this paper.

Interpreting Language Model Hidden States at Scale Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:16.879794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:16:16.879794Z digest=sha256:21d697d1c6167a0a9ef803c0de99d769bb01d2e51c5cc7e89c5a26f10308cb3c