Pith. sign in

Paper Citation Record · LEDGER

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

As of 13 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2412.17032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17032 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:55:14.902112Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:19:14.992316Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c3fa33c9-9fff-4ece-937b-ff73f8c67a48 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.648021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.648021Z digest=sha256:400105fb1fb286f2a8d3db6414fa302c2d44148e5743d1b08ce8a5b2c8146860

Observation 1ccf4b38-807d-498b-ae17-62afec84d6b8 · outbound

This paper cites Learning to Recover Reasoning Chains for Multi-Hop Question Answering via Cooperative Games.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Learning to Recover Reasoning Chains for Multi-Hop Question Answering via Cooperative Games

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:55:15.509178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.654283Z digest=sha256:17bc15666be81e0010e9deccea446512979c4c5df8a0242c4be014d4fbdeeaed

Observation 15774753-2e12-4a10-9856-cf60599fbd4a · outbound

This paper cites The Llama 3 Herd of Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.659377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.659377Z digest=sha256:1c08c2611b88f9d6eb925d5ac32eaf75fec1d1676f763ca4515558c4aa8b8240

Observation 04c5b463-0714-4264-a42a-5631638fdc3a · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.664341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.664341Z digest=sha256:6973b57ee7977286dd494585d0caea4f682b0fc0452866d11dbd872f8c8a9652

Observation 500c559b-11b4-4c74-8f0e-2d9a29db9ef0 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.669252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.669252Z digest=sha256:864819f441326c996c530395ef85434199b2f4cf341c19a0225b93e175ea823a

Observation 06437c8e-b0eb-469b-93d5-ccc896513f73 · outbound

This paper cites Qwen2.5-Coder Technical Report.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Qwen2.5-Coder Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.674445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.674445Z digest=sha256:9ec01ba21b267344e4b41dd1a2cdceca0d2aebfc59a15bea1a0bce3799ef7ab3

Observation 3b02736f-ad64-4f45-9448-9412ecd0ec97 · outbound

This paper cites Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.679583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.679583Z digest=sha256:a4d3bdf690d59bed1e2d87e9babb21516285b1263ddc44de44c5517d19dd74c3

Observation cd17f135-a0c9-4e03-b441-1586e96e3c1f · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.684533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.684533Z digest=sha256:3700173cc0a22688a3c5360ea81d31f91c5d307d3ae5bd3393c27c46f707242f

Observation 5c031ec2-67df-42d5-852b-7bb713bb8a1a · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.815764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.688361Z digest=sha256:7be9fb1592ce84edde7fb0ed2d2a45bdcaceb88ecbcd2fdf9130c19d999320f1

Observation fbad84af-98c5-481e-bdbb-9f151e4b04d1 · outbound

This paper cites Mistral 7B.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.692321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.692321Z digest=sha256:129eba34d87897a8df8fa3777f7efa6f9c455fcf56f5533750883ac8e46144bb

Observation 3a00bcdd-b499-4755-a9fd-2bb4ff41f467 · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.696605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.696605Z digest=sha256:a96ff9b654fe5bfb12d87919ee97740030be4cce84db240b32053cb2a483c78a

Observation 8851b154-f772-4be2-92ea-668b71633222 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.802024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.700990Z digest=sha256:b30a1e7c656eb03c80177ec282d2245dcfff2df416b712060cb6e70e150f7525

Observation 0955d653-7425-44d3-94e7-ab48c022ca2f · outbound

This paper cites Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.705536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.705536Z digest=sha256:f7b1e8f0530af2e2a61e9b274cf3d8e9379fdecaf288569144de3d106af9b5ae

Observation d43c95b4-fa1f-4398-8c3a-b04ee565d55b · outbound

This paper cites Gonzalez, Haotong Zhang, and Ion Stoica.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Gonzalez, Haotong Zhang, and Ion Stoica

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.709849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.709849Z digest=sha256:429bbb8b7ae0ea45078fcaeaa6bd978ab62ae68e11a409745e630d6052463cb7

Observation bb878223-5a60-4e33-b24d-03755f0d0a05 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.714204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.714204Z digest=sha256:6ec2581ab1155c721038ed71f1e9ab55186758d8d701fe3dd158ed2fbae3d3bf

Observation 02b048db-2b15-4816-8818-97996dcbb148 · outbound

This paper cites Large Language Models with Controllable Working Memory.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Large Language Models with Controllable Working Memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.719947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.719947Z digest=sha256:f8c084c951c3312f6a704eb93efd021993240db601a959c921c1a56ff850c0ab

Observation 55862471-7f16-4987-a417-96e3dfdfbf31 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.770632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.724903Z digest=sha256:444e312f581881a8591ca432c31fd626ea636210141ed725b2ea080c7f656bbb

Observation c8441089-dcbf-443e-ace2-e2863a413b03 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.757538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.732120Z digest=sha256:09fa1972decc9b6713c02b01eea87c1784e37ebd4cd909c9bbcdcb9626c13da6

Observation 6e0d918c-de0a-447f-9e7d-c2d6a31eef0f · outbound

This paper cites Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:55:15.381405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.735869Z digest=sha256:e5e2f0368df2658c65637a9d7774cbd89117dc6315167e4e906cc8d01b72dc84

Observation ef7d9c5f-36a5-43c4-8394-17cb944a296f · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.743955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.739708Z digest=sha256:4c2f7f0555054ca4a124d8bc01774967b2649c5de4f455393a1412127fc0f514

Observation 0fa58127-4eac-4b8f-ac01-1566fc47bb09 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.730181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.743987Z digest=sha256:18d3b920072b024e32dd49a7ad3b90206b0e2f876dfa8500dd12f1fbf1335256

Observation 2c7b9b02-5790-4afd-8d9b-fe30b38ab018 · outbound

This paper cites Large Dual Encoders Are Generalizable Retrievers.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Large Dual Encoders Are Generalizable Retrievers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.748838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.748838Z digest=sha256:42510d1ce31f4d364a481ac2796a3797214094c0be375c1bf6fd0f9c5e21b09c

Observation 3ec105f1-b80f-4935-bc89-fca35185e6d7 · outbound

This paper cites Guo, and Xueqi Cheng.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Guo, and Xueqi Cheng

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.716433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.754170Z digest=sha256:77525f219236560205555ca3821718bcea49c6b92f3f11e0f3994267b377855b

Observation a05b6d6b-a9dc-4058-8d9b-a2fb603eff54 · outbound

This paper cites Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.758418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.758418Z digest=sha256:4d80f578cbd253384de6a10741b9ca7c0802e91d8cc808efb2ebe43fd98d9dcb

Observation 7e030baa-9c5a-4ac3-a45e-0d9f5340b4e0 · outbound

This paper cites Robertson and Hugo Zaragoza.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Robertson and Hugo Zaragoza

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.762778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.762778Z digest=sha256:934470bf2356e6e84e845acc1cd8cc9809c5b9c0db3da6ece0956c3fc3144a85

Observation 1afe389f-61dd-4943-85fd-8433934924f2 · outbound

This paper cites Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.767197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.767197Z digest=sha256:92bc84486eae5a1e55f9299aa64dd4347ac195ead1333ac79f68f2855080649f

Observation 219b7b2a-d79e-4d49-a115-73da98a3c1d0 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.694290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.771553Z digest=sha256:4af1aeb7505ce2e564a98986c53af92f31017c3f1a73bf6e1dc126c4763e9585

Observation 968fa1e3-3d56-43a0-af76-4db2c8f9dc94 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.681754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.776241Z digest=sha256:640c7f073c7ec403220c2107f4f13065813d02be0c7efcf81bdb50c958d4f14d

Observation 5ad18645-bd1e-461f-8481-c266fa38f5a0 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.780421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.780421Z digest=sha256:aa2b393299019fd23ad37caa618c016fc7992b65d2baa55bfd993043773597dd

Observation 2632d840-75c1-46a7-80a1-555e5e21e3dd · outbound

This paper cites One Embedder, Any Task: Instruction-Finetuned Text Embeddings.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge One Embedder, Any Task: Instruction-Finetuned Text Embeddings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.784568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.784568Z digest=sha256:5c45b0ed15778931acd17bcd6535bcc02fea0eca55df2b9ed769ba3dc6883c8c

Observation 550c42f2-7140-49fe-bea9-de6241d3e40c · outbound

This paper cites Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.789259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.789259Z digest=sha256:c23061ab6f40935add093e1ecba9deb391dc6bbd24a839c9817b4687205cc92f

Observation 2c5cc6f1-7512-4a28-93d6-cc3f8b5542ca · outbound

This paper cites MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.793714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.793714Z digest=sha256:7507aa4045a0520c2e559e70a2c05be77b29b1f78fb0aa06fff4f6f759a9b447

Observation bd6bcbcd-51f6-4416-ba5d-e533e5e6260f · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Gemma 2: Improving Open Language Models at a Practical Size

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.798033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.798033Z digest=sha256:0faf3327a16b11e9a0a333e3fac9926bef1d8f3df63e8e163a8f6867cf9c1659

Observation f6b3bcba-d562-4e67-904f-6eae0c253f65 · outbound

This paper cites Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.668466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.802177Z digest=sha256:7e5068a03c01c20ac52ef9d484e55408c291b9e5acc158e596c5912a090f9503

Observation 8f51e6fb-cd0f-4d0a-8aa2-06f12269e604 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.805847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.805847Z digest=sha256:ee117c0aedd8a9e67f603dedb12f94ad83a7719d88f7aa816fab4c86bf1f9fbb

Observation 034e0e7b-aedc-43f0-9c26-b3fce9532a38 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.809422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.809422Z digest=sha256:c6b7eaf1d939c9e8b26e64d0b06ef5d3a970056133b2f5f8c68d017d5cbe70bb

Observation dac56400-d0a1-4af8-9c5c-c8eca28261bd · outbound

This paper cites Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.814055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.814055Z digest=sha256:f76ec185cbd1c3f8fb5a1dd09601d7b8d29e497fa2aa0db5771d86a974f3fb3b

Observation 06dc3a3d-0616-4137-9848-230da6870e7c · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.645885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.818452Z digest=sha256:92d0c4eb22c51c8f14f1044d3ee3eb437988a6e5b7c0015b2df3547cde608741

Observation c120a452-40e0-4e3c-b2a9-211633a44dbf · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.632440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.823098Z digest=sha256:acbf433e920d32a0b92fbd5fe7e2118d36a8a80fd0e26ec36e74956af603e55a

Observation d376a72f-5f67-4854-83dc-c55de4e41a2f · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.827330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.827330Z digest=sha256:7a637509a126069bd890598ee08e452bb104b0c12a54bc4b2945765e5241c852

Observation 96b9467a-0c67-42a6-bbeb-f1efb39208df · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.831416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.831416Z digest=sha256:96354226589170ace99898dffd82a9c131817741f9313d7b7cf164ec2f9c88e0

Observation b0cfbe42-b200-4ed5-a25b-255c5e4a5cf4 · outbound

This paper cites Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.836046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.836046Z digest=sha256:581869f2d736a716f8015d48085aa150cfd7afee9c3e7260bfa61348f7698c39

Observation b86cd3b5-149c-4dd1-aa52-7d208e5707f3 · outbound

This paper cites LM-Cocktail: Resilient Tuning of Language Models via Model Merging.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge LM-Cocktail: Resilient Tuning of Language Models via Model Merging

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.841141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.841141Z digest=sha256:97eaeb8bf0fe8a0b5ec4b78ff33ee825a3da9e5bd6b03c7cc24417ad3e2d46a0

Observation 2e64a726-be0e-4ea9-8590-bacd53e343a1 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-11T05:55:14.939744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.845683Z digest=sha256:8adebee73177e98deaba0d43dfb390f22c4b9cdadb756eb1ee283c853e21d96e

Observation 59b3c53c-3d12-4743-acb5-a38b61c4ffb7 · outbound

This paper cites Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.850001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.850001Z digest=sha256:38734d242abc459e4b05e152ebf4e471aff95b539706218c23ac2e1b7627b53a

Observation 79a03d29-3262-41a0-b30c-5d87dfe99e0d · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.609094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.854280Z digest=sha256:92c598a70cdd15d017951354564716928bb0104068824491dc84dfd7daeefbaa

Observation 69b744c9-ad5f-4f5b-ae1d-60f4bf7c28c2 · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.595699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.858421Z digest=sha256:e83a9a4c181e9be5c622e58b709d2a6a1feb1b4f66c9c53352ef12f40db6310b

Observation 07a26606-07bf-4f5e-a3d6-63e8dd4602e7 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.582715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.862741Z digest=sha256:05fd81c40fbf8c84f6b073496077f9d3dcff32fcce81a8a6959396d4a69f2a4d

Observation cb29f5ca-c30b-4377-8b3e-2a55bd819688 · outbound

This paper cites Yu, Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Yu, Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.569316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.867077Z digest=sha256:a2b2924d080cf11c82ddd1ace3d8ca2d43fadb1b41b1ef86e16ec9c832ce5e5e

Observation 0004aca3-7805-4f9e-966e-2ea550cf9e33 · outbound

This paper cites RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.871403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.871403Z digest=sha256:935708646c0bcac13f19f50af04d4ae41e5239350e2d0be25c1ba42bc1b2a909

Observation e7161233-33ae-4eeb-af2f-c2a0691163f1 · outbound

This paper cites Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.875132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.875132Z digest=sha256:534961b627778cb44f5a200dd51f440a5b19cbd13b967ef29229c3668bcd1e30

Observation 45b8792c-1047-4b93-9b26-99ed99da4b89 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.555225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.878899Z digest=sha256:581d0f8e52ee3492b9286f6e7c64227df3db2459670bb553703732275702277e

Observation 4954c2a4-3254-45a3-9c22-9d1d0b6196c0 · outbound

This paper cites Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.883477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.883477Z digest=sha256:9e9ea2ab83e37e02b6e18d8f4226f14f53c7ca0dbc6359daf58991d00fb51b49

Observation d3217fc6-19cf-4db7-b636-3b5574086b6c · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.888105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.888105Z digest=sha256:9d2afab0307b74596ca518ce36abb78000fb4ca5391a95d5e835cf4e2f439149

Observation c182206e-d2c1-45f7-a7a2-daa558bbfe16 · outbound

This paper cites EfficientRAG: Efficient Retriever for Multi-Hop Question Answering.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge EfficientRAG: Efficient Retriever for Multi-Hop Question Answering

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.892561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.892561Z digest=sha256:e59a7e0c8f9d3252a599ae62e5e077df5dc1bcfece9cc2ed9e4f744af7eb4d81

Observation 0b4d1837-57d9-4514-a222-7d94ae615d56 · outbound

This paper cites online" 'onlinestring :=.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge online" 'onlinestring :=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.897152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.897152Z digest=sha256:41e71a4f9ab9002245df001e429eb77e8f37d3a27c4db097c39d1d1f2fddfaea

Observation 4af56952-683f-4c5e-a41d-42f68a081186 · outbound

This paper cites write newline.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.902112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.902112Z digest=sha256:0f85de09adc25bb45443b6a571dd97b87f56f8a938f292935163933c0a959b13

Pith citing papers

Observation 0eee85ad-94c3-4cb4-99d2-2ce6b8a33120 · inbound

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation cites this paper.

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:14.992316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:14.992316Z digest=sha256:22c05dd8830e679648dc8e3e43d340dbad537a945e60c5c36c525e3cf6a8f5bc

Observation 24d508a4-30e9-4d4d-8781-b77d9e3ff792 · inbound

Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs cites this paper.

Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:38.663920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:38.663920Z digest=sha256:e65045df90cfe5b9320994303b0ff0d06b58b75100d4f8bb5c903d3aa7067522

Observation 9c9cb728-9e24-481d-9046-4501a99e1a5d · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:56.135207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:6ec2a97bc38ea5684b758eb456c17525a7859e2d0a08bfca2a16f189526b1bc4

Observation e47dcf84-bf68-4a0b-ac11-8d4ce0bb4789 · inbound

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents cites this paper.

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.385466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T10:01:45.332920Z digest=sha256:da96d8dea2001b7326795a7da2db5384c55eac9cc8f0f60eb3c41e58cf9893ab