Pith. sign in

Automatic Question-Answer Generation for Long-Tail Knowledge

1 Pith paper cite this work, alongside 1 external citations. Polarity classification is still indexing.

1 Pith paper citing it
1 external citations · Pith
abstract

Pretrained Large Language Models (LLMs) have gained significant attention for addressing open-domain Question Answering (QA). While they exhibit high accuracy in answering questions related to common knowledge, LLMs encounter difficulties in learning about uncommon long-tail knowledge (tail entities). Since manually constructing QA datasets demands substantial human resources, the types of existing QA datasets are limited, leaving us with a scarcity of datasets to study the performance of LLMs on tail entities. In this paper, we propose an automatic approach to generate specialized QA datasets for tail entities and present the associated research challenges. We conduct extensive experiments by employing pretrained LLMs on our newly generated long-tail QA datasets, comparing their performance with and without external resources including Wikipedia and Wikidata knowledge graphs.

fields

cs.CL 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality cs.CL · 2026-02-15 · conditional · none · ref 2024 · internal anchor

    Frontier LLMs complete 95–98% of tested Wikipedia facts when given their original source context, but directly answer only about two-thirds to three-quarters of plain questions about those same facts — recall, not storage, is the reported bottleneck.