Pith. sign in

REVIEW 1 cited by

Improving Scientific Document Retrieval with Concept Coverage-based Query Set Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11181 v1 pith:DHAV5EQG submitted 2025-02-16 cs.IR cs.AI

classification cs.IRcs.AI
keywords queriesquerygenerationccqgenconceptsdocumentconceptcoverage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In specialized fields like the scientific domain, constructing large-scale human-annotated datasets poses a significant challenge due to the need for domain expertise. Recent methods have employed large language models to generate synthetic queries, which serve as proxies for actual user queries. However, they lack control over the content generated, often resulting in incomplete coverage of academic concepts in documents. We introduce Concept Coverage-based Query set Generation (CCQGen) framework, designed to generate a set of queries with comprehensive coverage of the document's concepts. A key distinction of CCQGen is that it adaptively adjusts the generation process based on the previously generated queries. We identify concepts not sufficiently covered by previous queries, and leverage them as conditions for subsequent query generation. This approach guides each new query to complement the previous ones, aiding in a thorough understanding of the document. Extensive experiments demonstrate that CCQGen significantly enhances query quality and retrieval performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval

    cs.IR 2025-05 conditional novelty 5.0 of 10

    CoRank reranks scientific documents by first scoring 200 candidates from compact LLM-extracted features and then refining the top 20 with full text, improving average nDCG@10 from 50.6 to 55.5.

Pith tools