REVIEW 6 cited by
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Though current long-context large language models (LLMs) have demonstrated impressive capacities in answering user questions based on extensive text, the lack of citations in their responses makes user verification difficult, leading to concerns about their trustworthiness due to their potential hallucinations. In this work, we aim to enable long-context LLMs to generate responses with fine-grained sentence-level citations, improving their faithfulness and verifiability. We first introduce LongBench-Cite, an automated benchmark for assessing current LLMs' performance in Long-Context Question Answering with Citations (LQAC), revealing considerable room for improvement. To this end, we propose CoF (Coarse to Fine), a novel pipeline that utilizes off-the-shelf LLMs to automatically generate long-context QA instances with precise sentence-level citations, and leverage this pipeline to construct LongCite-45k, a large-scale SFT dataset for LQAC. Finally, we train LongCite-8B and LongCite-9B using the LongCite-45k dataset, successfully enabling their generation of accurate responses and fine-grained sentence-level citations in a single output. The evaluation results on LongBench-Cite show that our trained models achieve state-of-the-art citation quality, surpassing advanced proprietary models including GPT-4o.
Forward citations
Cited by 6 Pith papers
-
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
Ref-Long is a new long-context referencing benchmark on which all 13 tested LCLMs perform poorly, revealing a capability gap that simple retrieval benchmarks miss.
-
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
In multilingual retrieval-augmented generation, models cite English evidence more accurately than translated evidence, and this language preference can outweigh document relevance.
-
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
SelfCite uses context-ablation probability differences as a self-supervised reward to improve LLM sentence-level citations, raising LongBench-Cite citation F1 from 73.8 to 79.1.
-
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
MiniLongBench, a 237-sample compression of LongBench, is claimed to reproduce model rankings with a 0.97 Spearman correlation at 4.5% of the evaluation cost.
-
heiDS at ArchEHR-QA 2025: From Fixed-k to Query-dependent-k for Retrieval Augmented Generation
Query-dependent-k ranked list truncation performs comparably to, but not consistently better than, fixed-k retrieval in an attributed clinical RAG pipeline, with the best system still below the organizer baseline.
-
Position Paper: Metadata Enrichment Model: Integrating Neural Networks and Semantic Knowledge Graphs for Cultural Heritage Applications
A conceptual metadata enrichment framework integrating iterative vision analysis with LLM-driven decisions and RDF knowledge graphs is proposed, with a small annotated incunabula dataset released.
Discussion (0). Continue with ORCID to comment.