Pith. sign in

REVIEW 4 cited by

Evaluation of Retrieval-Augmented Generation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.07437 v2 pith:WL7PBAC5 submitted 2024-05-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords benchmarksevaluationgenerationchallengescurrentmetricsretrievalretrieval-augmented
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Retrieval-Augmented Generation (RAG) has recently gained traction in natural language processing. Numerous studies and real-world applications are leveraging its ability to enhance generative models through external information retrieval. Evaluating these RAG systems, however, poses unique challenges due to their hybrid structure and reliance on dynamic knowledge sources. To better understand these challenges, we conduct A Unified Evaluation Process of RAG (Auepora) and aim to provide a comprehensive overview of the evaluation and benchmarks of RAG systems. Specifically, we examine and compare several quantifiable metrics of the Retrieval and Generation components, such as relevance, accuracy, and faithfulness, within the current RAG benchmarks, encompassing the possible output and ground truth pairs. We then analyze the various datasets and metrics, discuss the limitations of current benchmarks, and suggest potential directions to advance the field of RAG benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DocNavRAG: Document-Structured Graph RAG with Stateful Evidence Construction for Complex Document Question Answering

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A training-free agentic graph RAG system that navigates document hierarchies and cross-region links, improving answer quality by 7.8% and context sufficiency by 17.7% over the strongest baseline across four CDQA benchmarks.

  2. Quasiparticle interference in LiFeAs: Signature of inelastic tunneling through spin fluctuations

    cond-mat.supr-con 2025-08 unverdicted novelty 6.0 of 10

    Replica QPI features in LiFeAs are attributed to inelastic tunneling through spin fluctuations at 8 to 10 meV, matching neutron scattering.

  3. RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RAG-Zeval uses rule-guided RL with ranking rewards on synthetic responses to train a 7B model that evaluates RAG faithfulness and correctness competitively with 70B-class judges.

  4. Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A knowledge graph based method that selects salient entity relationship triples produced question-answer pairs that subject matter experts rated higher than a paragraph-based method.

Pith tools