Pith. sign in

REVIEW 2 cited by

RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13998 v2 pith:VTEON2KI submitted 2024-07-19 cs.CL cs.AI

RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering

classification cs.CL cs.AI
keywords answersrag-qaarenalfrqalong-formshortansweringaugmented
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Question answering based on retrieval augmented generation (RAG-QA) is an important research topic in NLP and has a wide range of real-world applications. However, most existing datasets for this task are either constructed using a single source corpus or consist of short extractive answers, which fall short of evaluating large language model (LLM) based RAG-QA systems on cross-domain generalization. To address these limitations, we create Long-form RobustQA (LFRQA), a new dataset comprising human-written long-form answers that integrate short extractive answers from multiple documents into a single, coherent narrative, covering 26K queries and large corpora across seven different domains. We further propose RAG-QA Arena by directly comparing model-generated answers against LFRQA's answers using LLMs as evaluators. We show via extensive experiments that RAG-QA Arena and human judgments on answer quality are highly correlated. Moreover, only 41.3% of the most competitive LLM's answers are preferred to LFRQA's answers, demonstrating RAG-QA Arena as a challenging evaluation platform for future research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis

    cs.AI 2026-07 conditional novelty 6.0

    DOSA's tree-guided self-regressive chunking improves multi-page document hierarchy parsing by up to 4 F1 and 19 TEDS over prior methods on DocHieNet.

  2. AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

    cs.CL 2026-07 conditional novelty 4.0

    RAG with top-3 chunk retrieval lifts smaller LLMs on Indian legal QA (Llama2-70B: 45.7% to 51.7% on AIBE) but often hurts large models, and under the study's own rating protocol some AI answers outscored the reference...