Schema normalization repairs schema-drift failures in an agentic RAG benchmark (0.000 to 0.913) but does not recover stale, missing, denied, or wrong-session evidence, so reliability should be evaluated layer by layer.
Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Performance evaluation of Retrieval-Augmented Generation (RAG) systems within enterprise environments is governed by multi-dimensional and composite factors extending far beyond simple final accuracy checks. These factors include reasoning complexity, retrieval difficulty, the diverse structure of documents, and stringent requirements for operational explainability. Existing academic benchmarks fail to systematically diagnose these interlocking challenges, resulting in a critical gap where models achieving high performance scores fail to meet the expected reliability in practical deployment. To bridge this discrepancy, this research proposes a multi-dimensional diagnostic framework by defining a four-axis difficulty taxonomy and integrating it into an enterprise RAG benchmark to diagnose potential system weaknesses.
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
Schema normalization repairs schema-drift failures in an agentic RAG benchmark (0.000 to 0.913) but does not recover stale, missing, denied, or wrong-session evidence, so reliability should be evaluated layer by layer.