REVIEW 4 cited by
Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rapid increase in the number of parameters in large language models (LLMs) has significantly increased the cost involved in fine-tuning and retraining LLMs, a necessity for keeping models up to date and improving accuracy. Retrieval-Augmented Generation (RAG) offers a promising approach to improving the capabilities and accuracy of LLMs without the necessity of retraining. Although RAG eliminates the need for continuous retraining to update model data, it incurs a trade-off in the form of slower model inference times. Resultingly, the use of RAG in enhancing the accuracy and capabilities of LLMs often involves diverse performance implications and trade-offs based on its design. In an effort to begin tackling and mitigating the performance penalties associated with RAG from a systems perspective, this paper introduces a detailed taxonomy and characterization of the different elements within the RAG ecosystem for LLMs that explore trade-offs within latency, throughput, and memory. Our study reveals underlying inefficiencies in RAG for systems deployment, that can result in TTFT latencies that are twice as long and unoptimized datastores that consume terabytes of storage.
Forward citations
Cited by 4 Pith papers
-
Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture
Retrieval augmentation improves mental-health chatbot intent classification for 4 of 6 tested LLMs, mainly by catching more high-risk cases, at the cost of more false alarms.
-
NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems
A rule-guided self-generated fine-tuning method reduces verbal confidence miscalibration (ECE) in RAG question-answering by roughly 0.1 absolute across four open-weight models.
-
Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
Cosmos offloads full ANNS execution into compute-capable CXL memory devices with rank-level parallel distance calculation and adjacency-aware cluster placement, improving simulated QPS by up to 6.72x.
-
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
A purported impossibility theorem for LLM services reduces to the paper's own assumption that reasoning and authenticity necessarily consume extra inference budget.
Discussion (0). Continue with ORCID to comment.