Pith. sign in

REVIEW 4 cited by

Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.11854 v1 pith:FQOGW3TO submitted 2024-12-16 cs.AR

classification cs.AR
keywords llmsaccuracymodelretrainingsystemstrade-offscapabilitiesgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid increase in the number of parameters in large language models (LLMs) has significantly increased the cost involved in fine-tuning and retraining LLMs, a necessity for keeping models up to date and improving accuracy. Retrieval-Augmented Generation (RAG) offers a promising approach to improving the capabilities and accuracy of LLMs without the necessity of retraining. Although RAG eliminates the need for continuous retraining to update model data, it incurs a trade-off in the form of slower model inference times. Resultingly, the use of RAG in enhancing the accuracy and capabilities of LLMs often involves diverse performance implications and trade-offs based on its design. In an effort to begin tackling and mitigating the performance penalties associated with RAG from a systems perspective, this paper introduces a detailed taxonomy and characterization of the different elements within the RAG ecosystem for LLMs that explore trade-offs within latency, throughput, and memory. Our study reveals underlying inefficiencies in RAG for systems deployment, that can result in TTFT latencies that are twice as long and unoptimized datastores that consume terabytes of storage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Retrieval augmentation improves mental-health chatbot intent classification for 4 of 6 tested LLMs, mainly by catching more high-risk cases, at the cost of more false alarms.

  2. NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A rule-guided self-generated fine-tuning method reduces verbal confidence miscalibration (ECE) in RAG question-answering by roughly 0.1 absolute across four open-weight models.

  3. Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search

    cs.AR 2025-05 conditional novelty 6.0 of 10

    Cosmos offloads full ANNS execution into compute-capable CXL memory devices with rank-level parallel distance calculation and adjacency-aware cluster placement, improving simulated QPS by up to 6.72x.

  4. BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning

    cs.LG 2025-07 reject novelty 2.0 of 10

    A purported impossibility theorem for LLM services reduces to the paper's own assumption that reasoning and authenticity necessarily consume extra inference budget.

Pith tools