Pith. sign in

REVIEW 4 major objections 5 minor 121 references

Accelerating Retrieval-Augmented Generation

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read IKS, a CXL near-memory accelerator, makes exact nearest-neighbor search 13–28x faster and cuts RAG inference time by up to 26x.

desk verdict RAG profiling is solid; IKS speedups are plausible but unvalidated and inconsistently reported. read the letter →

arxiv 2412.15246 v1 pith:75JZUD72 submitted 2024-12-14 cs.CL cs.AIcs.ARcs.DCcs.IR

classification cs.CLcs.AIcs.ARcs.DCcs.IR
keywords retrieval-augmentedgenerationexactnearestneighborsearchnear-memoryaccelerationCXLmemoryexpandervectordatabaseapproximatelargelanguagemodelsdenseretrieval
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the retrieval phase, not the generative model, is the main latency and quality bottleneck in retrieval-augmented generation (RAG), and that exact nearest-neighbor search (ENNS) is the right retrieval regime despite being memory-bandwidth-hungry. It shows that high-quality exact retrieval lets an application send fewer documents to the language model at the same accuracy, while approximate search needs many more documents and ends up slower end-to-end. To make exact search affordable at scale, the authors design the Intelligent Knowledge Store (IKS), a type-2 CXL memory expander with near-memory accelerators attached to LPDDR5X packages. They report that IKS performs exact top-K search over a 512GB vector database 13.4–27.9x faster than a Sapphire Rapids CPU, cutting end-to-end inference time by 1.7–26.3x in representative question-answering pipelines. If these numbers hold, exact retrieval becomes a practical, cost-effective choice for large RAG deployments rather than a luxury.

What carries the argument

The load-bearing object is the Intelligent Knowledge Store (IKS), a type-2 CXL device that attaches eight LPDDR5X packages, each with a Near-Memory Accelerator (NMA) chip containing 64 processing engines, each combining a 68-MAC dot-product unit with a hardware Top-K unit that keeps an ordered 32-score list. The mechanism that makes the offload cheap is a cache-coherent CPU–accelerator interface built on the CXL.cache protocol: the host writes an offload context into coherent context buffers, rings a doorbell, and blocks with umwait(), while NMAs poll the same doorbell, read the query vectors, and write back partial top-K lists without DMA setup, interrupts, or kernel involvement. Vector layout also carries the argument: embedding vectors are stored column-major in blocks of 68 so the MAC array reads one dimension from 68 vectors per cycle and saturates the 136 GB/s LPDDR5X bandwidth, while the CPU reduces partial lists only at the end.

What would settle it

Take a real CXL type-2 device with LPDDR5X NMAs (or a cycle-accurate FPGA emulation) and measure exact top-32 retrieval over a 512GB corpus: if the latency is not close to the simulated 470.6 ms, or if the speedup over the optimized Faiss CPU baseline falls below 13.4x, the central performance claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the retrieval phase, not generation, should dominate the design of RAG systems, and that exact nearest-neighbor search can be made both fast and cheap with the right near-memory hardware. It shows that ENNS with a small K is on the Pareto frontier of accuracy and throughput; approximate search schemes such as HNSW must supply more documents to the LLM to match accuracy, which erases their search-time advantage. To make ENNS practical, IKS places 64 processing engines per LPDDR5X package, keeps a 32-entry ordered score list in hardware, and uses CXL.cache-coherent doorbells so the CPU and accelerators communicate with negligible overhead. On a 512GB vector database, IKS reports 470.6 ms retrieval latency, 13.4–27.9x faster than the Sapphire Rapids CPU baseline, and cuts end-to-end time-to-interactive by 1.7–26.3x for FiDT5, Llama-8B, and Llama-70B pipelines.

Load-bearing premise

Every reported speedup rests on the assumption that the authors' cycle-approximate simulator faithfully predicts the timing of real IKS hardware, since no physical IKS device was measured.

Editorial extensions

If this is right

  • RAG applications can use exact top-K retrieval with K as small as 1–4 and match or beat approximate-search accuracy, reducing generation cost and time-to-first-token.
  • Four IKS units can cover a 2TB corpus, scaling exact ENNS nearly linearly, with host-side top-K aggregation adding only tens of microseconds.
  • IKS's internal LPDDR5X can be disaggregated as CXL memory for co-running applications, so accelerator memory is not stranded when idle.
  • For datasets where ANNS cannot prune more than a small fraction of the corpus, exact search on IKS can beat ANNS on both accuracy and latency.
  • Because IKS always returns 32 candidates, varying K from 1 to 32 changes generation time but not retrieval time, which stays flat at about 470.6 ms for a 512GB corpus.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the simulator's timing is confirmed on real silicon, exact retrieval could displace ANNS in high-recall RAG serving, and the cost argument against GPUs for this memory-bound workload would strengthen because LPDDR5X is cheaper per byte than HBM.
  • The cache-coherent doorbell/umwait interface is general: the same CXL.cache mechanism could offload other memory-bound kernels that tolerate software-managed coherence, such as scans or scatter/gather operations.
  • A natural test not in the paper is combining IKS with early termination or coarse-grained pruning so exhaustive search stops once top-K is stable, reducing bandwidth interference for co-running memory-expander users without sacrificing accuracy.
  • The Pareto argument depends on the generative model's sensitivity to noisy documents; for models fine-tuned on retrieved contexts or tasks where K is large, ANNS may close the gap, so the claimed advantage should be rechecked per application.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies the retrieval phase of Retrieval-Augmented Generation (RAG) and argues that exact nearest neighbor search (ENNS) can be preferable to approximate search (ANNS) for end-to-end RAG accuracy and latency, because exact retrieval lets the system use a smaller document list K without losing accuracy. Based on this observation, the authors design Intelligent Knowledge Store (IKS), a type-2 CXL memory expander with near-memory accelerators (NMAs) attached to LPDDR5X packages, and implement a cache-coherent offload interface via CXL.cache. The paper reports a cycle-approximate simulator of IKS, RTL synthesis results for the NMA, and uses this to claim 13.4–27.9x faster ENNS over a 512 GB corpus relative to an Intel Sapphire Rapids CPU baseline, translating to 1.7–26.3x lower end-to-end RAG inference time across FiDT5, Llama-8B, and Llama-70B applications. The paper also profiles the RAG pipeline, shows that the CPU retrieval phase dominates time-to-interactive for large corpora, and argues that GPUs are expensive and inefficient for memory-bound ENNS.

Significance. If the performance and cost claims hold, this would be a valuable contribution to the systems community: it makes a concrete case for a CXL-attached, memory-expander-style accelerator for exact vector search, with a plausible hardware/software interface design grounded in RTL synthesis and a scale-out NMA architecture. The profiling section provides a useful quantitative view of RAG bottlenecks, and the observation that high-quality ENNS can reduce generation time by lowering K is an interesting and falsifiable claim. The authors ship public artifacts (simulator and modified Faiss) and report area, power, and cost analyses, which are strengths. The central reservation is that all IKS speedups come from a simulator whose key timing assumptions are not validated against hardware, and the headline speedup numbers are not stated consistently across the abstract and conclusion. With additional validation and sensitivity analysis, the core idea could be significant; in its current form the quantitative claims are not yet fully supported.

major comments (4)
  1. [Abstract vs. §10 (Conclusion)] The headline performance numbers are internally inconsistent. The abstract reports 13.4–27.9x ENNS speedup and 1.7–26.3x end-to-end inference speedup, while the conclusion states 18–52x ENNS speedup and 2.0–49x end-to-end speedup over the same 512 GB corpus. One of these ranges is not computed from the data in the paper, and this instability makes it impossible for a reader to know the central quantitative claim. The authors must reconcile these ranges and ensure that every stated speedup is directly traceable to a table or figure.
  2. [§6.1, Table 3, Fig. 9] All IKS retrieval times are produced by a cycle-approximate simulator rather than by IKS hardware or a validated FPGA prototype. The claim that IKS outperforms CPU and GPU therefore rests entirely on the simulator's assumptions for LPDDR5X bandwidth efficiency, CXL.cache traffic and latency, NMA timing, and the umwait/doorbell overhead. The paper provides no sensitivity analysis for these parameters, and no comparison between the simulator and any real CXL device or memory-side accelerator. Given that the reported 512 GB retrieval time of 470.6 ms is close to the pure bandwidth bound of 512 GB / 1.09 TB/s, the headline speedup is effectively a bandwidth-ratio claim; this needs to be demonstrated with a validated model, or the claims need to be weakened accordingly.
  3. [§3.4, §6.2, Fig. 4] The CPU baseline appears to be tuned, but the paper does not demonstrate that it is truly bandwidth-limited. The CPU reaches only about 31% of its DDR5 peak bandwidth in Fig. 4, and the authors argue this is structural. However, the only CPU ENNS implementation described is Faiss with a OneMKL BLAS backend, with the corpus block size increased to 16384; no attempt is shown to use an AVX-512 or AMX-optimized ENNS kernel that might saturate memory bandwidth more effectively. A saturating kernel could reduce the reported speedup to roughly 4–6x, which changes the paper's central conclusion. The authors should include a saturated-bandwidth ENNS implementation or explicitly bound how much of the speedup is an artifact of the particular BLAS configuration.
  4. [§7.1, Fig. 9] The comparison of IKS against GPU claims that IKS outperforms 1 H100 by 2.6x (batch 1) and 4.6x (batch 16) for a 50 GB corpus, and attributes this to inefficient top-K and low GPU memory bandwidth utilization. This is a strong claim that depends on the same simulator while the GPU numbers are measured. Since the IKS numbers are entirely modeled, the paper should at minimum include a sensitivity analysis on the effective LPDDR5X bandwidth utilization and the NMA clock frequency, and should state clearly which performance components are measured versus modeled. Without that, the relative IKS-vs-GPU comparison is not yet supported.
minor comments (5)
  1. [§10] The conclusion says 'Intel Sapphire Rapids accelerators,' which should be 'Intel Sapphire Rapids CPUs' to match the rest of the paper.
  2. [§5.3, Fig. 6] The 12-step transaction list in Figure 6 is helpful but the figure's step numbering is partially redundant with the arrows; a simpler one-line-per-step table would improve readability.
  3. [§5.1] The claim that the ×2 PCIe uplink oversubscription is 'neither a bottleneck for acceleration mode nor memory expander mode' is stated without a quantitative demonstration; providing a short bandwidth-accounting table would strengthen the argument.
  4. [§6.1] The simulator description in Appendix A is referenced, but the appendix in the provided text is only a checklist; the full simulator documentation should be integrated into the main artifact description or pointed to more explicitly in the repository.
  5. [§7.2, Fig. 10] The reported per-application end-to-end speedup ranges (e.g., 5.6–25.6x for FiDT5) are not presented in a table that would allow the reader to reproduce the minimum and maximum values; adding such a table would clarify how the abstract's global 1.7–26.3x range is derived.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: IKS speedups come from an independent cycle-approximate simulator and measured CPU/GPU baselines, with only minor non-load-bearing self-citations.

full rationale

I walked the paper's derivation chain from the RAG profiling observation to the IKS speedup claims. The observation that exact retrieval can reduce end-to-end latency while preserving accuracy is an empirical result from measured FiDT5, Llama-8B, and Llama-70B pipelines (Figure 2), not an assumption built into the conclusion. The IKS retrieval times in Table 3 are produced by the simulator described in Section 6.1, which uses timing parameters from RTL synthesis, LPDDR5X access timing, PCIe/CXL timing, and measured software overhead; the CPU and GPU baselines are measured on real systems (Table 2). I found no equation in the paper that defines an IKS retrieval time in terms of the speedup it is later claimed to produce, and no fitted parameter is renamed as a prediction. The paper's own limitations—no hardware/FPGA validation, no sensitivity analysis for CXL.cache latency, and evaluation limited to at most four IKS units—are correctness and validation risks, not circularity. Self-citations such as SmartDIMM and XFM appear in the Section 4 design rationale for choosing CXL over DIMM-based near-memory processing, but the headline 13.4-27.9x result does not depend on those citations; removing them would not change the simulator-based comparison. The inconsistency between the abstract's and conclusion's speedup ranges is a reporting inconsistency, not a definitional reduction. Under the stated rubric, this paper is self-contained against external baselines, so the circularity score is low.

Assumptions & free parameters 4 free parameters · 4 assumptions · 1 invented entities

The performance and cost claims depend on simulator timing assumptions, hand-chosen design parameters (68 MACs, 64 PEs, K=32), and tuned ANNS baselines. There are no statistical fits, but several domain assumptions lack independent validation.

free parameters (4)
  • MAC units per dot-product unit (NMA) = 68
    Chosen so 68 MACs at 1 GHz saturate the 136 GB/s LPDDR5X channel bandwidth. The speedup result depends on this compute-to-bandwidth ratio.
  • Processing engines per NMA = 64
    Supports batch sizes up to 64; the paper notes only one PE is used at batch size 1, so throughput is flat below 64.
  • Hardware top-K list size = 32
    The NMA returns 32 scores regardless of the requested K, which caps the maximum context size and shapes generation time comparisons.
  • HNSW index hyperparameters = M=32, efConstruction=128, efSearch=2048/10000
    Tuned by the authors to create strong ANNS baselines; the ANNS versus ENNS comparison depends on these choices.
assumptions (4)
  • domain assumption CXL.cache can implement the coherent doorbell and context-buffer interface with the modeled latency.
    Section 5.3 describes the interface; the only experimental evidence is a two-socket CPU test, not a real CXL device.
  • domain assumption The cycle-approximate simulator predicts real IKS performance.
    Section 6.1 states all IKS results come from the simulator; no silicon validation is shown.
  • domain assumption Generation accuracy depends primarily on retrieval recall, making ENNS with small K Pareto-optimal.
    Section 3.2 supports this on Natural Questions with three generators; the paper generalizes beyond this dataset.
  • domain assumption LPDDR5X bit flips are acceptable for ENNS, so ECC can be omitted.
    Section 5.1 argues similarity search is resilient to rare bit flips; this is untested in the evaluation.
invented entities (1)
  • Intelligent Knowledge Store (IKS)
    purpose: Type-2 CXL memory expander with scale-out near-memory accelerators for exact nearest neighbor search over large vector databases.
    Evaluated only via cycle-approximate simulation and RTL synthesis; no fabricated device or external measurement is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Accelerating Retrieval-Augmented Generation." pith.science (2026). https://pith.science/paper/75JZUD72

@misc{pith2026241215246,
  author       = {Pith},
  title        = {Pith review of: Accelerating Retrieval-Augmented Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/75JZUD72}},
  note         = {Machine review of arXiv:2412.15246}
}
read the original abstract

An evolving solution to address hallucination and enhance accuracy in large language models (LLMs) is Retrieval-Augmented Generation (RAG), which involves augmenting LLMs with information retrieved from an external knowledge source, such as the web. This paper profiles several RAG execution pipelines and demystifies the complex interplay between their retrieval and generation phases. We demonstrate that while exact retrieval schemes are expensive, they can reduce inference time compared to approximate retrieval variants because an exact retrieval model can send a smaller but more accurate list of documents to the generative model while maintaining the same end-to-end accuracy. This observation motivates the acceleration of the exact nearest neighbor search for RAG. In this work, we design Intelligent Knowledge Store (IKS), a type-2 CXL device that implements a scale-out near-memory acceleration architecture with a novel cache-coherent interface between the host CPU and near-memory accelerators. IKS offers 13.4-27.9x faster exact nearest neighbor search over a 512GB vector database compared with executing the search on Intel Sapphire Rapids CPUs. This higher search performance translates to 1.7-26.3x lower end-to-end inference time for representative RAG applications. IKS is inherently a memory expander; its internal DRAM can be disaggregated and used for other applications running on the server to prevent DRAM, which is the most expensive component in today's servers, from being stranded.

Figures

Figures reproduced from arXiv: 2412.15246 by the authors.

Figure 1
Figure 1. Overview of the Retrieval-Augmented Generation (RAG) pipeline. 1 Introduction State-of-the-art natural language processing systems heavily rely on large language models (LLMs)–deep Transformer net￾works [93] with hundreds of millions of parameters. There is much evidence that information presented in the LLM train￾ing corpora is “memorized” in the LLM parameters, forming a parametric knowledge base that the model de… view at source ↗
Figure 2
Figure 2. Generation accuracy vs. throughput (Queries/sec) of representative RAG applications for various retrieval algorithms and document counts (K). The corpus size is set to 50 GB and batch size to 16. over during the search, but for ANNS, the index can be more complex. For example, HNSW stores embedding vectors in a graph-based data structure [52]. To evaluate ANNS, we use the state-of-the-art HNSW [52] ANNS algorithm, a… view at source ↗
Figure 3
Figure 3. Latency breakdown of FiDT5, Llama-8B, Llama-70B for various val￾ues of K, corpus sizes. All configurations use batch size 1. Retrieval is ENNS and runs on CPU, generation runs on a single NVIDIA H100 (SXM) for all generative models. The value in each bar shows the absolute retrieval time. Batch Size 1 16 Corpus Size 50 GB 512 GB 50 GB 512 GB CPU 1 1 1 1 AMX 1.05 1.02 1.10 1.09 GPU 5.2 36.9 6.0 43.7 [PITH_FULL_IMAGE… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Roofline model for ENNS using Batch Size 1 and 16. See Section 6 for the experimental setup. 4 Case for Near-Memory ENNS Acceleration ENNS is characterized by the following features: • ENNS operations exhibit no data reuse for pair-wise similarity score calculations be…
Figure 5
Figure 5. Figure 5: (a) IKS internal DRAM, scratchpad spaces, and configuration registers are mapped to the host address space. The scratchpad and configuration register address ranges are labeled as Context Buffers (CB). (b) IKS is a compute-enabled CXL memory expander that includes eigh…
Figure 6
Figure 6. Figure 6: CPU-IKS interface through cache coherent CXL interconnect. buffers. Next, the host process writes into a doorbell regis￾ter, which is mapped to a cache line shared by NMAs. NMAs poll on the doorbell register, and as soon as there is a change, the offload starts (step 4…
Figure 7
Figure 7. Figure 7: Data layout inside each LPDDR5X package. The host CPU communi￾cates the base address "B", vector dimension "VD", and the number of vectors "N" to the NMAs for each offload. Four embedding vectors (EVs) are high￾lighted in this layout. (Batch Size = 4) Byte Offset Physi…
Figure 8
Figure 8. Figure 8: Data layout inside the query scratchpads mapped to host memory address at query scratchpad base address "QS_B". As we increase the batch size, more query scratchpads are populated with distinct query vectors. Platform Parameter Description CPU CPU model Intel Xeon 4416…
Figure 9
Figure 9. Figure 9: Comparison of ENNS retrieval time for CPU, AMX, GPU (1, 2, 4, and 8 devices), and IKS (1, and 4 devices) for various corpus sizes. The absence of bars in specific GPU and IKS configurations indicates that the corpus exceeds the capacity of the accelerator memory. The Y…
Figure 10
Figure 10. Figure 10: Inference time breakdown of CPU vs. IKS retrieval for FiDT5, Llama-8B, and Llama-70B. Generative model runs on GPU. than 64. As shown, the purposefully built NMA logic for ENNS enables 1 IKS unit to outperform 1 GPU for a 50 GB corpus for batch sizes 1 and 16 by 2.6× …
Figure 11
Figure 11. Figure 11: Comparison of accuracy and throughput of FiDT5, Llama-8B, and Llama-70B for various configurations. ANNS-2 is an HNSW index with M, efConstruction, and efSearch of 32, 128, and 2048, respectively. configurations) running on CPU, and RAG with ENNS run￾ning on IKS. The …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 17 canonical work pages

  1. [1]

    Gulavani, and Ramachandran Ramjee

    Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwa- tra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. Microsoft Research Blog (September 2023). https://doi.org/10.48550/arXiv.2308.16369

  2. [2]

    Yeonchan Ahn, Sang-Goo Lee, Junho Shim, and Jaehui Park. 2022. Retrieval-Augmented Response Generation for Knowledge-Grounded Conversation in the Wild. IEEE Access 10 (2022), 131374–131385. https://doi.org/10.1109/ACCESS.2022.3228964

  3. [3]

    Meta AI. 2024. Llama 3. Online; accessed 2024-12-13. https://llama.meta.com/llama3/

  4. [4]

    Mohammad Alian, Seung Won Min, Hadi Asgharimoghaddam, Ashutosh Dhar, Dong Kai Wang, Thomas Roewer, Adam McPad- den, Oliver O’Halloran, Deming Chen, Jinjun Xiong, et al . 2018. Application-Transparent Near-Memory Processing Architecture with Memory Channel Network. In 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 80...

  5. [5]

    Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2023. Llm in a flash: Efficient large language model inference with limited memory. arXiv preprint arXiv:2312.11514 (2023). https://doi.org/10.48550/arXiv.2312.11514

  6. [6]

    Martin Aumüller, Erik Bernhardsson, and Alexander Faithfull

  7. [7]

    Artem Babenko and Victor Lempitsky. 2012. The inverted multi-index. In 2012 IEEE Conference on Computer Vision and Pattern Recognition. 3069–3076. https://doi.org/10.1109/CVPR.2012.6248038

  8. [8]

    Giovanni Bonetta, Rossella Cancelliere, Ding Liu, and Paul Vozila

Show all 121 references
  1. [9]

    Francesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando, Raffaele Perego, and Salvatore Trani. 2024. Early Exit Strategies for Approximate k-NN Search in Dense Retrieval. In Proceedings of the 33rd ACM International Conference on Information and Knowledge ...

  2. [10]

    Deng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Wai Lam, and Shuming Shi. 2019. Skeleton-to-Response: Dialogue Generation Guided by Retrieval Memory. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Hu...

  3. [11]

    Wenhu Chen, Hexiang Hu, Xi Chen, Pat Verga, and William Cohen

  4. [12]

    Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W. Cohen. 2022. Re-Imagen: Retrieval-Augmented Text-to- Image Generator. ArXiv abs/2209.14491 (2022). https: //api.semanticscholar.org/CorpusID:252596087

  5. [13]

    Dally, Yatish Turakhia, and Song Han

    William J. Dally, Yatish Turakhia, and Song Han. 2020. Domain- specific hardware accelerators. Commun. ACM 63, 7 (2020), 48–57. https://doi.org/10.1145/3361682

  6. [14]

    Feyza Duman Keles, Pruthuvi Mahesakya Wijewardena, and Chinmay Hegde. 2023. On The Computational Complexity of Self-Attention. In Proceedings of The 34th International Conference on Algorithmic Learning Theory (Proceedings of Machine Learning Research, Vol. 201), Shipra Agrawa...

  7. [15]

    Zhengcong Fei. 2021. Memory-Augmented Image Captioning. Proceedings of the AAAI Conference on Artificial Intelligence 35, 2 (May 2021), 1317–1324. https://doi.org/10.1609/aaai.v35i2.16220

  8. [16]

    Amin Firoozshahian, Joel Coburn, Roman Levenstein, Rakesh Nattoji, Ashwin Kamath, Olivia Wu, Gurdeepak Grewal, Harish Aepala, Bhasker Jakka, Bob Dreyer, Adam Hutchin, Utku Diril, Krishnakumar Nair, Ehsan K. Aredestani, Martin Schatz, Yuchen Hao, Rakesh Komu- ravelli, Kunming H...

  9. [17]

    In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2023. Prompt cache: Modular attention reuse for low-latency inference. arXiv preprint arXiv:2311.04934 (2023). https://doi.org/10.48550/arXiv.2311.04934

  10. [18]

    Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. 2018. Search engine guided neural machine translation. In Proceed- ings of the AAAI Conference on Artificial Intelligence , Vol. 32. https://doi.org/10.1609/aaai.v32i1.12013

  11. [19]

    Liangke Gui, Borui Wang, Qiuyuan Huang, Alexander Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2022. KAT: A Knowledge Augmented Transformer for Vision-and-Language. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguisti...

  12. [20]

    Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang

    Tatsunori B. Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang. 2018. A Retrieve-and-Edit Framework for Predicting Struc- tured Outputs. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18). Curran Associa...

  13. [21]

    Qiuxiang He, Guoping Huang, Qu Cui, Li Li, and Lemao Liu. 2021. Fast and Accurate Neural Machine Translation with Translation Memory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natur...

  14. [22]

    Sebastian Hofstätter, Jiecao Chen, Karthik Raman, and Hamed Zamani

  15. [23]

    Mark Horowitz. 2014. 1.1 Computing’s energy problem (and what we can do about it). In 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) . 10–14. https://doi.org/10.1109/ISSCC.2014.6757323

  16. [24]

    Mohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H Loh, and Adwait Jog. 2021. Analyzing and leveraging decoupled L1 caches in GPUs. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 467–478. https://doi.org/10.1109/HPCA516...

  17. [25]

    Gautier Izacard and Edouard Grave. 2021. Distilling Knowl- edge from Reader to Retriever for Question Answering. In International Conference on Learning Representations . https://openreview.net/forum?id=NTEz-6wysdb

  18. [26]

    Gautier Izacard and Edouard Grave. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Paola Merlo, Jorg Tiede...

  19. [27]

    Hervé Jégou, Matthijs Douze, and Jeff Johnson. 2017. Faiss: A Library for Efficient Similarity Search. Engineering at Meta. https://engineering.fb.com/2017/03/29/data-infrastructure/faiss-a- library-for-efficient-similarity-search/ Accessed: 2024-12-13

  20. [28]

    Herve Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, 1 (2011), 117–128. https://doi.org/10.1109/TPAMI.2010.57

  21. [29]

    Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022. Deduplicat- ing training data mitigates privacy risks in language models. In International Conference on Machine Learning. PMLR, 10697–10707. https://proceedings.mlr.press/v162/kandpal22a.html

  22. [30]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)...

  23. [31]

    Smith, Yejin Choi, and Kentaro Inui

    Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Velocity Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2024. REALTIME QA: what’s the answer right now?. InProceedings of the 37th International Conference on Neural Informati...

  24. [32]

    Amirhossein Kazemnejad, Mohammadreza Salehi, and Mahdieh Soleymani Baghshah. 2020. Paraphrase Generation by Learning How to Edit from Samples. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schlute...

  25. [33]

    Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, YeonGon Cho, Jin Hyun Kim, Yongsuk Kwon, KyungSoo Kim, Jin Jung, Ilkwon Yun, Sung Joo Park, Hyunsun Park, Joonho Song, Jeonghyeon Cho, Kyomin Sohn, Nam Sung Kim, and Hsien-Hsin S. Lee. 2022. ...

  26. [34]

    Ben Keller, Rangharajan Venkatesan, Steve Dai, Stephen G Tell, Brian Zimmer, Charbel Sakr, William J Dally, C Thomas Gray, and Brucek Khailany. 2023. A 95.6-TOPS/W deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm. IEEE Journal of Solid-Stat...

  27. [35]

    Jin Hyun Kim, Shin-Haeng Kang, Sukhan Lee, Hyeonsu Kim, Yuhwan Ro, Seungwon Lee, David Wang, Jihyun Choi, Jinin So, YeonGon Cho, JoonHo Song, Jeonghyeon Cho, Kyomin Sohn, and Nam Sung Kim. 2022. Aquabolt-XL HBM2-PIM, LPDDR5-PIM With In-Memory Processing, and AXDIMM With Accele...

  28. [36]

    To Eun Kim, Alireza Salemi, Andrew Drozdov, Fernando Diaz, and Hamed Zamani. 2024. Retrieval-Enhanced Ma- chine Learning: Synthesis and Opportunities. arXiv (2024). https://doi.org/10.48550/arXiv.2407.12982 arXiv:2407.12982

  29. [37]

    Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani

    Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra, Alireza Salemi, Ryan A. Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani. 2024. LongLaMP: A Benchmark for Personalized ...

  30. [38]

    Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  31. [39]

    Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain ques- tion answering. arXiv preprint arXiv:1906.00300 (2019). https://doi.org/10.48550/arXiv.1906.00300

  32. [40]

    Lee, and Tae Jun Ham

    Yejin Lee, Hyunji Choi, Sunhong Min, Hyunseung Lee, Sangwon Beak, Dawoon Jeong, Jae W. Lee, and Tae Jun Ham. 2022. ANNA: Specialized Architecture for Approximate Nearest Neighbor Search. In2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 169–...

  33. [41]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela

  34. [42]

    Andersen, and Yuxiong He

    Conglong Li, Minjia Zhang, David G. Andersen, and Yuxiong He

  35. [43]

    Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D

    Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bianchini. 2023. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. In Procee...

  36. [44]

    Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu

  37. [45]

    Wen Li, Ying Zhang, Yifang Sun, Wei Wang, Mingjie Li, Wenjie Zhang, and Xuemin Lin. 2020. Approximate Nearest Neighbor Search on High Dimensional Data — Experiments, Analyses, and Improvement. IEEE Transactions on Knowledge and Data Engineering 32, 8 (2020), 1475–1488. https:/...

  38. [46]

    In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20)

    Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages. https://doi.org/1...

  39. [47]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Associa- tion for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013

  40. [48]

    InProceedings of the 2020 ACM SIGMOD International Conference on Management of Data (Portland, OR, USA) (SIGMOD ’20)

    Improving Approximate Nearest Neighbor Search through Learned Adaptive Early Termination. InProceedings of the 2020 ACM SIGMOD International Conference on Management of Data (Portland, OR, USA) (SIGMOD ’20). Association for Computing Machinery, New York, NY, USA, 2539–2554. ht...

  41. [49]

    Shangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow, and Yang Liu

  42. [50]

    Locuza. 2022. Die Analysis: Samsung Exynos 2200 with RDNA2 Graph- ics. Online; accessed 2024-12-13. https://locuza.substack.com/p/die- analysis-samsung-exynos-2200 Derrick Quinn et al

  43. [51]

    ArXiv abs/2202.01110 (2022)

    A Survey on Retrieval-Augmented Text Generation. ArXiv abs/2202.01110 (2022). https://api.semanticscholar.org/CorpusID: 246472929

  44. [52]

    Malkov and D

    Yu A. Malkov and D. A. Yashunin. 2020. Efficient and Ro- bust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 4 (apr 2020), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473

  45. [53]

    Pierre Lienhart. 2024. LLM Inference Series: 4. KV caching, a deeper look. Pierre Leinhart (Medium) (Jan 2024). https://medium.com/@plienhar/llm-inference-series-4-kv- caching-a-deeper-look-4ba9a77746c8

  46. [54]

    Timothy Prickett Morgan. 2024. He Who Can Pay Top Dollar For HBM Memory Controls AI Training. The Next Platform (2024). https://www.nextplatform.com/2024/02/27/he-who-can-pay-top- dollar-for-hbm-memory-controls-ai-training/ Accessed: 2024-06-23

  47. [55]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978 (2023). https://doi.org/10.48550/arXiv.2306.00978

  48. [56]

    OpenAI. 2023. ChatGPT plugins. OpenAI Blog (2023). https://openai.com/blog/chatgpt-plugins

  49. [57]

    In International Conference on Learning Representations

    Retrieval-Augmented Generation for Code Summarization via Hybrid GNN. In International Conference on Learning Representations. https://openreview.net/forum?id=zv-typ1gPxA

  50. [58]

    Md Rizwan Parvez, Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval Augmented Code Generation and Summarization. InFindings of the Association for Computational Linguistics: EMNLP 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Sco...

  51. [59]

    Loh, Natalie Enright Jerger, Ajaykumar Kannan, and Yasuko Eckert

    Gabriel H. Loh, Natalie Enright Jerger, Ajaykumar Kannan, and Yasuko Eckert. 2015. Interconnect-Memory Challenges for Multi-chip, Silicon Interposer Systems. In Proceedings of the 2015 International Symposium on Memory Systems (Washington DC, DC, USA)(MEMSYS ’15). Association ...

  52. [60]

    Dylan Patel and Jeremie Eliahou Ontiveros. 2024. CXL Is Dead In The AI Era. Online; accessed 2024-12-13. https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era

  53. [61]

    Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024. SpecInfer: Accelerating Large Language Model Serving w...

  54. [62]

    Neel Patel, Amin Mamandipoor, Derrick Quinn, and Mohammad Alian. 2023. XFM: Accelerated Software-Defined Far Memory. In Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture (Toronto, ON, Canada) (MICRO ’23) . Asso- ciation for Computing Machiner...

  55. [63]

    Samuel Naffziger, Kevin Lepak, Milam Paraschou, and Ma- hesh Subramony. 2020. 2.2 AMD Chiplet Architecture for High-Performance Server and Desktop Products. In 2020 IEEE International Solid- State Circuits Conference - (ISSCC) . 44–45. https://doi.org/10.1109/ISSCC19947.2020.9063103

  56. [64]

    Hao Peng, Ankur Parikh, Manaal Faruqui, Bhuwan Dhingra, and Dipanjan Das. 2019. Text Generation with Exemplar-based Adaptive Decoding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog...

  57. [66]

    Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a Benchmark for Knowledge Intensive Language Task...

  58. [67]

    Dylan Patel. 2022. Apple M2 Die Shot and Architecture Analysis – Big Cost Increase And A15 Based IP.SemiAnalysis (June 2022). https: //www.semianalysis.com/p/apple-m2-die-shot-and-architecture

  59. [68]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. The Journal of Machine Learning Research 21, 1, Article 140 (j...

  60. [69]

    Patel, A

    N. Patel, A. Mamandipoor, M. Nouri, and M. Alian. 2024. SmartDIMM: In-Memory Acceleration of Upper Layer Protocols. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE Computer Society, Los Alamitos, CA, USA, 312–329. https://doi.org/10....

  61. [70]

    Rita Ramos, Desmond Elliott, and Bruno Martins. 2023. Retrieval- augmented Image Captioning. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Andreas Vlachos and Isabelle Augenstein (Eds.). Association for Computati...

  62. [71]

    Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting . In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE Co...

  63. [73]

    Fabio Petroni, Aleksandra Piktus, and Angela Fan. 2020. Introducing KILT, a new unified benchmark for knowledge-intensive NLP tasks. Online; accessed 2024-11-22. https://ai.meta.com/blog/introducing- kilt-a-new-unified-benchmark-for-knowledge-intensive-nlp- tasks/

  64. [74]

    Alireza Salemi, Mahta Rafiee, and Hamed Zamani. 2023. Pre-Training Multi-Modal Dense Retrievers for Outside-Knowledge Visual Ques- tion Answering. InProceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (Taipei, Taiwan)(ICTIR ’23). Assoc...

  65. [75]

    Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2021 Conference of the North Amer...

  66. [77]

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang

  67. [78]

    Alireza Salemi and Hamed Zamani. 2024. Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval- Augmented Large Language Models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington ...

  68. [79]

    Satyabrata Sarangi and Bevan Baas. 2021. DeepScaleTool: A Tool for the Accurate Estimation of Technology Scaling in the Deep-Submicron Era. In 2021 IEEE International Symposium on Circuits and Systems (ISCAS). 1–5. https://doi.org/10.1109/ISCAS51556.2021.9401196

  69. [80]

    Alireza Salemi, Juan Altmayer Pizzorno, and Hamed Zamani

  70. [81]

    In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23)

    A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question Answering. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for Computing Mac...

  71. [82]

    Haihao Shen, Hanwen Chang, Bo Dong, Yu Luo, and Hengyu Meng. 2023. Efficient llm inference on cpus. arXiv preprint (2023). https://doi.org/10.48550/arXiv.2311.00502

  72. [83]

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024. LaMP: When Large Language Models Meet Per- sonalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martins, ...

  73. [84]

    Joonseop Sim, Soohong Ahn, Taeyoung Ahn, Seungyong Lee, Myunghyun Rhee, Jooyoung Kim, Kwangsik Shin, Donguk Moon, Euiseok Kim, and Kyoung Park. 2023. Computational CXL-Memory Solution for Accelerating Memory-Intensive Ap- plications. IEEE Computer Architecture Letters 22, 1 (2...

  74. [85]

    Alireza Salemi and Hamed Zamani. 2024. Comparing Retrieval-Augmentation and Parameter-Efficient Fine-Tuning for Privacy-Preserving Personalization of Large Language Models. arXiv:2409.09510 [cs.CL] https://arxiv.org/abs/2409.09510

  75. [86]

    Heidi Steen and Dan Wahlin. 2023. Retrieval Augumented Generation Overview. Microsoft Learn (2023). https://learn.microsoft.com/en- us/azure/search/retrieval-augmented-generation-overview

  76. [87]

    Alireza Salemi and Hamed Zamani. 2024. Learning to Rank for Multiple Retrieval-Augmented Models through Iterative Utility Max- imization. arXiv:2410.09942 [cs.CL] https://arxiv.org/abs/2410.09942

  77. [88]

    Su, Samuel Naffziger, and Mark Papermaster

    Lisa T. Su, Samuel Naffziger, and Mark Papermaster. 2017. Multi-chip technologies to unleash computing performance gains over the next decade. In 2017 IEEE International Electron Devices Meeting (IEDM) . 1.1.1–1.1.8. https://doi.org/10.1109/IEDM.2017.8268306

  78. [89]

    Yixuan Su, David Vandyke, Simon Baker, Yan Wang, and Nigel Collier. 2021. Keep the Primary, Rewrite the Secondary: A Two-Stage Approach for Paraphrase Generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, Wenjie Li...

  79. [90]

    Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cuc- chiara. 2022. Retrieval-Augmented Transformer for Image Captioning. In Proceedings of the 19th International Conference on Content-Based Multimedia Indexing (Graz, Austria) (CBMI ’22). Association for Computing Machin...

  80. [91]

    Schuh, Arvind Krishnamurthy, David Culler, Henry M

    Henry N. Schuh, Arvind Krishnamurthy, David Culler, Henry M. Levy, Luigi Rizzo, Samira Khan, and Brent E. Stephens. 2024. CC-NIC: a Cache-Coherent Interface to the NIC. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages an...

  81. [92]

    Zhiliang Tian, Wei Bi, Xiaopeng Li, and Nevin L. Zhang. 2019. Learn- ing to Abstract for Memory-augmented Conversational Response Generation. In Proceedings of the 57th Annual Meeting of the Associ- ation for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màr...

  82. [93]

    Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen...

  83. [94]

    Mengzhao Wang, Xiaoliang Xu, Qiang Yue, and Yuxiang Wang

  84. [95]

    Shamane Siriwardhana, Rivindu Weerasekera, Elliott Wen, Tharindu Kaluarachchi, Rajib Rana, and Suranga Nanayakkara. 2023. Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering. Transactions of the Association for Comp...

  85. [96]

    Zelin Wang, Ping Gong, Yibo Zhang, Jihao Gu, and Xuanyuan Yang

  86. [97]

    Jovan Stojkovic, Esha Choukse, Chaojie Zhang, Inigo Goiri, and Josep Torrellas. 2024. Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference.arXiv preprint arXiv:2403.20306 (2024). arXiv:2403.20306 [cs.AI] https://doi.org/10.48550/arXiv.2403.20306

  87. [98]

    WikiChip. 2024. Mask / Reticle. Online; accessed 2024-12-13. https://en.wikichip.org/wiki/mask

  88. [99]

    Yu Wu, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li, and Ming Zhou. 2019. Response Generation by Context-Aware Prototype Edit- ing. Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 7281–7288. https://doi.org/10.1609/aaai.v33i01.33017281

  89. [100]

    Gemini Team. 2024. Gemini: A Family of Highly Capable Multimodal Models. arXiv (2024). arXiv:2312.11805 [cs.CL] https://doi.org/10.48550/arXiv.2312.11805

  90. [101]

    David Thulke, Nico Daheim, Christian Dugast, and Hermann Ney

  91. [102]

    arXiv preprint arXiv:2102.04643 (2021)

    Efficient retrieval augmented generation from unstructured knowledge for task-oriented dialog. arXiv preprint arXiv:2102.04643 (2021). https://doi.org/10.48550/arXiv.2102.04643

  92. [103]

    Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective Retrieval Augmented Generation. arXiv (2024). https://doi.org/10.48550/arXiv.2401.15884 arXiv:2401.15884 [cs.CL]

  93. [104]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin

  94. [106]

    Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024. The Shift from Models to Compound AI Systems. Online; accessed 2024-12-13. https://bair.berkeley.edu/blog...

  95. [107]

    Pro- ceedings of the VLDB Endowment 14, 11 (jul 2021), 1964–1978

    A comprehensive survey and experimental comparison of graph-based approximate nearest neighbor search. Pro- ceedings of the VLDB Endowment 14, 11 (jul 2021), 1964–1978. https://doi.org/10.14778/3476249.3476255

  96. [108]

    Yitu Wang, Shiyu Li, Qilin Zheng, Linghao Song, Zongwang Li, Andrew Chang, Hai "Helen" Li, and Yiran Chen. 2024. NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neigh- bor Search through Near Data Processing. In Proceedings of the 39th Annual International Sym...

  97. [109]

    Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura. 2018. Guiding Neural Machine Translation with Retrieved Translation Pieces. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: H...

  98. [110]

    In Natural Language Processing and Chinese Computing , Fei Liu, Nan Duan, Qingting Xu, and Yu Hong (Eds.)

    Retrieval-Augmented Knowledge-Intensive Dialogue. In Natural Language Processing and Chinese Computing , Fei Liu, Nan Duan, Qingting Xu, and Yu Hong (Eds.). Springer Nature Switzerland, Cham, 16–28. https://doi.org/10.48550/arXiv.2005.11401

  99. [111]

    Jason Weston, Emily Dinan, and Alexander Miller. 2018. Retrieve and Refine: Improved Sequence Generation Models For Dialogue. In Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI, Aleksandr Chuklin, Jeff Dalton, Ju...

  100. [112]

    Yun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu, Liangchen Luo, Lei Meng, Bang Liu, and Jindong Chen. 2024. Accelerating Inference of Retrieval- Augmented Generation via Sparse Context Selection. arXiv (2024). https://doi.org/10.48550/arXiv.2405.16178

  101. [114]

    Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien De- mouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In Inter- national Conference on Machine Learning . PMLR, 38087–38099. https://doi.org/10.5555/361840...

  102. [115]

    Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=NG7sS51zVF

  103. [116]

    Jitao Xu, Josep Crego, and Jean Senellart. 2020. Boosting Neural Machine Translation with Similar Translations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Ass...

  104. [118]

    Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023. Inference with reference: Lossless acceleration of large language models. arXiv preprint arXiv:2304.04487 (2023). https://doi.org/10.48550/arXiv.2304.04487

  105. [121]

    Hamed Zamani and Michael Bendersky. 2024. Stochastic RAG: End-to- End Retrieval-Augmented Generation through Expected Utility Maxi- mization. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) ...

  106. [122]

    Hamed Zamani, Fernando Diaz, Mostafa Dehghani, Donald Metzler, and Michael Bendersky. 2022. Retrieval-Enhanced Machine Learning. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22). Asso...

  107. [124]

    Yunan Zhang, Shige Liu, and Jianguo Wang. [n. d.]. Are There Funda- mental Limitations in Supporting Vector Data Management in Rela- tional Databases? A Case Study of PostgreSQL.Preprint ([n. d.]). https: //www.cs.purdue.edu/homes/csjgwang/pubs/ICDE24_VecDB.pdf Accepted for pu...

  108. [125]

    Zhe Zhou, Cong Li, Fan Yang, and Guangyu Sun. 2023. DIMM- Link: Enabling Efficient Inter-DIMM Communication for Near- Memory Processing. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 302–316. https://doi.org/10.1109/HPCA56546.2023.1007...

  109. [2016]

    In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Jian Su, Kevin Duh, and Xavier Carreras (Eds.)

    SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392. https://...

  110. [2017]

    In Advances in Neural Information Processing Systems, I

    Attention is All you Need. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/ 2017/file/3f5...

  111. [2020]

    Information Systems 87 (2020), 101374

    ANN-Benchmarks: A benchmarking tool for approximate nearest neighbor algorithms. Information Systems 87 (2020), 101374. https://doi.org/10.48550/arXiv.1807.05614

  112. [2021]

    The International FLAIRS Conference Proceedings 34, 1 (April 2021)

    Retrieval-Augmented Transformer-XL for Close-Domain Dialog Generation. The International FLAIRS Conference Proceedings 34, 1 (April 2021). https://doi.org/10.32473/flairs.v34i1.128369

  113. [2022]

    InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.)

    MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Li...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.