REVIEW 4 major objections 5 minor 121 references
Accelerating Retrieval-Augmented Generation
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read IKS, a CXL near-memory accelerator, makes exact nearest-neighbor search 13–28x faster and cuts RAG inference time by up to 26x.
desk verdict RAG profiling is solid; IKS speedups are plausible but unvalidated and inconsistently reported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Intelligent Knowledge Store (IKS), a type-2 CXL device that attaches eight LPDDR5X packages, each with a Near-Memory Accelerator (NMA) chip containing 64 processing engines, each combining a 68-MAC dot-product unit with a hardware Top-K unit that keeps an ordered 32-score list. The mechanism that makes the offload cheap is a cache-coherent CPU–accelerator interface built on the CXL.cache protocol: the host writes an offload context into coherent context buffers, rings a doorbell, and blocks with umwait(), while NMAs poll the same doorbell, read the query vectors, and write back partial top-K lists without DMA setup, interrupts, or kernel involvement. Vector layout also carries the argument: embedding vectors are stored column-major in blocks of 68 so the MAC array reads one dimension from 68 vectors per cycle and saturates the 136 GB/s LPDDR5X bandwidth, while the CPU reduces partial lists only at the end.
What would settle it
Take a real CXL type-2 device with LPDDR5X NMAs (or a cycle-accurate FPGA emulation) and measure exact top-32 retrieval over a 512GB corpus: if the latency is not close to the simulated 470.6 ms, or if the speedup over the optimized Faiss CPU baseline falls below 13.4x, the central performance claim fails.
Extended reading notes
Core claim
The paper's central claim is that the retrieval phase, not generation, should dominate the design of RAG systems, and that exact nearest-neighbor search can be made both fast and cheap with the right near-memory hardware. It shows that ENNS with a small K is on the Pareto frontier of accuracy and throughput; approximate search schemes such as HNSW must supply more documents to the LLM to match accuracy, which erases their search-time advantage. To make ENNS practical, IKS places 64 processing engines per LPDDR5X package, keeps a 32-entry ordered score list in hardware, and uses CXL.cache-coherent doorbells so the CPU and accelerators communicate with negligible overhead. On a 512GB vector database, IKS reports 470.6 ms retrieval latency, 13.4–27.9x faster than the Sapphire Rapids CPU baseline, and cuts end-to-end time-to-interactive by 1.7–26.3x for FiDT5, Llama-8B, and Llama-70B pipelines.
Load-bearing premise
Every reported speedup rests on the assumption that the authors' cycle-approximate simulator faithfully predicts the timing of real IKS hardware, since no physical IKS device was measured.
Editorial extensions
If this is right
- RAG applications can use exact top-K retrieval with K as small as 1–4 and match or beat approximate-search accuracy, reducing generation cost and time-to-first-token.
- Four IKS units can cover a 2TB corpus, scaling exact ENNS nearly linearly, with host-side top-K aggregation adding only tens of microseconds.
- IKS's internal LPDDR5X can be disaggregated as CXL memory for co-running applications, so accelerator memory is not stranded when idle.
- For datasets where ANNS cannot prune more than a small fraction of the corpus, exact search on IKS can beat ANNS on both accuracy and latency.
- Because IKS always returns 32 candidates, varying K from 1 to 32 changes generation time but not retrieval time, which stays flat at about 470.6 ms for a 512GB corpus.
Reading between the lines
- If the simulator's timing is confirmed on real silicon, exact retrieval could displace ANNS in high-recall RAG serving, and the cost argument against GPUs for this memory-bound workload would strengthen because LPDDR5X is cheaper per byte than HBM.
- The cache-coherent doorbell/umwait interface is general: the same CXL.cache mechanism could offload other memory-bound kernels that tolerate software-managed coherence, such as scans or scatter/gather operations.
- A natural test not in the paper is combining IKS with early termination or coarse-grained pruning so exhaustive search stops once top-K is stable, reducing bandwidth interference for co-running memory-expander users without sacrificing accuracy.
- The Pareto argument depends on the generative model's sensitivity to noisy documents; for models fine-tuned on retrieved contexts or tasks where K is large, ANNS may close the gap, so the claimed advantage should be rechecked per application.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the retrieval phase of Retrieval-Augmented Generation (RAG) and argues that exact nearest neighbor search (ENNS) can be preferable to approximate search (ANNS) for end-to-end RAG accuracy and latency, because exact retrieval lets the system use a smaller document list K without losing accuracy. Based on this observation, the authors design Intelligent Knowledge Store (IKS), a type-2 CXL memory expander with near-memory accelerators (NMAs) attached to LPDDR5X packages, and implement a cache-coherent offload interface via CXL.cache. The paper reports a cycle-approximate simulator of IKS, RTL synthesis results for the NMA, and uses this to claim 13.4–27.9x faster ENNS over a 512 GB corpus relative to an Intel Sapphire Rapids CPU baseline, translating to 1.7–26.3x lower end-to-end RAG inference time across FiDT5, Llama-8B, and Llama-70B applications. The paper also profiles the RAG pipeline, shows that the CPU retrieval phase dominates time-to-interactive for large corpora, and argues that GPUs are expensive and inefficient for memory-bound ENNS.
Significance. If the performance and cost claims hold, this would be a valuable contribution to the systems community: it makes a concrete case for a CXL-attached, memory-expander-style accelerator for exact vector search, with a plausible hardware/software interface design grounded in RTL synthesis and a scale-out NMA architecture. The profiling section provides a useful quantitative view of RAG bottlenecks, and the observation that high-quality ENNS can reduce generation time by lowering K is an interesting and falsifiable claim. The authors ship public artifacts (simulator and modified Faiss) and report area, power, and cost analyses, which are strengths. The central reservation is that all IKS speedups come from a simulator whose key timing assumptions are not validated against hardware, and the headline speedup numbers are not stated consistently across the abstract and conclusion. With additional validation and sensitivity analysis, the core idea could be significant; in its current form the quantitative claims are not yet fully supported.
major comments (4)
- [Abstract vs. §10 (Conclusion)] The headline performance numbers are internally inconsistent. The abstract reports 13.4–27.9x ENNS speedup and 1.7–26.3x end-to-end inference speedup, while the conclusion states 18–52x ENNS speedup and 2.0–49x end-to-end speedup over the same 512 GB corpus. One of these ranges is not computed from the data in the paper, and this instability makes it impossible for a reader to know the central quantitative claim. The authors must reconcile these ranges and ensure that every stated speedup is directly traceable to a table or figure.
- [§6.1, Table 3, Fig. 9] All IKS retrieval times are produced by a cycle-approximate simulator rather than by IKS hardware or a validated FPGA prototype. The claim that IKS outperforms CPU and GPU therefore rests entirely on the simulator's assumptions for LPDDR5X bandwidth efficiency, CXL.cache traffic and latency, NMA timing, and the umwait/doorbell overhead. The paper provides no sensitivity analysis for these parameters, and no comparison between the simulator and any real CXL device or memory-side accelerator. Given that the reported 512 GB retrieval time of 470.6 ms is close to the pure bandwidth bound of 512 GB / 1.09 TB/s, the headline speedup is effectively a bandwidth-ratio claim; this needs to be demonstrated with a validated model, or the claims need to be weakened accordingly.
- [§3.4, §6.2, Fig. 4] The CPU baseline appears to be tuned, but the paper does not demonstrate that it is truly bandwidth-limited. The CPU reaches only about 31% of its DDR5 peak bandwidth in Fig. 4, and the authors argue this is structural. However, the only CPU ENNS implementation described is Faiss with a OneMKL BLAS backend, with the corpus block size increased to 16384; no attempt is shown to use an AVX-512 or AMX-optimized ENNS kernel that might saturate memory bandwidth more effectively. A saturating kernel could reduce the reported speedup to roughly 4–6x, which changes the paper's central conclusion. The authors should include a saturated-bandwidth ENNS implementation or explicitly bound how much of the speedup is an artifact of the particular BLAS configuration.
- [§7.1, Fig. 9] The comparison of IKS against GPU claims that IKS outperforms 1 H100 by 2.6x (batch 1) and 4.6x (batch 16) for a 50 GB corpus, and attributes this to inefficient top-K and low GPU memory bandwidth utilization. This is a strong claim that depends on the same simulator while the GPU numbers are measured. Since the IKS numbers are entirely modeled, the paper should at minimum include a sensitivity analysis on the effective LPDDR5X bandwidth utilization and the NMA clock frequency, and should state clearly which performance components are measured versus modeled. Without that, the relative IKS-vs-GPU comparison is not yet supported.
minor comments (5)
- [§10] The conclusion says 'Intel Sapphire Rapids accelerators,' which should be 'Intel Sapphire Rapids CPUs' to match the rest of the paper.
- [§5.3, Fig. 6] The 12-step transaction list in Figure 6 is helpful but the figure's step numbering is partially redundant with the arrows; a simpler one-line-per-step table would improve readability.
- [§5.1] The claim that the ×2 PCIe uplink oversubscription is 'neither a bottleneck for acceleration mode nor memory expander mode' is stated without a quantitative demonstration; providing a short bandwidth-accounting table would strengthen the argument.
- [§6.1] The simulator description in Appendix A is referenced, but the appendix in the provided text is only a checklist; the full simulator documentation should be integrated into the main artifact description or pointed to more explicitly in the repository.
- [§7.2, Fig. 10] The reported per-application end-to-end speedup ranges (e.g., 5.6–25.6x for FiDT5) are not presented in a table that would allow the reader to reproduce the minimum and maximum values; adding such a table would clarify how the abstract's global 1.7–26.3x range is derived.
Circularity Check
No significant circularity: IKS speedups come from an independent cycle-approximate simulator and measured CPU/GPU baselines, with only minor non-load-bearing self-citations.
full rationale
I walked the paper's derivation chain from the RAG profiling observation to the IKS speedup claims. The observation that exact retrieval can reduce end-to-end latency while preserving accuracy is an empirical result from measured FiDT5, Llama-8B, and Llama-70B pipelines (Figure 2), not an assumption built into the conclusion. The IKS retrieval times in Table 3 are produced by the simulator described in Section 6.1, which uses timing parameters from RTL synthesis, LPDDR5X access timing, PCIe/CXL timing, and measured software overhead; the CPU and GPU baselines are measured on real systems (Table 2). I found no equation in the paper that defines an IKS retrieval time in terms of the speedup it is later claimed to produce, and no fitted parameter is renamed as a prediction. The paper's own limitations—no hardware/FPGA validation, no sensitivity analysis for CXL.cache latency, and evaluation limited to at most four IKS units—are correctness and validation risks, not circularity. Self-citations such as SmartDIMM and XFM appear in the Section 4 design rationale for choosing CXL over DIMM-based near-memory processing, but the headline 13.4-27.9x result does not depend on those citations; removing them would not change the simulator-based comparison. The inconsistency between the abstract's and conclusion's speedup ranges is a reporting inconsistency, not a definitional reduction. Under the stated rubric, this paper is self-contained against external baselines, so the circularity score is low.
Assumptions & free parameters
free parameters (4)
- MAC units per dot-product unit (NMA) =
68
- Processing engines per NMA =
64
- Hardware top-K list size =
32
- HNSW index hyperparameters =
M=32, efConstruction=128, efSearch=2048/10000
assumptions (4)
- domain assumption CXL.cache can implement the coherent doorbell and context-buffer interface with the modeled latency.
- domain assumption The cycle-approximate simulator predicts real IKS performance.
- domain assumption Generation accuracy depends primarily on retrieval recall, making ENNS with small K Pareto-optimal.
- domain assumption LPDDR5X bit flips are acceptable for ENNS, so ECC can be omitted.
invented entities (1)
-
Intelligent Knowledge Store (IKS)
Cite this review
Pith. "Pith review of Accelerating Retrieval-Augmented Generation." pith.science (2026). https://pith.science/paper/75JZUD72
@misc{pith2026241215246,
author = {Pith},
title = {Pith review of: Accelerating Retrieval-Augmented Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/75JZUD72}},
note = {Machine review of arXiv:2412.15246}
}
read the original abstract
An evolving solution to address hallucination and enhance accuracy in large language models (LLMs) is Retrieval-Augmented Generation (RAG), which involves augmenting LLMs with information retrieved from an external knowledge source, such as the web. This paper profiles several RAG execution pipelines and demystifies the complex interplay between their retrieval and generation phases. We demonstrate that while exact retrieval schemes are expensive, they can reduce inference time compared to approximate retrieval variants because an exact retrieval model can send a smaller but more accurate list of documents to the generative model while maintaining the same end-to-end accuracy. This observation motivates the acceleration of the exact nearest neighbor search for RAG. In this work, we design Intelligent Knowledge Store (IKS), a type-2 CXL device that implements a scale-out near-memory acceleration architecture with a novel cache-coherent interface between the host CPU and near-memory accelerators. IKS offers 13.4-27.9x faster exact nearest neighbor search over a 512GB vector database compared with executing the search on Intel Sapphire Rapids CPUs. This higher search performance translates to 1.7-26.3x lower end-to-end inference time for representative RAG applications. IKS is inherently a memory expander; its internal DRAM can be disaggregated and used for other applications running on the server to prevent DRAM, which is the most expensive component in today's servers, from being stranded.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Gulavani, and Ramachandran Ramjee
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwa- tra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. Microsoft Research Blog (September 2023). https://doi.org/10.48550/arXiv.2308.16369
-
[2]
Yeonchan Ahn, Sang-Goo Lee, Junho Shim, and Jaehui Park. 2022. Retrieval-Augmented Response Generation for Knowledge-Grounded Conversation in the Wild. IEEE Access 10 (2022), 131374–131385. https://doi.org/10.1109/ACCESS.2022.3228964
arXiv 2022
-
[3]
Meta AI. 2024. Llama 3. Online; accessed 2024-12-13. https://llama.meta.com/llama3/
2024
-
[4]
Mohammad Alian, Seung Won Min, Hadi Asgharimoghaddam, Ashutosh Dhar, Dong Kai Wang, Thomas Roewer, Adam McPad- den, Oliver O’Halloran, Deming Chen, Jinjun Xiong, et al . 2018. Application-Transparent Near-Memory Processing Architecture with Memory Channel Network. In 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 80...
arXiv 2018
-
[5]
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2023. Llm in a flash: Efficient large language model inference with limited memory. arXiv preprint arXiv:2312.11514 (2023). https://doi.org/10.48550/arXiv.2312.11514
-
[6]
Martin Aumüller, Erik Bernhardsson, and Alexander Faithfull
-
[7]
Artem Babenko and Victor Lempitsky. 2012. The inverted multi-index. In 2012 IEEE Conference on Computer Vision and Pattern Recognition. 3069–3076. https://doi.org/10.1109/CVPR.2012.6248038
arXiv 2012
-
[8]
Giovanni Bonetta, Rossella Cancelliere, Ding Liu, and Paul Vozila
Show all 121 references
-
[9]
Francesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando, Raffaele Perego, and Salvatore Trani. 2024. Early Exit Strategies for Approximate k-NN Search in Dense Retrieval. In Proceedings of the 33rd ACM International Conference on Information and Knowledge ...
2024
-
[10]
Deng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Wai Lam, and Shuming Shi. 2019. Skeleton-to-Response: Dialogue Generation Guided by Retrieval Memory. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Hu...
2019 doi
-
[11]
Wenhu Chen, Hexiang Hu, Xi Chen, Pat Verga, and William Cohen
-
[12]
Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W. Cohen. 2022. Re-Imagen: Retrieval-Augmented Text-to- Image Generator. ArXiv abs/2209.14491 (2022). https: //api.semanticscholar.org/CorpusID:252596087
2022 arXiv
-
[13]
Dally, Yatish Turakhia, and Song Han
William J. Dally, Yatish Turakhia, and Song Han. 2020. Domain- specific hardware accelerators. Commun. ACM 63, 7 (2020), 48–57. https://doi.org/10.1145/3361682
2020 doi
-
[14]
Feyza Duman Keles, Pruthuvi Mahesakya Wijewardena, and Chinmay Hegde. 2023. On The Computational Complexity of Self-Attention. In Proceedings of The 34th International Conference on Algorithmic Learning Theory (Proceedings of Machine Learning Research, Vol. 201), Shipra Agrawa...
2023
-
[15]
Zhengcong Fei. 2021. Memory-Augmented Image Captioning. Proceedings of the AAAI Conference on Artificial Intelligence 35, 2 (May 2021), 1317–1324. https://doi.org/10.1609/aaai.v35i2.16220
2021 doi
-
[16]
Amin Firoozshahian, Joel Coburn, Roman Levenstein, Rakesh Nattoji, Ashwin Kamath, Olivia Wu, Gurdeepak Grewal, Harish Aepala, Bhasker Jakka, Bob Dreyer, Adam Hutchin, Utku Diril, Krishnakumar Nair, Ehsan K. Aredestani, Martin Schatz, Yuchen Hao, Rakesh Komu- ravelli, Kunming H...
2023
- [17]
-
[18]
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. 2018. Search engine guided neural machine translation. In Proceed- ings of the AAAI Conference on Artificial Intelligence , Vol. 32. https://doi.org/10.1609/aaai.v32i1.12013
2018 doi
-
[19]
Liangke Gui, Borui Wang, Qiuyuan Huang, Alexander Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2022. KAT: A Knowledge Augmented Transformer for Vision-and-Language. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguisti...
2022 doi
-
[20]
Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang
Tatsunori B. Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang. 2018. A Retrieve-and-Edit Framework for Predicting Struc- tured Outputs. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18). Curran Associa...
2018
-
[21]
Qiuxiang He, Guoping Huang, Qu Cui, Li Li, and Lemao Liu. 2021. Fast and Accurate Neural Machine Translation with Translation Memory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natur...
2021 doi
-
[22]
Sebastian Hofstätter, Jiecao Chen, Karthik Raman, and Hamed Zamani
-
[23]
Mark Horowitz. 2014. 1.1 Computing’s energy problem (and what we can do about it). In 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) . 10–14. https://doi.org/10.1109/ISSCC.2014.6757323
2014
-
[24]
Mohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H Loh, and Adwait Jog. 2021. Analyzing and leveraging decoupled L1 caches in GPUs. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 467–478. https://doi.org/10.1109/HPCA516...
2021
-
[25]
Gautier Izacard and Edouard Grave. 2021. Distilling Knowl- edge from Reader to Retriever for Question Answering. In International Conference on Learning Representations . https://openreview.net/forum?id=NTEz-6wysdb
2021
-
[26]
Gautier Izacard and Edouard Grave. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Paola Merlo, Jorg Tiede...
2021 doi
-
[27]
Hervé Jégou, Matthijs Douze, and Jeff Johnson. 2017. Faiss: A Library for Efficient Similarity Search. Engineering at Meta. https://engineering.fb.com/2017/03/29/data-infrastructure/faiss-a- library-for-efficient-similarity-search/ Accessed: 2024-12-13
2017
-
[28]
Herve Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, 1 (2011), 117–128. https://doi.org/10.1109/TPAMI.2010.57
2011 doi
-
[29]
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022. Deduplicat- ing training data mitigates privacy risks in language models. In International Conference on Machine Learning. PMLR, 10697–10707. https://proceedings.mlr.press/v162/kandpal22a.html
2022
-
[30]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)...
2020 doi
-
[31]
Smith, Yejin Choi, and Kentaro Inui
Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Velocity Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2024. REALTIME QA: what’s the answer right now?. InProceedings of the 37th International Conference on Neural Informati...
2024
-
[32]
Amirhossein Kazemnejad, Mohammadreza Salehi, and Mahdieh Soleymani Baghshah. 2020. Paraphrase Generation by Learning How to Edit from Samples. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schlute...
2020 doi
-
[33]
Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, YeonGon Cho, Jin Hyun Kim, Yongsuk Kwon, KyungSoo Kim, Jin Jung, Ilkwon Yun, Sung Joo Park, Hyunsun Park, Joonho Song, Jeonghyeon Cho, Kyomin Sohn, Nam Sung Kim, and Hsien-Hsin S. Lee. 2022. ...
2022
-
[34]
Ben Keller, Rangharajan Venkatesan, Steve Dai, Stephen G Tell, Brian Zimmer, Charbel Sakr, William J Dally, C Thomas Gray, and Brucek Khailany. 2023. A 95.6-TOPS/W deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm. IEEE Journal of Solid-Stat...
2023
-
[35]
Jin Hyun Kim, Shin-Haeng Kang, Sukhan Lee, Hyeonsu Kim, Yuhwan Ro, Seungwon Lee, David Wang, Jihyun Choi, Jinin So, YeonGon Cho, JoonHo Song, Jeonghyeon Cho, Kyomin Sohn, and Nam Sung Kim. 2022. Aquabolt-XL HBM2-PIM, LPDDR5-PIM With In-Memory Processing, and AXDIMM With Accele...
2022
- [36]
-
[37]
Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani
Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra, Alireza Salemi, Ryan A. Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani. 2024. LongLaMP: A Benchmark for Personalized ...
-
[38]
Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019 doi
- [39]
-
[40]
Lee, and Tae Jun Ham
Yejin Lee, Hyunji Choi, Sunhong Min, Hyunseung Lee, Sangwon Beak, Dawoon Jeong, Jae W. Lee, and Tae Jun Ham. 2022. ANNA: Specialized Architecture for Approximate Nearest Neighbor Search. In2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 169–...
2022
-
[41]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela
-
[42]
Andersen, and Yuxiong He
Conglong Li, Minjia Zhang, David G. Andersen, and Yuxiong He
-
[43]
Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D
Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bianchini. 2023. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. In Procee...
2023
-
[44]
Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu
-
[45]
Wen Li, Ying Zhang, Yifang Sun, Wei Wang, Mingjie Li, Wenjie Zhang, and Xuemin Lin. 2020. Approximate Nearest Neighbor Search on High Dimensional Data — Experiments, Analyses, and Improvement. IEEE Transactions on Knowledge and Data Engineering 32, 8 (2020), 1475–1488. https:/...
2020
-
[46]
In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages. https://doi.org/1...
-
[47]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Associa- tion for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013
2004
-
[48]
InProceedings of the 2020 ACM SIGMOD International Conference on Management of Data (Portland, OR, USA) (SIGMOD ’20)
Improving Approximate Nearest Neighbor Search through Learned Adaptive Early Termination. InProceedings of the 2020 ACM SIGMOD International Conference on Management of Data (Portland, OR, USA) (SIGMOD ’20). Association for Computing Machinery, New York, NY, USA, 2539–2554. ht...
2020
-
[49]
Shangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow, and Yang Liu
-
[50]
Locuza. 2022. Die Analysis: Samsung Exynos 2200 with RDNA2 Graph- ics. Online; accessed 2024-12-13. https://locuza.substack.com/p/die- analysis-samsung-exynos-2200 Derrick Quinn et al
2022
-
[51]
ArXiv abs/2202.01110 (2022)
A Survey on Retrieval-Augmented Text Generation. ArXiv abs/2202.01110 (2022). https://api.semanticscholar.org/CorpusID: 246472929
2022 arXiv
-
[52]
Malkov and D
Yu A. Malkov and D. A. Yashunin. 2020. Efficient and Ro- bust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 4 (apr 2020), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473
2020
-
[53]
Pierre Lienhart. 2024. LLM Inference Series: 4. KV caching, a deeper look. Pierre Leinhart (Medium) (Jan 2024). https://medium.com/@plienhar/llm-inference-series-4-kv- caching-a-deeper-look-4ba9a77746c8
2024
-
[54]
Timothy Prickett Morgan. 2024. He Who Can Pay Top Dollar For HBM Memory Controls AI Training. The Next Platform (2024). https://www.nextplatform.com/2024/02/27/he-who-can-pay-top- dollar-for-hbm-memory-controls-ai-training/ Accessed: 2024-06-23
2024
- [55]
-
[56]
OpenAI. 2023. ChatGPT plugins. OpenAI Blog (2023). https://openai.com/blog/chatgpt-plugins
2023
-
[57]
In International Conference on Learning Representations
Retrieval-Augmented Generation for Code Summarization via Hybrid GNN. In International Conference on Learning Representations. https://openreview.net/forum?id=zv-typ1gPxA
-
[58]
Md Rizwan Parvez, Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval Augmented Code Generation and Summarization. InFindings of the Association for Computational Linguistics: EMNLP 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Sco...
2021 doi
-
[59]
Loh, Natalie Enright Jerger, Ajaykumar Kannan, and Yasuko Eckert
Gabriel H. Loh, Natalie Enright Jerger, Ajaykumar Kannan, and Yasuko Eckert. 2015. Interconnect-Memory Challenges for Multi-chip, Silicon Interposer Systems. In Proceedings of the 2015 International Symposium on Memory Systems (Washington DC, DC, USA)(MEMSYS ’15). Association ...
2015
-
[60]
Dylan Patel and Jeremie Eliahou Ontiveros. 2024. CXL Is Dead In The AI Era. Online; accessed 2024-12-13. https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era
2024
-
[61]
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024. SpecInfer: Accelerating Large Language Model Serving w...
2024
-
[62]
Neel Patel, Amin Mamandipoor, Derrick Quinn, and Mohammad Alian. 2023. XFM: Accelerated Software-Defined Far Memory. In Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture (Toronto, ON, Canada) (MICRO ’23) . Asso- ciation for Computing Machiner...
2023
-
[63]
Samuel Naffziger, Kevin Lepak, Milam Paraschou, and Ma- hesh Subramony. 2020. 2.2 AMD Chiplet Architecture for High-Performance Server and Desktop Products. In 2020 IEEE International Solid- State Circuits Conference - (ISSCC) . 44–45. https://doi.org/10.1109/ISSCC19947.2020.9063103
2020
-
[64]
Hao Peng, Ankur Parikh, Manaal Faruqui, Bhuwan Dhingra, and Dipanjan Das. 2019. Text Generation with Exemplar-based Adaptive Decoding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog...
2019 doi
-
[66]
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a Benchmark for Knowledge Intensive Language Task...
2021
-
[67]
Dylan Patel. 2022. Apple M2 Die Shot and Architecture Analysis – Big Cost Increase And A15 Based IP.SemiAnalysis (June 2022). https: //www.semianalysis.com/p/apple-m2-die-shot-and-architecture
2022
-
[68]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. The Journal of Machine Learning Research 21, 1, Article 140 (j...
2020
-
[69]
Patel, A
N. Patel, A. Mamandipoor, M. Nouri, and M. Alian. 2024. SmartDIMM: In-Memory Acceleration of Upper Layer Protocols. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE Computer Society, Los Alamitos, CA, USA, 312–329. https://doi.org/10....
2024
-
[70]
Rita Ramos, Desmond Elliott, and Bruno Martins. 2023. Retrieval- augmented Image Captioning. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Andreas Vlachos and Isabelle Augenstein (Eds.). Association for Computati...
2023 doi
-
[71]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting . In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE Co...
2024
-
[73]
Fabio Petroni, Aleksandra Piktus, and Angela Fan. 2020. Introducing KILT, a new unified benchmark for knowledge-intensive NLP tasks. Online; accessed 2024-11-22. https://ai.meta.com/blog/introducing- kilt-a-new-unified-benchmark-for-knowledge-intensive-nlp- tasks/
2020
-
[74]
Alireza Salemi, Mahta Rafiee, and Hamed Zamani. 2023. Pre-Training Multi-Modal Dense Retrievers for Outside-Knowledge Visual Ques- tion Answering. InProceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (Taipei, Taiwan)(ICTIR ’23). Assoc...
2023
-
[75]
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2021 Conference of the North Amer...
2021
-
[77]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang
-
[78]
Alireza Salemi and Hamed Zamani. 2024. Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval- Augmented Large Language Models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington ...
2024
-
[79]
Satyabrata Sarangi and Bevan Baas. 2021. DeepScaleTool: A Tool for the Accurate Estimation of Technology Scaling in the Deep-Submicron Era. In 2021 IEEE International Symposium on Circuits and Systems (ISCAS). 1–5. https://doi.org/10.1109/ISCAS51556.2021.9401196
2021
-
[80]
Alireza Salemi, Juan Altmayer Pizzorno, and Hamed Zamani
-
[81]
In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23)
A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question Answering. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for Computing Mac...
- [82]
-
[83]
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024. LaMP: When Large Language Models Meet Per- sonalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martins, ...
2024 doi
-
[84]
Joonseop Sim, Soohong Ahn, Taeyoung Ahn, Seungyong Lee, Myunghyun Rhee, Jooyoung Kim, Kwangsik Shin, Donguk Moon, Euiseok Kim, and Kyoung Park. 2023. Computational CXL-Memory Solution for Accelerating Memory-Intensive Ap- plications. IEEE Computer Architecture Letters 22, 1 (2...
2023
-
[85]
Alireza Salemi and Hamed Zamani. 2024. Comparing Retrieval-Augmentation and Parameter-Efficient Fine-Tuning for Privacy-Preserving Personalization of Large Language Models. arXiv:2409.09510 [cs.CL] https://arxiv.org/abs/2409.09510
2024 arXiv
-
[86]
Heidi Steen and Dan Wahlin. 2023. Retrieval Augumented Generation Overview. Microsoft Learn (2023). https://learn.microsoft.com/en- us/azure/search/retrieval-augmented-generation-overview
2023
-
[87]
Alireza Salemi and Hamed Zamani. 2024. Learning to Rank for Multiple Retrieval-Augmented Models through Iterative Utility Max- imization. arXiv:2410.09942 [cs.CL] https://arxiv.org/abs/2410.09942
2024 arXiv
-
[88]
Su, Samuel Naffziger, and Mark Papermaster
Lisa T. Su, Samuel Naffziger, and Mark Papermaster. 2017. Multi-chip technologies to unleash computing performance gains over the next decade. In 2017 IEEE International Electron Devices Meeting (IEDM) . 1.1.1–1.1.8. https://doi.org/10.1109/IEDM.2017.8268306
2017
-
[89]
Yixuan Su, David Vandyke, Simon Baker, Yan Wang, and Nigel Collier. 2021. Keep the Primary, Rewrite the Secondary: A Two-Stage Approach for Paraphrase Generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, Wenjie Li...
2021 doi
-
[90]
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cuc- chiara. 2022. Retrieval-Augmented Transformer for Image Captioning. In Proceedings of the 19th International Conference on Content-Based Multimedia Indexing (Graz, Austria) (CBMI ’22). Association for Computing Machin...
2022
-
[91]
Schuh, Arvind Krishnamurthy, David Culler, Henry M
Henry N. Schuh, Arvind Krishnamurthy, David Culler, Henry M. Levy, Luigi Rizzo, Samira Khan, and Brent E. Stephens. 2024. CC-NIC: a Cache-Coherent Interface to the NIC. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages an...
2024
-
[92]
Zhiliang Tian, Wei Bi, Xiaopeng Li, and Nevin L. Zhang. 2019. Learn- ing to Abstract for Memory-augmented Conversational Response Generation. In Proceedings of the 57th Annual Meeting of the Associ- ation for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màr...
2019 doi
-
[93]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen...
2021 doi
-
[94]
Mengzhao Wang, Xiaoliang Xu, Qiang Yue, and Yuxiang Wang
-
[95]
Shamane Siriwardhana, Rivindu Weerasekera, Elliott Wen, Tharindu Kaluarachchi, Rajib Rana, and Suranga Nanayakkara. 2023. Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering. Transactions of the Association for Comp...
2023 doi
-
[96]
Zelin Wang, Ping Gong, Yibo Zhang, Jihao Gu, and Xuanyuan Yang
- [97]
-
[98]
WikiChip. 2024. Mask / Reticle. Online; accessed 2024-12-13. https://en.wikichip.org/wiki/mask
2024
-
[99]
Yu Wu, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li, and Ming Zhou. 2019. Response Generation by Context-Aware Prototype Edit- ing. Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 7281–7288. https://doi.org/10.1609/aaai.v33i01.33017281
2019 doi
- [100]
-
[101]
David Thulke, Nico Daheim, Christian Dugast, and Hermann Ney
- [102]
- [103]
-
[104]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin
-
[106]
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024. The Shift from Models to Compound AI Systems. Online; accessed 2024-12-13. https://bair.berkeley.edu/blog...
2024
-
[107]
Pro- ceedings of the VLDB Endowment 14, 11 (jul 2021), 1964–1978
A comprehensive survey and experimental comparison of graph-based approximate nearest neighbor search. Pro- ceedings of the VLDB Endowment 14, 11 (jul 2021), 1964–1978. https://doi.org/10.14778/3476249.3476255
2021
- [108]
-
[109]
Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura. 2018. Guiding Neural Machine Translation with Retrieved Translation Pieces. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: H...
2018 doi
-
[110]
In Natural Language Processing and Chinese Computing , Fei Liu, Nan Duan, Qingting Xu, and Yu Hong (Eds.)
Retrieval-Augmented Knowledge-Intensive Dialogue. In Natural Language Processing and Chinese Computing , Fei Liu, Nan Duan, Qingting Xu, and Yu Hong (Eds.). Springer Nature Switzerland, Cham, 16–28. https://doi.org/10.48550/arXiv.2005.11401
-
[111]
Jason Weston, Emily Dinan, and Alexander Miller. 2018. Retrieve and Refine: Improved Sequence Generation Models For Dialogue. In Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI, Aleksandr Chuklin, Jeff Dalton, Ju...
2018 doi
- [112]
-
[114]
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien De- mouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In Inter- national Conference on Machine Learning . PMLR, 38087–38099. https://doi.org/10.5555/361840...
2023
-
[115]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=NG7sS51zVF
2024
-
[116]
Jitao Xu, Josep Crego, and Jean Senellart. 2020. Boosting Neural Machine Translation with Similar Translations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Ass...
2020 doi
- [118]
-
[121]
Hamed Zamani and Michael Bendersky. 2024. Stochastic RAG: End-to- End Retrieval-Augmented Generation through Expected Utility Maxi- mization. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) ...
2024
-
[122]
Hamed Zamani, Fernando Diaz, Mostafa Dehghani, Donald Metzler, and Michael Bendersky. 2022. Retrieval-Enhanced Machine Learning. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22). Asso...
2022
-
[124]
Yunan Zhang, Shige Liu, and Jianguo Wang. [n. d.]. Are There Funda- mental Limitations in Supporting Vector Data Management in Rela- tional Databases? A Case Study of PostgreSQL.Preprint ([n. d.]). https: //www.cs.purdue.edu/homes/csjgwang/pubs/ICDE24_VecDB.pdf Accepted for pu...
-
[125]
Zhe Zhou, Cong Li, Fan Yang, and Guangyu Sun. 2023. DIMM- Link: Enabling Efficient Inter-DIMM Communication for Near- Memory Processing. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 302–316. https://doi.org/10.1109/HPCA56546.2023.1007...
2023
-
[2016]
In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Jian Su, Kevin Duh, and Xavier Carreras (Eds.)
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392. https://...
2016 doi
-
[2017]
In Advances in Neural Information Processing Systems, I
Attention is All you Need. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/ 2017/file/3f5...
2017
- [2020]
-
[2021]
The International FLAIRS Conference Proceedings 34, 1 (April 2021)
Retrieval-Augmented Transformer-XL for Close-Domain Dialog Generation. The International FLAIRS Conference Proceedings 34, 1 (April 2021). https://doi.org/10.32473/flairs.v34i1.128369
2021 doi
-
[2022]
InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.)
MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Li...
2022 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.