REVIEW 1 major objections 5 minor 97 references
D-NOVA claims RAG vector retrieval can be executed entirely inside 3D NAND flash arrays, using a new threshold-sensing metric that eliminates off-array re-ranking.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:41 UTC pith:DCQZFWSN
load-bearing objection A credible new in-array search mechanism with a thorough simulation study, but the headline gains rest on an unvalidated multi-wordline sensing mode. the 1 major comments →
D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
D-NOVA's central claim is that similarity search over dense vectors can be decomposed into a sequence of binary threshold checks that a NAND string performs natively: with multiple wordlines driven at query-dependent voltages, a string conducts only if every selected cell's threshold voltage lies below its applied voltage. The DTS metric alternates this upper-bound check with a complement-domain check, producing a tight score while staying fully digital. The architecture maps each embedding along a vertical string, stores a 4-bit complement copy, accumulates per-group pass/fail results as a 'deficit' in the existing page-buffer latches, and forwards only compact scores and metadata to the co
What carries the argument
Dual-Bound Tight Similarity Sensing (DTS), a digital distance proxy built from two complement upper-bound checks. UBS drives each of m wordlines at q_i + alpha and marks a group pass only if all cells conduct; Comp UBS does the same on 4-bit complement values (r' = 15 - r), replacing the loose lower-bound check that would otherwise cause false positives. The serial NAND string implements these as a single on/off decision, and per-window score deficits (S_max - S) are accumulated in the 4-bit data latches of the page buffer, avoiding wide adders. Multi-wordline activation and stage-dependent m (e.g., 1–2 for centroid/re-ranking, 4–8 for coarse search) provide the parallelism knob.
Load-bearing premise
The load-bearing premise is that real 3D NAND arrays can simultaneously drive multiple wordlines at distinct query-dependent voltages, with the serial string conduction faithfully implementing the per-cell threshold AND; the paper validates this with RC/SPICE simulation and prior patents, not with silicon measurement of this exact mode.
What would settle it
Program a real QLC die with known INT4 threshold levels, apply m=4–8 wordline voltages at q_i+alpha with alpha=2 under retention noise, and measure string-level conduction against the idealized AND; if per-cell sensing accuracy collapses below usable recall or multi-WL activation requires timing margins that erase the latency gain, the central claim fails.
If this is right
- Raw embedding vectors and partial scores never leave the NAND array; only compact candidate metadata crosses to the controller, removing the re-ranking memory wall.
- The IVF pipeline's three stages run in-storage: centroid selection, coarse Top-K2, and fine Top-K1, with controller sorting accounting for under 0.1% of latency in the final stage.
- DTS sensing replaces iterative multi-level QLC reads (7–15 sense operations) with a fixed small number of binary senses, giving SLC-like read latency at QLC density.
- Stage-aware m-adaptive sensing lets the coarse second stage run at m=8 with recall degradation confined to a few percent, since first and third stages stay narrow.
- A retrieval-aware adapter trained on DTS-mined hard negatives closes most of the recall gap to FP32 IVF, with only 0.5–1% query-encoding overhead.
Where Pith is reading between the lines
- If the multi-wordline sensing mode is validated on silicon, DTS-style bound checking could generalize beyond RAG to other in-storage filtering and similarity-join workloads that can tolerate a threshold score.
- The adapter recipe suggests a general pattern: when moving a continuous metric onto discrete in-memory hardware, a small learned projection can absorb the metric mismatch without changing storage-side logic.
- A testable scaling question the paper leaves open is how DTS recall behaves for very large nprobe and more heterogeneous embeddings; the current recall gap is measured at fixed K2=1000, K1=100.
- The reported energy benefits rely on WL toggling dominating energy; real dies with different peripheral costs could shift the balance, so per-die measurements of multi-WL activation energy are the next check.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. D-NOVA proposes an in-storage retrieval accelerator for RAG that executes the IVF pipeline inside 3D NAND. Embedding vectors are stored as INT4 values in QLC threshold-voltage levels along vertical NAND strings; the query is applied as wordline voltages, and strings are sensed in pass/fail mode to implement a new distance metric, Dual-Bound Tight Similarity Sensing (DTS). DTS combines an upper-bound check with a complement-domain upper-bound check, accumulates per-string score deficits in the page-buffer data latches, and transfers only compact scores and metadata to the SSD controller. A lightweight contrastive adapter, trained offline with DTS-mined hard negatives, maps encoder outputs into a DTS-friendly space. The three IVF stages (centroid search, coarse Top-K2, fine Top-K1 re-ranking) are claimed to run entirely within the NAND array. Evaluations on six datasets and two encoders with an in-house simulator report up to 41.7x lower latency and 71x lower energy than a CPU IVF baseline, and 4.3-12.1x higher throughput and 1.1-1.26x lower energy than the REIS in-storage baseline, with noise-aware Spectre simulations and m-sweep sensitivity studies.
Significance. The core idea is consequential: if the DTS sensing mode is physically realizable, D-NOVA would remove the re-ranking data-movement bottleneck that dominates REIS, and it would be a rare example of executing a full IVF retrieval pipeline inside NAND. The work also has genuine strengths: strict digital sensing avoids analog current accumulation; the contrastive-adapter co-design is a legitimate way to bridge a discrete threshold metric and pretrained embeddings; the stage-aware m-adaptive sensing and locality-aware mapping are well-motivated; and the noise simulations and m-sweeps are unusually thorough. However, the significance is conditional on three load-bearing issues: the unvalidated assumption of simultaneous per-WL query-adaptive voltages, an internal inconsistency in the DTS scoring equation and the Smax/4-bit deficit argument, and the underspecified adapter evaluation protocol. The paper does not release its simulator or code, which limits reproducibility of the headline numbers.
major comments (1)
- [§4.2, §5.1, Table 3] The adapter evaluation is underspecified. The paper does not state whether the queries used to train the adapter and to mine DTS hard negatives are disjoint from the evaluation queries. If DTS retrieval results and ground-truth labels from the reported benchmarks are used during training, the 'near-FP32-IVF' recall and the adapter gains in Table 3 and Fig. 11 are partly fitted and are not an independent accuracy measurement. Please specify the train/validation/test split for each benchmark, the negative-pool source and size, the InfoNCE temperature value, and report recall on held-out queries. If a disjoint split is already used, state it explicitly; if not, the accuracy claims must be re-evaluated.
minor comments (5)
- [§4.2 / Fig. 9] The InfoNCE loss equation and surrounding text are garbled with unicode artifacts, and Fig. 9 appears twice with different text. The equation is reconstructible, but the manuscript must be cleaned before publication.
- [§3.3 / Fig. 8] Figure 8 is duplicated with corrupted labels. Please replace the second copy and ensure all callouts match the figure contents.
- [§5.2 / Table 1] The simulator configuration lists the QLC threshold window and PTM models, but does not give the 16-level VTH distribution parameters or the exact RC model inputs. Adding the full configuration would help reproducibility.
- [§6.7] The statement that prior work [58] validates activation of up to 48 wordlines should be qualified: that work uses a uniform read voltage for bulk bitwise operations, which does not validate per-WL query-dependent voltages. The current sentence overstates the support provided by [58].
- [§6.6] The end-to-end latency/energy results do not state whether the host-side query adapter inference is included. The paper says adapter overhead is 0.5-1% of encoding time; please state explicitly whether it is counted in the reported numbers.
Circularity Check
No significant circularity: DTS is a newly introduced construction, the adapter is a co-design/training component rather than a disguised prediction, and headline speedups rest on an external feasibility assumption, not on a definitional identity.
full rationale
D-NOVA's central derivation is not circular. The DTS metric is explicitly defined as a new threshold-based score (UBS and Comp UBS) tailored for NAND string sensing, and the paper states: "D-NOVA does not implement cosine similarity or L2 distance inside the NAND array; instead, each IVF stage uses DTS as a threshold-based retrieval score." Thus DTS is not a renamed version of cosine/L2, and no equation in the paper reduces to its own input by construction. The contrastive adapter is trained offline using DTS-mined hard negatives and InfoNCE loss; the reported Recall@100 is measured against ground-truth retrieval quality, not against the DTS score itself. Training a query adapter to improve a metric-specific search is a legitimate co-design, not a fitted parameter being relabeled as a prediction. The multi-WL activation support is cited from external patents and Flash-Cosmos, and the paper acknowledges its own validation is via Spectre/RC simulations plus prior silicon for bulk bitwise operations. Even if distinct per-WL query-dependent voltages are not directly validated in real silicon, that is a feasibility gap, not a circular reduction. The only self-citations (FeNOMS, Proxima) appear in background or routing references and are not load-bearing for the central claim. The paper is self-contained in its evaluation against an FP32-IVF CPU baseline and the REIS in-storage baseline, and the reported speedups follow from the assumed DTS sensing mechanism rather than from a definitional tautology.
Axiom & Free-Parameter Ledger
free parameters (3)
- DTS tolerance alpha =
2 (default; swept 1–3)
- Stage-wise parallel sensing degrees m=(m1,m2,m3) =
(2,4,2), (2,8,2), (4,8,4)
- Adapter InfoNCE temperature and negative-pool size =
not reported
axioms (4)
- domain assumption 3D NAND supports activating multiple wordlines simultaneously with distinct read voltages while preserving binary string conductance logic.
- domain assumption INT4 embedding dimensions can be stored as ordered V_TH levels and queried by applying voltage q_i+alpha per WL without significant inter-cell interference.
- ad hoc to paper DTS score (threshold-bound pass count) is a sufficient ranking function for RAG after query-side adaptation.
- domain assumption Adapter trained with DTS-mined hard negatives generalizes to unseen queries/datasets.
read the original abstract
Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. However, its dense vector retrieval introduces significant latency and energy overhead, becoming the primary performance bottleneck. Although recent in-storage accelerators aim to reduce data movement, they still rely on host or embedded processors outside the memory, where nearly 70% of the total retrieval time is spent. As a result, they cannot fully overcome the bandwidth limitations, leading to yet another memory bottleneck. To tackle these limitations, we present D-NOVA, a hardware-software co-designed in-storage retrieval accelerator. D-NOVA executes an inverted file (IVF)-based hierarchical retrieval pipeline by deeply embedding the search functionality directly into the NAND memory array. This is achieved by incorporating a new distance metric, Dual-Bound Tight Similarity Sensing (DTS), which is specifically tailored for searching within the NAND string. In addition, we introduce a lightweight contrastive adapter that maps embedding vectors into a DTS-friendly domain, recovering near-software recall while improving performance and energy efficiency. D-NOVA is up to 41.7x faster and 71x more energy-efficient than a CPU baseline, and achieves 12.13x higher throughput while being up to 1.26x more energy-efficient than state-of-the-art in-storage RAG accelerators, demonstrating the potential of fully in-storage vector search for scalable RAG acceleration.
Figures
Reference graph
Works this paper leans on
-
[1]
Advanced Micro Devices, Inc. 2025. AMD uProf: Performance Analysis Tool. https://www.amd.com/en/developer/uprof.html
2025
-
[2]
Mohiuddin Ahmed, Raihan Seraj, and Syed Mohammed Shamsul Islam. 2020. The k-means algorithm: A comprehensive survey and performance evaluation. Electronics9, 8 (2020), 1295
2020
-
[3]
Toluwalope Ajayi et al. 2019. OpenROAD: Toward a Self-Driving, Open-Source Digital Layout Implementation Tool Chain. InDAC
2019
-
[4]
2023.AMD EPYC™9554
AMD. 2023.AMD EPYC™9554. https://www.amd.com/en/products/processors/ server/epyc/4th-generation-9004-and-8004-series/amd-epyc-9554.html Ac- cessed: 2025-11-16
2023
-
[5]
AMD Adaptive & Embedded Computing Group and Samsung. 2023. SmartSSD®Computational Storage Drive: Product Brief. https: //www.xilinx.com/publications/product-briefs/xilinx-smartssd-computational- storage-drive-product-brief.pdf. Accessed Nov. 2025
2023
-
[6]
Woorham Bae, Sung-Yong Cho, and Deog-Kyoon Jeong. 2021. A 1.93-pJ/Bit PCI Express Gen4 PHY Transmitter with On-Chip Supply Regulators in 28 nm CMOS. Electronics(2021). https://api.semanticscholar.org/CorpusID:234325987
2021
-
[7]
Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas
Rajeev Balasubramonian, Andrew B. Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories.ACM Transactions on Architecture and Code Optimization14, 2 (2017). doi:10.1145/3092639 Accessed: 2025-11-16
-
[8]
Haratsch, Yixin Luo, and Onur Mutlu
Yu Cai, Saugata Ghose, Erich F. Haratsch, Yixin Luo, and Onur Mutlu. 2017. Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives.Proc. IEEE105, 9 (2017), 1666–1704. doi:10.1109/JPROC.2017.2713127
arXiv 2017
-
[9]
Jianlyu Chen, Nan Wang, Chaofan Li, Bo Wang, Shitao Xiao, Han Xiao, Hao Liao, Defu Lian, and Zheng Liu. 2025. AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Tah...
2025
-
[10]
Kangqi Chen, Rakesh Nadig, Manos Frouzakis, Nika Mansouri Ghiasi, Yu Liang, Haiyu Mao, Jisung Park, Mohammad Sadrosadati, and Onur Mutlu. 2025. REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Pro- cessing. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing ...
arXiv 2025
-
[11]
Mingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang, Kuai Yu, Lei Dai, Ziming Yuan, Ying Wang, Lei Zhang, Huawei Li, and Xiaowei Li. 2025. DRIM- ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’25). Associat...
arXiv 2025
-
[12]
Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. In35th Conference on Neural Information Processing Systems (NeurIPS 2021)
2021
-
[13]
Sitian Chen, Amelie Chi Zhou, Yucheng Shi, Yusen Li, and Xin Yao. 2024. Me- mANNS: Enhancing Billion-Scale ANNS Efficiency with Practical PIM Hardware. arXiv:2410.23805. https://arxiv.org/abs/2410.23805
Pith/arXiv arXiv 2024
-
[15]
Hwanheechan Choi, Hyungjun Jo, Sangmin Ahn, Insang Han, and Hyungcheol Shin. 2026. Machine learning-based prediction of the impact of random grain boundary Z-interference on Vt distribution in 3-D NAND flash memory.Journal of Computational Electronics25, 1 (2026), 16
2026
-
[16]
Myungjun Chun, Jaeyong Lee, Sanggu Lee, Myungsuk Kim, and Jihong Kim
-
[17]
Lawrence T Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric. 2016. ASAP7: A 7-nm finFET predictive process design kit.Microelectronics Journal53 (July 2016), 105–115
2016
-
[18]
Cohere. 2023. wikipedia-2023-11-embed-multilingual-v3. https://huggingface. co/datasets/Cohere/wikipedia-2023-11-embed-multilingual-v3. Hugging Face Datasets
2023
-
[19]
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2025. THE FAISS LIBRARY.IEEE Transactions on Big Data(2025), 1–17. doi:10.1109/ TBDATA.2025.3618474
arXiv 2025
-
[20]
2021.PM9A3 NVMe PCIe SSD
Samsung Electronics. 2021.PM9A3 NVMe PCIe SSD. https://semiconductor. samsung.com/ssd/datacenter-ssd/pm9a3/ Accessed: 2025-11-16
2021
-
[21]
Keming Fan, Ashkan Moradifirouzabadi, Xiangjin Wu, Zheyu Li, Flavio Ponzina, Anton Persson, Eric Pop, Tajana Rosing, and Mingu Kang. 2024. SpecPCM: A Low-Power PCM-Based In-Memory Computing Accelerator for Full-Stack Mass Spectrometry Analysis.IEEE Journal on Exploratory Solid-State Computational Devices and Circuits10 (2024), 161–169. doi:10.1109/JXCDC.2...
arXiv 2024
-
[22]
Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, and Jie Wu
-
[23]
Hakim Hafidi, Mounir Ghogho, Philippe Ciblat, and Ananthram Swami. 2022. Negative sampling strategies for contrastive self-supervised learning of graph representations.Signal Processing190 (2022), 108310
2022
-
[24]
Tsutomu Higuchi, Takuyo Kodama, Koji Kato, Ryo Fukuda, Naoya Tokiwa, Mitsuhiro Abe, Teruo Takagiwa, Yuki Shimizu, Junji Musha, Katsuaki Sakurai, Jumpei Sato, Tetsuaki Utsumi, Kazuhide Yoneya, Yasuhiro Suematsu, Toshifumi Hashimoto, Takeshi Hioka, Kosuke Yanagidaira, Masatsugu Kojima, Junya Mat- suno, Kei Shiraishi, Kensuke Yamamoto, Shintaro Hayashi, Tomo...
arXiv 2021
-
[25]
Charles AR Hoare. 1962. Quicksort.The computer journal5, 1 (1962), 10–16
1962
-
[26]
Po-Kai Hsu, Weihong Xu, Tajana Rosing, and Shimeng Yu. 2023. An in-storage processing architecture with 3d nand heterogeneous integration for spectra open modification search. InProceedings of the International Symposium on Memory Systems. 1–7
2023
-
[27]
Zhengding Hu, Vibha Murthy, Zaifeng Pan, Wanlu Li, Xiaoyi Fang, Yufei Ding, and Yuke Wang. 2025. HedraRAG: Co-Optimizing Generation and Retrieval for Heterogeneous RAG Workflows. InProceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. 623–638
2025
-
[28]
Bongjoon Hyun, Taehun Kim, Dongjae Lee, and Minsoo Rhu. 2024. Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 263–279. doi:10.1109/HPCA57654.2024.00029
arXiv 2024
-
[29]
2023.LPDDR4X/LPDDR4 SDRAM MT53E768M64D4, MT53E1536M64D8, MT53E768M32D2, MT53E1536M32D4 Data Sheet
Micron Technology Inc. 2023.LPDDR4X/LPDDR4 SDRAM MT53E768M64D4, MT53E1536M64D8, MT53E768M32D2, MT53E1536M32D4 Data Sheet. https://www. 13 MICRO 2026, October 31–November 04, 2026, Athens, Greece Chang Eun Song, Sumukh Pinge, Tianqi Zhang, Sung Eun Kim, Tajana S Rosing, and Mingu Kang mouser.com/datasheet/2/671/z4bm_embedded_lpddr4x_lpddr4-3193428.pdf Rev....
2023
-
[30]
Yeonwoo Jeong, Hyunji Cho, Kyuri Park, Youngjae Kim, and Sungyong Park. 2025. CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases. arXiv preprint arXiv:2509.18670(2025)
arXiv 2025
-
[31]
Wenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso, Amir Yazdan- bakhsh, and Vidushi Dadu. 2025. Rago: Systematic performance optimization for retrieval-augmented generation serving. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 974–989
2025
-
[32]
Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Shufan Liu, Xuanzhe Liu, and Xin Jin. 2024. Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Systems(2024)
2024
-
[33]
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs.IEEE Transactions on Big Data7, 3 (2019), 535–547
2019
-
[34]
Norman P Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, et al. 2021. Ten lessons from three generations shaped google’s tpuv4i: Industrial product. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 1–14
2021
-
[35]
Ali Khakifirooz, Sriram Balasubrahmanyam, Richard Fastow, Kristopher H Gaewsky, Chang Wan Ha, Rezaul Haque, Owen W Jungroth, Steven Law, Alias- gar S Madraswala, Binh Ngo, et al. 2021. 30.2 a 1tb 4b/cell 144-tier floating-gate 3d-nand flash memory with 40mb/s program throughput and 13.8 gb/mm 2 bit density. In2021 IEEE International Solid-State Circuits C...
2021
-
[36]
Hyun-Jin Kim, Jeong-Don Lim, Jang-Woo Lee, Dae-Hoon Na, Joon-Ho Shin, Chae-Hoon Kim, Seung-Woo Yu, Ji-Yeon Shin, Seon-Kyoo Lee, Devraj Rajagopal, et al. 2015. 7.6 1gb/s 2tb nand flash multi-chip package with frequency-boosting interface chip. In2015 IEEE International Solid-State Circuits Conference-(ISSCC) Digest of Technical Papers. IEEE, 1–3
2015
-
[37]
Ji-Hoon Kim, Yeo-Reum Park, Jaeyoung Do, Soo-Young Ji, and Joo-Young Kim
-
[38]
Kana Kudo, Yuta Aiba, Kazuma Hasegawa, Xu Li, Yuichi Sano, and Tomoya Sanuki. 2025. Energy-Efficient In-Memory Computing using 3D Flash Memory with Sequential Multi-Block Activation and Current Control Cell (CC cell). In 2025 IEEE International Memory Workshop (IMW). 1–4. doi:10.1109/IMW61990. 2025.11026979
arXiv 2025
-
[39]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural Questions: A Benchmark for Question Answering Research.Tr...
-
[40]
Comput.72, 1 (2022), 278–290
Accelerating large-scale graph-based nearest neighbor search on a compu- tational storage platform.IEEE Trans. Comput.72, 1 (2022), 278–290
2022
-
[41]
Joo Hwan Lee, Hui Zhang, Veronica Lagrange, Praveen Krishnamoorthy, Xi- aodong Zhao, and Yang Seok Ki. 2020. SmartSSD: FPGA Accelerated Near-Storage Data Analytics on SSD.IEEE Computer Architecture Letters19, 2 (2020), 110–113. doi:10.1109/LCA.2020.3009347
arXiv 2020
-
[42]
Kyungmin Lee, Gunwook Yoon, Seung Jae Baik, and Myounggon Kang. 2026. Low-Power Stack-Level Programming Enabled by Optimized Dummy Word Line Voltage in 3-D NAND Flash Memory.IEEE Journal of the Electron Devices Society 14 (2026), 102–106. doi:10.1109/JEDS.2026.3659350
arXiv 2026
-
[43]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles. 611–626
2023
-
[44]
Nancy Leong, Sachit Chandra, and Hounien Chen. 2008. Random cache read using a double memory. US Patent 7,423,915
2008
-
[45]
Yinan Li, Bailu Ding, Ziyun Wei, Lukas M Maas, Momin Al-Ghosien, Spyros Blanas, Nicolas Bruno, Carlo Curino, Matteo Interlandi, Craig Peeper, et al. 2025. Scaling GPU-Accelerated Databases beyond GPU Memory Size.Proceedings of the VLDB Endowment18, 11 (2025), 4518–4531
2025
-
[46]
Peter Wung Lee. 2015. NAND array hierarchical bit-line structures for multiple word-line and all-bit-line simultaneous erase, erase-verify, program, program-verify, and read operations. https://patents.google.com/patent/ WO2015013689A2/en
2015
-
[48]
Hosam M Mahmoud, Reza Modarres, and Robert T Smythe. 1995. Analysis of quickselect: An algorithm for order statistics.RAIRO-Theoretical Informatics and Applications29, 4 (1995), 255–276
1995
-
[49]
Arm Ltd. 2016. Cortex-R8. https://www.arm.com/products/silicon-ip-cpu/cortex- r/cortex-r8. Accessed: 2025-11-16
2016
-
[50]
Conrado Martínez and Salvador Roura. 2001. Optimal sampling strategies in quicksort and quickselect.SIAM J. Comput.31, 3 (2001), 683–705
2001
-
[51]
Micron Technology
Inc. Micron Technology. 2025.DDR4 SDRAM. https://www.micron.com/products/ memory/dram-components/ddr4-sdram Accessed: 2025-11-16
2025
-
[52]
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Finan- cial Opinion Mining and Question Answering. InCompanion Proceedings of the The Web Conference 2018. Lyon, France, 1941–1942. doi:10.1145/3184558.3192301
arXiv 2018
-
[53]
Seock-Hwan Noh, Hoyeon Lee, Junkyum Kim, Junsu Im, Jay H Park, Sungjin Lee, Sam H Noh, Yeseong Kim, and Jaeha Kung. 2025. Flexible In-NAND Cryp- tographic Processing for Secure Flash Storage.arXiv preprint arXiv:2508.03866 (2025)
arXiv 2025
-
[54]
Yoshiaki Ogura et al . 2003. Semiconductor memory device and method for selecting multiple word lines. https://patents.google.com/patent/JP2003222422A Laid-open patent application
2003
-
[55]
Daehoon Na, Jang-woo Lee, Seon-Kyoo Lee, Hwasuk Cho, Junha Lee, Manjae Yang, Eunjin Song, Anil Kavala, Tongsung Kim, Dong-Su Jang, et al . 2021. A 1.8-Gb/s/pin 16-Tb NAND flash memory multi-chip package with F-chip for high-performance and high-capacity storage.IEEE Journal of Solid-State Circuits 56, 4 (2021), 1129–1140
2021
-
[56]
Nikolaos Papandreou, Haralampos Pozidis, Nikolas Ioannou, Thomas Parnell, Roman Pletka, Milos Stanisavljevic, Radu Stoica, Sasa Tomic, Patrick Breen, Gary Tressler, et al. 2020. Open block characterization and read voltage calibration of 3D QLC NAND flash. In2020 IEEE International Reliability Physics Symposium (IRPS). IEEE, 1–6
2020
-
[57]
Krishna Parat and Chuck Dennison. 2015. A floating gate based 3D NAND technology with CMOS under array. In2015 IEEE International Electron Devices Meeting (IEDM). IEEE, 3–3
2015
-
[58]
Hiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang, Tamas Feher, and Yong Wang. 2024. Cagra: Highly parallel graph construction and approximate nearest neighbor search for gpus. In2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 4236–4247
2024
-
[59]
Jisung Park, Myungsuk Kim, Myoungjun Chun, Lois Orosa, Jihong Kim, and Onur Mutlu. 2021. Reducing solid-state drive read latency by optimizing read-retry. InProceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems(Virtual, USA)(ASPLOS ’21). Association for Computing Machinery, New York, ...
arXiv 2021
-
[60]
Advait Parulekar, Liam Collins, Karthikeyan Shanmugam, Aryan Mokhtari, and Sanjay Shakkottai. 2023. Infonce loss provably learns cluster-preserving rep- resentations. InThe Thirty Sixth Annual Conference on Learning Theory. PMLR, 1914–1961
2023
-
[61]
Jisung Park, Roknoddin Azizi, Geraldo F Oliveira, Mohammad Sadrosadati, Rakesh Nadig, David Novo, Juan Gómez-Luna, Myungsuk Kim, and Onur Mutlu
-
[62]
In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
Flash-cosmos: In-flash bulk bitwise operations using inherent computation capability of nand flash memory. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 937–955
-
[63]
Derrick Quinn, Mohammad Nouri, Neel Patel, John Salihu, Alireza Salemi, Sukhan Lee, Hamed Zamani, and Mohammad Alian. 2025. Accelerating retrieval- augmented generation. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume
2025
-
[64]
D. J. Rosenkrantz, R. E. Stearns, and P. M. Lewis. 1977. An Analysis of Several Heuristics for the Traveling Salesman Problem.SIAM J. Comput.6, 3 (1977), 563–581. doi:10.1137/0206041
doi:10.1137/0206041 1977
-
[65]
Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H Pantha, Po-Kai Hsu, Zheyu Li, Weihong Xu, Zihan Xia, Flavio Ponzina, et al . 2025. FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) Flash. In2025 IEEE/ACM International Conference On Computer Aide...
2025
-
[66]
Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023. FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction. InProceedings of the 50th Annual International Symposium on Computer Architecture(Orlando, FL, USA)(ISCA ’23). Association for Computing Machiner...
arXiv 2023
-
[67]
Edward Suh, and Udit Gupta
Michael Shen, Muhammad Umar, Kiwan Maeng, G. Edward Suh, and Udit Gupta
-
[68]
Joobo Shim, Jaewon Oh, Hongchan Roh, Jaeyoung Do, and Sang-Won Lee. 2025. Turbocharging Vector Databases Using Modern SSDs.Proc. VLDB Endow.18, 11 (July 2025), 4710–4722. doi:10.14778/3749646.3749724
arXiv 2025
-
[69]
Sayed Ahmad Salehi. 2022. In-memory Bulk Bitwise Logic Operation for Multi- level Cell Non-volatile Memories. InProceedings of the 2022 International Sympo- sium on Memory Systems. 1–5
2022
-
[70]
Eran Sharon et al. 2014. Simultaneous sensing of multiple word-lines and detec- tion of NAND failures. https://patents.google.com/patent/EP2737487A1/en
2014
-
[71]
Chang Eun Song, Priyansh Bhatnagar, Zihan Xia, Nam Sung Kim, Tajana S Rosing, and Mingu Kang. 2025. Hybrid SLC-MLC RRAM Mixed-Signal Processing-in- Memory Architecture for Transformer Acceleration via Gradient Redistribution. InProceedings of the 52nd Annual International Symposium on Computer Architec- ture. 1155–1170
2025
-
[72]
InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25)
Hermes: Algorithm-System Co-design for Efficient Retrieval-Augmented Generation At-Scale. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing Machinery, New York, NY, USA, 958–973. doi:10.1145/3695053.3731076 14 D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Sim...
arXiv 2026
-
[73]
Chang Eun Song, Ashkan Moradifirouzabadi, Tajana Rosing, and Mingu Kang
-
[74]
Wonbo Shim, Hongwu Jiang, Xiaochen Peng, and Shimeng Yu. 2021. Architectural Design of 3D NAND Flash based Compute-in-Memory for Inference Engine. In Proceedings of the International Symposium on Memory Systems(Washington, DC, USA)(MEMSYS ’20). Association for Computing Machinery, New York, NY, USA, 77–85. doi:10.1145/3422575.3422779
arXiv 2021
-
[75]
Tinku Singh, Durgesh Kumar Srivastava, and Alok Aggarwal. 2017. A novel approach for CPU utilization on a multicore paradigm using parallel quicksort. In 2017 3rd International Conference on Computational Intelligence & Communication Technology (CICT). IEEE, 1–6
2017
-
[76]
Yoshiki Takai, Mamoru Fukuchi, Reika Kinoshita, Chihiro Matsui, and Ken Takeuchi. 2019. Analysis on heterogeneous SSD configuration with quadruple- level cell (QLC) NAND flash memory. In2019 IEEE 11th International Memory Workshop (IMW). IEEE, 1–4
2019
-
[77]
Chang Eun Song, Yidong Li, Amardeep Ramnani, Pulkit Agrawal, Purvi Agrawal, Sung-Joon Jang, Sang-Seol Lee, Tajana Rosing, and Mingu Kang. 2024. 52.5 TOPS/W 1.7 GHz Reconfigurable XGBoost Inference Accelerator Based on Modular-Unit-Tree with Dynamic Data and Compute Gating. In2024 IEEE Custom Integrated Circuits Conference (CICC). IEEE, 1–2
2024
-
[78]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal
-
[79]
Bing Tian, Haikun Liu, Zhuohui Duan, Xiaofei Liao, Hai Jin, and Yu Zhang. 2024. Scalable billion-point approximate nearest neighbor search using SmartSSDs. In Proceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference (Santa Clara, CA, USA)(USENIX ATC’24). USENIX Association, USA, Article 69, 16 pages
2024
-
[80]
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. MPNet: Masked and Permuted Pre-training for Language Understanding. InAdvances in Neural Information Processing Systems 33 (NeurIPS 2020). https://proceedings.neurips. cc/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf
2020
-
[81]
Stillmaker and B
A. Stillmaker and B. Baas. 2017. Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm.Integration, the VLSI Jour- nal58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration. TechScale/
2017
-
[82]
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.