REVIEW 3 major objections 5 minor 125 references
XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read XBOF claims that a JBOF with half the per-SSD computing resources can match a fully equipped JBOF's throughput, by having idle SSDs lend their processors and DRAM over cache-coherent CXL to handle busy SSDs' metadata work.
desk verdict A genuinely new idea for inter-SSD compute harvesting over CXL, well evaluated, but the processor-harvesting path has a real crash-consistency gap and the headline numbers rest on a self-built simulator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the combination of disaggregated SSD internals with a cache-coherent CXL 3.0 fabric. Each SSD is split into a compute-end (processor and DRAM, running firmware such as address translation) and a data-end (flash channels, DMA engine, data buffer), and the two parts are independently visible to peers. Idle resource descriptors written into globally shared DRAM let a busy SSD atomically claim a lender's processor or DRAM segments. Processor harvesting works by NVMe queue-pair binding: a borrower queue pair is paired with a lender's shadow queue pair, and a weighted-round-robin load-balance formula splits redirected commands in proportion to both sides' processor utilization. Lender-side execution of borrower firmware is possible only because CXL keeps the borrower's address-translation mapping directory and table coherent in global fabric-attached memory, so the lender can read and update them with ordinary load/store operations. DRAM harvesting uses an online miss-ratio curve predictor to decide which segments can be lent or borrowed, plus redo logs flushed back to the borrower to protect offsite metadata; the paper reports these operations cost hundreds of nanoseconds, small relative to flash I/O.
What would settle it
Measure the round-trip latency and coherence-invalidation cost of a cache-coherent CXL peer-to-peer access between two real SSD controllers, then replay a read-dominated production trace on a two-SSD XBOF prototype: if borrower throughput falls below a conventional fully equipped JBOF by more than the claimed negligible margin, or if lender throughput drops well beyond the reported ~1.3%, the central claim is falsified.
Extended reading notes
Core claim
The paper claims that inter-SSD resource sharing over a cache-coherent fabric can substitute for per-SSD computing resources. XBOF disaggregates each SSD into a compute-end (processor and DRAM, which run firmware such as address translation) and a data-end (flash channels, DMA engine, data buffer), and exposes them separately over CXL. When one SSD's processor is saturated, typically during read bursts, the modified host NVMe (standard SSD command protocol) driver redirects a fraction of its I/O commands to an idle SSD's shadow queue pair; the lender uses the borrower's mapping tables, kept coherent in the global CXL memory space, to do command parsing and address translation, then sends DMA and flash operations back to the borrower's data-end for the actual data movement. When an SSD's onboard DRAM cannot hold enough of its mapping table, it caches table segments in a lender's DRAM with a log-based crash-consistency protocol. The evaluation claims that with half the per-SSD processors and DRAM, XBOF matches the throughput of a conventional fully equipped JBOF across microbenchmarks and production traces, improves SSD resource utilization by 50.4%, and saves 19.0% of SSD BOM cost, while the lender's own performance loss is about 1.3%.
Load-bearing premise
The whole benefit rests on the assumption that real CXL cache-coherent device-to-device access between SSDs has roughly the latency, bandwidth, and coherence cost that the paper's simulator assigns it; if remote mapping-table access and cache invalidation are several times slower on real hardware, the harvested throughput advantage shrinks toward the overhead.
Editorial extensions
If this is right
- A JBOF with the same peak throughput can be built for roughly 19% lower SSD bill-of-material cost, so cloud storage suppliers can either serve more capacity per dollar or meet the same performance targets with cheaper hardware.
- Read-dominated workloads, which data-redirection harvesting cannot accelerate, become shareable: the lender helps with metadata processing while data still flows from the borrower's own flash.
- Lender reclamation no longer requires copying written data back, so the extra writes and SSD wear that plague data-redirection harvesting disappear.
- Because resource management is decentralized through globally visible descriptors, the JBOF host's weak CPU need not become the bottleneck of the whole enclosure.
- The changes stay inside the NVMe driver and SSD firmware, so mainstream operating systems and applications see a normal NVMe device.
Reading between the lines
- The same compute/data disaggregation could be applied one level up: with CXL's multi-level switching, an idle rack's SSD processors and DRAM could serve a burst on another rack, effectively pooling metadata processing across the data center rather than just one enclosure.
- If real CXL silicon arrives with higher coherence-invalidation costs than the simulator assigns, the margin that should be tested first is the sub-microsecond remote mapping-table access; the paper's FPGA prototype cannot yet bound that cost because it states that public CXL 3.0 hardware is unavailable.
- The mechanism is not limited to SSDs: any metadata-hungry device with bursty compute demand, such as computational storage drives, smart NICs, or memory-semantic SSDs, could borrow peer compute for address translation or index lookup while its own data path stays local.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes XBOF, a JBOF design that disaggregates each SSD into a compute-end (processor and DRAM) and a data-end (flash and DMA), and uses CXL 3.0 cache-coherent peer-to-peer communication so that idle SSDs can lend their processors and DRAM to busy SSDs. The claimed benefits are a 50.4% improvement in SSD resource utilization, a 19.0% reduction in BOM cost, and performance comparable to a conventional JBOF while using only half the per-SSD computing resources. The evaluation combines SimpleSSD-based simulation with an integrated CXL simulator (ESF), microbenchmarks and production traces, sensitivity studies, a 10-run complex scenario, and a NUMA-based end-to-end prototype.
Significance. If the central claim holds, XBOF is a significant result: it offers a concrete path to reduce JBOF BOM costs substantially while preserving performance, and it is one of the first system designs to exploit CXL 3.0 Type-2 coherent peer-to-peer access for storage resource harvesting. The paper has notable strengths: the evaluation is broad and uses third-party components where possible (SimpleSSD, public traces), the design narrative is internally consistent, and the authors provide a hardware prototype of part of the firmware path. The headline percentages are simulation outputs rather than fitted constants, and the workloads are public traces, which makes the evaluation largely falsifiable. The main weaknesses are the PLP gap in the processor-harvesting path and the dependence of the result on CXL hardware parameters that have not yet been validated on real Type-2 CXL 3.0 hardware.
major comments (3)
- [§4.4] Section 4.4 describes the processor-harvesting path, in which the lender executes FTL address translation and updates the borrower's mapping table via CXL load/store. The crash-consistency mechanism in §4.5 covers only the DRAM-harvesting path, where offsite metadata in the lender's DRAM is protected with log pages and cacheline flush instructions (DCCSW). The processor-harvesting path has no analogous flush or fence before I/O completion, so dirty mapping lines held in the lender's write-back cache are not reachable by the borrower's power-loss protection circuit. On borrower power loss, the completed flash write can be orphaned or the mapping can revert, violating the PLP requirement that §4.5 itself states. This is a load-bearing correctness gap for enterprise deployment and must be addressed in the design.
- [§4.6, §5.1, §5.4] The evaluation of the core benefit depends on the ESF CXL simulator (ref [4]) for the latency and coherence cost of Type-2 CXL peer-to-peer access, and §4.6 acknowledges that no CXL 3.0 hardware is publicly available. The central claim of XBOF rests on sub-microsecond coherent remote access and small lock/invalidation overhead, yet §5.4 varies only processor cores, DRAM capacity, and borrower/lender ratios; it never varies CXL latency, bandwidth, or coherence overhead. Please add a sensitivity study sweeping these parameters and, if possible, validate ESF against a real CXL 3.0 Type-2 device or against published hardware measurements of such devices.
- [§5.2, Abstract] The abstract claims a 50.4% resource-utilization improvement as a general result, but the 50.4% figure in §5.2 is reported for a single workload (256 KB sequential read, Figure 9c). Similarly, the claim of 'comparable performance to Conv in all workloads' is presented as an average over many workloads, but Figure 11 shows per-workload differences without confidence intervals or statistical significance. Please report per-workload deltas and variance, and qualify the utilization and performance claims so they are not over-generalized.
minor comments (5)
- [§3.2] The challenge numbering is confusing: 'Challenge 2 and 3.1' and 'Challenge 3.2' suggest that Challenge 3 has sub-challenges, but the list in §3.1 presents only three numbered challenges. Please renumber for consistency.
- [§5.2] The BOM cost saving of 19.0% rests on assumed market prices and a 10% premium for CXL-enabled controllers and DRAM taken from prior work (ref [95]). A short sensitivity analysis over these price inputs would materially strengthen the cost-efficiency claim.
- [§4.4] The load-balance formula uses variables such as W_shadowSQ and W_borrowerSQ but the notation is not fully explained in the text, making the formula hard to reproduce. Please expand the variable definitions.
- [Figure 9c] Figure 9c is difficult to read: the x-axis is labeled 'Time (ms)', but the harvesting-start marker and the relationship between the time axis and the workload throughput curves are not clearly explained in the caption.
- [General] The paper does not include an artifact availability statement. Given the 18K LOC simulator extension and the 1K/2K LOC driver and firmware modifications, releasing these artifacts would substantially improve reproducibility.
Circularity Check
Performance-parity claim rests on the same team's CXL simulator (ESF); otherwise the evaluation is self-contained.
-
self citation load bearing
[Section 4.6 (Simulator) and Section 5.3 (Extra latency); Ref [4]]
"Due to the lack of publicly available CXL 3.0 hardware, we prototype the firmware-side modification of our disaggregated SSD designs on DaisyPlus OpenSSD board [45, 96]. ... To evaluate the CXL fabrics, we integrate ESF [4], a cycle-accurate CXL simulator, which can accurately model the features defined in CXL 3.0 standard [15]."
Ref [4] (ESF) shares authors with this paper (An, Yi, Li, Wang, Luo, and Zhang). The paper's conclusion that XBOF 'takes minor Inter-SSD time (up to 2.9%)' because CXL 'delivers sub-microsecond remote access' (Section 5.3) is generated by this same-team simulator; no public CXL 3.0 hardware is used, and the NUMA emulator (Section 5.6) does not exercise CXL cache coherence. Thus the central premise of the performance-parity claim is supported by a self-citation that is not independently verified. This is load-bearing, but it is not a by-construction fit, so the rest of the evaluation retains independent content.
full rationale
The evaluation is largely self-contained: SimpleSSD is a third-party simulator, all workloads are public traces, and the cost model uses stated market prices plus an external 10% CXL-premium estimate [95]; none of the headline numbers are fitted constants or renamed inputs. The 50.4% utilization number measures the borrowed cycles directly, and the 19.0% cost saving is arithmetic from stated prices, so neither is circular. The one load-bearing self-citation is ESF [4], the same team's CXL 3.0 simulator, used in Section 4.6 to produce the sub-microsecond remote-access cost on which 'negligible performance loss' rests. Since no CXL 3.0 hardware was available and the NUMA emulation does not test CXL coherence, the quantitative core of the central claim is self-referential. This raises the score to 4; it is not 6+ because the performance parity is a simulator output conditioned on that model, not a quantity equal to its inputs by construction. XHarvest [74] appears only as related work and is not load-bearing. The power-loss-protection gap in Section 4.4 is a correctness risk rather than a circular step and does not affect this score.
Assumptions & free parameters
free parameters (5)
- busy/underutilized watermark =
75%
- DRAM harvesting segment size =
2 MB default
- resource descriptor sync period =
10 ms
- BOM price inputs =
NAND $4.95/128GB, DRAM $7.2/GB, controller $48, other $6
- CXL controller/DRAM premium =
10% over Shrunk
assumptions (5)
- domain assumption CXL 3.0 provides cache-coherent peer-to-peer access among up to 4096 points with HDM-DB coherency and sub-microsecond remote latency
- domain assumption ESF accurately models CXL 3.0 features
- domain assumption SimpleSSD faithfully models SSD firmware, ARM processor, DRAM, and flash timing
- domain assumption SHARDS predicts the miss ratio curve accurately enough for DRAM lend/borrow decisions
- domain assumption Reader-writer locks over remote mapping tables preserve FTL consistency across SSDs
invented entities (3)
-
compute-end/data-end disaggregated SSD architecture
-
shadow QP I/O redirection in the NVMe driver
-
idle resource descriptor table in global fabric-attached memory
Cite this review
Pith. "Pith review of XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing." pith.science (2026). https://pith.science/paper/QJVCCDEW
@misc{pith2026250910251,
author = {Pith},
title = {Pith review of: XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing},
year = {2026},
howpublished = {\url{https://pith.science/paper/QJVCCDEW}},
note = {Machine review of arXiv:2509.10251}
}
read the original abstract
Enterprise SSDs integrate numerous computing resources (e.g., ARM processor and onboard DRAM) to satisfy the ever-increasing performance requirements of I/O bursts. While these resources substantially elevate the monetary costs of SSDs, the sporadic nature of I/O bursts causes severe SSD resource underutilization in just a bunch of flash (JBOF) level. Tackling this challenge, we propose XBOF, a cost-efficient JBOF design, which only reserves moderate computing resources in SSDs at low monetary cost, while achieving demanded I/O performance through efficient inter-SSD resource sharing. Specifically, XBOF first disaggregates SSD architecture into multiple disjoint parts based on their functionality, enabling fine-grained SSD internal resource management. XBOF then employs a decentralized scheme to manage these disaggregated resources and harvests the computing resources of idle SSDs to assist busy SSDs in handling I/O bursts. This idea is facilitated by the cache-coherent capability of Compute eXpress Link (CXL), with which the busy SSDs can directly utilize the harvested computing resources to accelerate metadata processing. The evaluation results show that XBOF improves SSD resource utilization by 50.4% and saves 19.0% monetary costs with a negligible performance loss, compared to existing JBOF designs.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[4]
Yuda An, Shushu Yi, Bo Mao, Qiao Li, Mingzhe Zhang, Ke Zhou, Nong Xiao, Guangyu Sun, Xiaolin Wang, Yingwei Luo, and Jie Zhang. 2024. A Novel Extensible Simulation Framework for CXL-Enabled Systems. arXiv preprint arXiv:2411.08312(2024)
work page Pith review arXiv 2024
-
[1]
Alibaba. 2018. Alibaba Open Cluster Trace Program. https://github.com/alibaba/clusterdata/blob/master/cluster- trace-v2018/trace_2018.md
2018
-
[2]
Alibaba. 2024. Alibaba Block Traces.https://github .com/alibaba/ block-traces
2024
-
[3]
Pradeep Ambati, Íñigo Goiri, Felipe Frujeri, Alper Gun, Ke Wang, Brian Dolan, Brian Corell, Sekhar Pasupuleti, Thomas Moscibroda, Sameh Elnikety, Marcus Fontoura, and Ricardo Bianchini. 2020. Pro- viding SLOs for Resource-HarvestingVMs in cloud platforms. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 735–751
2020
-
[5]
Erick Bauman, Gbadebo Ayoade, and Zhiqiang Lin. 2015. A sur- vey on hypervisor-based monitoring: approaches, applications, and evolutions.ACM Computing Surveys (CSUR)48, 1 (2015), 1–33
2015
-
[6]
Erick Bauman, Huibo Wang, Mingwei Zhang, and Zhiqiang Lin. 2018. Sgxelide: enabling enclave code secrecy via self-modification. In Proceedings of the 2018 International Symposium on Code Generation and Optimization. 75–86
2018
-
[7]
Matias Bjørling, Javier Gonzalez, and Philippe Bonnet. 2017. Light- NVM: The Linux Open-ChannelSSD Subsystem. In15th USENIX Conference on File and Storage Technologies (FAST 17). 359–374
2017
-
[8]
Irina Calciu, Dave Dice, Yossi Lev, Victor Luchangco, Virendra J Marathe, and Nir Shavit. 2013. NUMA-aware reader-writer locks. In Proceedings of the 18th ACM SIGPLAN symposium on Principles and practice of parallel programming. 157–166
2013
Show all 125 references
-
[9]
Karthik Chandrasekar, Christian Weis, Yonghui Li, Benny Akesson, Norbert Wehn, and Kees Goossens. 2012. DRAMPower: Open-source DRAM power & energy estimation tool.URL: http://www. drampower. info22 (2012)
2012
-
[10]
Peter M Chen, Edward K Lee, Garth A Gibson, Randy H Katz, and David A Patterson. 1994. RAID: High-performance, reliable sec- ondary storage.ACM Computing Surveys (CSUR)26, 2 (1994), 145– 185
1994
-
[11]
Tae-Sun Chung, Dong-Joo Park, Sangwon Park, Dong-Ho Lee, Sang- Won Lee, and Ha-Joo Song. 2009. A survey of flash translation layer. Journal of Systems Architecture55, 5-6 (2009), 332–343
2009
-
[12]
Gilberto Contreras and Margaret Martonosi. 2005. Power prediction for Intel XScale®processors using performance monitoring unit events. InProceedings of the 2005 international symposium on Low power electronics and design. 221–226
2005
-
[13]
Crucial. 2025. Crucial T705 PCIe 5.0 NVMe.https://www.crucial.com/ ssd/t705/ct2000t705ssd5a
2025
-
[14]
CXL. 2025. Compute Express Link.https://computeexpresslink .org
2025
-
[15]
CXL. 2025. Compute Express LinkTM (CXLTM) Specification 3.2.https://computeexpresslink .org/wp-content/uploads/2024/12/ CXL_3.2-Spec-Announcement_FINAL-1.pdf
2025
-
[16]
Christoffer Dall and Jason Nieh. 2014. KVM/ARM: the design and implementation of the linux ARM hypervisor.Acm Sigplan Notices 49, 4 (2014), 333–348
2014
-
[17]
Li Deng, Yu-Lin Ren, Fei Xu, Heng He, and Chao Li. 2018. Resource utilization analysis of Alibaba cloud. InIntelligent Computing Theories and Application: 14th International Conference, ICIC 2018, Wuhan, China, August 15-18, 2018, Proceedings, Part I 14. Springer, 183–194
2018
-
[18]
ARM Developer. 2025. DCCSW, Data Cache line Clean by Set/Way.https://developer.arm.com/documentation/ddi0601/2024- 12/AArch32-Instructions/DCCSW--Data-Cache-line-Clean-by- Set-Way
2025
-
[19]
ARM Developer. 2025. Interaction with the Performance Moni- toring Unit (PMU).https://developer .arm.com/documentation/ ddi0469/b/functional-description/operation/interaction-with-the- performance-monitoring-unit--pmu-
2025
-
[20]
ARM Developer. 2025. Overview of the Armv8 Archi- tecture.https://developer .arm.com/documentation/dui0801/l/ Overview-of-the-Armv8-Architecture
2025
-
[21]
Yaozu Dong, Zhao Yu, and Greg Rose. 2008. SR-IOV Networking in Xen: Architecture, Design and Implementation.. InWorkshop on I/O virtualization, Vol. 2. San Diego, CA, USA
2008
-
[22]
DRAMeXchange. 2024. The Global Price of NAND Flash and LPDDR. https://www.dramexchange.com/
2024
-
[23]
Zelin Du, Shaoqi Li, Zixuan Huang, Jin Xue, Kecheng Huang, Tianyu Wang, and Zili Shao. 2024. PipeSSD: A Lock-free Pipelined SSD Firmware Design for Multi-core Architecture. InProceedings of the 61st ACM/IEEE Design Automation Conference. 1–6
2024
-
[24]
facebook. 2025. db_bench.https://github .com/facebook/rocksdb/ wiki/Benchmarking-tools
2025
-
[25]
Facebook. 2025. Rocksdb.http://rocksdb.org/
2025
-
[26]
filebench. 2025. A Model Based File System Workload Generator. https://github.com/filebench/filebench
2025
-
[27]
Donghyun Gouk, Miryeong Kwon, Jie Zhang, Sungjoon Koh, Wonil Choi, Nam Sung Kim, Mahmut Kandemir, and Myoungsoo Jung. 2018. Amber: Enabling precise full-system simulation with detailed model- ing of all SSD resources. In2018 51st Annual IEEE/ACM International Symposium on Micr...
2018
-
[28]
Jim Gray, Paul McJones, Mike Blasgen, Bruce Lindsay, Raymond Lorie, Tom Price, Franco Putzolu, and Irving Traiger. 1981. The recovery manager of the System R database manager.ACM Computing Surveys (CSUR)13, 2 (1981), 223–242
1981
-
[29]
Zerui Guo, Hua Zhang, Chenxingyu Zhao, Yuebin Bai, Michael Swift, and Ming Liu. 2023. Leed: A low-power, fast persistent key-value store on smartnic jbofs. InProceedings of the ACM SIGCOMM 2023 Conference. 1012–1027
2023
-
[30]
Aayush Gupta, Youngjae Kim, and Bhuvan Urgaonkar. 2009. DFTL: a flash translation layer employing demand-based selective caching of page-level address mappings.Acm Sigplan Notices44, 3 (2009), 229–240
2009
-
[31]
Jian Huang, Anirudh Badam, Laura Caulfield, Suman Nath, Sudipta Sengupta, Bikash Sharma, and Moinuddin K Qureshi. 2017. FlashBlox: 12 Achieving Both Performance Isolation and Uniform Lifetime for Virtualized SSDs. In15th USENIX Conference on File and Storage Technologies (FAST...
2017
-
[32]
Intel. 2025. Intel®Xeon®Platinum 8562Y+ Processor. https://www.intel.com/content/www/us/en/products/sku/ 237558/intel-xeon-platinum-8562y-processor-60m-cache-2- 80-ghz/specifications.html
2025
-
[33]
Houxiang Ji, Srikar Vanavasam, Yang Zhou, Qirong Xia, Jinghan Huang, Yifan Yuan, Ren Wang, Pekon Gupta, Bhushan Chitlur, Ipoom Jeong, and Nam Sung Kim. 2024. Demystifying a CXL Type-2 Device: A Heterogeneous Cooperative Computing Perspective. In2024 57th IEEE/ACM International...
2024
-
[34]
Sheng Jiang and Ming Liu. 2025. Building an Elastic Block Storage over EBOFs Using Shadow Views. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). 1137–1153
2025
-
[35]
Tianyang Jiang, Guangyan Zhang, Zican Huang, Xiaosong Ma, Junyu Wei, Zhiyue Li, and Weimin Zheng. 2021. FusionRAID: Achieving Consistent Low Latency for Commodity SSD Arrays. In19th USENIX Conference on File and Storage Technologies (FAST 21). 355–370
2021
-
[36]
Myoungsoo Jung. 2022. Hello bytes, bye blocks: Pcie storage meets compute express link for memory expansion (cxl-ssd). InProceedings of the 14th ACM Workshop on Hot Topics in Storage and File Systems. 45–51
2022
-
[37]
Myoungsoo Jung, Jie Zhang, Ahmed Abulila, Miryeong Kwon, Narges Shahidi, John Shalf, Nam Sung Kim, and Mahmut Kandemir. 2017. SimpleSSD: Modeling solid state drives for holistic system simulation. IEEE Computer Architecture Letters17, 1 (2017), 37–41
2017
-
[38]
Luyi Kang, Yuqi Xue, Weiwei Jia, Xiaohao Wang, Jongryool Kim, Changhwan Youn, Myeong Joon Kang, Hyung Jin Lim, Bruce Jacob, and Jian Huang. 2021. Iceclave: A trusted execution environment for in-storage computing. InMICRO-54: 54th Annual IEEE/ACM Interna- tional Symposium on M...
2021
-
[39]
Swaroop Kavalanekar, Bruce Worthington, Qi Zhang, and Vishal Sharda. 2008. Characterization of storage workload traces from production windows servers. In2008 IEEE International Symposium on Workload Characterization. IEEE, 119–128
2008
-
[40]
Aleksandr Khasymski, M Mustafa Rafique, Ali R Butt, Sudharshan S Vazhkudai, and Dimitrios S Nikolopoulos. 2012. On the use of GPUs in realizing cost-effective distributed RAID. In2012 IEEE 20th Interna- tional Symposium on Modeling, Analysis and Simulation of Computer and Tele...
2012
-
[41]
Jiho Kim, Myoungsoo Jung, and John Kim. 2023. Decoupled SSD: Rethinking SSD Architecture through Network-based Flash Con- trollers. InProceedings of the 50th Annual International Symposium on Computer Architecture. 1–13
2023
-
[42]
Jiho Kim, Seokwon Kang, Yongjun Park, and John Kim. 2022. Net- worked SSD: Flash Memory Interconnection Network for High- Bandwidth SSD. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 388–403
2022
-
[43]
Sang-Hoon Kim, Jaehoon Shim, Euidong Lee, Seongyeop Jeong, Ilkueon Kang, and Jin-Soo Kim. 2023. NVMeVirt: A versatile software- defined virtual NVMe device. In21st USENIX Conference on File and Storage Technologies (FAST 23). 379–394
2023
-
[44]
Thomas Kim, Jekyeom Jeon, Nikhil Arora, Huaicheng Li, Michael Kaminsky, David G Andersen, Gregory R Ganger, George Amvrosiadis, and Matias Bjørling. 2023. RAIZN: Redundant Array of Independent Zoned Namespaces. InProceedings of the 28th ACM International Conference on Architec...
2023
-
[45]
Jaewook Kwak, Sangjin Lee, Kibin Park, Jinwoo Jeong, and Yong Ho Song. 2020. Cosmos+ openssd: Rapid prototype for flash storage systems.ACM Transactions on Storage (TOS)16, 3 (2020), 1–35
2020
-
[46]
Dongup Kwon, Junehyuk Boo, Dongryeong Kim, and Jangwoo Kim
-
[47]
Miryeong Kwon, Sangwon Lee, and Myoungsoo Jung. 2023. Cache in hand: Expander-driven cxl prefetcher for next generation cxl-ssd. InProceedings of the 15th ACM Workshop on Hot Topics in Storage and File Systems. 24–30
2023
-
[48]
James R Larus and Thomas Ball. 1994. Rewriting executable files to measure program behavior.Software: Practice and Experience24, 2 (1994), 197–218
1994
-
[49]
Chunghan Lee, Tatsuo Kumano, Tatsuma Matsuki, Hiroshi Endo, Naoto Fukumoto, and Mariko Sugawara. 2017. Understanding storage traffic characteristics on enterprise virtual desktop infrastructure. InProceedings of the 10th ACM International Systems and Storage Conference. 1–11
2017
-
[50]
Hill, Marcus Fontoura, and Ricardo Bian- chini
Huaicheng Li, Daniel S Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bian- chini. 2023. Pond: Cxl-based memory pooling systems for cloud platforms. InProc...
2023
-
[51]
Sheng Li, Jung Ho Ahn, Richard D Strong, Jay B Brockman, Dean M Tullsen, and Norman P Jouppi. 2009. McPAT: An integrated power, area, and timing modeling framework for multicore and manycore architectures. InProceedings of the 42nd annual ieee/acm international symposium on mi...
2009
-
[52]
Shaobo Li, Yirui Zhou, Hao Ren, and Jian Huang. 2025. Bytefs: Sys- tem support for (CXL-based) memory-semantic solid-state drives. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume
2025
-
[53]
Changyue Liao, Mo Sun, Zihan Yang, Kaiqi Chen, Binhang Yuan, Fei Wu, and Zeke Wang. 2024. Adding NVMe SSDs to Enable and Accelerate 100B Model Fine-tuning on a Single GPU.arXiv preprint arXiv:2403.06504(2024)
2024 arXiv
-
[54]
Linux. 2021. Linux v5.15.https://github .com/torvalds/linux/tree/ v5.15
2021
-
[55]
Linux. 2025. Lightnvm Driver.https://github .com/torvalds/linux/ tree/v5.10-rc3/drivers/lightnvm
2025
-
[56]
Linux. 2025. mdraid layer.https://github .com/torvalds/linux/tree/ master/drivers/md
2025
-
[57]
Linux. 2025. Memory Management.https://docs .kernel.org/admin- guide/mm/index.html
2025
-
[58]
Longsys. 2021. Application Documents for Initial Public Of- fering and Listing on GEM of Shenzhen Jiang Bolong Electron- ics Co. Initial Public Offering of Shares and Listing on GEM Board Application Documents Report in response to the audit in- quiry letter from.https://qccda...
2021
-
[59]
Sunilkumar S Manvi and Gopal Krishna Shyam. 2014. Resource management for Infrastructure as a Service (IaaS) in cloud computing: A survey.Journal of network and computer applications41 (2014), 424–440
2014
-
[60]
China Flash Market. 2024. The Price of NAND Flash and LPDDR. https://en.chinaflashmarket.com/
2024
-
[61]
Avantika Mathur, Mingming Cao, Suparna Bhattacharya, Andreas Dilger, Alex Tomas, and Laurent Vivier. 2007. The new ext4 filesystem: current status and future plans. InProceedings of the Linux symposium, Vol. 2. Citeseer, 21–33
2007
-
[62]
Sara McAllister, Benjamin Berg, Julian Tutuncu-Macias, Juncheng Yang, Sathya Gunasekar, Jimmy Lu, Daniel S Berger, Nathan Beck- mann, and Gregory R Ganger. 2021. Kangaroo: Caching billions of 13 Shushu Yi et al. tiny objects on flash. InProceedings of the ACM SIGOPS 28th sympo...
2021
-
[63]
Rino Micheloni, Alessia Marelli, Kam Eshghi, and G Wong. 2013. SSD market overview.Inside Solid State Drives (SSDs)(2013), 1–17
2013
-
[64]
Micron. 2025. Micron 9550 NVMe SSD.https://www .micron.com/ products/storage/ssd/data-center-ssd/9550-ssd
2025
-
[65]
Micron. 2025. Micron leads ecosystem: first to develop PCIe Gen6 data center SSD.https://www .micron.com/about/blog/storage/ssd/ micron-leads-ecosystem-first-to-develop-pcie-gen6-data-center- ssd
2025
-
[66]
Micron. 2025. Micron MT40A512M16TD-062E.https: //www.mouser.com/ProductDetail/Micron/MT40A512M16TD- 062E-AITR?qs=3Rah4i%252BhyCENjNab2Szuaw%3D%3D
2025
-
[67]
Jaehong Min, Ming Liu, Tapan Chugh, Chenxingyu Zhao, Andrew Wei, In Hwan Doh, and Arvind Krishnamurthy. 2021. Gimbal: en- abling multi-tenant storage disaggregation on SmartNIC JBOFs. In Proceedings of the 2021 ACM SIGCOMM 2021 Conference. 106–122
2021
-
[68]
Rakesh Nadig, Mohammad Sadrosadati, Haiyu Mao, Nika Man- souri Ghiasi, Arash Tavakkol, Jisung Park, Hamid Sarbazi-Azad, Juan Gómez Luna, and Onur Mutlu. 2023. Venice: Improving Solid- State Drive Parallelism at Low Cost via Conflict-Free Accesses. In Proceedings of the 50th An...
2023
-
[69]
Iyswarya Narayanan, Di Wang, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu, Badrid- dine Khessib, and Kushagra Vaid. 2016. SSD failures in datacenters: What? when? and why?. InProceedings of the 9th ACM International on Systems and ...
2016
-
[70]
Nvidia. 2025. NVIDIA BLUEFIELD-3 DPU.https://www .nvidia.com/ content/dam/en-zz/Solutions/Data-Center/documents/datasheet- nvidia-bluefield-3-dpu.pdf
2025
-
[71]
NVMe. 2022. NVM Command Set Specification.https: //nvmexpress.org/wp-content/uploads/NVM-Express-NVM- Command-Set-Specification-1.0c-2022.10.03-Ratified.pdf
2022
-
[72]
NVMe. 2025. NVM Express®Base Specification.https:// nvmexpress.org/specification/nvm-express-base-specification/
2025
-
[73]
ONFi. 2025. Open NAND Flash Interface Specifications.https:// onfi.org/specs.html
2025
-
[74]
Li Peng, Wenbo Wu, Shushu Yi, Xianzhang Chen, Chenxi Wang, Shengwen Liang, Zhe Wang, Nong Xiao, Qiao Li, Mingzhe Zhang, and Jie Zhang. 2025. XHarvest: Rethinking High-Performance and Cost-Efficient SSD Architecture with CXL-Driven Harvesting. In Proceedings of the 52nd Annual ...
2025
-
[75]
Phison. 2025. PS5021-E21T.https://www .phison.com/en/products/ ssd/ps5021-e21t
2025
-
[76]
Phison. 2025. PS5026-E26.https://www .phison.com/en/products/ ssd/ps5026-e26
2025
-
[77]
Benjamin Reidys, Jinghan Sun, Anirudh Badam, Shadi Noghabi, and Jian Huang. 2022. BlockFlex: Enabling Storage Harvesting with Software-Defined Flash in Modern Cloud Platforms. In16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 17–33
2022
-
[78]
Rusty Russell. 2008. virtio: towards a de-facto standard for virtual I/O devices.ACM SIGOPS Operating Systems Review42, 5 (2008), 95–103
2008
-
[79]
Samsung. 2020. Samsung 980Pro NVMe SSD.https: //www.samsung.com/us/computing/memory-storage/solid- state-drives/980-pro-pcie-4-0-nvme-ssd-1tb-mz-v8p1t0b-am/
2020
-
[80]
Samsung. 2023. Samsung PM1743.https:// semiconductor.samsung.com/ssd/enterprise-ssd/pm1743/
2023
-
[81]
Henry N Schuh, Arvind Krishnamurthy, David Culler, Henry M Levy, Luigi Rizzo, Samira Khan, and Brent E Stephens. 2024. CC-NIC: a Cache-Coherent Interface to the NIC. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and O...
2024
-
[82]
Kia Shakiba, Sari Sultan, and Michael Stumm. 2024. Kosmo: efficient online miss ratio curve generation for eviction policy evaluation. In 22nd USENIX Conference on File and Storage Technologies (FAST 24). 89–105
2024
-
[83]
Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. 2018. LegoOS: A disseminated, distributed OS for hardware resource disag- gregation. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 69–87
2018
-
[84]
Xuanhua Shi, Ming Li, Wei Liu, Hai Jin, Chen Yu, and Yong Chen
-
[85]
Junyi Shu, Ruidong Zhu, Yun Ma, Gang Huang, Hong Mei, Xuanzhe Liu, and Xin Jin. 2023. Disaggregated RAID Storage in Modern Datacenters. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3. 147–163
2023
-
[86]
SimpleSSD. 2020. SimpleSSD version 2.0: Open-Source Licensed Educational SSD Simulator for High-Performance Storage and Full- System Evaluations.https://github.com/SimpleSSD/SimpleSSD
2020
-
[87]
Ryan Smith. 2020. SSD Calculator Soruces.https: //www.soothsawyer.com/wp-content/uploads/2020/03/ Public_SSD_Cost_Calculator_Share.pdf
2020
-
[88]
SNIA. 2025. MSR Cambridge Traces.http://iotta .snia.org/traces/ block-io/388
2025
-
[89]
Yong Ho Song, Sanghyuk Jung, Sang-Won Lee, and Jin-Soo Kim. 2014. Cosmos openSSD: A PCIe-based open source SSD platform.Proc. Flash Memory Summit(2014), 1–30
2014
-
[90]
Kang-Deog Suh, Byung-Hoon Suh, Young-Ho Lim, Jin-Ki Kim, Young- Joon Choi, Yong-Nam Koh, Sung-Soo Lee, Suk-Chon Kwon, Byung- Soon Choi, Jin-Sun Yum, Jung-Hyuk Choi, Jang-Rae Kim, and Hyung- Kyu Lim. 1995. A 3.3 V 32 Mb NAND flash memory with incremental step pulse programming ...
1995
-
[91]
Jinghan Sun, Benjamin Reidys, Daixuan Li, Jichuan Chang, Marc Snir, and Jian Huang. 2025. Fleetio: Managing multi-tenant cloud storage with multi-agent reinforcement learning. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Langu...
2025
-
[92]
Xun Sun, Mingxing Zhang, Yingdi Shan, Kang Chen, Jinlei Jiang, and Yongwei Wu. 2025. Scalio: Scaling up DPU-basedJBOF Key-value Store with NVMe-oF Target Offload. In19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). 449–464
2025
-
[93]
Yan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Wang, Jung Ho Ahn, Tianyin Xu, and Nam Sung Kim. 2023. Demys- tifying cxl memory with genuine cxl-ready systems and devices. In Proceedings of th...
2023
-
[94]
SuperMicro. 2025. Storage SuperServer SSG-229J-5BU24JBF. https://www.supermicro.com/en/products/system/storage/2u/ssg- 229j-5bu24jbf
2025
-
[95]
Yupeng Tang, Ping Zhou, Wenhui Zhang, Henry Hu, Qirui Yang, Hao Xiang, Tongping Liu, Jiaxin Shan, Ruoyun Huang, Cheng Zhao, Cheng Chen, Hui Zhang, Fei Liu, Shuai Zhang, Xiaoning Ding, and Jianjun Chen. 2024. Exploring performance and cost optimization with asic-based cxl memor...
2024
-
[96]
CRZ Technology. 2025. DaisyPlus OpenSSD.https://www .crz- tech.com/crz/article/DaisyPlus/
2025
-
[97]
Maxio Technology. 2024. About Lianyun Technology (Hangzhou) Co. Initial Public Offering and Listing on Technology and Innovation Board (TECHNOLOGY) Response to the Audit Inquiry Letter on the Filing Documents.https://static .sse.com.cn/stock/disclosure/ 14 announcement/c/20240...
2024
-
[98]
TRENDFORCE. 2024. NAND Flash Price Trends.https:// www.trendforce.com/price/flash
2024
-
[99]
Shivani Tripathy and Manoranjan Satpathy. 2022. SSD internal cache management policies: A survey.Journal of Systems Architecture122 (2022), 102334
2022
-
[100]
Hung-Wei Tseng, Laura Grupp, and Steven Swanson. 2011. Under- standing the impact of power loss on flash memory. InProceedings of the 48th Design Automation Conference. 35–40
2011
-
[101]
Carl A Waldspurger, Nohhyun Park, Alexander Garthwaite, and Irfan Ahmad. 2015. Efficient MRC construction with SHARDS. In 13th USENIX Conference on File and Storage Technologies (FAST 15). 95–110
2015
-
[102]
Rui Wang, Yongkun Li, Hong Xie, Yinlong Xu, and John CS Lui
-
[103]
Yawen Wang, Kapil Arya, Marios Kogias, Manohar Vanga, Aditya Bhandari, Neeraja J Yadwadkar, Siddhartha Sen, Sameh Elnikety, Christos Kozyrakis, and Ricardo Bianchini. 2021. Smartharvest: Har- vesting idle cpus safely and efficiently in the cloud. InProceedings of the Sixteenth...
2021
-
[104]
Reinhold P Weicker. 1984. Dhrystone: a synthetic systems program- ming benchmark.Commun. ACM27, 10 (1984), 1013–1030
1984
-
[105]
In2020 USENIX Annual Technical Conference (USENIX ATC 20)
GraphWalker: An I/O-Efficient and Resource-Friendly Graph Analytic System for Fast and Scalable Random Walks. In2020 USENIX Annual Technical Conference (USENIX ATC 20). 559–571
-
[106]
Gala Yadgar, MOSHE Gabel, Shehbaz Jaffer, and Bianca Schroeder
-
[107]
Shao-Peng Yang, Minjae Kim, Sanghyun Nam, Juhyung Park, Jin- Yong Choi, Eyee Hyun Nam, Eunji Lee, Sungjin Lee, and Bryan S Kim
-
[108]
Jiwon Woo, Minwoo Ahn, Gyusun Lee, and Jinkyu Jeong. 2021. D2FQ: Device-Direct Fair Queueing for NVMeSSDs. In19th USENIX Confer- ence on File and Storage Technologies (FAST 21). 403–415
2021
-
[109]
Zhe Yang, Youyou Lu, Xiaojian Liao, Youmin Chen, Junru Li, Siyu He, and Jiwu Shu. 2023. 𝜆-IO: A Unified IO Stack for Computational Storage. In21st USENIX Conference on File and Storage Technologies (FAST 23). 347–362
2023
-
[110]
Min Ye, Qiao Li, Yina Lv, Jie Zhang, Tianyu Ren, Daniel Wen, Tei- Wei Kuo, and Chun Jason Xue. 2024. Achieving Near-Zero Read Retry for 3D NAND Flash Memory. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating...
2024
-
[111]
Shushu Yi, Shaocong Sun, Li Peng, Yingbo Sun, Ming-Chang Yang, Zhichao Cao, Qiao Li, Myoungsoo Jung, Ke Zhou, and Jie Zhang
-
[112]
Shushu Yi, Yanning Yang, Yunxiao Tang, Zixuan Zhou, Junzhe Li, Chen Yue, Myoungsoo Jung, and Jie Zhang. 2022. Scalaraid: Opti- mizing linux software raid system for next-generation storage. In Proceedings of the 14th ACM Workshop on Hot Topics in Storage and File Systems. 119–125
2022
-
[113]
Ziye Yang, Changpeng Liu, Yanbo Zhou, Xiaodong Liu, and Gang Cao. 2018. Spdk vhost-nvme: Accelerating i/os in virtual machines on nvme ssds via user space vhost target. In2018 IEEE 8th International Symposium on Cloud and Service Computing (SC2). IEEE, 67–76
2018
-
[114]
Haoyang Zhang, Yuqi Xue, Yirui Eric Zhou, Shaobo Li, and Jian Huang. 2025. SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design.arXiv preprint arXiv:2501.10682(2025)
2025 arXiv
-
[115]
Jie Zhang, Miryeong Kwon, Michael Swift, and Myoungsoo Jung
-
[116]
Jian Zhang, Yujie Ren, Marie Nguyen, Changwoo Min, and Sudarsun Kannan. 2024. OmniCache: Collaborative Caching for Near-storage Accelerators. In22nd USENIX Conference on File and Storage Tech- nologies (FAST 24). 35–50
2024
-
[117]
Yu Zhang, Ping Huang, Ke Zhou, Hua Wang, Jianying Hu, Yongguang Ji, and Bin Cheng. 2020. OSCA: An Online-Model Based Cache Allocation Scheme in Cloud Block Storage Systems. In2020 USENIX Annual Technical Conference (USENIX ATC 20). 785–798
2020
-
[118]
Mai Zheng, Joseph Tucek, Feng Qin, and Mark Lillibridge. 2013. Understanding the robustness of SSDs under power fault. In11th USENIX Conference on File and Storage Technologies (FAST 13). 271– 284. 15
2013
-
[119]
Sangjin Yoo and Dongkun Shin. 2020. Reinforcement Learning- BasedSLC Cache Technique for Enhancing SSD Write Performance. In12th USENIX Workshop on Hot Topics in Storage and File Systems (HotStorage 20)
2020
-
[122]
In18th USENIX Conference on File and Storage Technologies (FAST 20)
Scalable parallel flash firmware for many-core architectures. In18th USENIX Conference on File and Storage Technologies (FAST 20). 121–136
-
[2017]
In Proceedings of the international conference on supercomputing
Ssdup: a traffic-aware ssd burst buffer for hpc systems. In Proceedings of the international conference on supercomputing. 1–10
-
[2020]
In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)
FVM:FPGA-assisted Virtual Device Emulation for Fast, Scalable, and Flexible Storage Virtualization. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 955–971
-
[2021]
SSD-based workload characteristics and their performance implications.ACM Transactions on Storage (TOS)17, 1 (2021), 1–26
2021
-
[2023]
In2023 USENIX Annual Technical Conference (USENIX ATC 23)
Overcoming the Memory Wall with CXL-EnabledSSDs. In2023 USENIX Annual Technical Conference (USENIX ATC 23). 601–617
-
[2024]
InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles
BIZA: Design of Self-Governing Block-Interface ZNS AFA for Endurance and Performance. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles. 313–329
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.