REVIEW 4 major objections 4 minor 101 references
DMG: A Scalable and Efficient Memory-Disaggregated Graph Processing System
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper claims that DMG is the first practical graph processing system for disaggregated memory, one that scales across multiple compute and memory nodes while keeping compute-side caches at conventional DM sizes and delivering performan
desk verdict Genuinely new multi-CN/MN graph processing on DM, with a thorough ablation, but the MN-CPU assumption is tested only on 24-core EPYC nodes and could be the weak point. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is an adaptive per-vertex index and edge-store combined with a redistribute update path. The 32-byte per-vertex index lives in a single range-partitioned array on the memory pool; for low-degree vertices it stores the edge list inline (‘in-place’), turning two dependent remote reads into one, while for high-degree vertices it stores the edge-list address and compressed per-segment lengths (‘out-of-place’). The segment lengths let a node fetch only the portion of a hub’s edge list that falls in a given partition, and they let the update coordinator hand whole segments to the node that owns the destination attributes. This machinery converts RDMA’s IOPS bottleneck into fe
What would settle it
Run the same BFS and PageRank workloads with memory nodes limited to one or two low-power cores and measure both end-to-end time and memory-node CPU utilization during the densest iteration; if any memory node’s utilization saturates or the end-to-end time degrades disproportionately, the collaborative-update claim collapses.
Extended reading notes
Core claim
The paper’s core discovery is that the three obstacles to practical graph processing on DM—IOPS-limited remote reads, costly remote update propagation, and tail effects from hub vertices—can each be turned around by exploiting where data already resides. DMG stores vertex attributes once in a shared memory pool and gives each compute node a small cache covering its assigned chunk. For retrieval, a 32-byte per-vertex index entry either embeds the edge list of low-degree vertices or stores a compressed segment layout for high-degree vertices, so most vertex reads become one merged RDMA request instead of two dependent fine-grained ones. For updates, a collaborative scheme batches update candid
Load-bearing premise
The load-bearing premise is that a memory node’s scarce CPU can absorb the offloaded update work (ValRD) and RPC service without becoming a bottleneck—the testbed gives each memory node a 24-core server running only two threads, so the reported numbers depend on memory-node CPU being cheap.
Editorial extensions
If this is right
- Disaggregated-memory graph systems can store one copy of the graph in a shared memory pool and elastically add compute nodes without re-coupling memory, so tenants pay only for the resource they need.
- A conventional 1–2 GB compute-side cache is enough for billion-scale graphs, because each compute node caches only the attribute slice of its assigned chunk; aggregate compute-node memory stays a small fraction of memory-pool usage.
- Graph partitioning for load balancing can be redone in sub-seconds using tiny-chunk metadata, making repeated resource reconfiguration practical on DM.
- The computation-time gap versus a local-memory distributed system stands at about 40%, and the paper argues this gap will narrow as interconnect bandwidth rises toward 800 Gbps and beyond.
- The same storage and update techniques carry over to CXL-based memory pooling, because they reduce fine-grained remote accesses rather than relying on RDMA-specific behavior.
Reading between the lines
- The paper does not test memory nodes with truly scarce CPU: its memory nodes are 24-core servers running only two threads. If a real memory node has one or two weak cores, the offloaded ValRD update path could saturate that CPU, and the reported cache savings and speedups would shrink.
- The in-place index scheme exploits power-law degree distributions; on graphs with more uniform degrees, fewer vertices fit inline and the retrieval benefit should diminish. The paper includes a synthetic R-MAT graph but does not isolate this effect.
- The collaborative-update advantage depends on chunk locality; a graph whose vertex IDs are shuffled to destroy locality would send far more update candidates across the network. DMG does not report a degradation curve for such adversarial layouts.
- Because the loaded graph store and segment metadata are reusable across different compute-node counts, a natural extension is mid-query elastic resizing of the compute pool without reloading the graph; the paper does not implement this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. DMG proposes a graph processing system for memory-disaggregated (DM) architectures. It contributes three main designs: (i) a DM-friendly graph store that embeds low-degree edge lists in an enlarged index and applies merged/batched RDMA retrieval; (ii) an adaptive update coordinator that selects between collaborative update (with pass-by-reference and pass-by-value re-distribution) and direct remote update depending on update density; and (iii) a two-stage workload manager using coarse-grained tiny-chunk partitioning plus runtime hub re-scheduling. The paper claims this is the first practical multi-CN/multi-MN graph processing system on DM, achieving up to 4.9× speedup over FAM-Graph, up to 18.9× reduction in compute-side cache demand, and computation time within 40% of the local-memory distributed system Gemini. The evaluation includes internal ablations over the proposed components and comparisons with FAM-Graph, DMG-Base, and Gemini on four billion-scale graphs.
Significance. If the claims hold, this is a substantial contribution to the DM systems literature. The paper identifies an important practical gap in existing DM graph systems (single CN/MN and large compute-side caches) and proposes concrete mechanisms to address it. A particular strength is the internal ablation in §5.4: separating the effects of index embedding, retrieval optimizations, the two re-distribution modes, and runtime re-scheduling gives good evidence that the design choices, not just the overall architecture, drive the reported gains. The cache-efficiency result (orders of magnitude below FAM-Graph) and the startup-time comparison with Gemini are also compelling. However, the central scalability and performance claims rest on assumptions about memory-node CPU availability and on a fixed 4-MN testbed; these are not yet fully stress-tested, and the absence of repeated-run statistics makes some quantitative claims difficult to assess.
major comments (4)
- [§5.1/§5.4.2] MN CPU model: The system model in §2.2 assumes MNs have weak computation power (e.g., 1–2 CPU cores), but the testbed uses 24-core EPYC 7402P MNs (only two threads active). The ValRD mechanism (§4.3.1) is load-bearing: Fig. 25 shows +ValRD provides a large improvement over +RefRD. However, the MN CPU validation in Fig. 26 reports at most 44% usage of one EPYC core for TW; this does not establish that the same offload is sustainable on a real 1–2 core, lower-performance MN that also runs RPC-serving threads. I recommend adding an experiment that pins MN threads to one or two cores, or otherwise models weak MNs, and reports end-to-end performance and ValRD CPU/throughput under that constraint. If ValRD saturates, the fallback path could erode the reported speedups.
- [§5.2/§5.1] Memory-side scalability is not directly demonstrated. The abstract claims elastic scaling of both compute and memory, but all scaling experiments vary only the number of CNs (1, 2, 4) with a fixed 4-MN pool. I could not find an experiment that varies the number of MNs, or that increases graph size while scaling MNs. The clueweb12 result shows DMG can use a 4-MN pool where FAM-Graph cannot, but not that adding MNs elastically accommodates ever-increasing graphs. Please add a memory-scaling experiment (e.g., 1→4 MNs with fixed CN count, or a graph that grows with MN count) or temper the claim.
- [§5 (general methodology)] All reported numbers appear to be from single executions, with no error bars or variance information. This matters for the quantitative headline claims (4.9× speedup, 18.9× cache reduction) and for small differences such as the UK 0.91–1.19× speedup range in §5.2. I request repeated runs (at least 3–5 per configuration) with confidence intervals, or at minimum a statement of run-to-run variability, for the key figures (Figs. 15–19, 24, 25, 28).
- [§4.2.1/§5.4.1] Several design thresholds—32B index entry size, in-place degree ≤7, RS hub threshold >1024, 4 coroutines per thread, 1024 tiny-chunks—appear to be chosen based on the four benchmark graphs, and the index-size sweep (Fig. 22) is shown only for TW. Since the main comparisons use the same datasets, it is unclear whether these values are overfit to the testbed. Please add a sensitivity analysis for the most impactful thresholds across datasets and at least one other hardware configuration, or give an explicit argument that the thresholds are hardware- and dataset-independent.
minor comments (4)
- [Figs. 15–18] The captions do not define what the annotated ratios (e.g., '99x', '179x') refer to. Clarify whether they are speedups of DMG over DMG-Base or over FAM-Graph, and how FAM-Graph is plotted at 2 and 4 CNs when it supports only one CN.
- [Fig. 19] Please define 'per-CN cache usage' precisely (maximum RSS, allocated cache size, or measured working set) and state whether the FAM-Graph bar is for one CN only. This would help readers interpret the 18.9× claim.
- [§5.3] The statement that DMG achieves computation time 'within 40% of Gemini' is ambiguous: does it mean 40% slower, or 40% of Gemini's time? The text later says 'moderate computation overhead,' suggesting the former. Please rephrase.
- [Abstract and §5.3] There are small text issues: 'toDM-friendly' in the abstract and 'ontwitter-2010' in §5.3 are missing spaces. Also, the paper promises open-source code; please include the repository or artifact link at the final version.
Circularity Check
No load-bearing circularity: claims are measured against external baselines; self-citations and the MN-CPU setup are caveats, not circular inputs.
full rationale
DMG is an evaluation-driven systems paper; its central claims are established by direct measurements against external baselines (FAM-Graph, Gemini) and an internal ablation (DMG-Base) across four billion-scale graphs, not by a derivation in which an output is defined in terms of an input. The design constants (32B index, in-place degree threshold, RS threshold, 2GB cache) are presented as empirical tuning choices with parameter studies (Figs. 22-23) and ablations (Figs. 24-28), and the headline 4.9x/18.9x numbers are measured comparisons, not fitted outputs. Self-citations ([21] Aceso, [36] DEX, [74-76] Seraph/Oasis) support background claims or appear in related work; the one load-bearing assumption they touch—that MNs can execute offloaded ValRD work cheaply—is independently measured in Fig. 26, so the citations do not carry the argument. The stated limitations (§4.5.2 dual-mode selection restricted by hardware; §4.5.3 static graphs only) narrow scope but do not create circularity. The main caveat is external validity, not circularity: MNs are simulated on 24-core EPYC machines (§5.1) while §2.2 assumes 1-2 weak cores, so the ValRD CPU-usage numbers may overstate real-MN headroom; this affects how representative the results are, not whether the results were derived from their own assumptions.
Assumptions & free parameters
free parameters (5)
- Index entry size =
32 bytes
- In-place / ValRD degree threshold =
degree ≤ 7
- RS hub threshold =
degree > 1024
- Coroutines per thread =
4
- Tiny-chunk count =
1024
assumptions (5)
- domain assumption RDMA NIC throughput is IOPS-bound for requests < 4 KB
- domain assumption Real-world graphs exhibit power-law degree distributions
- domain assumption Chunk-based contiguous partitioning preserves locality in real graphs
- domain assumption DM compute nodes have 1-2 GB local cache and MNs have scarce CPU
- domain assumption DM pool is globally addressable via 16-bit MN ID + 48-bit offset
Cite this review
Pith. "Pith review of DMG: A Scalable and Efficient Memory-Disaggregated Graph Processing System." pith.science (2026). https://pith.science/paper/LKQIBDWU
@misc{pith2026260720881,
author = {Pith},
title = {Pith review of: DMG: A Scalable and Efficient Memory-Disaggregated Graph Processing System},
year = {2026},
howpublished = {\url{https://pith.science/paper/LKQIBDWU}},
note = {Machine review of arXiv:2607.20881}
}
read the original abstract
Traditional graph processing systems are built on monolithic servers, which couple a fixed ratio of compute and memory resources but often result in resource under-utilization in data centers. Although the disaggregated memory (DM) architecture has emerged to address this inefficiency, we identify that existing graph processing systems on DM remain highly impractical. They rely on unscalable architectures that fail to scale beyond a single memory node and a single compute node, and they require compute-side caches that are orders of magnitude larger than conventional practice in DM. To this end, this paper presents DMG, the first practical graph processing system on DM, which demonstrates superior system scalability and cache efficiency while delivering high performance. To improve efficiency of graph retrieval on DM, DMG proposes a DM-friendly graph store with retrieval optimizations. To mitigate costly update propagation, DMG presents an adaptive update coordinator that coordinates compute and memory nodes to perform update propagation with low overhead. To enable fast and effective load balancing, DMG employs a two-stage workload manager that includes a coarse-grained initial partitioning and a fine-grained runtime re-scheduling. Experimental results substantiate that compared with the state-of-the-art DM-based graph processing system, DMG can elastically scale up both compute and memory resources, delivering up to 4.9X better performance and accommodating graphs with ever-increasing sizes; meanwhile, it effectively tames the compute-side cache demands by up to 18.9X, positioning itself as a DM-ready solution in practice.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
[n. d.]. perftest: Infiniband Verbs Performance Tests. https://github.com/linux- rdma/perftest. https://github.com/linux-rdma/perftest
-
[2]
Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker
Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ouster- hout, Marcos K. Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker
-
[3]
Paolo Boldi and Sebastiano Vigna. 2004. The WebGraph Framework I: Com- pression Techniques. InProc. of the Thirteenth International World Wide Web Conference (WWW 2004). ACM Press, Manhattan, USA, 595–601
2004
-
[4]
Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli
Irina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli. 2021. Rethinking software runtimes for disag- gregated memory. InProceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems(Virtual, USA)(ASPLOS ’21). Association for Computing Machi...
2021
-
[5]
Deepayan Chakrabarti, Yiping Zhan, and Christos Faloutsos. [n. d.].R-MAT: A Recursive Model for Graph Mining. 442–446. arXiv:https://epubs.siam.org/doi/pdf/10.1137/1.9781611972740.43 doi:10.1137/1.9781611972740.43
-
[6]
Dechuang Chen, Sibo Wang, and Qintian Guo. 2025. ACGraph: An Efficient Asynchronous Out-of-Core Graph Processing Framework.Proc. ACM Manag. Data3, 6, Article 290 (Dec. 2025), 26 pages. doi:10.1145/3769755
-
[7]
Rong Chen, Jiaxin Shi, Yanzhe Chen, and Haibo Chen. 2015. PowerLyra: differen- tiated graph computation and partitioning on skewed graphs. InProceedings of the Tenth European Conference on Computer Systems(Bordeaux, France)(EuroSys ’15). New York, NY, USA, Article 1, 15 pages. doi:10.1145/2741948.2741970
arXiv 2015
-
[8]
Zheng Chen, Feng Zhang, JiaWei Guan, Jidong Zhai, Xipeng Shen, Huanchen Zhang, Wentong Shu, and Xiaoyong Du. 2023. CompressGraph: Efficient Parallel Graph Analytics with Rule-Based Compression.Proc. ACM Manag. Data1, 1, Article 4 (May 2023), 31 pages
2023
Show all 101 references
-
[9]
Pengjie Cui, Haotian Liu, Dong Jiang, Bo Tang, and Ye Yuan. 2025. Nezha: An Efficient Distributed Graph Processing System on Heterogeneous Hardware. Proc. ACM Manag. Data3, 1, Article 57 (Feb. 2025), 27 pages. doi:10.1145/3709707
2025 doi
-
[10]
Roshan Dathathri, Gurbinder Gill, Loc Hoang, Hoang-Vu Dang, Alex Brooks, Nikoli Dryden, Marc Snir, and Keshav Pingali. 2018. Gluon: A communication- optimizing substrate for distributed heterogeneous graph analytics. InProceed- ings of the 39th ACM SIGPLAN conference on progra...
2018
-
[11]
Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Pra...
2019
-
[12]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization.arXiv preprint arXiv:2404.16130 (2024)
2024 arXiv
-
[13]
Orri Erling, Alex Averbuch, Josep Larriba-Pey, Hassan Chafi, Andrey Gubichev, Arnau Prat, Minh-Duc Pham, and Peter Boncz. 2015. The LDBC Social Net- work Benchmark: Interactive Workload. InProceedings of the 2015 ACM SIGMOD International Conference on Management of Data(Melbou...
2015
-
[14]
Gurbinder Gill, Roshan Dathathri, Loc Hoang, and Keshav Pingali. 2018. A study of partitioning policies for graph analytics on large-scale distributed platforms. Proceedings of the VLDB Endowment12, 4 (2018), 321–334
2018
-
[15]
Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin
Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin
-
[16]
Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Sangwon Lee, and My- oungsoo Jung. 2023. Memory Pooling With CXL.IEEE Micro43, 2 (2023), 48–57
2023
-
[17]
Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, and Kang G Shin. 2017. Efficient memory disaggregation with infiniswap. In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). 649–667
2017
-
[18]
Hao Guo and Youyou Lu. 2025. Achieving Low-Latency Graph-Based Vector Search via Aligning Best-First Search Algorithm with SSD. In19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). USENIX Association, Boston, MA, 171–186
2025
-
[19]
Jing Guo, Zihao Chang, Sa Wang, Haiyang Ding, Yihui Feng, Liang Mao, and Yungang Bao. 2019. Who limits the resource efficiency of my datacenter: an anal- ysis of Alibaba datacenter traces. InProceedings of the International Symposium on Quality of Service(Phoenix, Arizona)(IWQ...
2019
-
[20]
Bernstein
Zhihan Guo, Xinyu Zeng, Kan Wu, Wuh-Chwen Hwang, Ziwei Ren, Xiangyao Yu, Mahesh Balakrishnan, and Philip A. Bernstein. 2022. Cornus: atomic commit for a cloud DBMS with storage disaggregation.Proc. VLDB Endow.16, 2 (Oct. 2022), 379–392
2022
-
[21]
Zhisheng Hu, Pengfei Zuo, Yizou Chen, Chao Wang, Junliang Hu, and Ming- Chang Yang. 2024. Aceso: Achieving Efficient Fault Tolerance in Memory- Disaggregated Key-Value Stores. InProceedings of the ACM SIGOPS 30th Sympo- sium on Operating Systems Principles(Austin, TX, USA)(SOS...
2024
-
[22]
Chengying Huan, Zhengyi Yang, Haoshen Yang, Shaonan Ma, Rong Gu, Fang Xi, Yongchao Liu, Guihai Chen, and Chen Tian. 2025. Gem: Scalable Monotonic Graph Processing Beyond Billion-Scale on a Single Machine.Proc. ACM Manag. Data3, 6, Article 330 (Dec. 2025), 30 pages. doi:10.1145/3769795
2025 doi
-
[23]
Jialiang Huang, MingXing Zhang, Teng Ma, Zheng Liu, Sixing Lin, Kang Chen, Jinlei Jiang, Xia Liao, Yingdi Shan, Ning Zhang, Mengting Lu, Tao Ma, Haifeng Gong, and YongWei Wu. 2024. TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Node...
2024
-
[24]
Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019. DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node. InAdvances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. ...
2019
-
[25]
Anuj Kalia, Michael Kaminsky, and David G Andersen. 2016. Design guide- lines for high performance{RDMA} systems. In2016 USENIX Annual Technical Conference (USENIX ATC 16). 437–450
2016
-
[26]
Pradeep Kumar and H Howie Huang. 2020. Graphone: A data store for real-time analytics on evolving graphs.ACM Transactions on Storage (TOS)15, 4 (2020), 1–40
2020
-
[27]
Haewoon Kwak, Changhyun Lee, Hosung Park, and Sue Moon. 2010. What is Twitter, a social network or a news media?. InProceedings of the 19th international conference on World wide web. 591–600
2010
-
[28]
Aapo Kyrola, Guy Blelloch, and Carlos Guestrin. 2012. GraphChi: Large-Scale Graph Computation on Just a PC. In10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12). Hollywood, CA, 31–46
2012
-
[29]
Aguilera, Kimberly Keeton, and Vijay Chidambaram
Sekwon Lee, Soujanya Ponnapalli, Sharad Singhal, Marcos K. Aguilera, Kimberly Keeton, and Vijay Chidambaram. 2022. DINOMO: An Elastic, Scalable, High- Performance Key-Value Store for Disaggregated Persistent Memory.Proc. VLDB Endow.15, 13 (Sept. 2022), 4023–4037. doi:10.14778/...
2022
-
[30]
Seung-Seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong, and Abhishek Bhattacharjee. 2021. MIND: In-Network Memory Management for Disaggregated Data Centers. InSOSP ’21: ACM SIGOPS 28th Symposium on Operating Systems Principles, Koblenz, Germany. ACM, 488–504
2021
-
[31]
Guoliang Li, Wengang Tian, Jinyu Zhang, Ronen Grosman, Zongchao Liu, and Sihao Li. 2024. GaussDB: A Cloud-Native Multi-Primary Database with Compute-Memory-Storage Disaggregation.Proc. VLDB Endow.17, 12 (Aug. 2024), 3786–3798. doi:10.14778/3685800.3685806
2024
-
[32]
Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D
Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bianchini. 2023. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. InProceed...
2023
-
[33]
Compute Express Link. 2023. Compute express link: The breakthrough cpu-to- device interconnect. https://www.computeexpresslink.org/
2023
-
[34]
Howie Huang
Hang Liu and H. Howie Huang. 2017. Graphene: Fine-Grained IO Management for Graph Computing. In15th USENIX Conference on File and Storage Technologies (FAST 17). Santa Clara, CA, 285–300
2017
-
[35]
Yushi Liu, Shixuan Sun, Zijun Li, Quan Chen, Sen Gao, Bingsheng He, Chao Li, and Minyi Guo. 2024. FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless Computing. InProceedings of the 29th ACM International Conference on Architectural Sup...
2024 doi
-
[36]
Baotong Lu, Kaisong Huang, Chieh-Jan Mike Liang, Tianzheng Wang, and Eric Lo. 2024. DEX: Scalable Range Indexing on Disaggregated Memory.Proc. VLDB Endow.17, 10 (Aug. 2024), 2603–2616
2024
-
[37]
Chengzhi Lu, Kejiang Ye, Guoyao Xu, Cheng-Zhong Xu, and Tongxin Bai. 2017. Imbalance in the cloud: An analysis on Alibaba cluster trace. In2017 IEEE Inter- national Conference on Big Data (Big Data). 2884–2892. doi:10.1109/BigData.2017. 8258257
2017 doi
-
[38]
Haodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan, Xiaofei Liao, Hai Jin, and Yu Zhang. 2025. Fast distributed transactions for RDMA-based disaggregated memory. InProceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference(Boston, MA, USA)(USENIX ATC ’25). U...
2025
-
[39]
Lyu, and Yang- fan Zhou
Xuchuan Luo, Jiacheng Shen, Pengfei Zuo, Xin Wang, Michael R. Lyu, and Yang- fan Zhou. 2024. CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles(Austin, TX, USA)(SOSP...
2024
-
[40]
Lyu, and Yangfan Zhou
Xuchuan Luo, Pengfei Zuo, Jiacheng Shen, Jiazhen Gu, Xin Wang, Michael R. Lyu, and Yangfan Zhou. 2023. SMART: A High-Performance Adaptive Radix Tree for Disaggregated Memory. In17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). Boston, MA
2023
-
[41]
Peter Macko, Virendra J Marathe, Daniel W Margo, and Margo I Seltzer. 2015. Llama: Efficient graph analytics using large multiversioned arrays. In2015 IEEE 31st International Conference on Data Engineering. IEEE, 363–374
2015
-
[42]
Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.IEEE transactions on pattern analysis and machine intelligence42, 4 (2018), 824–836
2018
-
[43]
Kanaujia, and Prakash Chauhan
Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agar- wal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit O. Kanaujia, and Prakash Chauhan. 2023. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory. InProceedings of the 28th ACM ...
2023
-
[44]
Donald Nguyen, Andrew Lenharth, and Keshav Pingali. 2013. A lightweight infrastructure for graph analytics. InProceedings of the twenty-fourth ACM sym- posium on operating systems principles. 456–471
2013
-
[45]
Vlad Nitu, Boris Teabe, Alain Tchana, Canturk Isci, and Daniel Hagimont. 2018. Welcome to zombieland: practical and energy-efficient memory disaggregation in a datacenter. InProceedings of the Thirteenth EuroSys Conference(Porto, Portugal) (EuroSys ’18). Association for Comput...
2018
-
[46]
Prashant Pandey, Brian Wheatman, Helen Xu, and Aydin Buluc. 2021. Terrace: A hierarchical graph container for skewed dynamic graphs. InProceedings of the 2021 international conference on management of data. 1372–1385
2021
-
[47]
Xi Pang and Jianguo Wang. 2024. Understanding the Performance Implications of the Design Principles in Storage-Disaggregated Databases.Proc. ACM Manag. Data2, 3, Article 180 (May 2024), 26 pages. doi:10.1145/3654983
2024 doi
-
[48]
Feng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia, Zuoning Chen, and Yongwei Wu. 2024. Scaling Up Memory Disaggregated Applications with SMART. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volum...
2024
-
[49]
Jie Ren, Minjia Zhang, and Dong Li. 2020. HM-ANN: efficient billion-point nearest neighbor search on heterogeneous memory. InProceedings of the 34th International Conference on Neural Information Processing Systems(Vancouver, BC, Canada)(NIPS ’20). Curran Associates Inc., Red ...
2020
-
[50]
André Ryser, Alberto Lerner, Alex Forencich, and Philippe Cudré-Mauroux. 2022. D-RDMA: Bringing Zero-Copy RDMA to Database Systems. In12th Conference on Innovative Data Systems Research, CIDR 2022, Chaminade, CA, USA, January 9-12,
2022
-
[51]
Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. 2018. LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). Carlsbad, CA
2018
-
[52]
Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Yuxin Su, Jiazhen Gu, Hao Feng, Yangfan Zhou, and Michael R Lyu. 2023. Ditto: An elastic and adaptive memory- disaggregated caching system. InProceedings of the 29th Symposium on Operating Systems Principles. 675–691
2023
-
[53]
Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Tianyi Yang, Yuxin Su, Yangfan Zhou, and Michael R. Lyu. 2023. FUSEE: A Fully Memory-Disaggregated Key-Value Store. In21st USENIX Conference on File and Storage Technologies (FAST 23). Santa Clara, CA
2023
-
[54]
Blelloch
Julian Shun and Guy E. Blelloch. 2013. Ligra: a lightweight graph processing framework for shared memory. InProceedings of the 18th ACM SIGPLAN sym- posium on Principles and practice of parallel programming. New York, NY, USA, 135–146
2013
-
[55]
Jixian Su, Chiyu Hao, Shixuan Sun, Hao Zhang, Sen Gao, Jiaxin Jiang, Yao Chen, Chenyi Zhang, Bingsheng He, and Minyi Guo. 2025. Revisiting the Design of In-Memory Dynamic Graph Storage.Proc. ACM Manag. Data3, 1, Article 70 (Feb. 2025), 27 pages. doi:10.1145/3709720
2025 doi
-
[56]
John Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng, Guanzhou Hu, Zhihao Jia, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2021. Dorylus: Affordable, Scalable, and Accurate GNN Training with Dis- tributed CPU Servers and Serverless Threads. In...
2021
-
[57]
Haque, Zhijing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes
Muhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque, Zhijing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes. 2020. Borg: the next generation. InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Association for Co...
2020
-
[58]
Infiniband trade association. 2023. InfiniBand. https://www.infinibandta.org/
2023
-
[59]
Chenxi Wang, Haoran Ma, Shi Liu, Yifan Qiao, Jonathan Eyolfson, Christian Navasca, Shan Lu, and Guoqing Harry Xu. 2022. MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime. In16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2...
2022
-
[60]
Chenxi Wang, Yifan Qiao, Haoran Ma, Shi Liu, Wenguang Chen, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2023. Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory. In20th USENIX Symposium on Networked Systems Design and Implementation, NSDI 202...
2023
-
[61]
Jing Wang, Chao Li, Taolei Wang, Lu Zhang, Pengyu Wang, Junyi Mei, and Minyi Guo. 2022. Excavating the potential of graph workload on rdma-based far memory architecture. In2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1029–1039
2022
-
[62]
Qing Wang, Youyou Lu, and Jiwu Shu. 2022. Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory. InProceedings of the 2022 International Conference on Management of Data(Philadelphia, PA, USA)(SIG- MOD ’22). New York, NY, USA
2022
-
[63]
TamerÖzsu, and Walid G
Ruihong Wang, Chuqing Gao, Jianguo Wang, Prishita Kadam, M. TamerÖzsu, and Walid G. Aref. 2024. Optimizing LSM-based indexes for disaggregated memory. DMG : A Scalable and Efficient Memory-Disaggregated Graph Processing System The VLDB Journal33, 6 (June 2024), 1813–1836. doi:...
2024 doi
-
[65]
Rui Wang, Weixu Zong, Shuibing He, Xinyu Chen, Zhenxin Li, and Zheng Dang
-
[66]
Yuke Wang, Boyuan Feng, Zheng Wang, Tong Geng, Kevin Barker, Ang Li, and Yufei Ding. 2023. MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms. In17th USENIX Symposium on Operating Systems Design an...
2023
-
[67]
Zilong Wang, Xinchen Wan, Luyang Li, Yijun Sun, Peng Xie, Xin Wei, Qingsong Ning, Junxue Zhang, and Kai Chen. 2024. Fast, Scalable, and Accurate Rate Limiter for RDMA NICs. InProceedings of the ACM SIGCOMM 2024 Conference, ACM SIGCOMM 2024, Sydney, NSW, Australia, August 4-8, ...
2024
-
[68]
Xingda Wei, Zhiyuan Dong, Rong Chen, and Haibo Chen. 2018. Deconstructing RDMA-enabled Distributed Transactions: Hybrid is Better!. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 233–251
2018
-
[69]
Marcel Weisgut, Daniel Ritter, Pinar Tözün, Lawrence Benson, and Tilmann Rabl
-
[70]
Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. 2022. Graph Neural Networks in Recommender Systems: A Survey.ACM Comput. Surv.55, 5, Article 97 (Dec. 2022), 37 pages
2022
-
[71]
Haoxuan Xie, Junfeng Liu, Siqiang Luo, and Kai Wang. 2026. RadixGraph: A Fast, Space-Optimized Data Structure for Dynamic Graph Storage (Extended Version). arXiv:2601.01444 [cs.DB] https://arxiv.org/abs/2601.01444
2026
-
[72]
Bin Yan, Youyou Lu, Qing Wang, Minhui Xie, and Jiwu Shu. 2023. Patronus: High-Performance and Protective Remote Memory. In21st USENIX Conference on File and Storage Technologies (FAST 23). USENIX Association, Santa Clara, CA, 315–330. https://www.usenix.org/conference/fast23/p...
2023
-
[73]
Ke Yang, MingXing Zhang, Kang Chen, Xiaosong Ma, Yang Bai, and Yong Jiang
-
[74]
Tsun-Yu Yang, Yizou Chen, Yuhong Liang, and Ming-Chang Yang. 2024. Ser- aph: Towards Scalable and Efficient Fully-external Graph Computation via On- demand Processing. In22nd USENIX Conference on File and Storage Technologies (FAST 24). Santa Clara, CA, 373–387
2024
-
[75]
Tsun-Yu Yang, Yizou Chen, Yuhong Liang, and Ming-Chang Yang. 2025. Lever- aging On-demand Processing to Co-optimize Scalability and Efficiency for Fully- external Graph Computation.ACM Trans. Storage21, 2, Article 11 (Feb. 2025), 31 pages. doi:10.1145/3701037
2025 doi
-
[76]
Tsun-Yu Yang, Yi Li, Yizou Chen, Bingzhe Li, and Ming-Chang Yang. 2025. Oa- sis: An Out-of-core Approximate Graph System via All-Distances Sketches. In 23rd USENIX Conference on File and Storage Technologies (FAST 25). USENIX Association, Santa Clara, CA, 523–537
2025
-
[77]
Xinjun Yang, Yingqiang Zhang, Hao Chen, Feifei Li, Gerry Fan, Yang Kong, Bo Wang, Jing Fang, Yuhui Wang, Tao Huang, Wenpu Hu, Jim Kao, and Jianping Jiang. 2025. Unlocking the Potential of CXL for Disaggregated Memory in Cloud- Native Databases. InCompanion of the 2025 Internat...
2025
-
[78]
Xinjun Yang, Yingqiang Zhang, Hao Chen, Feifei Li, Bo Wang, Jing Fang, Chuan Sun, and Yuhui Wang. 2024. PolarDB-MP: A Multi-Primary Cloud- Native Database via Disaggregated Shared Memory. InCompanion of the 2024 International Conference on Management of Data(Santiago AA, Chile...
2024
-
[79]
Yiwei Yang, Pooneh Safayenikoo, Jiacheng Ma, Tanvir Ahmed Khan, and An- drew Quinn. 2023. CXLMemSim: A pure software simulated CXL. mem for performance characterization.arXiv preprint arXiv:2303.06153(2023)
2023 arXiv
-
[80]
Peiqi Yin, Xiao Yan, Shiyuan Deng, Hui Li, Yifan Zhu, Xiangyu Zhi, Jingqi Mao, Ran Xu, Wenliang Zhang, and James Cheng. 2026. DistVS: Large-scale Vector Search with Compute-Memory Disaggregation. In23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26)....
2026
-
[81]
Song Yu, Shufeng Gong, Qian Tao, Sijie Shen, Yanfeng Zhang, Wenyuan Yu, Pengxi Liu, Zhixin Zhang, Hongfu Li, Xiaojian Luo, Ge Yu, and Jingren Zhou
-
[82]
Xiangyao Yu. 2025. Disaggregation: A New Architecture for Cloud Databases. Proc. VLDB Endow.18, 12 (Aug. 2025), 5527–5530. doi:10.14778/3750601.3760520
2025
-
[83]
Daniel Zahka and Ada Gavrilovska. 2022. FAM-Graph: Graph analytics on disag- gregated memory. In2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 81–92
2022
-
[84]
Shaoxun Zeng, Xiaojian Liao, Hao Guo, and Youyou Lu. 2024. Volley: Accelerat- ing Write-Read Orders in Disaggregated Storage. InProceedings of the Nineteenth European Conference on Computer Systems(Athens, Greece)(EuroSys ’24). Asso- ciation for Computing Machinery, New York, ...
2024
-
[85]
Hantian Zha, Teng Ma, Baotong Lu, Yuansen Wang, Dongbiao He, Yuanhui Luo, Dafang Zhang, Yunpeng Chai, Yuxing Chen, and Anqun Pan. 2025. Shard: A Scalable and Resize-Optimized Hash Index on Disaggregated Memory.Proc. VLDB Endow.19, 4 (Dec. 2025), 684–697. doi:10.14778/3785297.3785309
2025
-
[86]
Chenzi Zhang, Fan Wei, Qin Liu, Zhihao Gavin Tang, and Zhenguo Li. 2017. Graph Edge Partitioning via Neighborhood Heuristic. InProceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax, NS, Canada)(KDD ’17). Association for Com...
2017
-
[87]
Ming Zhang, Yu Hua, and Zhijun Yang. 2024. Motor: Enabling Multi-Versioning for Distributed Transactions on Disaggregated Memory. In18th USENIX Sympo- sium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA. USENIX Association, 801–819
2024
-
[88]
ACM Manag
LSMGraph: A High-Performance Dynamic Graph Storage System with Multi-Level CSR.Proc. ACM Manag. Data2, 6, Article 243 (Dec. 2024), 28 pages. doi:10.1145/3698818
2024 doi
-
[89]
Qizhen Zhang, Xinyi Chen, Sidharth Sankhe, Zhilei Zheng, Ke Zhong, Sebastian Angel, Ang Chen, Vincent Liu, and Boon Thau Loo. 2022. Optimizing data- intensive systems in disaggregated data centers with teleport. InProceedings of the 2022 International Conference on Management ...
2022
-
[90]
Wenqian Zhang, Zhengyi Yang, Dong Wen, Wentao Li, Wenjie Zhang, and Xuemin Lin. 2025. Accelerating Core Decomposition in Billion-Scale Hy- pergraphs.Proc. ACM Manag. Data3, 1, Article 6 (Feb. 2025), 27 pages. doi:10.1145/3709656
2025 doi
-
[91]
Yang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao, James Mickens, Minlan Yu, Chris Kennelly, Paul Turner, David E. Culler, Henry M. Levy, and Amin Vahdat. 2022. Carbink: Fault-Tolerant Far Memory. In16th USENIX Symposium on Operating Systems Design and Implementation, OSDI...
2022
-
[92]
Xiaowei Zhu, Wenguang Chen, Weimin Zheng, and Xiaosong Ma. 2016. Gemini: A{Computation-Centric} distributed graph processing system. In12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 301–316
2016
-
[93]
Xiaowei Zhu, Wentao Han, and Wenguang Chen. 2015. GridGraph: Large-Scale Graph Processing on a Single Machine Using 2-Level Hierarchical Partitioning. In2015 USENIX Annual Technical Conference (USENIX ATC 15). Santa Clara, CA, 375–386
2015
-
[94]
Zhiting Zhu, Newton Ni, Yibo Huang, Yan Sun, Zhipeng Jia, Nam Sung Kim, and Emmett Witchel. 2024. Lupin: Tolerating Partial Failures in a CXL Pod. In Proceedings of the 2nd Workshop on Disruptive Memory Systems(Austin, TX, USA) (DIMES ’24). Association for Computing Machinery,...
2024
-
[95]
Mingxing Zhang, Teng Ma, Jinqi Hua, Zheng Liu, Kang Chen, Ning Ding, Fan Du, Jinlei Jiang, Tao Ma, and Yongwei Wu. 2023. Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared Memory. InProceedings of the 29th Symposium on Operating Systems Princ...
2023
-
[102]
Tobias Ziegler, Sumukha Tumkur Vani, Carsten Binnig, Rodrigo Fonseca, and Tim Kraska. 2019. Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks. InProceedings of the 2019 International Conference on Management of Data(Amsterdam, Netherlands)(SIGMOD...
2019
-
[2012]
In10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12)
PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs. In10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12). Hollywood, CA, 17–30
-
[2019]
InProceedings of the 27th ACM Symposium on Operating Systems Principles(Huntsville, Ontario, Canada)(SOSP ’19)
KnightKing: a fast distributed graph random walk engine. InProceedings of the 27th ACM Symposium on Operating Systems Principles(Huntsville, Ontario, Canada)(SOSP ’19). Association for Computing Machinery, New York, NY, USA, 524–537
-
[2020]
InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20)
Can far memory improve job throughput?. InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Association for Computing Machinery, New York, NY, USA, Article 14, 16 pages
-
[2022]
https://www.cidrdb.org/cidr2022/papers/p77-ryser.pdf
www.cidrdb.org. https://www.cidrdb.org/cidr2022/papers/p77-ryser.pdf
-
[2024]
In2024 USENIX Annual Technical Conference (USENIX ATC 24)
Efficient Large Graph Processing with Chunk-Based Graph Representation Model. In2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 1239–1255
-
[2025]
VLDB Endow.18, 9 (May 2025), 3119–3133
CXL Memory Performance for In-Memory Data Processing.Proc. VLDB Endow.18, 9 (May 2025), 3119–3133. doi:10.14778/3746405.3746432
2025
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.