Pith. sign in

REVIEW 4 major objections 4 minor 101 references

DMG: A Scalable and Efficient Memory-Disaggregated Graph Processing System

T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This paper claims that DMG is the first practical graph processing system for disaggregated memory, one that scales across multiple compute and memory nodes while keeping compute-side caches at conventional DM sizes and delivering performan

desk verdict Genuinely new multi-CN/MN graph processing on DM, with a thorough ablation, but the MN-CPU assumption is tested only on 24-core EPYC nodes and could be the weak point. read the letter →

arxiv 2607.20881 v1 pith:LKQIBDWU submitted 2026-07-23 cs.DB cs.DC

classification cs.DBcs.DC
keywords graphprocessingdisaggregatedmemoryRDMACXLcacheefficiencypartitioningloadbalancingstore
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Traditional distributed graph platforms couple CPU and memory in fixed-ratio servers, wasting resources; existing attempts to run graph analytics on disaggregated memory are stuck on one compute node and one memory node and demand tens of gigabytes of compute-side cache. DMG claims to be the first practical disaggregated-memory graph system that scales out both compute and memory elastically, needs only about 1–2 GB of cache per compute node, and still runs billion-edge graphs within roughly 40% of a local-memory distributed system’s computation time. It achieves this with a DM-friendly graph store that inlines short edge lists into the index, an adaptive update coordinator that moves update work to whichever node holds the destination vertex’s attributes, and a two-stage workload manager that partitions fast at startup and re-schedules hub vertices at runtime. If these claims hold, disaggregated memory stops being a toy for graph analytics and becomes a way to allocate graph memory independently of compute.

What carries the argument

The central mechanism is an adaptive per-vertex index and edge-store combined with a redistribute update path. The 32-byte per-vertex index lives in a single range-partitioned array on the memory pool; for low-degree vertices it stores the edge list inline (‘in-place’), turning two dependent remote reads into one, while for high-degree vertices it stores the edge-list address and compressed per-segment lengths (‘out-of-place’). The segment lengths let a node fetch only the portion of a hub’s edge list that falls in a given partition, and they let the update coordinator hand whole segments to the node that owns the destination attributes. This machinery converts RDMA’s IOPS bottleneck into fe

What would settle it

Run the same BFS and PageRank workloads with memory nodes limited to one or two low-power cores and measure both end-to-end time and memory-node CPU utilization during the densest iteration; if any memory node’s utilization saturates or the end-to-end time degrades disproportionately, the collaborative-update claim collapses.

Watch

Extended reading notes

Core claim

The paper’s core discovery is that the three obstacles to practical graph processing on DM—IOPS-limited remote reads, costly remote update propagation, and tail effects from hub vertices—can each be turned around by exploiting where data already resides. DMG stores vertex attributes once in a shared memory pool and gives each compute node a small cache covering its assigned chunk. For retrieval, a 32-byte per-vertex index entry either embeds the edge list of low-degree vertices or stores a compressed segment layout for high-degree vertices, so most vertex reads become one merged RDMA request instead of two dependent fine-grained ones. For updates, a collaborative scheme batches update candid

Load-bearing premise

The load-bearing premise is that a memory node’s scarce CPU can absorb the offloaded update work (ValRD) and RPC service without becoming a bottleneck—the testbed gives each memory node a 24-core server running only two threads, so the reported numbers depend on memory-node CPU being cheap.

Editorial extensions

If this is right

  • Disaggregated-memory graph systems can store one copy of the graph in a shared memory pool and elastically add compute nodes without re-coupling memory, so tenants pay only for the resource they need.
  • A conventional 1–2 GB compute-side cache is enough for billion-scale graphs, because each compute node caches only the attribute slice of its assigned chunk; aggregate compute-node memory stays a small fraction of memory-pool usage.
  • Graph partitioning for load balancing can be redone in sub-seconds using tiny-chunk metadata, making repeated resource reconfiguration practical on DM.
  • The computation-time gap versus a local-memory distributed system stands at about 40%, and the paper argues this gap will narrow as interconnect bandwidth rises toward 800 Gbps and beyond.
  • The same storage and update techniques carry over to CXL-based memory pooling, because they reduce fine-grained remote accesses rather than relying on RDMA-specific behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not test memory nodes with truly scarce CPU: its memory nodes are 24-core servers running only two threads. If a real memory node has one or two weak cores, the offloaded ValRD update path could saturate that CPU, and the reported cache savings and speedups would shrink.
  • The in-place index scheme exploits power-law degree distributions; on graphs with more uniform degrees, fewer vertices fit inline and the retrieval benefit should diminish. The paper includes a synthetic R-MAT graph but does not isolate this effect.
  • The collaborative-update advantage depends on chunk locality; a graph whose vertex IDs are shuffled to destroy locality would send far more update candidates across the network. DMG does not report a degradation curve for such adversarial layouts.
  • Because the loaded graph store and segment metadata are reusable across different compute-node counts, a natural extension is mid-query elastic resizing of the compute pool without reloading the graph; the paper does not implement this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. DMG proposes a graph processing system for memory-disaggregated (DM) architectures. It contributes three main designs: (i) a DM-friendly graph store that embeds low-degree edge lists in an enlarged index and applies merged/batched RDMA retrieval; (ii) an adaptive update coordinator that selects between collaborative update (with pass-by-reference and pass-by-value re-distribution) and direct remote update depending on update density; and (iii) a two-stage workload manager using coarse-grained tiny-chunk partitioning plus runtime hub re-scheduling. The paper claims this is the first practical multi-CN/multi-MN graph processing system on DM, achieving up to 4.9× speedup over FAM-Graph, up to 18.9× reduction in compute-side cache demand, and computation time within 40% of the local-memory distributed system Gemini. The evaluation includes internal ablations over the proposed components and comparisons with FAM-Graph, DMG-Base, and Gemini on four billion-scale graphs.

Significance. If the claims hold, this is a substantial contribution to the DM systems literature. The paper identifies an important practical gap in existing DM graph systems (single CN/MN and large compute-side caches) and proposes concrete mechanisms to address it. A particular strength is the internal ablation in §5.4: separating the effects of index embedding, retrieval optimizations, the two re-distribution modes, and runtime re-scheduling gives good evidence that the design choices, not just the overall architecture, drive the reported gains. The cache-efficiency result (orders of magnitude below FAM-Graph) and the startup-time comparison with Gemini are also compelling. However, the central scalability and performance claims rest on assumptions about memory-node CPU availability and on a fixed 4-MN testbed; these are not yet fully stress-tested, and the absence of repeated-run statistics makes some quantitative claims difficult to assess.

major comments (4)
  1. [§5.1/§5.4.2] MN CPU model: The system model in §2.2 assumes MNs have weak computation power (e.g., 1–2 CPU cores), but the testbed uses 24-core EPYC 7402P MNs (only two threads active). The ValRD mechanism (§4.3.1) is load-bearing: Fig. 25 shows +ValRD provides a large improvement over +RefRD. However, the MN CPU validation in Fig. 26 reports at most 44% usage of one EPYC core for TW; this does not establish that the same offload is sustainable on a real 1–2 core, lower-performance MN that also runs RPC-serving threads. I recommend adding an experiment that pins MN threads to one or two cores, or otherwise models weak MNs, and reports end-to-end performance and ValRD CPU/throughput under that constraint. If ValRD saturates, the fallback path could erode the reported speedups.
  2. [§5.2/§5.1] Memory-side scalability is not directly demonstrated. The abstract claims elastic scaling of both compute and memory, but all scaling experiments vary only the number of CNs (1, 2, 4) with a fixed 4-MN pool. I could not find an experiment that varies the number of MNs, or that increases graph size while scaling MNs. The clueweb12 result shows DMG can use a 4-MN pool where FAM-Graph cannot, but not that adding MNs elastically accommodates ever-increasing graphs. Please add a memory-scaling experiment (e.g., 1→4 MNs with fixed CN count, or a graph that grows with MN count) or temper the claim.
  3. [§5 (general methodology)] All reported numbers appear to be from single executions, with no error bars or variance information. This matters for the quantitative headline claims (4.9× speedup, 18.9× cache reduction) and for small differences such as the UK 0.91–1.19× speedup range in §5.2. I request repeated runs (at least 3–5 per configuration) with confidence intervals, or at minimum a statement of run-to-run variability, for the key figures (Figs. 15–19, 24, 25, 28).
  4. [§4.2.1/§5.4.1] Several design thresholds—32B index entry size, in-place degree ≤7, RS hub threshold >1024, 4 coroutines per thread, 1024 tiny-chunks—appear to be chosen based on the four benchmark graphs, and the index-size sweep (Fig. 22) is shown only for TW. Since the main comparisons use the same datasets, it is unclear whether these values are overfit to the testbed. Please add a sensitivity analysis for the most impactful thresholds across datasets and at least one other hardware configuration, or give an explicit argument that the thresholds are hardware- and dataset-independent.
minor comments (4)
  1. [Figs. 15–18] The captions do not define what the annotated ratios (e.g., '99x', '179x') refer to. Clarify whether they are speedups of DMG over DMG-Base or over FAM-Graph, and how FAM-Graph is plotted at 2 and 4 CNs when it supports only one CN.
  2. [Fig. 19] Please define 'per-CN cache usage' precisely (maximum RSS, allocated cache size, or measured working set) and state whether the FAM-Graph bar is for one CN only. This would help readers interpret the 18.9× claim.
  3. [§5.3] The statement that DMG achieves computation time 'within 40% of Gemini' is ambiguous: does it mean 40% slower, or 40% of Gemini's time? The text later says 'moderate computation overhead,' suggesting the former. Please rephrase.
  4. [Abstract and §5.3] There are small text issues: 'toDM-friendly' in the abstract and 'ontwitter-2010' in §5.3 are missing spaces. Also, the paper promises open-source code; please include the repository or artifact link at the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No load-bearing circularity: claims are measured against external baselines; self-citations and the MN-CPU setup are caveats, not circular inputs.

full rationale

DMG is an evaluation-driven systems paper; its central claims are established by direct measurements against external baselines (FAM-Graph, Gemini) and an internal ablation (DMG-Base) across four billion-scale graphs, not by a derivation in which an output is defined in terms of an input. The design constants (32B index, in-place degree threshold, RS threshold, 2GB cache) are presented as empirical tuning choices with parameter studies (Figs. 22-23) and ablations (Figs. 24-28), and the headline 4.9x/18.9x numbers are measured comparisons, not fitted outputs. Self-citations ([21] Aceso, [36] DEX, [74-76] Seraph/Oasis) support background claims or appear in related work; the one load-bearing assumption they touch—that MNs can execute offloaded ValRD work cheaply—is independently measured in Fig. 26, so the citations do not carry the argument. The stated limitations (§4.5.2 dual-mode selection restricted by hardware; §4.5.3 static graphs only) narrow scope but do not create circularity. The main caveat is external validity, not circularity: MNs are simulated on 24-core EPYC machines (§5.1) while §2.2 assumes 1-2 weak cores, so the ValRD CPU-usage numbers may overstate real-MN headroom; this affects how representative the results are, not whether the results were derived from their own assumptions.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The paper's central claims rely mainly on engineering assumptions about RDMA behavior, graph structure, and DM hardware conventions rather than unproved mathematical axioms. No new physical or formal entities are introduced. The principal 'costs' are empirical design constants fit to the benchmark graphs.

free parameters (5)
  • Index entry size = 32 bytes
    Chosen to balance performance and memory efficiency (Figure 22) on benchmark graphs; a fitted design constant.
  • In-place / ValRD degree threshold = degree ≤ 7
    Threshold below which edges are embedded in the index or passed by value; tuned to power-law graphs in the evaluation.
  • RS hub threshold = degree > 1024
    Degree above which a vertex's edge list is split into segments for runtime re-scheduling; chosen empirically (§4.4.2).
  • Coroutines per thread = 4
    Default concurrency level for hiding RDMA latency; set by the system and not swept in the evaluation.
  • Tiny-chunk count = 1024
    Granularity for coarse-grained initial partitioning; a fixed system constant chosen for fast partition times.
assumptions (5)
  • domain assumption RDMA NIC throughput is IOPS-bound for requests < 4 KB
    Used to justify increasing request size via in-place index and merged/batched reads; supported by Figure 6 but treated as a hardware property.
  • domain assumption Real-world graphs exhibit power-law degree distributions
    Underlies in-place scheme, RefRD/ValRD split, and hub-vertex RS; the four benchmark graphs may not cover all graph types.
  • domain assumption Chunk-based contiguous partitioning preserves locality in real graphs
    Adopted from Gemini [92]; DMG's cache-efficiency argument depends on neighbors being clustered within chunks.
  • domain assumption DM compute nodes have 1-2 GB local cache and MNs have scarce CPU
    Standard DM convention from prior work; the testbed uses 24-core machines as MNs, so the scarcity is not fully reproduced.
  • domain assumption DM pool is globally addressable via 16-bit MN ID + 48-bit offset
    Addressing scheme from prior DM systems [40,62]; assumed to hold in the target deployment.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DMG: A Scalable and Efficient Memory-Disaggregated Graph Processing System." pith.science (2026). https://pith.science/paper/LKQIBDWU

@misc{pith2026260720881,
  author       = {Pith},
  title        = {Pith review of: DMG: A Scalable and Efficient Memory-Disaggregated Graph Processing System},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LKQIBDWU}},
  note         = {Machine review of arXiv:2607.20881}
}
read the original abstract

Traditional graph processing systems are built on monolithic servers, which couple a fixed ratio of compute and memory resources but often result in resource under-utilization in data centers. Although the disaggregated memory (DM) architecture has emerged to address this inefficiency, we identify that existing graph processing systems on DM remain highly impractical. They rely on unscalable architectures that fail to scale beyond a single memory node and a single compute node, and they require compute-side caches that are orders of magnitude larger than conventional practice in DM. To this end, this paper presents DMG, the first practical graph processing system on DM, which demonstrates superior system scalability and cache efficiency while delivering high performance. To improve efficiency of graph retrieval on DM, DMG proposes a DM-friendly graph store with retrieval optimizations. To mitigate costly update propagation, DMG presents an adaptive update coordinator that coordinates compute and memory nodes to perform update propagation with low overhead. To enable fast and effective load balancing, DMG employs a two-stage workload manager that includes a coarse-grained initial partitioning and a fine-grained runtime re-scheduling. Experimental results substantiate that compared with the state-of-the-art DM-based graph processing system, DMG can elastically scale up both compute and memory resources, delivering up to 4.9X better performance and accommodating graphs with ever-increasing sizes; meanwhile, it effectively tames the compute-side cache demands by up to 18.9X, positioning itself as a DM-ready solution in practice.

Figures

Figures reproduced from arXiv: 2607.20881 by the authors.

Figure 1
Figure 1. Data structures and workflow in graph processing. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The Disaggregated Memory Architecture. The disaggregated memory (DM) architecture has recently been proposed for cloud infrastructure [45, 51] to address the resource inefficiency in traditional data centers built on monolithic servers, which couple different resource components together [19, 37, 57]. As shown in [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. An example of vertex replication in distributed [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Different system architectures and data layouts for [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: CDF of access granu￾larity for graph retrieval dur￾ing BFS on twitter-2010 8 32 128 512 2k 8k 32k Access Granularity (Bytes) 1 10 IOPS (Million/sec) IOPS bound IOPS Bandwidth 1 10 100 Bandwidth (Gbps) Bandwidth bound [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 9
Figure 9. Figure 9: The Overview of DMG. 4.2 DM-friendly Graph Store The storage format fundamentally determines the efficiency of storing and accessing graph data on DM [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 11
Figure 11. Figure 11: Merged and batched retrieval of graph data. [PITH_FULL_IMAGE:figures/full_fig_p007_11.png]
Figure 12
Figure 12. Figure 12: Different strategies in DMG to perform update [PITH_FULL_IMAGE:figures/full_fig_p007_12.png]
Figure 14
Figure 14. Figure 14: Structure of Batched Re-Distribution (RD) Buffer [PITH_FULL_IMAGE:figures/full_fig_p008_14.png]
Figure 15
Figure 15. Figure 15: Computation time on twitter-2010. 1 2 4 Number of CN (a) BFS 2 1 2 2 2 3 2 4 2 5 Comp. Time (s) FAM-Graph 1 2 4 Number of CN (b) CC 2 3 2 4 2 5 2 6 DMG-Base 1 2 4 Number of CN (c) PR 2 2 2 3 2 4 2 5 2 6 DMG [PITH_FULL_IMAGE:figures/full_fig_p010_15.png]
Figure 16
Figure 16. Figure 16: Computation time on uk-2007-05. 1 2 4 Number of CN (a) BFS 2 3 2 4 2 5 2 6 2 7 Comp. Time (s) 9.1x 12x FAM-Graph 1 2 4 Number of CN (b) CC 2 4 2 5 2 6 2 7 10x 16x DMG-Base 1 2 4 Number of CN (c) PR 2 6 2 7 2 8 2 9 2 10 16x 23x DMG [PITH_FULL_IMAGE:figures/full_fig_p0…
Figure 19
Figure 19. Figure 19: Cache demands in each CN. social network graph, rmat-29 (R29) as a synthetic graph gener￾ated by R-MAT [5] with default parameters, uk-2007-05 (UK) and clueweb12 (CW) [3] as web crawler graphs. We evaluate three rep￾resentative graph processing workloads: breadth-firs…
Figure 20
Figure 20. Figure 20: Startup time (s) of DMG and Gemini for clueweb12 with 4 servers/CNs. 10 1 10 2 10 3 Mem Usage (GB) 7 7 7 128 256 512 10 2 10 3 Mem Usage (GB) Gemini OOM Gemini OOM 195 195 196 512 1 2 4 CN/Server Num (a) twitter-2010 0 1 2 Comp Time (s) 1 2 4 CN/Server Num (b) clueweb…
Figure 24
Figure 24. Figure 24: Design effect of retrieval optimizations. [PITH_FULL_IMAGE:figures/full_fig_p012_24.png]
Figure 25
Figure 25. Figure 25: Design effect of two types of RD. BFS CC PR (a) twitter-2010 0 20 40 60 1 Core CPU Usage Rate (%) 21.8 44.4 30.2 44.1 36.2 36.2 BFS CC PR (b) clueweb12 0 10 20 30 3.86 8.67 7.24 13.9 0.86 0.86 Total Exec Densest Iter [PITH_FULL_IMAGE:figures/full_fig_p012_25.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

101 extracted references · 9 canonical work pages

  1. [1]

    [n. d.]. perftest: Infiniband Verbs Performance Tests. https://github.com/linux- rdma/perftest. https://github.com/linux-rdma/perftest

  2. [2]

    Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker

    Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ouster- hout, Marcos K. Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker

  3. [3]

    Paolo Boldi and Sebastiano Vigna. 2004. The WebGraph Framework I: Com- pression Techniques. InProc. of the Thirteenth International World Wide Web Conference (WWW 2004). ACM Press, Manhattan, USA, 595–601

  4. [4]

    Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli

    Irina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli. 2021. Rethinking software runtimes for disag- gregated memory. InProceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems(Virtual, USA)(ASPLOS ’21). Association for Computing Machi...

  5. [5]

    Deepayan Chakrabarti, Yiping Zhan, and Christos Faloutsos. [n. d.].R-MAT: A Recursive Model for Graph Mining. 442–446. arXiv:https://epubs.siam.org/doi/pdf/10.1137/1.9781611972740.43 doi:10.1137/1.9781611972740.43

  6. [6]

    Dechuang Chen, Sibo Wang, and Qintian Guo. 2025. ACGraph: An Efficient Asynchronous Out-of-Core Graph Processing Framework.Proc. ACM Manag. Data3, 6, Article 290 (Dec. 2025), 26 pages. doi:10.1145/3769755

  7. [7]

    Rong Chen, Jiaxin Shi, Yanzhe Chen, and Haibo Chen. 2015. PowerLyra: differen- tiated graph computation and partitioning on skewed graphs. InProceedings of the Tenth European Conference on Computer Systems(Bordeaux, France)(EuroSys ’15). New York, NY, USA, Article 1, 15 pages. doi:10.1145/2741948.2741970

  8. [8]

    Zheng Chen, Feng Zhang, JiaWei Guan, Jidong Zhai, Xipeng Shen, Huanchen Zhang, Wentong Shu, and Xiaoyong Du. 2023. CompressGraph: Efficient Parallel Graph Analytics with Rule-Based Compression.Proc. ACM Manag. Data1, 1, Article 4 (May 2023), 31 pages

Show all 101 references
  1. [9]

    Pengjie Cui, Haotian Liu, Dong Jiang, Bo Tang, and Ye Yuan. 2025. Nezha: An Efficient Distributed Graph Processing System on Heterogeneous Hardware. Proc. ACM Manag. Data3, 1, Article 57 (Feb. 2025), 27 pages. doi:10.1145/3709707

  2. [10]

    Roshan Dathathri, Gurbinder Gill, Loc Hoang, Hoang-Vu Dang, Alex Brooks, Nikoli Dryden, Marc Snir, and Keshav Pingali. 2018. Gluon: A communication- optimizing substrate for distributed heterogeneous graph analytics. InProceed- ings of the 39th ACM SIGPLAN conference on progra...

  3. [11]

    Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Pra...

  4. [12]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization.arXiv preprint arXiv:2404.16130 (2024)

  5. [13]

    Orri Erling, Alex Averbuch, Josep Larriba-Pey, Hassan Chafi, Andrey Gubichev, Arnau Prat, Minh-Duc Pham, and Peter Boncz. 2015. The LDBC Social Net- work Benchmark: Interactive Workload. InProceedings of the 2015 ACM SIGMOD International Conference on Management of Data(Melbou...

  6. [14]

    Gurbinder Gill, Roshan Dathathri, Loc Hoang, and Keshav Pingali. 2018. A study of partitioning policies for graph analytics on large-scale distributed platforms. Proceedings of the VLDB Endowment12, 4 (2018), 321–334

  7. [15]

    Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin

    Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin

  8. [16]

    Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Sangwon Lee, and My- oungsoo Jung. 2023. Memory Pooling With CXL.IEEE Micro43, 2 (2023), 48–57

  9. [17]

    Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, and Kang G Shin. 2017. Efficient memory disaggregation with infiniswap. In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). 649–667

  10. [18]

    Hao Guo and Youyou Lu. 2025. Achieving Low-Latency Graph-Based Vector Search via Aligning Best-First Search Algorithm with SSD. In19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). USENIX Association, Boston, MA, 171–186

  11. [19]

    Jing Guo, Zihao Chang, Sa Wang, Haiyang Ding, Yihui Feng, Liang Mao, and Yungang Bao. 2019. Who limits the resource efficiency of my datacenter: an anal- ysis of Alibaba datacenter traces. InProceedings of the International Symposium on Quality of Service(Phoenix, Arizona)(IWQ...

  12. [20]

    Bernstein

    Zhihan Guo, Xinyu Zeng, Kan Wu, Wuh-Chwen Hwang, Ziwei Ren, Xiangyao Yu, Mahesh Balakrishnan, and Philip A. Bernstein. 2022. Cornus: atomic commit for a cloud DBMS with storage disaggregation.Proc. VLDB Endow.16, 2 (Oct. 2022), 379–392

  13. [21]

    Zhisheng Hu, Pengfei Zuo, Yizou Chen, Chao Wang, Junliang Hu, and Ming- Chang Yang. 2024. Aceso: Achieving Efficient Fault Tolerance in Memory- Disaggregated Key-Value Stores. InProceedings of the ACM SIGOPS 30th Sympo- sium on Operating Systems Principles(Austin, TX, USA)(SOS...

  14. [22]

    Chengying Huan, Zhengyi Yang, Haoshen Yang, Shaonan Ma, Rong Gu, Fang Xi, Yongchao Liu, Guihai Chen, and Chen Tian. 2025. Gem: Scalable Monotonic Graph Processing Beyond Billion-Scale on a Single Machine.Proc. ACM Manag. Data3, 6, Article 330 (Dec. 2025), 30 pages. doi:10.1145/3769795

  15. [23]

    Jialiang Huang, MingXing Zhang, Teng Ma, Zheng Liu, Sixing Lin, Kang Chen, Jinlei Jiang, Xia Liao, Yingdi Shan, Ning Zhang, Mengting Lu, Tao Ma, Haifeng Gong, and YongWei Wu. 2024. TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Node...

  16. [24]

    Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019. DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node. InAdvances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. ...

  17. [25]

    Anuj Kalia, Michael Kaminsky, and David G Andersen. 2016. Design guide- lines for high performance{RDMA} systems. In2016 USENIX Annual Technical Conference (USENIX ATC 16). 437–450

  18. [26]

    Pradeep Kumar and H Howie Huang. 2020. Graphone: A data store for real-time analytics on evolving graphs.ACM Transactions on Storage (TOS)15, 4 (2020), 1–40

  19. [27]

    Haewoon Kwak, Changhyun Lee, Hosung Park, and Sue Moon. 2010. What is Twitter, a social network or a news media?. InProceedings of the 19th international conference on World wide web. 591–600

  20. [28]

    Aapo Kyrola, Guy Blelloch, and Carlos Guestrin. 2012. GraphChi: Large-Scale Graph Computation on Just a PC. In10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12). Hollywood, CA, 31–46

  21. [29]

    Aguilera, Kimberly Keeton, and Vijay Chidambaram

    Sekwon Lee, Soujanya Ponnapalli, Sharad Singhal, Marcos K. Aguilera, Kimberly Keeton, and Vijay Chidambaram. 2022. DINOMO: An Elastic, Scalable, High- Performance Key-Value Store for Disaggregated Persistent Memory.Proc. VLDB Endow.15, 13 (Sept. 2022), 4023–4037. doi:10.14778/...

  22. [30]

    Seung-Seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong, and Abhishek Bhattacharjee. 2021. MIND: In-Network Memory Management for Disaggregated Data Centers. InSOSP ’21: ACM SIGOPS 28th Symposium on Operating Systems Principles, Koblenz, Germany. ACM, 488–504

  23. [31]

    Guoliang Li, Wengang Tian, Jinyu Zhang, Ronen Grosman, Zongchao Liu, and Sihao Li. 2024. GaussDB: A Cloud-Native Multi-Primary Database with Compute-Memory-Storage Disaggregation.Proc. VLDB Endow.17, 12 (Aug. 2024), 3786–3798. doi:10.14778/3685800.3685806

  24. [32]

    Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D

    Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bianchini. 2023. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. InProceed...

  25. [33]

    Compute Express Link. 2023. Compute express link: The breakthrough cpu-to- device interconnect. https://www.computeexpresslink.org/

  26. [34]

    Howie Huang

    Hang Liu and H. Howie Huang. 2017. Graphene: Fine-Grained IO Management for Graph Computing. In15th USENIX Conference on File and Storage Technologies (FAST 17). Santa Clara, CA, 285–300

  27. [35]

    Yushi Liu, Shixuan Sun, Zijun Li, Quan Chen, Sen Gao, Bingsheng He, Chao Li, and Minyi Guo. 2024. FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless Computing. InProceedings of the 29th ACM International Conference on Architectural Sup...

  28. [36]

    Baotong Lu, Kaisong Huang, Chieh-Jan Mike Liang, Tianzheng Wang, and Eric Lo. 2024. DEX: Scalable Range Indexing on Disaggregated Memory.Proc. VLDB Endow.17, 10 (Aug. 2024), 2603–2616

  29. [37]

    Chengzhi Lu, Kejiang Ye, Guoyao Xu, Cheng-Zhong Xu, and Tongxin Bai. 2017. Imbalance in the cloud: An analysis on Alibaba cluster trace. In2017 IEEE Inter- national Conference on Big Data (Big Data). 2884–2892. doi:10.1109/BigData.2017. 8258257

  30. [38]

    Haodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan, Xiaofei Liao, Hai Jin, and Yu Zhang. 2025. Fast distributed transactions for RDMA-based disaggregated memory. InProceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference(Boston, MA, USA)(USENIX ATC ’25). U...

  31. [39]

    Lyu, and Yang- fan Zhou

    Xuchuan Luo, Jiacheng Shen, Pengfei Zuo, Xin Wang, Michael R. Lyu, and Yang- fan Zhou. 2024. CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles(Austin, TX, USA)(SOSP...

  32. [40]

    Lyu, and Yangfan Zhou

    Xuchuan Luo, Pengfei Zuo, Jiacheng Shen, Jiazhen Gu, Xin Wang, Michael R. Lyu, and Yangfan Zhou. 2023. SMART: A High-Performance Adaptive Radix Tree for Disaggregated Memory. In17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). Boston, MA

  33. [41]

    Peter Macko, Virendra J Marathe, Daniel W Margo, and Margo I Seltzer. 2015. Llama: Efficient graph analytics using large multiversioned arrays. In2015 IEEE 31st International Conference on Data Engineering. IEEE, 363–374

  34. [42]

    Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.IEEE transactions on pattern analysis and machine intelligence42, 4 (2018), 824–836

  35. [43]

    Kanaujia, and Prakash Chauhan

    Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agar- wal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit O. Kanaujia, and Prakash Chauhan. 2023. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory. InProceedings of the 28th ACM ...

  36. [44]

    Donald Nguyen, Andrew Lenharth, and Keshav Pingali. 2013. A lightweight infrastructure for graph analytics. InProceedings of the twenty-fourth ACM sym- posium on operating systems principles. 456–471

  37. [45]

    Vlad Nitu, Boris Teabe, Alain Tchana, Canturk Isci, and Daniel Hagimont. 2018. Welcome to zombieland: practical and energy-efficient memory disaggregation in a datacenter. InProceedings of the Thirteenth EuroSys Conference(Porto, Portugal) (EuroSys ’18). Association for Comput...

  38. [46]

    Prashant Pandey, Brian Wheatman, Helen Xu, and Aydin Buluc. 2021. Terrace: A hierarchical graph container for skewed dynamic graphs. InProceedings of the 2021 international conference on management of data. 1372–1385

  39. [47]

    Xi Pang and Jianguo Wang. 2024. Understanding the Performance Implications of the Design Principles in Storage-Disaggregated Databases.Proc. ACM Manag. Data2, 3, Article 180 (May 2024), 26 pages. doi:10.1145/3654983

  40. [48]

    Feng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia, Zuoning Chen, and Yongwei Wu. 2024. Scaling Up Memory Disaggregated Applications with SMART. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volum...

  41. [49]

    Jie Ren, Minjia Zhang, and Dong Li. 2020. HM-ANN: efficient billion-point nearest neighbor search on heterogeneous memory. InProceedings of the 34th International Conference on Neural Information Processing Systems(Vancouver, BC, Canada)(NIPS ’20). Curran Associates Inc., Red ...

  42. [50]

    André Ryser, Alberto Lerner, Alex Forencich, and Philippe Cudré-Mauroux. 2022. D-RDMA: Bringing Zero-Copy RDMA to Database Systems. In12th Conference on Innovative Data Systems Research, CIDR 2022, Chaminade, CA, USA, January 9-12,

  43. [51]

    Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. 2018. LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). Carlsbad, CA

  44. [52]

    Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Yuxin Su, Jiazhen Gu, Hao Feng, Yangfan Zhou, and Michael R Lyu. 2023. Ditto: An elastic and adaptive memory- disaggregated caching system. InProceedings of the 29th Symposium on Operating Systems Principles. 675–691

  45. [53]

    Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Tianyi Yang, Yuxin Su, Yangfan Zhou, and Michael R. Lyu. 2023. FUSEE: A Fully Memory-Disaggregated Key-Value Store. In21st USENIX Conference on File and Storage Technologies (FAST 23). Santa Clara, CA

  46. [54]

    Blelloch

    Julian Shun and Guy E. Blelloch. 2013. Ligra: a lightweight graph processing framework for shared memory. InProceedings of the 18th ACM SIGPLAN sym- posium on Principles and practice of parallel programming. New York, NY, USA, 135–146

  47. [55]

    Jixian Su, Chiyu Hao, Shixuan Sun, Hao Zhang, Sen Gao, Jiaxin Jiang, Yao Chen, Chenyi Zhang, Bingsheng He, and Minyi Guo. 2025. Revisiting the Design of In-Memory Dynamic Graph Storage.Proc. ACM Manag. Data3, 1, Article 70 (Feb. 2025), 27 pages. doi:10.1145/3709720

  48. [56]

    John Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng, Guanzhou Hu, Zhihao Jia, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2021. Dorylus: Affordable, Scalable, and Accurate GNN Training with Dis- tributed CPU Servers and Serverless Threads. In...

  49. [57]

    Haque, Zhijing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes

    Muhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque, Zhijing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes. 2020. Borg: the next generation. InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Association for Co...

  50. [58]

    Infiniband trade association. 2023. InfiniBand. https://www.infinibandta.org/

  51. [59]

    Chenxi Wang, Haoran Ma, Shi Liu, Yifan Qiao, Jonathan Eyolfson, Christian Navasca, Shan Lu, and Guoqing Harry Xu. 2022. MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime. In16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2...

  52. [60]

    Chenxi Wang, Yifan Qiao, Haoran Ma, Shi Liu, Wenguang Chen, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2023. Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory. In20th USENIX Symposium on Networked Systems Design and Implementation, NSDI 202...

  53. [61]

    Jing Wang, Chao Li, Taolei Wang, Lu Zhang, Pengyu Wang, Junyi Mei, and Minyi Guo. 2022. Excavating the potential of graph workload on rdma-based far memory architecture. In2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1029–1039

  54. [62]

    Qing Wang, Youyou Lu, and Jiwu Shu. 2022. Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory. InProceedings of the 2022 International Conference on Management of Data(Philadelphia, PA, USA)(SIG- MOD ’22). New York, NY, USA

  55. [63]

    TamerÖzsu, and Walid G

    Ruihong Wang, Chuqing Gao, Jianguo Wang, Prishita Kadam, M. TamerÖzsu, and Walid G. Aref. 2024. Optimizing LSM-based indexes for disaggregated memory. DMG : A Scalable and Efficient Memory-Disaggregated Graph Processing System The VLDB Journal33, 6 (June 2024), 1813–1836. doi:...

  56. [65]

    Rui Wang, Weixu Zong, Shuibing He, Xinyu Chen, Zhenxin Li, and Zheng Dang

  57. [66]

    Yuke Wang, Boyuan Feng, Zheng Wang, Tong Geng, Kevin Barker, Ang Li, and Yufei Ding. 2023. MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms. In17th USENIX Symposium on Operating Systems Design an...

  58. [67]

    Zilong Wang, Xinchen Wan, Luyang Li, Yijun Sun, Peng Xie, Xin Wei, Qingsong Ning, Junxue Zhang, and Kai Chen. 2024. Fast, Scalable, and Accurate Rate Limiter for RDMA NICs. InProceedings of the ACM SIGCOMM 2024 Conference, ACM SIGCOMM 2024, Sydney, NSW, Australia, August 4-8, ...

  59. [68]

    Xingda Wei, Zhiyuan Dong, Rong Chen, and Haibo Chen. 2018. Deconstructing RDMA-enabled Distributed Transactions: Hybrid is Better!. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 233–251

  60. [69]

    Marcel Weisgut, Daniel Ritter, Pinar Tözün, Lawrence Benson, and Tilmann Rabl

  61. [70]

    Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. 2022. Graph Neural Networks in Recommender Systems: A Survey.ACM Comput. Surv.55, 5, Article 97 (Dec. 2022), 37 pages

  62. [71]

    Haoxuan Xie, Junfeng Liu, Siqiang Luo, and Kai Wang. 2026. RadixGraph: A Fast, Space-Optimized Data Structure for Dynamic Graph Storage (Extended Version). arXiv:2601.01444 [cs.DB] https://arxiv.org/abs/2601.01444

  63. [72]

    Bin Yan, Youyou Lu, Qing Wang, Minhui Xie, and Jiwu Shu. 2023. Patronus: High-Performance and Protective Remote Memory. In21st USENIX Conference on File and Storage Technologies (FAST 23). USENIX Association, Santa Clara, CA, 315–330. https://www.usenix.org/conference/fast23/p...

  64. [73]

    Ke Yang, MingXing Zhang, Kang Chen, Xiaosong Ma, Yang Bai, and Yong Jiang

  65. [74]

    Tsun-Yu Yang, Yizou Chen, Yuhong Liang, and Ming-Chang Yang. 2024. Ser- aph: Towards Scalable and Efficient Fully-external Graph Computation via On- demand Processing. In22nd USENIX Conference on File and Storage Technologies (FAST 24). Santa Clara, CA, 373–387

  66. [75]

    Tsun-Yu Yang, Yizou Chen, Yuhong Liang, and Ming-Chang Yang. 2025. Lever- aging On-demand Processing to Co-optimize Scalability and Efficiency for Fully- external Graph Computation.ACM Trans. Storage21, 2, Article 11 (Feb. 2025), 31 pages. doi:10.1145/3701037

  67. [76]

    Tsun-Yu Yang, Yi Li, Yizou Chen, Bingzhe Li, and Ming-Chang Yang. 2025. Oa- sis: An Out-of-core Approximate Graph System via All-Distances Sketches. In 23rd USENIX Conference on File and Storage Technologies (FAST 25). USENIX Association, Santa Clara, CA, 523–537

  68. [77]

    Xinjun Yang, Yingqiang Zhang, Hao Chen, Feifei Li, Gerry Fan, Yang Kong, Bo Wang, Jing Fang, Yuhui Wang, Tao Huang, Wenpu Hu, Jim Kao, and Jianping Jiang. 2025. Unlocking the Potential of CXL for Disaggregated Memory in Cloud- Native Databases. InCompanion of the 2025 Internat...

  69. [78]

    Xinjun Yang, Yingqiang Zhang, Hao Chen, Feifei Li, Bo Wang, Jing Fang, Chuan Sun, and Yuhui Wang. 2024. PolarDB-MP: A Multi-Primary Cloud- Native Database via Disaggregated Shared Memory. InCompanion of the 2024 International Conference on Management of Data(Santiago AA, Chile...

  70. [79]

    Yiwei Yang, Pooneh Safayenikoo, Jiacheng Ma, Tanvir Ahmed Khan, and An- drew Quinn. 2023. CXLMemSim: A pure software simulated CXL. mem for performance characterization.arXiv preprint arXiv:2303.06153(2023)

  71. [80]

    Peiqi Yin, Xiao Yan, Shiyuan Deng, Hui Li, Yifan Zhu, Xiangyu Zhi, Jingqi Mao, Ran Xu, Wenliang Zhang, and James Cheng. 2026. DistVS: Large-scale Vector Search with Compute-Memory Disaggregation. In23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26)....

  72. [81]

    Song Yu, Shufeng Gong, Qian Tao, Sijie Shen, Yanfeng Zhang, Wenyuan Yu, Pengxi Liu, Zhixin Zhang, Hongfu Li, Xiaojian Luo, Ge Yu, and Jingren Zhou

  73. [82]

    Xiangyao Yu. 2025. Disaggregation: A New Architecture for Cloud Databases. Proc. VLDB Endow.18, 12 (Aug. 2025), 5527–5530. doi:10.14778/3750601.3760520

  74. [83]

    Daniel Zahka and Ada Gavrilovska. 2022. FAM-Graph: Graph analytics on disag- gregated memory. In2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 81–92

  75. [84]

    Shaoxun Zeng, Xiaojian Liao, Hao Guo, and Youyou Lu. 2024. Volley: Accelerat- ing Write-Read Orders in Disaggregated Storage. InProceedings of the Nineteenth European Conference on Computer Systems(Athens, Greece)(EuroSys ’24). Asso- ciation for Computing Machinery, New York, ...

  76. [85]

    Hantian Zha, Teng Ma, Baotong Lu, Yuansen Wang, Dongbiao He, Yuanhui Luo, Dafang Zhang, Yunpeng Chai, Yuxing Chen, and Anqun Pan. 2025. Shard: A Scalable and Resize-Optimized Hash Index on Disaggregated Memory.Proc. VLDB Endow.19, 4 (Dec. 2025), 684–697. doi:10.14778/3785297.3785309

  77. [86]

    Chenzi Zhang, Fan Wei, Qin Liu, Zhihao Gavin Tang, and Zhenguo Li. 2017. Graph Edge Partitioning via Neighborhood Heuristic. InProceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax, NS, Canada)(KDD ’17). Association for Com...

  78. [87]

    Ming Zhang, Yu Hua, and Zhijun Yang. 2024. Motor: Enabling Multi-Versioning for Distributed Transactions on Disaggregated Memory. In18th USENIX Sympo- sium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA. USENIX Association, 801–819

  79. [88]

    ACM Manag

    LSMGraph: A High-Performance Dynamic Graph Storage System with Multi-Level CSR.Proc. ACM Manag. Data2, 6, Article 243 (Dec. 2024), 28 pages. doi:10.1145/3698818

  80. [89]

    Qizhen Zhang, Xinyi Chen, Sidharth Sankhe, Zhilei Zheng, Ke Zhong, Sebastian Angel, Ang Chen, Vincent Liu, and Boon Thau Loo. 2022. Optimizing data- intensive systems in disaggregated data centers with teleport. InProceedings of the 2022 International Conference on Management ...

  81. [90]

    Wenqian Zhang, Zhengyi Yang, Dong Wen, Wentao Li, Wenjie Zhang, and Xuemin Lin. 2025. Accelerating Core Decomposition in Billion-Scale Hy- pergraphs.Proc. ACM Manag. Data3, 1, Article 6 (Feb. 2025), 27 pages. doi:10.1145/3709656

  82. [91]

    Yang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao, James Mickens, Minlan Yu, Chris Kennelly, Paul Turner, David E. Culler, Henry M. Levy, and Amin Vahdat. 2022. Carbink: Fault-Tolerant Far Memory. In16th USENIX Symposium on Operating Systems Design and Implementation, OSDI...

  83. [92]

    Xiaowei Zhu, Wenguang Chen, Weimin Zheng, and Xiaosong Ma. 2016. Gemini: A{Computation-Centric} distributed graph processing system. In12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 301–316

  84. [93]

    Xiaowei Zhu, Wentao Han, and Wenguang Chen. 2015. GridGraph: Large-Scale Graph Processing on a Single Machine Using 2-Level Hierarchical Partitioning. In2015 USENIX Annual Technical Conference (USENIX ATC 15). Santa Clara, CA, 375–386

  85. [94]

    Zhiting Zhu, Newton Ni, Yibo Huang, Yan Sun, Zhipeng Jia, Nam Sung Kim, and Emmett Witchel. 2024. Lupin: Tolerating Partial Failures in a CXL Pod. In Proceedings of the 2nd Workshop on Disruptive Memory Systems(Austin, TX, USA) (DIMES ’24). Association for Computing Machinery,...

  86. [95]

    Mingxing Zhang, Teng Ma, Jinqi Hua, Zheng Liu, Kang Chen, Ning Ding, Fan Du, Jinlei Jiang, Tao Ma, and Yongwei Wu. 2023. Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared Memory. InProceedings of the 29th Symposium on Operating Systems Princ...

  87. [102]

    Tobias Ziegler, Sumukha Tumkur Vani, Carsten Binnig, Rodrigo Fonseca, and Tim Kraska. 2019. Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks. InProceedings of the 2019 International Conference on Management of Data(Amsterdam, Netherlands)(SIGMOD...

  88. [2012]

    In10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12)

    PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs. In10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12). Hollywood, CA, 17–30

  89. [2019]

    InProceedings of the 27th ACM Symposium on Operating Systems Principles(Huntsville, Ontario, Canada)(SOSP ’19)

    KnightKing: a fast distributed graph random walk engine. InProceedings of the 27th ACM Symposium on Operating Systems Principles(Huntsville, Ontario, Canada)(SOSP ’19). Association for Computing Machinery, New York, NY, USA, 524–537

  90. [2020]

    InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20)

    Can far memory improve job throughput?. InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Association for Computing Machinery, New York, NY, USA, Article 14, 16 pages

  91. [2022]

    https://www.cidrdb.org/cidr2022/papers/p77-ryser.pdf

    www.cidrdb.org. https://www.cidrdb.org/cidr2022/papers/p77-ryser.pdf

  92. [2024]

    In2024 USENIX Annual Technical Conference (USENIX ATC 24)

    Efficient Large Graph Processing with Chunk-Based Graph Representation Model. In2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 1239–1255

  93. [2025]

    VLDB Endow.18, 9 (May 2025), 3119–3133

    CXL Memory Performance for In-Memory Data Processing.Proc. VLDB Endow.18, 9 (May 2025), 3119–3133. doi:10.14778/3746405.3746432

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.