Pith. sign in

REVIEW 3 major objections 5 minor 125 references

Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read SINLK claims node-grained, leaf-tracked placement can raise tree-index throughput on CXL heterogeneous memory by up to 71%.

desk verdict A credible CXL-tiering scheme with a real-hardware evaluation, but the leaf-centric hotness assumption is unproven for scans and failed lookups, so the headline overclaims generality. read the letter →

arxiv 2507.18559 v1 pith:ZVAWQYSX submitted 2025-07-24 cs.OS

classification cs.OS
keywords heterogeneousmemoryCXLtieringtree-structuredindexesB+treeradixtreedataplacementhotpathmigrationwatermarkcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Tree-structured indexes such as B+ trees and radix trees lose roughly 70% of their throughput when three quarters of the tree's memory lands on CXL-attached memory. This paper argues that the right remedy is to place data at the granularity of tree nodes rather than memory pages, and to let the tree's own structure decide what goes into fast memory: upper-level nodes first, and then the complete root-to-leaf paths that lead to frequently read leaves. The proposed scheme, SINLK, tracks access frequency only at leaves, migrates entire hot paths upward and cold nodes downward while preserving a single fast-to-slow boundary on every path, and adjusts its own allocation and migration parameters from real-time fast-memory usage. Integrated into one B+ tree and one radix tree with less than three percent internal code changes, SINLK reports up to 71% higher throughput and 81% lower P99 latency than state-of-the-art page-tiering and HM-optimized index schemes on YCSB and real-world workloads. If correct, the result means a small fast-memory tier can deliver most of the performance of a fully fast index.

What carries the argument

Three mechanisms carry the argument. Leaf-centric access tracking records a per-access frequency only for the destination leaf (two metadata bytes), identifying hot paths by their leaves; the paper measures this at a 5–7% slowdown versus roughly 60% for per-node tracking. Structure-aware migration is built on the single-boundary invariant that every fast-memory node's ancestors are also in fast memory, so each root-to-leaf path has at most one fast-to-slow transition; promotion walks from a hot leaf upward through its slow ancestors, while demotion removes a node only after all its children are already in slow memory and obeys a level cutoff L_demote. The hyper watermark mechanism ties these together, adjusting the allocation level L_fast, the hot and cold percentile thresholds, and L_demote asymmetrically when fast-memory usage crosses high (95%) or low (85%) watermarks so that allocation, promotion, and demotion all push fast-memory usage toward the same stable point.

What would settle it

Run SINLK on a workload of many short range scans over scattered cold keys that share a small set of internal subtrees, so internal nodes are hot while every individual leaf is cold; if the fraction of accesses served by unpromoted internal nodes exceeds the paper's reported 0.9% false-negative bound, the leaf proxy fails and a variant that also tracks internal nodes will measurably outperform SINLK.

Watch

Extended reading notes

Core claim

SINLK is a node-grained, tree-structure-aware data placement scheme for CXL-based heterogeneous memory. Its central claim is that placement decisions for a tree index on a fast/slow memory pair should follow the tree's own units: nodes are the unit of management, upper levels are inherently hotter than lower levels, and access happens along root-to-leaf paths. From this it follows that hot nodes should live in fast memory, that hotness should be measured by how often leaves are accessed rather than by instrumenting every node, and that a leaf's entire slow ancestor chain should be promoted together so that every fast node keeps all of its ancestors in fast memory, yielding the single-boundary structure where each root-to-leaf path crosses the fast/slow boundary at most once. The paper further claims that a coordinated hyper-watermark controller, which adjusts allocation depth, hot and cold thresholds, and demotion depth from current fast-memory usage, is what prevents both fast-memory exhaustion and burst migrations. On a real CXL platform with YCSB and production block traces, the scheme reports throughput gains up to 71% and P99 latency reductions up to 81% compared with page-level placement and static HM-optimized indexes.

Load-bearing premise

The whole hot-path mechanism rests on one proxy: a path is hot exactly when its leaf is frequently accessed, so tracking only leaf counts is enough to know which ancestors deserve fast memory.

Editorial extensions

If this is right

  • Fast memory can be provisioned at modest fractions (10–20%) of total index memory and still capture most of a fully fast index's performance for skewed workloads.
  • The single-boundary invariant turns placement into a cut of the tree, bounding every root-to-leaf path to at most one slow-memory segment and making access latency more predictable.
  • The framework transfers across index shapes: both a B+ tree and a radix tree were adapted with under 3% internal code modification, and the same machinery is proposed for multi-tier hierarchies by applying it to adjacent memory pairs.
  • Tail latency improves along with throughput because the watermark controller prevents burst demotions; P99 latency drops up to 81% on real-world traces.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the leaf-centric hot-path proxy could be tested on other index families with value-carrying leaves, such as skip lists or learned indexes; a positive result would extend the method to most in-memory index shapes.
  • Beyond the paper, the paper's own discussion anticipates false negatives when infrequent leaves share a hot ancestor; a stress workload of many short scans over cold leaves would directly quantify this gap and could motivate hybrid leaf-plus-internal-node tracking.
  • Beyond the paper, the hyper watermark mechanism is a feedback controller on fast-memory occupancy, so the same coordination logic could inform OS-level CXL tiering for objects other than tree nodes.
  • Beyond the paper, recovery after a hot-region shift (roughly 17 seconds in the microbenchmark) depends on worker wake-up intervals; event-triggered migration would likely be needed for faster-changing cloud workloads.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes SINLK, a node-grained, tree-structure-aware data placement scheme for tree indexes on CXL heterogeneous memory. SINLK tracks access hotness only at leaf nodes, uses layer-aware allocation to keep upper-level nodes in fast memory, migrates entire hot paths and cold subtrees in a structure-aware way, and coordinates allocation and migration with a hyper watermark mechanism based on fast memory usage. The scheme is integrated into Masstree and ART with small code changes and evaluated on a real CXL 1.1 platform against MEMTIS, TPP, Caption, and PACTree-based baselines under YCSB, synthetic skewed-partition, and Alibaba block-trace workloads. The paper reports up to 71% throughput improvement and up to 81% P99 latency reduction relative to these baselines.

Significance. The paper makes a useful and timely contribution: it identifies a granularity mismatch between page-level CXL tiering and tree nodes, and it proposes a concrete node-grained alternative with a plausible design rationale (layer principle, path principle, single-boundary structure). The evaluation is on real CXL hardware, includes a factor analysis isolating each technique, and includes sensitivity analysis for dynamic workloads, worker wake-up intervals, and fast-memory ratios; these are clear strengths. The central claim, however, is supported only under the assumption that leaf access frequency identifies hot paths, which is not validated for scans and failed lookups, and the headline numbers come from single-point comparisons without released code or error bars. If those gaps are closed, the result would be a solid systems contribution.

major comments (3)
  1. [§4.1, §4.2, §7] The load-bearing assumption of the design is that hot paths are exactly the paths to frequently accessed leaves. This is stated in §4.1 ('determining hot paths is based on leaf nodes' access frequency') and inherited by promotion (§4.2.1) and demotion (§4.2.2). The assumption is exact only for successful point lookups and updates. For failed lookups in a radix tree, the search can terminate at an internal node without reaching a leaf, so a hot internal node with a cold subtree never triggers promotion. For range scans, the paper never specifies whether a scan increments one leaf's counter or every scanned leaf's counter; the first choice hides the hot descent path, and the second choice can mark a broad band of individually cold leaves as hot and exhaust fast memory. The authors' own YCSB-E result in Figure 14(b) (9.5% improvement over baseline, versus 31–84% elsewhere) is consistent with this failure mode. Section 7 bounds false-negative internal nodes only for a Zipfian B+tree under point accesses (<0.9%) and does not cover failed lookups, scans, or radix-tree prefix aborts. Because allocation, promotion, and demotion all depend on this proxy, the abstract's generality claim is not yet supported for workloads containing these access types.
  2. [§6.1, Figures 13–22] The abstract's quantitative claims ('up to 71% throughput, up to 81% P99 latency') are based on single-point measurements. No error bars, confidence intervals, or run-to-run variance are reported in §6, and the code is not released. For a systems paper whose contribution is empirical, this makes the magnitude of the claimed improvement difficult to verify, especially because throughput and tail latency on a 28-thread, 32 GiB CXL setup are sensitive to allocation placement and background-worker scheduling. I would like to see at least 3–5 runs per configuration with error bars on the headline figures, and release of the SINLK framework and integration code to enable reproducibility.
  3. [§4.3, §6.3] The hyper watermark mechanism uses two fixed thresholds, U_high=95% and U_low=85% (§4.3.1, §5), but the sensitivity analysis in §6.3 varies only worker wake-up intervals and maximum fast memory usage; it does not vary U_high and U_low. Since the stability claim rests on these thresholds, the paper should show that throughput and latency are insensitive to reasonable choices of U_high/U_low (e.g., 85/75, 90/80, 95/85, 98/90). Similarly, P_hot and P_cold are said to be initialized from the maximum fast memory usage, but the initialization formula is not given, so the reader cannot assess the sensitivity of the histogram-based classification to these values.
minor comments (5)
  1. [Title, §5] The running head and several passages use 'S INLK' with an unintended space (e.g., the title and Section 5); the spacing should be fixed throughout.
  2. [§2.2, §6.2.2] Figure cross-references are inconsistent: §2.2 cites 'Figure 22' where Figure 2 is meant, and §6.2.2 cites 'Figure 21(a)' and 'Figure 22(a)' when the throughput figures in that section are Figures 13 and 14.
  3. [§6.2.1] The SINLK-Prophet comparison for With Insert is not apples-to-apples because Prophet uses 1.3–1.5× more fast memory than the SINLK limit; the text acknowledges this, but the reader should be told explicitly that the claimed 'similar to Prophet' result excludes that workload.
  4. [§6.4] The statement that SINLK's throughput at 48 threads is '28.6–38.1× that of a single thread' should be accompanied by the baseline's single-thread-to-multi-thread scaling, so the reader can separate SINLK's scaling from Masstree's inherent scaling.
  5. [Table 2] The 'Run Time Ratio' values in Table 2 are parts per thousand; this is stated in the body but should also appear in the table caption, as the caption alone is ambiguous.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: SINLK's claims rest on external baselines and ablation experiments, not on definitions that presuppose the conclusions.

full rationale

The paper's central contribution is an empirical placement scheme evaluated against external baselines (MEMTIS, TPP, Caption, PACTree, Masstree) on real CXL hardware, and its component techniques are isolated through factor analysis (Figure 17) rather than assumed. The leaf-centric tracking assumption equates hot paths with paths to hot leaves, but this is an explicitly stated design proxy and robustness limitation (Section 7 bounds false negatives only for Zipfian point accesses); it is not a step in which a fitted parameter is renamed as a prediction or in which the derivation reduces to its own inputs. The only apparent self-citation, XIndex [78], is used solely as an example of skewed real-world access patterns and is not load-bearing. Performance numbers such as the 71% throughput and 81% P99 latency improvements are measured against independent systems, so the headline claims are not forced by construction. No circular step meeting the required quote-and-reduction standard was found.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests on several hand-chosen system parameters and assumptions about tree access patterns. The parameters are not fitted to the reported speedups and some are subject to sensitivity analysis, but they are not derived from first principles. The most fragile assumption is that leaf access frequency is a sufficient proxy for path hotness, which the paper only partially validates.

free parameters (6)
  • high watermark U_high = 95%
    Hand-set threshold for triggering aggressive demotion and parameter adjustment (Sections 4.3 and 5). Not derived from workload data.
  • low watermark U_low = 85%
    Hand-set threshold for conservative promotion adjustment (Sections 4.3 and 5).
  • L_fast initial value and adjustment bounds = not specified
    Highest level allocated in fast memory; dynamically adjusted by the watermark maintainer, with limits based on maximum fast memory usage (Sections 4.1 and 4.3).
  • L_demote = not specified
    Lowest level eligible for demotion; dynamically adjusted under memory pressure (Sections 4.2.2 and 4.3).
  • P_hot and P_cold = initialized based on maximum fast memory usage
    Percentage thresholds for hot and cold leaf classification; adjusted by the hyper watermark mechanism (Sections 4.2.3 and 4.3).
  • worker wake-up intervals = 500 ms trigger, 2000 ms cooler, 100 ms watermark maintainer
    Configuration constants in Section 5; sensitivity tested but fixed for the main results.
assumptions (5)
  • domain assumption CXL-attached memory has a stable roughly 2x latency and roughly 60% bandwidth gap versus DRAM, and this gap dominates the migration and caching trade-off.
    Measured on one CXL 1.1 platform (Figure 1c) and assumed representative of CXL-HM, including future CXL switches (Section 2.1).
  • domain assumption Upper tree levels are accessed far more often than lower levels, so fast-memory placement of upper nodes is beneficial.
    Validated by access-count distributions and placement experiments (Section 3.1.1, Figure 6), but treated as a general tree property.
  • domain assumption Hot paths can be identified from leaf-node access frequency alone; internal nodes inherit hotness from leaves.
    Foundation of leaf-centric tracking (Section 4.1); the false-negative analysis in Section 7 only covers one Zipfian B+tree case.
  • domain assumption Workloads are skewed so that a small fraction of keys receives most accesses.
    Used in the microbenchmark (90% of requests to 5% of keys) and YCSB Zipfian default; real-world traces are selected for large working sets (Section 6.1).
  • domain assumption Maintaining a single fast/slow boundary along each root-to-leaf path is beneficial and should be preserved during migration.
    Motivated by an access-ratio argument and Figure 7; used as a design invariant in Section 4.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK." pith.science (2026). https://pith.science/paper/ZVAWQYSX

@misc{pith2026250718559,
  author       = {Pith},
  title        = {Pith review of: Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZVAWQYSX}},
  note         = {Machine review of arXiv:2507.18559}
}
read the original abstract

On heterogeneous memory (HM) where fast memory (i.e., CPU-attached DRAM) and slow memory (e.g., remote NUMA memory, RDMA-connected memory, Persistent Memory (PM)) coexist, optimizing the placement of tree-structure indexes (e.g., B+tree) is crucial to achieving high performance while enjoying memory expansion. Nowadays, CXL-based heterogeneous memory (CXL-HM) is emerging due to its high efficiency and memory semantics. Prior tree-structure index placement schemes for HM cannot effectively boost performance on CXL-HM, as they fail to adapt to the changes in hardware characteristics and semantics. Additionally, existing CXL-HM page-level data placement schemes are not efficient for tree-structure indexes due to the granularity mismatch between the tree nodes and the page. In this paper, we argue for a CXL native, tree-structure aware data placement scheme to optimize tree-structure indexes on CXL-HM. Our key insight is that the placement of tree-structure indexes on CXL-HM should match the tree's inherent characteristics with CXL-HM features. We present SINLK, a tree-structure aware, node-grained data placement scheme for tree-structure indexes on CXL-HM. With SINLK, developers can easily adapt existing tree-structure indexes to CXL-HM. We have integrated the B+tree and radix tree with SINLK to demonstrate its effectiveness. Evaluations show that SINLK improves throughput by up to 71% and reduces P99 latency by up to 81% compared with state-of-the-art data placement schemes (e.g., MEMTIS) and HM-optimized tree-structure indexes in YCSB and real-world workloads.

Figures

Figures reproduced from arXiv: 2507.18559 by the authors.

Figure 1
Figure 1. CXL-based heterogeneous memory system. (a,b) System architecture. (c) Comparison of fast and slow mem￾ory latency and bandwidth tested by Intel MLC [63] on our CXL1.1 platform. tree. Evaluation on a real CXL platform with various work￾loads demonstrates the advantages of SINLK. 2 Background and Motivation 2.1 CXL-based Heterogeneous Memory CXL [49, 59] is an emerging interconnect technology no￾table for its memory e… view at source ↗
Figure 5
Figure 5. Node access count distribution across pages [PITH_FULL_IMAGE:figures/full_fig_p003_5.png] view at source ↗
Figure 3
Figure 3. Comparison of SMART, PACTree, and ART on CXL-HM. PAC stands for PACTree. All indexes have a fast memory usage ratio of 20%, except ART-F at 40%. PAC, ART, and ART-F use the same memory allocation policy as [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Performance breakdown per operation type. [PITH_FULL_IMAGE:figures/full_fig_p003_4.png]
Figure 6
Figure 6. Figure 6: Tput. of different node placement strategies. [PITH_FULL_IMAGE:figures/full_fig_p004_6.png]
Figure 8
Figure 8. Figure 8: Tput. at different hot path ratios in fast mem [PITH_FULL_IMAGE:figures/full_fig_p004_8.png]
Figure 9
Figure 9. Figure 9: Architecture and interactions of SINLK.PQ stands for promotion queue, DQ stands for demotion queue. access frequency of leaf nodes (❷), laying the groundwork for hot path and cold node identification. The background module consists of several background workers to cond…
Figure 10
Figure 10. Figure 10: Promotion and demotion procedures. When allocating new nodes or changing the node’s position by auto-balancing, the decision on the node’s placement is based on its current level 𝑙 and its parent’s type. If 𝑙 < 𝐿𝑓 𝑎𝑠𝑡 and its parent is in fast memory, the node is allo…
Figure 12
Figure 12. Figure 12: System parameter adjustments and the col [PITH_FULL_IMAGE:figures/full_fig_p007_12.png]
Figure 13
Figure 13. Figure 13: Performance of SINLK (S-Masstree) and compared systems in the microbenchmark. Update Heavy Read Mostly Read Only Read Latest Read Modify Write 0 10 20 30 40 50 60 Throughput (Mop/s) 58.5% 68.7% 83.6% 31.7% 68.6% PAC-L Baseline Caption (a) S-ART Update Heavy Read Mostl…
Figure 14
Figure 14. Figure 14: Throughput of SINLK and compared systems in the macrobenchmark. • Does SINLK perform well in various scenarios? (§6.2) • Can SINLK maintain good performance in dynamic chang￾ing workloads and different initial system settings? (§6.3) • How scalable is SINLK when varyi…
Figure 15
Figure 15. Figure 15: Sensitivity analysis of SINLK (S-Masstree). 1 4 8 162428324048 Thread Count 0 10 20 30 Throughput (Mop/s) Update Heavy Baseline 1 4 8 162428324048 Thread Count 0 10 20 30 40 Read Mostly SINLK 1 4 8 162428324048 Thread Count 0 10 20 30 40 50 Read Only 1 4 8 16242832404…
Figure 16
Figure 16. Figure 16: Scalability of SINLK (S-Masstree) in microbenchmark. ting, we further evaluate SINLK’s performance across various data scales and fast memory usage later to demonstrate its performance stability (§6.3.3, §6.4 part1). As shown in [PITH_FULL_IMAGE:figures/full_fig_p010…
Figure 18
Figure 18. Figure 18: Impact of key sizes (Micro Update Heavy). Read Avg Read P90 Read P99 Write Avg Write P90 Write P99 0 1 2 3 Latency ( s) Passive & Separated Mgmt. SINLK [PITH_FULL_IMAGE:figures/full_fig_p011_18.png]
Figure 20
Figure 20. Figure 20: Fast memory usage ratio variation over time. [PITH_FULL_IMAGE:figures/full_fig_p011_20.png]
Figure 21
Figure 21. Figure 21: Performance comparison between S-ART and compared systems with Alibaba Block Traces. [PITH_FULL_IMAGE:figures/full_fig_p013_21.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

125 extracted references · 80 canonical work pages

  1. [1]

    Cache crafti- ness for fast multicore key-value storage

    Yandong Mao, Eddie Kohler, and Robert Tappan Morris. Cache crafti- ness for fast multicore key-value storage. In Proceedings of the 7th ACM european conference on Computer Systems , pages 183–196, Bern, Switzerland, 2012. ACM

  2. [2]

    Wormhole: A fast ordered index for in-memory data management

    Xingbo Wu, Fan Ni, and Song Jiang. Wormhole: A fast ordered index for in-memory data management. In Proceedings of the Fourteenth EuroSys Conference 2019, pages 1–16, 2019

  3. [3]

    Cuckoo Trie: exploiting memory- level parallelism for efficient dram indexing

    Adar Zeitak and Adam Morrison. Cuckoo Trie: exploiting memory- level parallelism for efficient dram indexing. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles , pages 147–162, 2021

  4. [4]

    HOT: A height optimized trie index for main-memory data- base systems

    Robert Binna, Eva Zangerle, Martin Pichl, Günther Specht, and Viktor Leis. HOT: A height optimized trie index for main-memory data- base systems. In Proceedings of the 2018 International Conference on Management of Data, pages 521–534, 2018

  5. [5]

    FAST: fast architecture sensitive tree search on modern CPUs and GPUs

    Changkyu Kim, Jatin Chhugani, Nadathur Satish, Eric Sedlar, An- thony D Nguyen, Tim Kaldewey, Victor W Lee, Scott A Brandt, and Pradeep Dubey. FAST: fast architecture sensitive tree search on modern CPUs and GPUs. In Proceedings of the 2010 ACM SIGMOD International Conference on Management of data , pages 339–350, 2010

  6. [6]

    Making B+-trees cache conscious in main memory

    Jun Rao and Kenneth A Ross. Making B+-trees cache conscious in main memory. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data , pages 475–486, 2000

  7. [7]

    The Bw-Tree: A B-tree for new hardware platforms

    Justin J Levandoski, David B Lomet, and Sudipta Sengupta. The Bw-Tree: A B-tree for new hardware platforms. In 2013 IEEE 29th International Conference on Data Engineering (ICDE) , pages 302–313. IEEE, 2013

  8. [8]

    Occualizer: Optimistic concur- rent search trees from sequential code

    Tomer Shanny and Adam Morrison. Occualizer: Optimistic concur- rent search trees from sequential code. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , pages 321–337, 2022

Show all 125 references
  1. [9]

    Andersen

    Ziqi Wang, Andrew Pavlo, Hyeontaek Lim, Viktor Leis, Huanchen Zhang, Michael Kaminsky, and David G. Andersen. Building a Bw- Tree Takes More Than Just Buzz Words. In Proceedings of the 2018 International Conference on Management of Data , SIGMOD ’18, page 473–488, New York, NY...

  2. [10]

    The adaptive radix tree: ARTful indexing for main-memory databases

    Viktor Leis, Alfons Kemper, and Thomas Neumann. The adaptive radix tree: ARTful indexing for main-memory databases. In 2013 IEEE 29th International Conference on Data Engineering (ICDE) , pages 38–49. IEEE, 2013

  3. [11]

    The ART of practical synchronization

    Viktor Leis, Florian Scheibner, Alfons Kemper, and Thomas Neumann. The ART of practical synchronization. In Proceedings of the 12th Inter- national Workshop on Data Management on New Hardware , DaMoN ’16, New York, NY, USA, 2016. Association for Computing Machinery

  4. [12]

    Speedy transactions in multicore in-memory databases

    Stephen Tu, Wenting Zheng, Eddie Kohler, Barbara Liskov, and Samuel Madden. Speedy transactions in multicore in-memory databases. In Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles, SOSP ’13, page 18–32, New York, NY, USA, 2013. Association for C...

  5. [13]

    Andersen, Andrew Pavlo, Michael Kamin- sky, Lin Ma, and Rui Shen

    Huanchen Zhang, David G. Andersen, Andrew Pavlo, Michael Kamin- sky, Lin Ma, and Rui Shen. Reducing the Storage Overhead of Main- Memory OLTP Databases with Hybrid Indexes. In Proceedings of the 2016 International Conference on Management of Data , SIGMOD ’16, page 1567–1581, ...

  6. [14]

    https: //hyper-db.de/, 2024

    HyPer – a hybrid OLTP&OLAP high performance DBMS. https: //hyper-db.de/, 2024

  7. [15]

    https://www.memsql.com/, 2024

    MemSQL. https://www.memsql.com/, 2024

  8. [16]

    https://www.sap.com/products/hana.html, 2024

    SAP HANA. https://www.sap.com/products/hana.html, 2024

  9. [17]

    https://www.oracle.com/database/technologies/related/ timesten.html, 2024

    TimesTen: Fastest OLTP database, ultra high availability, elastic scalability. https://www.oracle.com/database/technologies/related/ timesten.html, 2024

  10. [18]

    https://www.voltdb.com/, 2024

    VoltDB. https://www.voltdb.com/, 2024

  11. [19]

    KVell: the design and implementation of a fast persistent key-value store

    Baptiste Lepers, Oana Balmau, Karan Gupta, and Willy Zwaenepoel. KVell: the design and implementation of a fast persistent key-value store. In Proceedings of the 27th ACM Symposium on Operating Sys- tems Principles, SOSP ’19, page 447–461, New York, NY, USA, 2019. Association ...

  12. [20]

    Andersen, and Michael Kamin- sky

    Hyeontaek Lim, Bin Fan, David G. Andersen, and Michael Kamin- sky. SILT: a memory-efficient, high-performance key-value store. In Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles, SOSP ’11, page 1–13, New York, NY, USA, 2011. Association for Comput...

  13. [21]

    Xin, Cheng Lian, Yin Huai, Davies Liu, Joseph K

    Michael Armbrust, Reynold S. Xin, Cheng Lian, Yin Huai, Davies Liu, Joseph K. Bradley, Xiangrui Meng, Tomer Kaftan, Michael J. Franklin, Ali Ghodsi, and Matei Zaharia. Spark SQL: Relational Data Processing in Spark. In Proceedings of the 2015 ACM SIGMOD International Conferenc...

  14. [22]

    Index- Accelerated Pattern Matching in Event Stores

    Michael Körber, Nikolaus Glombiewski, and Bernhard Seeger. Index- Accelerated Pattern Matching in Event Stores. In Proceedings of the 2021 International Conference on Management of Data , SIGMOD ’21, page 1023–1036, New York, NY, USA, 2021. Association for Comput- ing Machinery

  15. [23]

    https://druid.apache.org/, 2025

    Apache Druid. https://druid.apache.org/, 2025

  16. [24]

    https: //flink.apache.org/, 2025

    Apache Flink®: Stateful Computations over Data Streams. https: //flink.apache.org/, 2025

  17. [25]

    https://www.elastic.co/ elasticsearch, 2025

    Elasticsearch: The heart of the Elastic Stack. https://www.elastic.co/ elasticsearch, 2025

  18. [26]

    Meet the walkers: accelerating index traversals for in-memory databases

    Onur Kocberber, Boris Grot, Javier Picorel, Babak Falsafi, Kevin Lim, 13 and Parthasarathy Ranganathan. Meet the walkers: accelerating index traversals for in-memory databases. In Proceedings of the 46th Annual IEEE/ACM International Symposium on Microarchitecture , MICRO-46, ...

  19. [27]

    Adaptive Hybrid Indexes

    Christoph Anneser, Andreas Kipf, Huanchen Zhang, Thomas Neu- mann, and Alfons Kemper. Adaptive Hybrid Indexes. In Proceedings of the 2022 International Conference on Management of Data , SIG- MOD ’22, pages 1626–1639, New York, NY, USA, 2022. Association for Computing Machinery

  20. [28]

    Andersen, Michael Kamin- sky, Kimberly Keeton, and Andrew Pavlo

    Huanchen Zhang, Xiaoxuan Liu, David G. Andersen, Michael Kamin- sky, Kimberly Keeton, and Andrew Pavlo. Order-Preserving Key Compression for In-Memory Search Trees. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data , SIG- MOD ’20, pages 1601–...

  21. [29]

    Workload analysis of a large-scale key-value store

    Berk Atikoglu, Yuehai Xu, Eitan Frachtenberg, Song Jiang, and Mike Paleczny. Workload analysis of a large-scale key-value store. In Proceedings of the 12th ACM SIGMETRICS/PERFORMANCE joint in- ternational conference on Measurement and Modeling of Computer Systems, pages 53–64, 2012

  22. [30]

    Andersen, and Michael Kamin- sky

    Hyeontaek Lim, Dongsu Han, David G. Andersen, and Michael Kamin- sky. MICA: a holistic approach to fast in-memory key-value stor- age. In Proceedings of the 11th USENIX Conference on Networked Systems Design and Implementation , NSDI’14, pages 429–444, USA,

  23. [31]

    HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVM

    Amanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez, and Simon Peter. HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVM. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles , SOSP ’21, page 392–407, New York, NY, USA, 2...

  24. [32]

    In-Memory Big Data Management and Processing: A Survey

    Hao Zhang, Gang Chen, Beng Chin Ooi, Kian-Lee Tan, and Meihui Zhang. In-Memory Big Data Management and Processing: A Survey. IEEE Transactions on Knowledge and Data Engineering , 27(7):1920– 1948, 2015

  25. [33]

    Frans Kaashoek

    Zviad Metreveli, Nickolai Zeldovich, and M. Frans Kaashoek. CPHASH: a cache-partitioned hash table. In Proceedings of the 17th ACM SIGPLAN Symposium on Principles and Practice of Parallel Pro- gramming, PPoPP ’12, page 319–320, New York, NY, USA, 2012. As- sociation for Comput...

  26. [34]

    Using Elimina- tion and Delegation to Implement a Scalable NUMA-Friendly Stack

    Irina Calciu, Justin Gottschlich, and Maurice Herlihy. Using Elimina- tion and Delegation to Implement a Scalable NUMA-Friendly Stack. In 5th USENIX Workshop on Hot Topics in Parallelism (HotPar 13) , San Jose, CA, June 2013. USENIX Association

  27. [35]

    Aguilera

    Irina Calciu, Siddhartha Sen, Mahesh Balakrishnan, and Marcos K. Aguilera. Black-box Concurrent Data Structures for NUMA Architec- tures. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, ASPL...

  28. [36]

    Designing distributed tree-based index structures for fast rdma-capable networks

    Ziegler, Tobias and Tumkur Vani, Sumukha and Binnig, Carsten and Fonseca, Rodrigo and Kraska, Tim. Designing distributed tree-based index structures for fast rdma-capable networks. InProceedings of the 2019 International Conference on Management of Data , SIGMOD ’19, page 741–...

  29. [37]

    Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory

    Qing Wang, Youyou Lu, and Jiwu Shu. Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory. In Proceedings of the 2022 International Conference on Management of Data , SIG- MOD ’22, page 1033–1048, New York, NY, USA, 2022. Association for Computing Machinery

  30. [38]

    Lyu, and Yangfan Zhou

    Xuchuan Luo, Pengfei Zuo, Jiacheng Shen, Jiazhen Gu, Xin Wang, Michael R. Lyu, and Yangfan Zhou. SMART: A high-performance adaptive radix tree for disaggregated memory. In 17th USENIX Sym- posium on Operating Systems Design and Implementation (OSDI 23) , pages 553–571, Boston,...

  31. [39]

    Lyu, and Yangfan Zhou

    Xuchuan Luo, Jiacheng Shen, Pengfei Zuo, Xin Wang, Michael R. Lyu, and Yangfan Zhou. CHIME: A cache-efficient and high-performance hybrid index on disaggregated memory. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles , SOSP ’24, page 110–126, Ne...

  32. [40]

    DEX: Scalable Range Indexing on Disaggregated Memory

    Baotong Lu, Kaisong Huang, Chieh-Jan Mike Liang, Tianzheng Wang, and Eric Lo. DEX: Scalable Range Indexing on Disaggregated Memory. Proc. VLDB Endow., 17(10):2603–2616, August 2024

  33. [41]

    ROLEX: a scalable RDMA-oriented learned key-value store for disag- gregated memory systems

    Pengfei Li, Yu Hua, Pengfei Zuo, Zhangyu Chen, and Jiajie Sheng. ROLEX: a scalable RDMA-oriented learned key-value store for disag- gregated memory systems. In Proceedings of the 21st USENIX Confer- ence on File and Storage Technologies , FAST’23, USA, 2023. USENIX Association

  34. [42]

    NV-Tree: reducing consistency cost for NVM-based single level systems

    Jun Yang, Qingsong Wei, Cheng Chen, Chundong Wang, Khai Leong Yong, and Bingsheng He. NV-Tree: reducing consistency cost for NVM-based single level systems. In Proceedings of the 13th USENIX Conference on File and Storage Technologies , FAST’15, page 167–181, USA, 2015. USENIX...

  35. [43]

    FPTree: A Hybrid SCM-DRAM Persistent and Concurrent B-Tree for Storage Class Memory

    Ismail Oukid, Johan Lasperas, Anisoara Nica, Thomas Willhalm, and Wolfgang Lehner. FPTree: A Hybrid SCM-DRAM Persistent and Concurrent B-Tree for Storage Class Memory. In Proceedings of the 2016 International Conference on Management of Data , SIGMOD ’16, pages 371–386, New Yo...

  36. [44]

    Hyun Lim, Hyunsub Song, Beomseok Nam, and Sam H

    Se Kwon Lee, K. Hyun Lim, Hyunsub Song, Beomseok Nam, and Sam H. Noh. WORT: Write Optimal Radix Tree for Persistent Memory Storage Systems. In 15th USENIX Conference on File and Storage Technologies (FAST 17), pages 257–270, Santa Clara, CA, February

  37. [45]

    DPTree: differential indexing for persistent memory

    Xinjing Zhou, Lidan Shou, Ke Chen, Wei Hu, and Gang Chen. DPTree: differential indexing for persistent memory. Proc. VLDB Endow. , 13(4):421–434, dec 2019

  38. [46]

    uTree: a persistent B+-tree with low tail latency

    Youmin Chen, Youyou Lu, Kedong Fang, Qing Wang, and Jiwu Shu. uTree: a persistent B+-tree with low tail latency. Proc. VLDB Endow., 13(12):2634–2648, jul 2020

  39. [47]

    LB+Trees: optimizing persistent index performance on 3DXPoint memory

    Jihang Liu, Shimin Chen, and Lujun Wang. LB+Trees: optimizing persistent index performance on 3DXPoint memory. Proc. VLDB Endow., 13(7):1078–1090, mar 2020

  40. [48]

    NBTree: a lock-free PM-friendly persistent B+-tree for eADR-enabled PM systems

    Bowen Zhang, Shengan Zheng, Zhenlin Qi, and Linpeng Huang. NBTree: a lock-free PM-friendly persistent B+-tree for eADR-enabled PM systems. Proc. VLDB Endow., 15(6):1187–1200, feb 2022

  41. [49]

    Compute express link (CXL)

    Compute Express Link. Compute express link (CXL). https:// computeexpresslink.org/, February 2024

  42. [50]

    https://semiconductor

    CXL Memory Module - Box (CMM-B). https://semiconductor. samsung.com/news-events/tech-blog/cxl-memory-module-box- cmm-b/, 2024

  43. [51]

    TPP: Transpar- ent Page Placement for CXL-Enabled Tiered-Memory

    Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit Kanaujia, and Prakash Chauhan. TPP: Transpar- ent Page Placement for CXL-Enabled Tiered-Memory. In Proceedings of the 28th ACM Internat...

  44. [52]

    MEMTIS: Efficient Memory Tiering with Dynamic Page Clas- sification and Page Size Determination

    Taehyung Lee, Sumit Kumar Monga, Changwoo Min, and Young Ik Eom. MEMTIS: Efficient Memory Tiering with Dynamic Page Clas- sification and Page Size Determination. In Proceedings of the 29th Symposium on Operating Systems Principles , pages 17–34, 2023

  45. [53]

    Nomad: Non-Exclusive Memory Tiering via Trans- actional Page Migration

    Lingfeng Xiang, Zhen Lin, Weishu Deng, Hui Lu, Jia Rao, Yifan Yuan, and Ren Wang. Nomad: Non-Exclusive Memory Tiering via Trans- actional Page Migration. In 18th USENIX Symposium on Operating 14 Systems Design and Implementation (OSDI 24) , pages 19–35, Santa Clara, CA, July 2...

  46. [54]

    Midhul Vuppalapati and Rachit Agarwal. Tiered Memory Manage- ment: Access Latency is the Key! In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles , SOSP ’24, page 79–94, New York, NY, USA, 2024. Association for Computing Machin- ery

  47. [55]

    NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering, 2024

    Zhe Zhou, Yiqi Chen, Tao Zhang, Yang Wang, Ran Shu, Shuotao Xu, Peng Cheng, Lei Qu, Yongqiang Xiong, Jie Zhang, and Guangyu Sun. NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering, 2024

  48. [56]

    Cooper, Adam Silberstein, Erwin Tam, Raghu Ramakrishnan, and Russell Sears

    Brian F. Cooper, Adam Silberstein, Erwin Tam, Raghu Ramakrishnan, and Russell Sears. Benchmarking Cloud Serving Systems with YCSB. In Proceedings of the 1st ACM Symposium on Cloud Computing , SoCC ’10, pages 143–154. Association for Computing Machinery, 2010

  49. [57]

    Madhava Krishnan, Xinwei Fu, Sanidhya Kashyap, and Changwoo Min

    Wook-Hee Kim, R. Madhava Krishnan, Xinwei Fu, Sanidhya Kashyap, and Changwoo Min. PACTree: A High Performance Persistent Range Index Using PAC Guidelines. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles , SOSP ’21, pages 424–439, New York, NY, U...

  50. [58]

    https://github.com/alibaba/block-traces, 2020

    alibaba/block-traces. https://github.com/alibaba/block-traces, 2020

  51. [59]

    An introduction to the compute express link (CXL) interconnect

    Debendra Das Sharma, Robert Blankenship, and Daniel Berger. An introduction to the compute express link (CXL) interconnect. ACM Comput. Surv., 56(11), July 2024

  52. [60]

    https://computeexpresslink.org/wp-content/ uploads/2024/02/CXL-1.0-Specification.pdf , 2024

    CXL 1.0 specification. https://computeexpresslink.org/wp-content/ uploads/2024/02/CXL-1.0-Specification.pdf , 2024

  53. [61]

    https://computeexpresslink.org/wp-content/ uploads/2024/02/CXL-2.0-Specification.pdf , 2024

    CXL 2.0 specification. https://computeexpresslink.org/wp-content/ uploads/2024/02/CXL-2.0-Specification.pdf , 2024

  54. [62]

    https://computeexpresslink.org/wp-content/ uploads/2024/02/CXL-3.0-Specification.pdf , 2024

    CXL 3.0 specification. https://computeexpresslink.org/wp-content/ uploads/2024/02/CXL-3.0-Specification.pdf , 2024

  55. [63]

    https://www.intel

    Intel® Memory Latency Checker V3.11. https://www.intel. com/content/www/us/en/developer/articles/tool/intelr-memory- latency-checker.html, 2021

  56. [64]

    CXL Switch for Scalable & Composable Memory Pool- ing/Sharing

    JP Jiang. CXL Switch for Scalable & Composable Memory Pool- ing/Sharing. https://files.futurememorystorage.com/proceedings/ 2024/20240807_CXLT-202-1_Jiang.pdf, August 2024

  57. [65]

    [PATCH] mm: mempolicy: N:M interleave policy for tiered memory nodes

    Johannes Weiner. [PATCH] mm: mempolicy: N:M interleave policy for tiered memory nodes. https://lore.kernel.org/linux-mm/YqD0% 2FtzFwXvJ1gK6@cmpxchg.org/T/, 2022

  58. [66]

    WASP: Workload-Aware Self- Replicating Page-Tables for NUMA Servers

    Hongliang Qu and Zhibin Yu. WASP: Workload-Aware Self- Replicating Page-Tables for NUMA Servers. In Proceedings of the 29th ACM International Conference on Architectural Support for Pro- gramming Languages and Operating Systems, Volume 2 , ASPLOS ’24, page 1233–1249, New York,...

  59. [67]

    NUMASK: high performance scalable skip list for NUMA

    Henry Daly, Ahmed Hassan, Michael F Spear, and Roberto Palmieri. NUMASK: high performance scalable skip list for NUMA. In 32nd In- ternational Symposium on Distributed Computing (DISC 2018). Schloss- Dagstuhl-Leibniz Zentrum für Informatik, 2018

  60. [68]

    An adaptive concurrent priority queue for NUMA architectures

    Foteini Strati, Christina Giannoula, Dimitrios Siakavaras, Georgios Goumas, and Nectarios Koziris. An adaptive concurrent priority queue for NUMA architectures. In Proceedings of the 16th ACM International Conference on Computing Frontiers, CF ’19, page 135–144, New York, NY, ...

  61. [69]

    Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated Memory

    Kai Lu, Siqi Zhao, Haikang Shan, Qiang Wei, Guokuan Li, Jiguang Wan, Ting Yao, Huatao Wu, and Daohui Wang. Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated Memory. ACM Trans. Archit. Code Optim., 21(3), September 2024

  62. [70]

    ROART: Range-query Opti- mized Persistent ART

    Shaonan Ma, Kang Chen, Shimin Chen, Mengxing Liu, Jianglang Zhu, Hongbo Kang, and Yongwei Wu. ROART: Range-query Opti- mized Persistent ART. In 19th USENIX Conference on File and Storage Technologies (FAST 21), pages 1–16. USENIX Association, February 2021

  63. [71]

    PLIN: a persistent learned index for non-volatile memory with high performance and instant recovery

    Zhou Zhang, Zhaole Chu, Peiquan Jin, Yongping Luo, Xike Xie, Shouhong Wan, Yun Luo, Xufei Wu, Peng Zou, Chunyang Zheng, Guoan Wu, and Andy Rudoff. PLIN: a persistent learned index for non-volatile memory with high performance and instant recovery. Proc. VLDB Endow., 16(2):243–...

  64. [72]

    Persistent B+-trees in non-volatile main memory

    Shimin Chen and Qin Jin. Persistent B+-trees in non-volatile main memory. Proc. VLDB Endow., 8(7):786–797, February 2015

  65. [73]

    Prism: Optimizing Key-Value Store for Modern Heterogeneous Storage Devices

    Yongju Song, Wook-Hee Kim, Sumit Kumar Monga, Changwoo Min, and Young Ik Eom. Prism: Optimizing Key-Value Store for Modern Heterogeneous Storage Devices. In Proceedings of the 28th ACM In- ternational Conference on Architectural Support for Programming Lan- guages and Operatin...

  66. [74]

    https: //github.com/begeekmyfriend/bplustree, 2014

    begeekmyfriend/bplustree: A minimal but extreme fast B+ tree in- dexing structure demo for billions of key-value storage. https: //github.com/begeekmyfriend/bplustree, 2014

  67. [75]

    https://github

    armon/libart: Adaptive Radix Trees implemented in C. https://github. com/armon/libart, 2013

  68. [76]

    Anti-caching: A new approach to database man- agement system architecture

    Justin DeBrabant, Andrew Pavlo, Stephen Tu, Michael Stonebraker, and Stan Zdonik. Anti-caching: A new approach to database man- agement system architecture. Proceedings of the VLDB Endowment , 6(14):1942–1953, 2013

  69. [77]

    Trekking through siberia: Managing cold data in a memory-optimized database

    Ahmed Eldawy, Justin Levandoski, and Per-Åke Larson. Trekking through siberia: Managing cold data in a memory-optimized database. Proceedings of the VLDB Endowment , 7(11):931–942, 2014

  70. [78]

    XIndex: a scalable learned index for multicore data storage

    Chuzhe Tang, Youyun Wang, Zhiyuan Dong, Gansen Hu, Zhaoguo Wang, Minjie Wang, and Haibo Chen. XIndex: a scalable learned index for multicore data storage. In Proceedings of the 25th ACM SIGPLAN symposium on principles and practice of parallel programming , pages 308–320, 2020

  71. [79]

    Juncheng Yang, Yao Yue, and K. V. Rashmi. A large scale analysis of hundreds of in-memory cache clusters at Twitter. pages 191–208, November 2020

  72. [80]

    https://github.com/ kohler/masstree-beta, 2012

    kohler/masstree-beta: Beta release of Masstree. https://github.com/ kohler/masstree-beta, 2012

  73. [81]

    Towards an adaptable systems architecture for memory tiering at warehouse-scale

    Padmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar, David Culler, Zhiyi Xu, Jianing Fan, Christopher Kennelly, Bill McCloskey, Danijela Mijailovic, et al. Towards an adaptable systems architecture for memory tiering at warehouse-scale. InProceedings of the 28th ACM Internat...

  74. [82]

    Software-defined far memory in warehouse-scale computers

    Andres Lagar-Cavilla, Junwhan Ahn, Suleiman Souhlal, Neha Agar- wal, Radoslaw Burny, Shakeel Butt, Jichuan Chang, Ashwin Chau- gule, Nan Deng, Junaid Shahid, et al. Software-defined far memory in warehouse-scale computers. In Proceedings of the Twenty-Fourth International Conf...

  75. [83]

    FlexMem: Adaptive Page Profiling and Migration for Tiered Memory

    Dong Xu, Junhee Ryu, Kwangsik Shin, Pengfei Su, and Dong Li. FlexMem: Adaptive Page Profiling and Migration for Tiered Memory. In 2024 USENIX Annual Technical Conference (USENIX ATC 24) , pages 817–833, Santa Clara, CA, July 2024. USENIX Association

  76. [84]

    Hankins and Jignesh M

    Richard A. Hankins and Jignesh M. Patel. Effect of node size on the performance of cache-conscious b+-trees. In Proceedings of the 2003 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems , SIGMETRICS ’03, page 283–294, New York, NY, USA, 20...

  77. [85]

    https://www.montage- tech.com/MXC, 2024

    CXL® Memory eXpander Controller (MXC). https://www.montage- tech.com/MXC, 2024

  78. [86]

    Zhichao Cao, Siying Dong, Sagar Vemuri, and David H. C. Du. Charac- terizing, modeling, and benchmarking RocksDB key-value workloads at facebook. In Proceedings of the 18th USENIX Conference on File and Storage Technologies, FAST’20, pages 209–224, USA, 2020. USENIX Association

  79. [87]

    Qiuping Wang, Jinhong Li, Patrick P. C. Lee, Tao Ouyang, Chao Shi, and Lilong Huang. Separating Data via Block Invalidation Time In- 15 ference for Write Amplification Reduction in Log-Structured Storage. In 20th USENIX Conference on File and Storage Technologies (FAST 22) , p...

  80. [88]

    Demystifying CXL memory with genuine CXL-ready systems and devices

    Yan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, et al. Demystifying CXL memory with genuine CXL-ready systems and devices. In Proceedings of the 56th Annual IEEE/ACM International Symposium on Micr...

  81. [89]

    Nimble page management for tiered memory systems

    Zi Yan, Daniel Lustig, David Nellans, and Abhishek Bhattacharjee. Nimble page management for tiered memory systems. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems , pages 331–345, 2019

  82. [90]

    What’s the story in EBS glory: evolutions and lessons in building cloud block store

    Weidong Zhang, Erci Xu, Qiuping Wang, Xiaolu Zhang, Yuesheng Gu, Zhenwei Lu, Tao Ouyang, Guanqun Dai, Wenwen Peng, Zhe Xu, Shuo Zhang, Dong Wu, Yilei Peng, Tianyun Wang, Haoran Zhang, Jiasheng Wang, Wenyuan Yan, Yuanyuan Dong, Wenhui Yao, Zhongjie Wu, Lingjun Zhu, Chao Shi, Yi...

  83. [91]

    https://www.intel

    Performance monitoring unit sharing guide. https://www.intel. com/content/www/us/en/content-details/727001/performance- monitoring-unit-sharing-guide.html , 2022

  84. [92]

    Amd uprof v4.0 user guide

    Advanced Micro Devices. Amd uprof v4.0 user guide. https://www.amd.com/content/dam/amd/en/documents/ developer/uprof-v4.0-gaGA-user-guide.pdf , 2022. Accessed: 2024-03-26

  85. [93]

    Neha Agarwal and Thomas F. Wenisch. Thermostat: Application- transparent page management for two-tiered main memory. In Pro- ceedings of the Twenty-Second International Conference on Architec- tural Support for Programming Languages and Operating Systems , ASPLOS ’17, pages 63...

  86. [94]

    https://git.kernel.org/pub/scm/linux/kernel/git/vishal/tiering.git/ log/?h=tiering-0.8, 2022

    kernel/git/vishal/tiering.git - Vishal Verma’s fork of linux.git. https://git.kernel.org/pub/scm/linux/kernel/git/vishal/tiering.git/ log/?h=tiering-0.8, 2022

  87. [95]

    Exploring the Design Space of Page Management for Multi-Tiered Memory Systems

    Jonghyeon Kim, Wonkyo Choe, and Jeongseob Ahn. Exploring the Design Space of Page Management for Multi-Tiered Memory Systems. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) , pages 715–728. USENIX Association, July 2021

  88. [96]

    HeteroOS: OS Design for Heterogeneous Memory Manage- ment in Datacenter

    Sudarsun Kannan, Ada Gavrilovska, Vishal Gupta, and Karsten Schwan. HeteroOS: OS Design for Heterogeneous Memory Manage- ment in Datacenter. In Proceedings of the 44th Annual International Symposium on Computer Architecture, ISCA ’17, pages 521–534, New York, NY, USA, 2017. As...

  89. [97]

    MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large Memory

    Jie Ren, Dong Xu, Junhee Ryu, Kwangsik Shin, Daewoo Kim, and Dong Li. MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large Memory. InProceedings of the Nineteenth European Conference on Computer Systems , EuroSys ’24, pages 803–817, New York, NY, USA, 2024. As...

  90. [98]

    Page migration support for disaggregated non-volatile memories

    Vamsee Reddy Kommareddy, Simon David Hammond, Clayton Hughes, Ahmad Samih, and Amro Awad. Page migration support for disaggregated non-volatile memories. In Proceedings of the Interna- tional Symposium on Memory Systems , MEMSYS ’19, pages 417–427, New York, NY, USA, 2019. Ass...

  91. [99]

    John, and Arkaprava Basu

    Jee Ho Ryoo, Lizy K. John, and Arkaprava Basu. A case for granu- larity aware page migration. In Proceedings of the 2018 International Conference on Supercomputing, ICS ’18, pages 352–362, New York, NY, USA, 2018. Association for Computing Machinery

  92. [100]

    Dancing in the dark: Profiling for tiered memory

    Jinyoung Choi, Sergey Blagodurov, and Hung-Wei Tseng. Dancing in the dark: Profiling for tiered memory. In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS) , pages 13–22. IEEE, 2021

  93. [101]

    Automatic numa balancing

    Rik van Riel and Vinod Chegu. Automatic numa balancing. In Red Hat Summit, 2014

  94. [102]

    Multi-clock: Dynamic tiering for hybrid memory systems

    Adnan Maruf, Ashikee Ghosh, Janki Bhimani, Daniel Campello, Andy Rudoff, and Raju Rangaswami. Multi-clock: Dynamic tiering for hybrid memory systems. In HPCA, pages 925–937, 2022

  95. [103]

    More cache for less cash: CXL Memory Bandwidth and Capacity Expansion in Software Caches

    Don Moon, Daniel Byrne, and Sounak Gupta. More cache for less cash: CXL Memory Bandwidth and Capacity Expansion in Software Caches. In 2023 OCP Global Summit - Server: Composable Memory System (CMS). OCP, 2023. Presented by Don Moon (SK hynix), Daniel Byrne (Intel), and Sounak...

  96. [104]

    https://github.com/ facebook/CacheLib/discussions/102, 2021

    Introducing new memory types to cachelib. https://github.com/ facebook/CacheLib/discussions/102, 2021

  97. [105]

    Data tiering in heterogeneous memory systems

    Subramanya R Dulloor, Amitabha Roy, Zheguang Zhao, Narayanan Sundaram, Nadathur Satish, Rajesh Sankaran, Jeff Jackson, and Karsten Schwan. Data tiering in heterogeneous memory systems. In Proceedings of the Eleventh European Conference on Computer Systems , pages 1–16, 2016

  98. [106]

    Unimem: runtime data man- agementon non-volatile memory-based heterogeneous main memory

    Kai Wu, Yingchao Huang, and Dong Li. Unimem: runtime data man- agementon non-volatile memory-based heterogeneous main memory. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , SC ’17, New York, NY, USA, 2017. Asso...

  99. [107]

    Opti- mal data placement for heterogeneous cache, memory, and storage systems

    Lei Zhang, Reza Karimi, Irfan Ahmad, and Ymir Vigfusson. Opti- mal data placement for heterogeneous cache, memory, and storage systems. Proceedings of the ACM on Measurement and Analysis of Computing Systems, 4(1):1–27, 2020

  100. [108]

    M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory Systems

    Yan Sun, Jongyul Kim, Dougas Yu, Jiyuan Zhang, Siyuan Chai, Michael Jaemin Kim, Hwayong Nam, Jaehyun Park, Eojin Na, Yifan Yuan, Ren Wang, Jung Ho Ahn, Tianyin Xu, and Nam Sung Kim. M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory Systems. In Proc...

  101. [109]

    Berger, Carl Waldspurger, Ryan Wee, Ishwar Agarwal, Rajat Agarwal, Frank Hady, Karthik Kumar, Mark D

    Yuhong Zhong, Daniel S. Berger, Carl Waldspurger, Ryan Wee, Ishwar Agarwal, Rajat Agarwal, Frank Hady, Karthik Kumar, Mark D. Hill, Mosharaf Chowdhury, and Asaf Cidon. Managing Memory Tiers with CXL in Virtualized Environments. In 18th USENIX Symposium on Operating Systems Des...

  102. [110]

    Johnny Cache: the End of DRAM Cache Conflicts (in Tiered Main Memory Systems)

    Baptiste Lepers and Willy Zwaenepoel. Johnny Cache: the End of DRAM Cache Conflicts (in Tiered Main Memory Systems). In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23), pages 519–534, Boston, MA, July 2023. USENIX Association

  103. [111]

    CXL and the Return of Scale-Up Database Engines

    Alberto Lerner and Gustavo Alonso. CXL and the Return of Scale-Up Database Engines. Proc. VLDB Endow., 17(10):2568–2575, August 2024

  104. [112]

    Enabling CXL Memory Expansion for In-Memory Database Management Systems

    Minseon Ahn, Andrew Chang, Donghun Lee, Jongmin Gim, Jungmin Kim, Jaemin Jung, Oliver Rebholz, Vincent Pham, Krishna Malladi, and Yang Seok Ki. Enabling CXL Memory Expansion for In-Memory Database Management Systems. In Proceedings of the 18th Interna- tional Workshop on Data ...

  105. [113]

    Elastic Use of Far Memory for In-Memory Database Management Systems

    Donghun Lee, Thomas Willhalm, Minseon Ahn, Suprasad Mutalik De- sai, Daniel Booss, Navneet Singh, Daniel Ritter, Jungmin Kim, and Oliver Rebholz. Elastic Use of Far Memory for In-Memory Database Management Systems. In Proceedings of the 19th International Work- shop on Data Ma...

  106. [114]

    Database Kernels: Seamless Integration of Database Sys- tems and Fast Storage via CXL

    Sangjin Lee, Alberto Lerner, Philippe Bonnet, and Philippe Cudré- Mauroux. Database Kernels: Seamless Integration of Database Sys- tems and Fast Storage via CXL. In CIDR, 2024

  107. [115]

    Pasha: An Efficient, Scalable Database Architecture for CXL Pods

    Yibo Huang, Newton Ni, Vijay Chidambaram, Emmett Witchel, and Dixin Tang. Pasha: An Efficient, Scalable Database Architecture for CXL Pods. 16

  108. [116]

    Patronus: high-performance and protective remote memory

    Bin Yan, Youyou Lu, Qing Wang, Minhui Xie, and Jiwu Shu. Patronus: high-performance and protective remote memory. In Proceedings of the 21st USENIX Conference on File and Storage Technologies , FAST’23, USA, 2023. USENIX Association

  109. [117]

    Marcos K. Aguilera, Nadav Amit, Irina Calciu, Xavier Deguillard, Jayneel Gandhi, Stanko Novaković, Arun Ramanathan, Pratap Sub- rahmanyam, Lalith Suresh, Kiran Tati, Rajesh Venkatasubramanian, and Michael Wei. Remote regions: a simple abstraction for remote memory. In 2018 USE...

  110. [118]

    Adaptive Placement for In-memory Storage Functions

    Ankit Bhardwaj, Chinmay Kulkarni, and Ryan Stutsman. Adaptive Placement for In-memory Storage Functions. In 2020 USENIX An- nual Technical Conference (USENIX ATC 20) , pages 127–141. USENIX Association, July 2020

  111. [119]

    Effectively prefetching remote memory with leap

    Hasan Al Maruf and Mosharaf Chowdhury. Effectively prefetching remote memory with leap. In Proceedings of the 2020 USENIX Confer- ence on Usenix Annual Technical Conference, USENIX ATC’20, USA,

  112. [120]

    UniMem: Redesigning Disaggregated Memory within A Unified Local-Remote Memory Hierarchy

    Yijie Zhong, Minqiang Zhou, Zhirong Shen, and Jiwu Shu. UniMem: Redesigning Disaggregated Memory within A Unified Local-Remote Memory Hierarchy. In 2024 USENIX Annual Technical Conference (USENIX ATC 24), pages 463–477, Santa Clara, CA, July 2024. USENIX Association

  113. [121]

    Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory

    Chenxi Wang, Yifan Qiao, Haoran Ma, Shi Liu, Wenguang Chen, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. Canvas: Isolated and Adaptive Swapping for Multi-Applications on Remote Memory. In 20th USENIX Symposium on Networked Systems Design and Imple- mentation (NSDI 23), p...

  114. [122]

    Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony

    Yifan Qiao, Chenxi Wang, Zhenyuan Ruan, Adam Belay, Qingda Lu, Yiying Zhang, Miryung Kim, and Guoqing Harry Xu. Hermit: Low-Latency, High-Throughput, and Transparent Remote Memory via Feedback-Directed Asynchrony. In 20th USENIX Symposium on Networked Systems Design and Implem...

  115. [123]

    Mira: A program- behavior-guided far memory system

    Zhiyuan Guo, Zijian He, and Yiying Zhang. Mira: A program- behavior-guided far memory system. In Proceedings of the 29th Sym- posium on Operating Systems Principles , pages 692–708, 2023

  116. [124]

    Aguilera, and Adam Belay

    Zhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, and Adam Belay. AIFM: high-performance, application-integrated far memory. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation, OSDI’20, USA, 2020. USENIX Association

  117. [125]

    Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Yuxin Su, Jiazhen Gu, Hao Feng, Yangfan Zhou, and Michael R. Lyu. Ditto: An Elastic and Adaptive Memory-Disaggregated Caching System. In Proceedings of the 29th Symposium on Operating Systems Principles , SOSP ’23, pages 675–691, New Yo...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.