Pith. sign in

REVIEW 5 major objections 5 minor 138 references

Valinor: Architectural Support for Fast, Energy-Efficient and Programmable Physical Memory Allocation

T0 review · 5 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Valinor makes page allocation both hardware-fast and fully programmable by executing OS-supplied allocation libraries on a small engine in the memory controller.

desk verdict Valinor's programmable page-allocation engine is a real new idea with an honest FPGA prototype, but the paper never measures the steady-state translation cost of its no-PTE fast path, so the headline speedup is narrower than it looks. read the letter →

arxiv 2607.14789 v1 pith:O3J4UWXR submitted 2026-07-16 cs.AR

classification cs.AR
keywords memoryallocationpagefaulthandlingprogrammablehardwareenginehardware-OScooperationserverlessworkloadsplacementpolicydrambankpressuresingle-ownershipprotocol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that physical memory allocation—the kernel routine that establishes virtual-to-physical mappings on demand—can be moved out of the software page-fault path and into hardware without sacrificing the ability to change allocation policy. Valinor adds a small programmable engine inside the memory controller that executes OS-supplied allocation libraries, so a minor page fault on a library-bound region is resolved entirely in hardware, avoiding the kernel trap, pipeline flushes, and lock contention that make Linux faults cost thousands of cycles. The authors argue this is significant for serverless functions and microservices, where minor faults can account for up to 54% of runtime and up to 40% of system energy. On a RISC-V FPGA prototype running Linux, they report 17x faster minor page-fault handling, 16% average end-to-end speedups, and up to 8% energy savings, with under 1.5% area overhead on an 8-core design. The paper's deeper claim is that hardware-speed allocation and software-programmable placement policies are not in conflict, and that the same substrate can express bank-pressure-aware, integrity-checking, speculative, and tier-aware allocation policies.

What carries the argument

The load-bearing mechanism is the programmable allocation engine (PAE): a lightweight, three-stage in-order RISC-V pipeline with 4 KB caches, placed in the memory controller and shared across cores. It executes trusted OS-supplied allocation libraries via a compact RV32I-compatible ISA extended with telemetry-read instructions. Two supporting structures carry the rest of the argument: the VMA filter, a small CAM-plus-properties cache that lets the page-table walker decide in hardware whether a faulting region is library-bound and with what permissions; and the single-ownership coordination protocol, which replaces cache coherence with an explicit, barrier-carrying handoff of mapping authorit

What would settle it

Run a multi-threaded test where one thread faults on a PAE-managed page while a second thread calls mprotect on the same region to revoke write permission; if the PAE's alloc commits a writable mapping after the mprotect completes, the fenced VMA-mutation protocol has a bug. Alternatively, measure whether PAE translate-routine traffic for a workload whose working set exceeds TLB capacity keeps minor-fault latency below the baseline kernel handler; if not, the 16% end-to-end speedup will not generalize to larger-footprint workloads.

Watch

Extended reading notes

Core claim

Valinor's central claim is that page allocation can be both hardware-fast and software-programmable. The paper introduces the programmable allocation engine (PAE), a small three-stage in-order pipeline placed in the memory controller, running compact 'allocation libraries' that the OS loads and configures per process or per memory object. When the page-table walker faults on a library-bound region, the request is forwarded to the PAE, which executes the library's translate and alloc routines and returns a page-table entry directly to the MMU, completing the allocation without a kernel trap. The paper reports that on a BOOM RISC-V soft-core running Linux this makes minor page-fault handling 1

Load-bearing premise

The end-to-end gains rest on the assumption that a PAE-created mapping need not be installed in the OS page table on the fast path: every later TLB miss on such a page must be serviced by the PAE's translate routine, and correctness depends on the single-ownership protocol whose own revision note admits 'remaining concurrency corner cases' were closed only by making two barrier rules explicit.

Editorial extensions

If this is right

  • Minor page faults on library-bound regions can be resolved without entering the kernel, removing pipeline flushes, TLB shootdowns, and zone-lock contention from the critical path.
  • Allocation policy becomes a loadable software artifact: the OS can bind different regions to different libraries, so placement can adapt to DRAM bank pressure, memory tiers, or integrity requirements without silicon redesign.
  • Because the PAE sits in the memory controller, libraries can act on live telemetry—such as per-bank conflict counters—that a software handler cannot cheaply observe, enabling interference-aware placement for co-located workloads.
  • PAE-created mappings are kept out of the OS page table on the fast path, so after TLB eviction the PAE's translate routine, not the kernel, services the miss; this keeps fault handling in hardware but makes the PAE a new element on the translation critical path.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same template—an OS-programmable engine executing kernel data-plane routines—could extend to other trap-heavy services, such as TLB shootdown handling or page-table walks for non-allocation faults; Valinor does not claim this, but its PAE is a natural incubator for it.
  • The fast path makes the PAE's metadata effectively a second page table. A stress test interleaving fork, munmap, and mprotect with PAE-managed allocations would expose whether the single-ownership invariant truly covers all interleavings; the paper's own revision note indicates 'remaining concurrency corner cases' were only recently closed by making two barrier rules explicit.
  • The reported energy savings come from a single in-order engine shared across cores; on many-core or multi-socket systems the PAE could become a contention point for TLB-refill traffic, so a larger-scale evaluation would be needed to confirm the per-allocation energy and latency claims.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. Valinor proposes a hardware-OS cooperative memory allocation substrate built around a programmable allocation engine (PAE) at the memory controller and a VMA filter (VMAF) in the MMU path. Applications or the OS bind memory regions to compact OS-supplied allocation libraries, and minor page faults for those regions are resolved by the PAE without trapping into the kernel. On the fast path, PAE-created mappings are cached in the TLB and recorded in library metadata but not installed in the OS page table; later TLB misses are serviced by the PAE's translate routine. The paper reports a 17× reduction in minor-fault handling time on a BOOM RISC-V FPGA prototype running Linux, 16% average end-to-end speedup, 5% average (up to 8%) energy savings, and <1.5% area overhead, together with a simulated study of six allocation libraries and four PAE microarchitectures.

Significance. If the measurements hold, Valinor would be a significant advance: it is the first page-granularity allocation substrate to combine programmable, policy-rich allocation with hardware-class fault handling, and it demonstrates policies—bank-pressure-aware placement, mapping-integrity checks, speculation, tier-aware placement—that fixed-function hardware cannot express. The FPGA prototype on a real BOOM/Linux system, the direct comparison against SW-only and Fixed-HW implementations, the pool-size sensitivity analysis (Fig. 12), and the six-library design-space exploration are concrete strengths that go well beyond a purely simulated proposal. The main caveats are that the no-PTE fast path moves post-allocation TLB refills onto a single shared PAE without being measured, and that the headline energy claims are not actually shown; both directly qualify the central end-to-end claims.

major comments (5)
  1. [§6.3.2] The fast path deliberately does not install PAE-created mappings into the OS page table, so any later TLB miss on a PAE-managed page must be serviced by the PAE's translate routine instead of a hardware page-table walk over a present PTE. The paper reports per-fault latency (§9.1) but never reports TLB-refill traffic, PAE queue occupancy, or PAE translate-hit latency relative to a baseline PTW. For workloads whose active footprint exceeds L2 TLB capacity and whose pages are reused, every reuse after TLB eviction must traverse VMAF+PAE through a single shared in-order engine (§9.3). The 17× fault-latency improvement is therefore not the right metric for the end-to-end claim; total translation cost per page over its lifetime is needed. Without this measurement, the 16% average speedup may not generalize beyond first-touch-dominated workloads.
  2. [Abstract; §3.3; §9.1] The abstract states that minor faults account for 'up to 40% of system energy,' but the measurements in §3.3 and Fig. 2 report up to 18% (and §1 also says 18%). This is a 2.2× discrepancy in a headline motivation. In addition, §9.1 says energy was measured with Synopsys but the figure is 'not shown due to space constraints,' so the claimed 5% average / up to 8% energy savings is not actually presented. Energy is central to the title and abstract; the authors should include the missing figure and reconcile the 40% vs 18% numbers.
  3. [§9.1] The text says Fixed-HW achieves an average speedup of 16% and is 'the fastest,' while Valinor also achieves an average speedup of 16% yet is 'only 3% slower on average' than Fixed-HW. If Valinor is 3% slower, its speedup should be approximately 12.5%, not 16%, unless the numbers are rounded to the nearest integer. This contradiction affects the central claim of matching fixed hardware. Please report precise per-benchmark and average values, either in the figure or in a table.
  4. [§6.5] The no-PTE fast path makes the PAE's metadata authoritative for PAE-owned regions, so correctness rests entirely on the single-ownership invariant. However, the protocol is described only in prose, and the section's own text says the 'revision makes the two barrier rules below explicit to close the remaining concurrency corner cases.' No formal model, proof, model-checking result, or concurrency stress test is provided for the drain and VMA-mutation fences. A missed race between drain and a fault, or between a VMA mutation and an in-flight alloc, could install a stale or conflicting mapping. Given that this is the correctness backbone of the architecture, a verification argument or litmus-test evidence is needed.
  5. [§8] The design-space exploration uses Virtuoso's MimicOS configured to 'mimic Linux's minor page-fault handling behavior.' Since Virtuoso is from the same group and no validation against the real kernel's fault path is reported, the simulated library-latency numbers may partly reflect the imitation's accuracy. Please provide a validation showing that MimicOS reproduces the §3 Linux fault-path characteristics (e.g., instruction counts, lock contention) for the evaluated workloads, or bound the sensitivity of Fig. 13 to the imitation.
minor comments (5)
  1. [Fig. 10; Fig. 12] Check axis and label typos: Fig. 10 appears to contain 'trnasnp' (likely 'transp'), and Fig. 12's y-axis is labeled 'Reduction in Page Fault Latency' but reaches 120% with reduction values above 100%, which is not meaningful as a reduction; clarify whether this is a speedup ratio.
  2. [§7.3] The tier-utility weights w_bw, w_cap, locality_value_hot/cold are free parameters, but Fig. 15 sweeps only the per-VMA hint fraction. Please report the default parameter values and provide at least a small sensitivity analysis, since the 2.6× latency result may depend on them.
  3. [§6.5] The manuscript contains self-referential revision language ('The submitted design already enforces... The revision makes...') and 'revision' phrasing that should be removed for archival publication; the final paper should present the protocol declaratively.
  4. [§6.4] The claim that the fork() drain cost is 'insignificant' because fork already walks the VMA tree is not self-evident for large PAE-owned segments; this cost is not measured. A quantitative statement of the drain overhead would strengthen the CoW/fork correctness argument.
  5. [§9.1] The energy methodology is too thin: it says only 'energy measurements with Synopsys.' Please specify what was measured (core, DRAM, full system), how the FPGA prototype's power was obtained, and how the 5% average / up-to-8% energy savings were derived.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; headline FPGA results are measured against real Linux; only minor non-load-bearing self-citation of the Virtuoso simulator.

full rationale

Valinor's headline numbers—17× fault-latency reduction, 16% end-to-end speedup, up to 8% energy savings, <1.5% area—come from the FPGA prototype running real Linux, not from a fitted model: §9.1 reports 'baseline Linux requires 5881 cycles per minor fault' vs 'Valinor ... still provides a 17.07x improvement over Linux', and Fig. 10 is explicitly the 'end-to-end performance speedup measured on the FPGA prototype'. These are empirical measurements with an independent hardware ground truth. The only self-citation in the evaluation chain is Virtuoso: 'We also implement Valinor in a full-system simulator (Virtuoso [10]+Sniper [11])' and 'We configured MimicOS (the lightweight OS modeled in Virtuoso) to mimic Linux's minor page-fault handling behavior'. Virtuoso is a same-group simulator, and the OS-imitation methodology is calibrated to the real-Xeon measurements of §3; it is used for design-space exploration and the six-library evaluation, not for the central FPGA speedup claims. The tier-aware placement result (§9.2, Fig. 15) shows the per-VMA hint steering pages toward fast tiers; this follows from the library's own utility definition (utility(t)=V(hint)·locality(t)−w_bw·bw_util(t)−w_cap·cap_util(t)), so the 'hint steers placement' observation is by construction a sanity check of the library, not an independent prediction—but it is not load-bearing for the paper's main claims, which are about the programmability of the PAE substrate. The §6.5 revision note ('the revision makes the two barrier rules below explicit to close the remaining concurrency corner cases') is a correctness/verification limitation, not a circularity, and it does not enter the numerical results. No target result is assumed to produce the numbers.

Assumptions & free parameters 2 free parameters · 4 assumptions · 2 invented entities

The central evaluation rests on the chosen pool size, the fidelity of the authors' own OS simulator, the unverified ownership protocol, and trust in PAE library code. None of these are fitted to produce the headline numbers, but they are hand-chosen or assumed.

free parameters (2)
  • Allocation pool size = 16 MB
    FPGA experiments set the short-lived allocation pool to 16 MB by default; Fig. 12 shows page-fault latency reduction depends strongly on pool size (4-16 MB), so headline speedups are tied to this hand-chosen configuration.
  • Tier-utility weights w_bw, w_cap, locality_value_hot/cold = not specified
    The tier-aware library's utility function in §7.3 uses hand-chosen weights; no sensitivity analysis is given, so the placement results of Fig. 15 are conditional on these values.
assumptions (4)
  • domain assumption Virtuoso's MimicOS faithfully mimics Linux minor page-fault handling behavior.
    Full-system evaluation (§8) configures MimicOS to mimic Linux's minor fault behavior; all simulator results inherit the accuracy of this imitation.
  • domain assumption The single-ownership invariant of §6.5 is sufficient for correctness and can be enforced without hardware cache coherence.
    The paper provides a prose protocol with barrier rules, not a formal proof or machine-checked model; the revision note admits that the submitted design left concurrency corner cases to be closed by the new rules.
  • domain assumption PAE library code is trusted and user-space cannot influence which library runs or corrupt its code.
    The security argument in §6.4 relies on OS-managed CSRs and on libraries executing inside the kernel trust domain; a library bug or compromised OS is declared out of scope.
  • domain assumption VMA lookup can be performed in hardware by the VMAF using a small CAM plus a reprogrammable vma_walk routine.
    The PTW fast path depends on VMAF correctly and cheaply discovering memory-object properties; this requires cooperation from kernel VMA data structures and the correctness of the vma_walk function.
invented entities (2)
  • Programmable Allocation Engine (PAE) independent evidence
    purpose: Executes OS-supplied allocation libraries (alloc/translate/free) in hardware, resolving page faults without trapping into the kernel.
    Implemented as a RISCV-MINI in-order core on an FPGA prototype and synthesized for area/power with Yosys; not validated in production silicon.
  • Virtual Memory Area Filter (VMAF) independent evidence
    purpose: Hardware filter in the page-table walker that identifies library-bound memory objects and extracts VMA properties, enabling hardware allocation eligibility checks.
    Part of the FPGA prototype and simulated design; correctness depends on OS VMA layouts and the vma_walk routine.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Valinor: Architectural Support for Fast, Energy-Efficient and Programmable Physical Memory Allocation." pith.science (2026). https://pith.science/paper/O3J4UWXR

@misc{pith2026260714789,
  author       = {Pith},
  title        = {Pith review of: Valinor: Architectural Support for Fast, Energy-Efficient and Programmable Physical Memory Allocation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O3J4UWXR}},
  note         = {Machine review of arXiv:2607.14789}
}
read the original abstract

Physical memory allocation establishes virtual-to-physical mappings on demand. In current systems, each minor page fault traps into the kernel and triggers pipeline flushes, stalls, and a long sequence of allocation steps that can cost tens of thousands of cycles. These overheads are increasingly significant for short-lived workloads such as serverless functions and microservices, where minor faults can account for up to 54% of runtime and up to 40% of system energy. Prior hardware allocation proposals avoid traps and context switches, but either sacrifice useful placement optimizations or rely on fixed-function logic that cannot adapt to new policies or changing hardware conditions. We present Valinor, a hardware-OS cooperative memory allocation substrate that combines software flexibility with hardware-class performance. Valinor introduces a programmable hardware allocation engine that executes compact OS-supplied allocation libraries at close to fixed-hardware speed. It supports diverse policies, including short-lived object allocators, integrity mechanisms, and hardware-telemetry-guided placement. We implement Valinor on a BOOM RISC-V soft core running Linux and in a full-system simulator. On real hardware, Valinor accelerates allocation by 17x, improves end-to-end performance by 16%, and reduces energy consumption by up to 8%. Full-system simulation further evaluates the programmable allocation engine and six allocation libraries, showing that Valinor provides hardware-class performance without sacrificing programmability.

Figures

Figures reproduced from arXiv: 2607.14789 by the authors.

Figure 1
Figure 1. Resolving an anonymous, minor page fault in Linux [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 3
Figure 3. CDF of execution latency of minor page faults. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Page fault latency overhead with and without THP [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (10 more)
Figure 5
Figure 5. Figure 5: CDF of number of instructions executed during [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Minor page fault latency breakdown. 4 Limitations of Prior Work Prior works proposed accelerating page allocation by offloading parts of memory management to hardware, avoiding software fault￾handling overheads [5–7]. Tirumalasetty et al. [7] extend the MMU with a hard…
Figure 8
Figure 8. Figure 8: Valinor’s architecture and integration with the OS. [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 7
Figure 7. Figure 7: Valinor’s software architecture (Low-latency allocation library example). [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 9
Figure 9. Figure 9: Valinor’s hardware architecture. 8 [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 12
Figure 12. Figure 12: Page fault latency reduction across different allo [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 11
Figure 11. Figure 11: Average number of cycles taken per page fault for [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 13
Figure 13. Figure 13: Page fault handling latency reduction across four allocation libraries and four PAE designs. Fixed HW is a fixed-function implementation of the same Valinor policy used in each group. Adaptive Policies Beyond Fixed HW. The previous experi￾ment holds the allocation pol…
Figure 15
Figure 15. Figure 15: Tier-aware placement as latency-sensitive VMA fraction varies (simulation). Bars show pages placed per tier (left axis). The red line shows mean allocation latency (right axis). The per-VMA intent hint steers placement from CXL (fraction 0) to Local DRAM (fraction 1.0…
Figure 14
Figure 14. Figure 14: Normalized execution time of co-located applica [PITH_FULL_IMAGE:figures/full_fig_p013_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

138 extracted references · 4 linked inside Pith

  1. [1]

    Memory Coherence in Shared Virtual Memory Systems

    Kai Li and Paul Hudak. Memory Coherence in Shared Virtual Memory Systems. InTOCS, 1989

  2. [2]

    Appel and Kai Li

    Andrew W. Appel and Kai Li. Virtual Memory Primitives for User Programs. InASPLOS, 1991

  3. [3]

    Machine-Independent Virtual Memory Management for Paged Uniprocessor and Multiprocessor Ar- chitectures

    Richard Rashid, Avadis Tevanian, Michael Young, David Golub, Robert Baron, David Black, William Bolosky, and Jonathan Chew. Machine-Independent Virtual Memory Management for Paged Uniprocessor and Multiprocessor Ar- chitectures. InOSR, 1987

  4. [4]

    Satyanarayanan, Henry H

    M. Satyanarayanan, Henry H. Mashburn, Puneet Kumar, David C. Steere, and James J. Kistler. Lightweight Recoverable Virtual Memory. InSOSP, 1993

  5. [5]

    Memento: Architectural Support for Ephemeral Mem- ory Management in Serverless Environments

    Ziqi Wang, Kaiyang Zhao, Pei Li, Andrew Jacob, Michael Kozuch, Todd Mowry, and Dimitrios Skarlatos. Memento: Architectural Support for Ephemeral Mem- ory Management in Serverless Environments. InMICRO, 2023

  6. [6]

    Lee, and Jinkyu Jeong

    Gyusun Lee, Wenjing Jin, Wonsuk Song, Jeonghun Gong, Jonghyun Bae, Tae Jun Ham, Jae W. Lee, and Jinkyu Jeong. A Case for Hardware-Based Demand Paging. InISCA, 2020

  7. [7]

    Reducing Minor Page Fault Overheads through Enhanced Page Walker

    Chandrahas Tirumalasetty, Chih Chieh Chou, Narasimha Reddy, Paul Gratz, and Ayman Abouelwafa. Reducing Minor Page Fault Overheads through Enhanced Page Walker. InTACO, 2022

  8. [8]

    Benchmarking, Analysis, and Optimization of Serverless Function Snap- shots

    Dmitrii Ustiugov, Plamen Petrov, Marios Kogias, Edouard Bugnion, and Boris Grot. Benchmarking, Analysis, and Optimization of Serverless Function Snap- shots. InASPLOS, 2021

Show all 138 references
  1. [9]

    Accessed: April 7th, 2026

    Xilinx ZCU 106 FPGA: https://www.amd.com/en/products/adaptive-socs-and- fpgas/evaluation-boards/zcu106.html. Accessed: April 7th, 2026

  2. [10]

    Nisa Bostanci, An- dreas Kosmas Kakolyris, Berkin K

    Konstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, An- dreas Kosmas Kakolyris, Berkin K. Konar, Rahul Bera, Mohammad Sadrosadati, Rakesh Kumar, Nandita Vijaykumar, and Onur Mutlu. Virtuoso: Enabling fast and accurate virtual memory research via an imitation-ba...

  3. [11]

    Carlson, Wim Heirman, and Lieven Eeckhout

    Trevor E. Carlson, Wim Heirman, and Lieven Eeckhout. Sniper: Exploring the Level of Abstraction for Scalable and Accurate Parallel Multi-Core Simulations. InSC, 2011

  4. [12]

    Shen and James L

    Kenneth K. Shen and James L. Peterson. A Weighted Buddy Method for Dynamic Storage Allocation. InCACM, 1974

  5. [13]

    Plimpton, Ron Brightwell, Courtenay Vaughan, Keith Underwood, and Mike Davis

    Steven J. Plimpton, Ron Brightwell, Courtenay Vaughan, Keith Underwood, and Mike Davis. A Simple Synchronous Distributed-Memory Algorithm for the HPCC RandomAccess Benchmark. InCluster, 2006

  6. [14]

    Tramm, Andrew R

    John R. Tramm, Andrew R. Siegel, Tanzima Islam, and Martin Schulz. XSBench - The Development and Verification of a Performance Abstraction for Monte Carlo Reactor Analysis. InPHYSOR, 2014

  7. [15]

    A Hybrid CNN–LSTM Model for Typhoon Formation Forecasting

    Chen Rui, Xiang Wang, Weimin Zhang, Xiaoyu Zhu, Aiping Li, and Chao Yang. A Hybrid CNN–LSTM Model for Typhoon Formation Forecasting. 2019

  8. [16]

    HPCG Benchmark: A New Metric for Ranking High Performance Computing Systems

    Jack Dongarra, Michael A Heroux, and Piotr Luszczek. HPCG Benchmark: A New Metric for Ranking High Performance Computing Systems. Techni- cal Report UT-EECS-15-736, Univ. of Tennessee Dept. of Electrical Engg. and Computer Sci., 2015

  9. [17]

    https://docs.mongodb.org/ manual/tutorial/transparent-huge-pages/, 2016

    MongoDB Recommends Disabling Huge Pages. https://docs.mongodb.org/ manual/tutorial/transparent-huge-pages/, 2016

  10. [18]

    http://antirez.com/news/52, 2016

    Redis Recommends Disabling Huge Pages. http://antirez.com/news/52, 2016

  11. [19]

    Distributed caching with memcached

    Brad Fitzpatrick. Distributed caching with memcached. InLinux Journal, 2004

  12. [20]

    http://www.tpc.org/tpcc/detail.asp

    TPC-C. http://www.tpc.org/tpcc/detail.asp

  13. [21]

    Transaction Processing Performance Council (TPC).TPC BENCHMARK H, 2.17.1 edition, 2014

  14. [22]

    Ligra: A Lightweight Graph Processing Frame- work for Shared Memory

    Julian Shun and Guy E Blelloch. Ligra: A Lightweight Graph Processing Frame- work for Shared Memory. InPPoPP, 2013

  15. [23]

    Graphicionado: A High-Performance and Energy-Efficient Acceler- ator for Graph Analytics

    Tae Jun Ham, Lisa Wu, Narayanan Sundaram, Nadathur Satish, and Margaret Martonosi. Graphicionado: A High-Performance and Energy-Efficient Acceler- ator for Graph Analytics. InMICRO, 2016

  16. [24]

    The PageR- ank Citation Ranking: Bringing Order to the Web

    Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. The PageR- ank Citation Ranking: Bringing Order to the Web. Technical report, Stanford InfoLab, 1999

  17. [25]

    GraphMineSuite: Enabling High-Performance and Programmable Graph Mining Algorithms with Set Algebra

    Maciej Besta, Zur Vonarburg-Shmaria, Yannick Schaffner, Leonardo Schwarz, Grzegorz Kwasniewski, Lukas Gianinazzi, Jakub Beranek, Kacper Janda, Tobias Holenstein, Sebastian Leisinger, Peter Tatkowski, Esref Ozdemir, Adrian Balla, Marcin Copik, Philipp Lindenberger, Pavel Kalvod...

  18. [26]

    Graph 500 Large-Scale Benchmarks

    Graph 500. Graph 500 Large-Scale Benchmarks. http://www.graph500.org/

  19. [27]

    Vasimud- din, Sanchit Misra, David Blaauw, Satish Narayanasamy, and Reetuparna Das

    Arun Subramaniyan, Yufeng Gu, Timothy Dunn, Somnath Paul, Md. Vasimud- din, Sanchit Misra, David Blaauw, Satish Narayanasamy, and Reetuparna Das. GenomicsBench: A Benchmark Suite for Genomics. InISPASS, 2021

  20. [28]

    Biobench: A Benchmark Suite of Bioinformatics Applications

    Kursad Albayraktaroglu, Aamer Jaleel, Xue Wu, Manoj Franklin, Bruce Ja- cob, Chau-Wen Tseng, and Donald Yeung. Biobench: A Benchmark Suite of Bioinformatics Applications. InISPASS, 2005

  21. [29]

    GRIM-Filter: Fast Seed Location Filtering in DNA Read Mapping Using Processing-in-Memory Technologies

    Jeremie S Kim, Damla Senol Cali, Hongyi Xin, Donghyuk Lee, Saugata Ghose, Mohammed Alser, Hasan Hassan, Oguz Ergin, Can Alkan, and Onur 14 Mutlu. GRIM-Filter: Fast Seed Location Filtering in DNA Read Mapping Using Processing-in-Memory Technologies. 2018

  22. [30]

    GenStore: A High-Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis

    Nika Mansouri Ghiasi, Jisung Park, Harun Mustafa, Jeremie Kim, Ataberk Olgun, Arvid Gollwitzer, Damla Senol Cali, Can Firtina, Haiyu Mao, Nour Almadhoun Alserr, et al. GenStore: A High-Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis. I...

  23. [31]

    Damla Senol Cali, Konstantinos Kanellopoulos, Joël Lindegger, Zülal Bingöl, Gurpreet S. Kalsi, Ziyi Zuo, Can Firtina, Meryem Banu Cavlak, Jeremie Kim, Nika Mansouri Ghiasi, Gagandeep Singh, Juan Gómez-Luna, Nour Almadhoun Alserr, Mohammed Alser, Sreenivas Subramoney, Can Alkan...

  24. [32]

    Mapping Single Molecule Sequencing Reads Using Basic Local Alignment with Successive Refinement (BLASR): Application and Theory

    Mark J Chaisson and Glenn Tesler. Mapping Single Molecule Sequencing Reads Using Basic Local Alignment with Successive Refinement (BLASR): Application and Theory. 2012

  25. [33]

    Aligning Sequence Reads, Clone Sequences and Assembly Contigs with BWA-MEM

    Heng Li. Aligning Sequence Reads, Clone Sequences and Assembly Contigs with BWA-MEM. arXiv:1303.3997 [q-bio.GN], 2013

  26. [34]

    Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform

    Heng Li and Richard Durbin. Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform. 2009

  27. [35]

    A Comparison of Seed-and- Extend Techniques in Modern DNA Read Alignment Algorithms

    Nauman Ahmed, Koen Bertels, and Zaid Al-Ars. A Comparison of Seed-and- Extend Techniques in Modern DNA Read Alignment Algorithms. InBIBM, 2016

  28. [36]

    Accelerating Genome Analysis: A Primer on an Ongoing Journey

    Mohammed Alser, Zülal Bingöl, Damla Senol Cali, Jeremie Kim, Saugata Ghose, Can Alkan, and Onur Mutlu. Accelerating Genome Analysis: A Primer on an Ongoing Journey. 2020

  29. [37]

    GateKeeper: A New Hardware Architecture for Accelerating Pre- Alignment in DNA Short Read Mapping

    Mohammed Alser, Hasan Hassan, Hongyi Xin, Oğuz Ergin, Onur Mutlu, and Can Alkan. GateKeeper: A New Hardware Architecture for Accelerating Pre- Alignment in DNA Short Read Mapping. 2017

  30. [38]

    Appuswamy, J

    R. Appuswamy, J. Fellay, and N. Chaturvedi. Sequence Alignment Through the Looking Glass. InIPDPSW, 2018

  31. [39]

    GenASM: A High-Performance, Low- Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis

    Damla Senol Cali, Gurpreet S Kalsi, Zülal Bingöl, Can Firtina, Lavanya Sub- ramanian, Jeremie S Kim, Rachata Ausavarungnirun, Mohammed Alser, Juan Gomez-Luna, Amirali Boroumand, et al. GenASM: A High-Performance, Low- Power Approximate String Matching Acceleration Framework fo...

  32. [40]

    GRIM- Filter: Fast Seed Filtering in Read Mapping using Emerging Memory Technolo- gies

    Jeremie S Kim, Damla Senol, Hongyi Xin, Donghyuk Lee, Saugata Ghose, Mo- hammed Alser, Hasan Hassan, Oguz Ergin, Can Alkan, and Onur Mutlu. GRIM- Filter: Fast Seed Filtering in Read Mapping using Emerging Memory Technolo- gies. arXiv:1708.04329 [q-bio.GN], 2017

  33. [41]

    GraphAligner: Rapid and Versatile Sequence-to-Graph Alignment

    Mikko Rautiainen and Tobias Marschall. GraphAligner: Rapid and Versatile Sequence-to-Graph Alignment. InGenome Biology, 2020

  34. [42]

    Architectural Implications of Function-as-a-Service Computing

    Mohammad Shahrad, Jonathan Balkind, and David Wentzlaff. Architectural Implications of Function-as-a-Service Computing. MICRO, 2019

  35. [43]

    Catalyzer: Sub-Millisecond Startup for Serverless Computing with Initialization-Less Booting

    Dong Du, Tianyi Yu, Yubin Xia, Binyu Zang, Guanglu Yan, Chenggang Qin, Qixuan Wu, and Haibo Chen. Catalyzer: Sub-Millisecond Startup for Serverless Computing with Initialization-Less Booting. ASPLOS, 2020

  36. [44]

    Wisefuse: Workload characterization and dag transformation for serverless workflows.Proceedings of the ACM on Measurement and Analysis of Computing Systems, 6(2):1–28, 2022

    Ashraf Mahgoub, Edgardo Barsallo Yi, Karthick Shankar, Eshaan Minocha, Sameh Elnikety, Saurabh Bagchi, and Somali Chaterji. Wisefuse: Workload characterization and dag transformation for serverless workflows.Proceedings of the ACM on Measurement and Analysis of Computing Syste...

  37. [45]

    Architectural implications of function-as-a-service computing

    Mohammad Shahrad, Jonathan Balkind, and David Wentzlaff. Architectural implications of function-as-a-service computing. InProceedings of the 52nd annual IEEE/ACM international symposium on microarchitecture, pages 1063– 1075, 2019

  38. [46]

    Xfaas: Hyperscale and low cost serverless functions at meta

    Alireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang, Abhigna Nagaraja, Neeraj Pathak, Girish Joshi, Carla Souza, Bo Huang, Wyatt Cook, et al. Xfaas: Hyperscale and low cost serverless functions at meta. InProceedings of the 29th Symposium on Operating Systems Pri...

  39. [47]

    Faster and cheaper serverless computing on harvested resources

    Yanqi Zhang, Íñigo Goiri, Gohar Irfan Chaudhry, Rodrigo Fonseca, Sameh Elnikety, Christina Delimitrou, and Ricardo Bianchini. Faster and cheaper serverless computing on harvested resources. InProceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles, pages 7...

  40. [48]

    Icebreaker: Warming serverless functions better with heterogeneity

    Rohan Basu Roy, Tirthak Patel, and Devesh Tiwari. Icebreaker: Warming serverless functions better with heterogeneity. InProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, pages 753–767, 2022

  41. [49]

    Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider

    Mohammad Shahrad, Rodrigo Fonseca, Íñigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini. Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. In2020 U...

  42. [50]

    Sebs: A serverless benchmark suite for function-as-a-service computing, 2021

    Marcin Copik, Grzegorz Kwasniewski, Maciej Besta, Michal Podstawski, and Torsten Hoefler. Sebs: A serverless benchmark suite for function-as-a-service computing, 2021

  43. [51]

    Faa $ t: A transparent auto-scaling cache for serverless applications

    Francisco Romero, Gohar Irfan Chaudhry, Íñigo Goiri, Pragna Gopa, Paul Ba- tum, Neeraja J Yadwadkar, Rodrigo Fonseca, Christos Kozyrakis, and Ricardo Bianchini. Faa $ t: A transparent auto-scaling cache for serverless applications. InProceedings of the ACM symposium on cloud c...

  44. [52]

    {SONIC}: Application-aware data passing for chained serverless applications

    Ashraf Mahgoub, Li Wang, Karthick Shankar, Yiming Zhang, Huangshi Tian, Subrata Mitra, Yuxing Peng, Hongqi Wang, Ana Klimovic, Haoran Yang, et al. {SONIC}: Application-aware data passing for chained serverless applications. In2021 USENIX Annual Technical Conference (USENIX ATC...

  45. [53]

    Dvfaas: Leveraging dvfs for faas workflows.IEEE Computer Architecture Letters, 2023

    Achilleas Tzenetopoulos, Dimosthenis Masouros, Dimitrios Soudris, and Sotirios Xydis. Dvfaas: Leveraging dvfs for faas workflows.IEEE Computer Architecture Letters, 2023

  46. [54]

    Sequence clock: A dynamic resource orchestrator for serverless architectures

    Ioannis Fakinos, Achilleas Tzenetopoulos, Dimosthenis Masouros, Sotirios Xy- dis, and Dimitrios Soudris. Sequence clock: A dynamic resource orchestrator for serverless architectures. In2022 IEEE 15th International Conference on Cloud Computing (CLOUD), pages 81–90, 2022

  47. [55]

    In2023 USENIX Annual Technical Conference (USENIX ATC 23), pages 879–896, 2023

    Ghazal Sadeghian, Mohamed Elsakhawy, Mohanna Shahrad, Joe Hattori, and Mohammad Shahrad.{UnFaaSener}: Latency and cost aware offloading of func- tions from serverless platforms. In2023 USENIX Annual Technical Conference (USENIX ATC 23), pages 879–896, 2023

  48. [56]

    {ORION} and the three rights: Sizing, bundling, and prewarming for serverless{DAGs}

    Ashraf Mahgoub, Edgardo Barsallo Yi, Karthick Shankar, Sameh Elnikety, So- mali Chaterji, and Saurabh Bagchi. {ORION} and the three rights: Sizing, bundling, and prewarming for serverless{DAGs}. In16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), ...

  49. [57]

    Stepconf: Slo-aware dynamic resource configuration for serverless function workflows

    Zhaojie Wen, Yishuo Wang, and Fangming Liu. Stepconf: Slo-aware dynamic resource configuration for serverless function workflows. InIEEE INFOCOM 2022-IEEE Conference on Computer Communications, pages 1868–1877. IEEE, 2022

  50. [58]

    Sizeless: Predicting the optimal size of serverless functions

    Simon Eismann, Long Bui, Johannes Grohmann, Cristina Abad, Nikolas Herbst, and Samuel Kounev. Sizeless: Predicting the optimal size of serverless functions. InProceedings of the 22nd International Middleware Conference, pages 248–259, 2021

  51. [59]

    Parrotfish: Parametric regression for optimizing serverless functions

    Arshia Moghimi, Joe Hattori, Alexander Li, Mehdi Ben Chikha, and Mohammad Shahrad. Parrotfish: Parametric regression for optimizing serverless functions. InProceedings of the 2023 ACM Symposium on Cloud Computing, pages 177–192, 2023

  52. [60]

    How does it function? characterizing long- term trends in production serverless workloads

    Artjom Joosen, Ahmed Hassan, Martin Asenov, Rajkarn Singh, Luke Darlow, Jianfeng Wang, and Adam Barker. How does it function? characterizing long- term trends in production serverless workloads. InProceedings of the 2023 ACM Symposium on Cloud Computing, pages 443–458, 2023

  53. [61]

    Cloud programming simplified: A berkeley view on serverless computing.arXiv preprint arXiv:1902.03383, 2019

    Eric Jonas, Johann Schleier-Smith, Vikram Sreekanti, Chia-Che Tsai, Anurag Khandelwal, Qifan Pu, Vaishaal Shankar, Joao Carreira, Karl Krauth, Neeraja Yadwadkar, et al. Cloud programming simplified: A berkeley view on serverless computing.arXiv preprint arXiv:1902.03383, 2019

  54. [62]

    Specfaas: Accelerating serverless applications with speculative function execution

    Jovan Stojkovic, Tianyin Xu, Hubertus Franke, and Josep Torrellas. Specfaas: Accelerating serverless applications with speculative function execution. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 814–827. IEEE, 2023

  55. [63]

    Kraken: Adaptive container provisioning for deploying dynamic dags in serverless platforms

    Vivek M Bhasi, Jashwant Raj Gunasekaran, Prashanth Thinakaran, Cyan Subhra Mishra, Mahmut Taylan Kandemir, and Chita Das. Kraken: Adaptive container provisioning for deploying dynamic dags in serverless platforms. InProceedings of the ACM Symposium on Cloud Computing, pages 15...

  56. [64]

    Triggerflow: trigger-based orchestration of serverless workflows

    Pedro García López, Aitor Arjona, Josep Sampé, Aleksander Slominski, and Lionel Villard. Triggerflow: trigger-based orchestration of serverless workflows. InProceedings of the 14th ACM international conference on distributed and event- based systems, pages 3–14, 2020

  57. [65]

    Server- less data analytics in the ibm cloud

    Josep Sampé, Gil Vernik, Marc Sánchez-Artigas, and Pedro García-López. Server- less data analytics in the ibm cloud. InProceedings of the 19th International Middleware Conference Industry, pages 1–8, 2018

  58. [66]

    Agile cold starts for scalable serverless

    Anup Mohan, Harshad Sane, Kshitij Doshi, Saikrishna Edupuganti, Naren Nayak, and Vadim Sukhomlinov. Agile cold starts for scalable serverless. In 11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19), 2019

  59. [67]

    Ecofaas: Rethinking the design of serverless environments for energy efficiency

    Jovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke, and Josep Torrellas. Ecofaas: Rethinking the design of serverless environments for energy efficiency. InProceedings of the 51st Annual International Symposium on Computer Architecture (ISCA’24), 2024

  60. [68]

    Incendio: priority-based scheduling for alleviating cold start in serverless computing.IEEE Transactions on Computers, 2024

    Xinquan Cai, Qianlong Sang, Chuang Hu, Yili Gong, Kun Suo, Xiaobo Zhou, and Dazhao Cheng. Incendio: priority-based scheduling for alleviating cold start in serverless computing.IEEE Transactions on Computers, 2024

  61. [69]

    Faas- batch: Boosting serverless efficiency with in-container parallelism and resource multiplexing.IEEE Transactions on Computers, 73(4):1071–1085, 2024

    Zhaorui Wu, Yuhui Deng, Yi Zhou, Jie Li, Shujie Pang, and Xiao Qin. Faas- batch: Boosting serverless efficiency with in-container parallelism and resource multiplexing.IEEE Transactions on Computers, 73(4):1071–1085, 2024

  62. [70]

    Jolteon: unleashing the promise of serverless for serverless workflows

    Zili Zhang, Chao Jin, and Xin Jin. Jolteon: unleashing the promise of serverless for serverless workflows. In21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 167–183, 2024

  63. [71]

    Towards elastic memory allocation of serverless functions in disaggre- gated memory systems

    Achilleas Tzenetopoulos, Dimosthenis Masouros, Dimitrios Soudris, and Sotirios Xydis. Towards elastic memory allocation of serverless functions in disaggre- gated memory systems. InProceedings of the 4th Workshop on Heterogeneous Composable and Disaggregated Systems, HCDS 2025...

  64. [72]

    Faasmem: Improving memory efficiency of serverless computing with memory pool architecture

    Chuhao Xu, Yiyu Liu, Zijun Li, Quan Chen, Han Zhao, Deze Zeng, Qian Peng, Xueqi Wu, Haifeng Zhao, Senbo Fu, and Minyi Guo. Faasmem: Improving memory efficiency of serverless computing with memory pool architecture. In Rajiv Gupta, Nael B. Abu-Ghazaleh, Madan Musuvathi, and Dan...

  65. [73]

    Codecrunch: Improving serverless performance via function compression and cost-aware warmup location optimization

    Rohan Basu Roy, Tirthak Patel, Rohan Garg, and Devesh Tiwari. Codecrunch: Improving serverless performance via function compression and cost-aware warmup location optimization. In Rajiv Gupta, Nael B. Abu-Ghazaleh, Madan Musuvathi, and Dan Tsafrir, editors,Proceedings of the 2...

  66. [74]

    Serverless cold starts and where to find them

    Artjom Joosen, Ahmed Hassan, Martin Asenov, Rajkarn Singh, Luke Darlow, Jianfeng Wang, Qiwen Deng, and Adam Barker. Serverless cold starts and where to find them. InProceedings of the Twentieth European Conference on Computer Systems, pages 938–953, 2025

  67. [75]

    The Architectural Implications of Cloud Microservices

    Yu Gan and Christina Delimitrou. The Architectural Implications of Cloud Microservices. 2018

  68. [76]

    An Open-Source Benchmark Suite for Microservices and Their Hardware-Software Implications for Cloud & Edge Systems

    Yu Gan, Yanqi Zhang, Dailun Cheng, Ankitha Shetty, Priyal Rathi, Nayan Katarki, Ariana Bruno, Justin Hu, Brian Ritchken, Brendon Jackson, Kelvin Hu, Meghna Pancholi, Yuan He, Brett Clancy, Chris Colen, Fukang Wen, Catherine Leung, Siyuan Wang, Leon Zaruvinsky, Mateo Espinosa, ...

  69. [77]

    Ursa: Lightweight resource management for cloud-native microservices

    Yanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety, and Christina Delimitrou. Ursa: Lightweight resource management for cloud-native microservices. InIEEE International Symposium on High-Performance Computer Architecture, HPCA 2024, Edinburgh, United Kingdom, March 2-6, 2024, page...

  70. [78]

    Rusty: Runtime interference-aware predictive monitoring for modern multi-tenant systems

    Dimosthenis Masouros, Sotirios Xydis, and Dimitrios Soudris. Rusty: Runtime interference-aware predictive monitoring for modern multi-tenant systems. IEEE Transactions on Parallel and Distributed Systems, 32(1):184–198, 2020

  71. [79]

    Caliper: Interference estimator for multi-tenant environments sharing architectural resources.ACM Transactions on Architecture and Code Optimization (TACO), 16(3):1–25, 2019

    Ram Srivatsa Kannan, Michael Laurenzano, Jeongseob Ahn, Jason Mars, and Lingjia Tang. Caliper: Interference estimator for multi-tenant environments sharing architectural resources.ACM Transactions on Architecture and Code Optimization (TACO), 16(3):1–25, 2019

  72. [80]

    In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pages 177–194, 2018

    Akshitha Sriraman and Thomas F Wenisch.{𝜇 Tune}:{Auto-Tuned} threading for{OLDI} microservices. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pages 177–194, 2018

  73. [81]

    Parties: Qos-aware resource partitioning for multiple interactive services

    Shuang Chen, Christina Delimitrou, and José F Martínez. Parties: Qos-aware resource partitioning for multiple interactive services. InProceedings of the Twenty-Fourth International Conference on Architectural Support for Program- ming Languages and Operating Systems, pages 107...

  74. [82]

    Clite: Efficient and qos-aware co-location of multiple latency-critical jobs for warehouse scale computers

    Tirthak Patel and Devesh Tiwari. Clite: Efficient and qos-aware co-location of multiple latency-critical jobs for warehouse scale computers. In2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), pages 193–206. IEEE, 2020

  75. [83]

    Understanding, predicting and scheduling serverless workloads under partial interference

    Laiping Zhao, Yanan Yang, Yiming Li, Xian Zhou, and Keqiu Li. Understanding, predicting and scheduling serverless workloads under partial interference. In Proceedings of the International conference for high performance computing, networking, storage and analysis, pages 1–15, 2021

  76. [84]

    Twig: Multi-agent task management for colocated latency-critical cloud services

    Rajiv Nishtala, Vinicius Petrucci, Paul Carpenter, and Magnus Sjalander. Twig: Multi-agent task management for colocated latency-critical cloud services. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), pages 167–179. IEEE, 2020

  77. [85]

    Sinan: Ml-based and qos-aware resource management for cloud microservices

    Yanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G Edward Suh, and Christina Delimitrou. Sinan: Ml-based and qos-aware resource management for cloud microservices. InProceedings of the 26th ACM international conference on architectural support for programming languages and operatin...

  78. [86]

    Ant-man: Towards agile power management in the microservice era

    Xiaofeng Hou, Chao Li, Jiacheng Liu, Lu Zhang, Yang Hu, and Minyi Guo. Ant-man: Towards agile power management in the microservice era. InSC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–14. IEEE, 2020

  79. [87]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient Memory Management for Large Language Model Serving with PagedAttention. InSOSP, 2023

  80. [88]

    LLaMA: Open and Efficient Foundation Language Models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and Efficient Foundation Language...

  81. [89]

    BAGEL: Bootstrapping Agents by Guiding Exploration with Language

    Shikhar Murty, Christopher Manning, Peter Shaw, Mandar Joshi, and Kenton Lee. BAGEL: Bootstrapping Agents by Guiding Exploration with Language. In ICML, 2024

  82. [90]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...

  83. [91]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need. InNIPS 2017

  84. [92]

    GPT-4 Technical Report

    OpenAI. GPT-4 Technical Report

  85. [93]

    throttll’em: Predictive gpu throttling for energy efficient llm inference serving

    Andreas Kosmas Kakolyris, Dimosthenis Masouros, Petros Vavaroutsos, Sotirios Xydis, and Dimitrios Soudris. throttll’em: Predictive gpu throttling for energy efficient llm inference serving. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA), p...

  86. [94]

    Dynamo: Accelerating language model inference with dynamic multi-token sampling.arXiv preprint arXiv:2405.00888, 2024

    Shikhar Tuli, Chi-Heng Lin, Yen-Chang Hsu, Niraj K Jha, Yilin Shen, and Hongxia Jin. Dynamo: Accelerating language model inference with dynamic multi-token sampling.arXiv preprint arXiv:2405.00888, 2024

  87. [95]

    {ServerlessLLM}:{Low-Latency} serverless in- ference for large language models

    Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. {ServerlessLLM}:{Low-Latency} serverless in- ference for large language models. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24), pages 135–153, 2024

  88. [96]

    Parsing gigabytes of JSON per second.The VLDB Journal, 28(6):941–960, 2019

    Geoff Langdale and Daniel Lemire. Parsing gigabytes of JSON per second.The VLDB Journal, 28(6):941–960, 2019

  89. [97]

    https://github.com/vhive-serverless/vSwarm

    vSwarm. https://github.com/vhive-serverless/vSwarm

  90. [98]

    1-bit ai infra: Part 1.1, fast and lossless bitnet b1.58 inference on cpus, 2024

    Jinheng Wang, Hansong Zhou, Ting Song, Shaoguang Mao, Shuming Ma, Hongyu Wang, Yan Xia, and Furu Wei. 1-bit ai infra: Part 1.1, fast and lossless bitnet b1.58 inference on cpus, 2024

  91. [99]

    Rapl: Memory power estimation and capping

    Howard David, Eugene Gorbatov, Ulf R Hanebutte, Rahul Khanna, and Christian Le. Rapl: Memory power estimation and capping. InProceedings of the 16th ACM/IEEE international symposium on Low power electronics and design, pages 189–194, 2010

  92. [100]

    Fast and efficient memory reclamation for serverless microvms

    (see arxiv:2411.12893). Fast and efficient memory reclamation for serverless microvms. 2024

  93. [101]

    ftrace - Function Tracer: https://docs.kernel.org/trace/ftrace.html

  94. [102]

    Coordinated bank and cache coloring for temporal protection of memory accesses

    Noriaki Suzuki, Hyoseung Kim, Dionisio de Niz, Bjorn Andersson, Lutz Wrage, Mark Klein, and Ragunathan Rajkumar. Coordinated bank and cache coloring for temporal protection of memory accesses. In2013 IEEE 16th International Conference on Computational Science and Engineering

  95. [103]

    Tmc: Near-optimal resource allocation for tiered-memory systems

    Yuanjiang Ni, Pankaj Mehra, Ethan Miller, and Heiner Litz. Tmc: Near-optimal resource allocation for tiered-memory systems. SoCC ’23, 2023

  96. [104]

    Tiered-Latency DRAM: A Low Latency and Low Cost DRAM Architecture

    Donghyuk Lee, Yoongu Kim, Vivek Seshadri, Jamie Liu, Lavanya Subramanian, and Onur Mutlu. Tiered-Latency DRAM: A Low Latency and Low Cost DRAM Architecture. InHPCA, 2013

  97. [105]

    Citadel: Rethinking mem- ory allocation to safeguard against inter-domain rowhammer exploits

    Anish Saxena, Walter Wang, and Alexandros Daglis. Citadel: Rethinking mem- ory allocation to safeguard against inter-domain rowhammer exploits. MICRO ’25, 2025

  98. [106]

    Aws lambda: Run code without thinking about servers or clusters

  99. [107]

    Borg, omega, and kubernetes.Queue, 14(1):70–93, 2016

    Brendan Burns, Brian Grant, David Oppenheimer, Eric Brewer, and John Wilkes. Borg, omega, and kubernetes.Queue, 14(1):70–93, 2016

  100. [108]

    TCMalloc: Thread-Caching Malloc

    Sanjay Ghemawat and Paul Menage. TCMalloc: Thread-Caching Malloc. http: //goog-perftools.sourceforge.net/doc/tcmalloc.html

  101. [109]

    https://docs.kernel.org/admin-guide/mm/concepts.html

    Anonymous Memory. https://docs.kernel.org/admin-guide/mm/concepts.html

  102. [110]

    https://www.kernel.org/

    The linux kernel 6.13.0. https://www.kernel.org/

  103. [111]

    Accessed: April 7th, 2026

    RISCV-MINI: https://github.com/ucb-bar/riscv-mini. Accessed: April 7th, 2026

  104. [112]

    Accessed: April 7th, 2026

    RISC-V architecture: https://riscv.org/specifications/ratified/. Accessed: April 7th, 2026

  105. [113]

    Ralph C. Merkle. A digital signature based on a conventional encryption function. InAdvances in Cryptology — CRYPTO ’87, volume 293 ofLecture Notes in Computer Science, pages 369–378. Springer, 1987

  106. [114]

    Linux (5.11.6) [operating system]

    Linus Torvalds. Linux (5.11.6) [operating system]. https://github.com/torvalds/ linux/releases/tag/

  107. [115]

    Nisa Bostanci, An- dreas Kosmas Kakolyris, Berkin Kerim Konar, Rahul Bera, Mohammad Sadrosa- dati, Rakesh Kumar, Nandita Vijaykumar, and Onur Mutlu

    Konstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, An- dreas Kosmas Kakolyris, Berkin Kerim Konar, Rahul Bera, Mohammad Sadrosa- dati, Rakesh Kumar, Nandita Vijaykumar, and Onur Mutlu. Virtuoso: Enabling fast and accurate virtual memory research via an imitati...

  108. [116]

    Design Compiler

    Synopsys, Inc. Design Compiler. Accessed: October 31th, 2021

  109. [117]

    Hycube: A cgra with reconfigurable single-cycle multi-hop interconnect

    Manupa Karunaratne, Aditi Kulkarni Mohite, Tulika Mitra, and Li-Shiuan Peh. Hycube: A cgra with reconfigurable single-cycle multi-hop interconnect. DAC ’17, 2017

  110. [118]

    Accelerating edge ai with morpher: An integrated design, compilation and simulation framework for cgras, 2023

    Dhananjaya Wijerathne, Zhaoying Li, and Tulika Mitra. Accelerating edge ai with morpher: An integrated design, compilation and simulation framework for cgras, 2023

  111. [119]

    https://github.com/YosysHQ/

    Yosys open synthesis suite. https://github.com/YosysHQ/. Accessed: April 7th, 2026

  112. [120]

    https://github.com/The-OpenROAD- Project/OpenROAD-flow-scripts.git

    Nangate freepdk45 open cell library. https://github.com/The-OpenROAD- Project/OpenROAD-flow-scripts.git. Accessed: April 7th, 2026

  113. [121]

    Chandrahas Tirumalasetty, Chih-Chieh Chou, A. L. Narasimha Reddy, Paul Gratz, and Ayman Abouelwafa. Reducing minor page fault overheads through 16 enhanced page walker.ACM Trans. Archit. Code Optim., 19(4):57:1–57:26, 2022

  114. [122]

    Clio: A Hardware-software co-designed Disaggregated Memory system

    Zhiyuan Guo, Yizhou Shan, Xuhao Luo, Yutong Huang, and Yiying Zhang. Clio: A Hardware-software co-designed Disaggregated Memory system. InASPLOS, 2022

  115. [123]

    URL: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2019/EECS- 2019-154.pdf

    Enabling efficient and transparent remote memory access in disaggregated datacenters. URL: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2019/EECS- 2019-154.pdf

  116. [124]

    Mallacc: Acceler- ating Memory Allocation

    Svilen Kanev, Sam Likun Xi, Gu-Yeon Wei, and David Brooks. Mallacc: Acceler- ating Memory Allocation. InASPLOS, 2017

  117. [125]

    H. Cam, M. Abd-El-Barr, and S.M. Sait. A high-performance hardware-efficient memory allocation technique and design. InProceedings 1999 IEEE Interna- tional Conference on Computer Design: VLSI in Computers and Processors (Cat. No.99CB37040), pages 274–276, 1999

  118. [126]

    Morris Chang and Edward F

    J. Morris Chang and Edward F. Gehringer. A High-Performance Memory Allocator for Object-Oriented Systems. InTC, 1996

  119. [127]

    Morris Chang, Witawas Srisa-An, and C-TD Lo

    J. Morris Chang, Witawas Srisa-An, and C-TD Lo. Architectural Support for Dynamic Memory Management. InICCD, 2000

  120. [128]

    A Page-Based Hybrid (Software- Hardware) Dynamic Memory Allocator

    Wentong Li, Saraju Mohanty, and Krishna Kavi. A Page-Based Hybrid (Software- Hardware) Dynamic Memory Allocator. InCAL, 2006

  121. [129]

    Fea- sibility of Decoupling Memory Management From the Execution Pipeline

    Wentong Li, Mehran Rezaei, Krishna Kavi, Afrin Naz, and Philip Sweany. Fea- sibility of Decoupling Memory Management From the Execution Pipeline. InJ. Syst. Archit., 2007

  122. [130]

    Seidl, and Mario Wolczko

    Greg Wright, Matthew L. Seidl, and Mario Wolczko. An object-aware memory architecture. Technical report, USA, 2005

  123. [131]

    Chang and E.F

    J.M. Chang and E.F. Gehringer. Evaluation of an object-caching coprocessor design for object-oriented systems. InProceedings of 1993 IEEE International Conference on Computer Design ICCD’93, pages 132–139, 1993

  124. [132]

    Flexible reference-counting-based hardware acceleration for garbage collection

    José A Joao, Onur Mutlu, and Yale N Patt. Flexible reference-counting-based hardware acceleration for garbage collection. InACM SIGARCH Computer Architecture News, volume 37, pages 418–428. ACM, 2009

  125. [133]

    Iris Bahar

    Samuel Thomas, Jiwon Choe, Ofir Gordon, Erez Petrank, Tali Moreshet, Maurice Herlihy, and R. Iris Bahar. Towards hardware accelerated garbage collection with near-memory processing. In2022 IEEE High Performance Extreme Com- puting Conference (HPEC), pages 1–6, 2022

  126. [134]

    Schmidt and Kelvin D

    William J. Schmidt and Kelvin D. Nilsen. Performance of a hardware-assisted real-time garbage collector. InProceedings of the Sixth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS VI, page 76–85, New York, NY, USA, 1994...

  127. [135]

    Wise, Brian Heck, Caleb Hess, Willie Hunt, and Eric Ost

    David S. Wise, Brian Heck, Caleb Hess, Willie Hunt, and Eric Ost. Research demonstration of a hardware reference-counting heap.Lisp Symb. Comput., 10(2):159–181, July 1997

  128. [136]

    M. Meyer. An On-Chip Garbage Collection Coprocessor for Embedded Real- Time Systems. InProceedings of the 11th IEEE International Conference on Embedded and Real-Time Computing Systems and Applications, pages 517–524, August 2005

  129. [137]

    Integrated hardware garbage collection.ACM Trans

    Andrés Amaya García, David May, and Ed Nutting. Integrated hardware garbage collection.ACM Trans. Embed. Comput. Syst., 20(5), July 2021

  130. [138]

    Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory

    Jaeyoung Jang, Jun Heo, Yejin Lee, Jaeyeon Won, Seonghak Kim, Sung Jun Jung, Hakbeom Jang, Tae Jun Ham, and Jae W Lee. Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory. InMICRO, 2019. 17

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.