REVIEW 3 major objections 3 minor 95 references
EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read EDM claims that moving the remote-memory protocol stack into the Ethernet PHY and adding a centralized in-switch scheduler cuts remote read and write latency to roughly 300 ns, beating RoCEv2 and TCP/IP by an order of magnitude and…
desk verdict Solid FPGA demonstration of a sub-300ns in-PHY memory fabric, but the high-load latency claims rest on a simulator with inconsistent parameters. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the Ethernet Physical Coding Sublayer (PCS), the part of the PHY that encodes data into 66-bit blocks, repurposed as the data path for memory messages via new /M*/, /N/, and /G/ block types. On the switch, the key mechanism is a centralized scheduler that implements priority-based Parallel Iterative Matching (PIM), the classic switch-scheduling algorithm that matches inputs to outputs iteratively, extended so each iteration takes a constant three clock cycles through ordered-list hardware and a priority encoder. The scheduler's grants create virtual circuits in the PHY between source and destination ports, which is what eliminates queuing, drops, and layer-2 forwarding at the switch for memory traffic.
What would settle it
Build or simulate a 512-port EDM switch at the claimed 3 GHz clock and measure whether a maximal matching is actually formed in about 27 cycles; if the matching latency exceeds the transmission time of the 128-byte or 256-byte chunk, the zero-queuing promise fails on the first incast, and average latency under 80 percent all-to-all load will exceed the claimed 1.3 times unloaded.
Extended reading notes
Core claim
The central claim is that the Ethernet MAC layer is not an unavoidable price of using Ethernet for remote memory access. By operating directly in the Physical Coding Sublayer, EDM carries memory messages as custom 66-bit PHY block types, reuses inter-frame gap bits, and achieves intra-frame preemption at 66-bit granularity, eliminating the 64-byte minimum frame, the IFG overhead, and head-of-line blocking caused by large non-memory frames. The companion claim is that a centralized scheduler implemented in the switch PHY can form a priority-based online maximal matching fast enough to reserve bandwidth in advance, so memory traffic is forwarded through virtual circuits with zero queuing and zero layer-2 processing delay. On the testbed, remote reads and writes complete through the fabric in about 300 ns, and in simulation the average completion time stays within 1.2 to 1.4 times the ideal for several disaggregated application workloads, outperforming reactive transports and CXL-style credit flow control under load.
Load-bearing premise
The scheduler's zero-queuing and high-utilization guarantees depend on it forming a maximal matching in about 3 times log(N) clock cycles at a clock rate fast enough that the chosen chunk size keeps every link busy; if computing the matching takes longer than sending a chunk, links go idle and queues build.
Editorial extensions
If this is right
- If EDM's numbers hold, remote memory reads and writes over Ethernet can run at latencies previously associated with PCIe-attached CXL, so memory disaggregation does not require a separate fabric.
- Memory traffic can share links with IP and storage traffic without sacrificing small-message latency, because 66-bit-granularity preemption bounds interference from large frames.
- Removing the TCP/IP or RDMA transport stack from the memory data path eliminates per-packet encapsulation and congestion-control delays, simplifying the host NIC for memory traffic.
- With a scheduler guaranteeing zero queuing, applications see predictable memory-access latency even under incast-like many-to-one loads.
- The ASIC synthesis suggests the scheduler fits within a modern switch chip: roughly 10 mm2 of area and about 1 MB of SRAM for a 512-port switch, which is a concrete scaling path beyond the two-port FPGA prototype.
Reading between the lines
- A testable extension the paper does not explore is whether the same PHY-level multiplexing can carry other latency-critical traffic, such as coherence or synchronization primitives, alongside memory and IP traffic; only memory messages and grants are defined.
- The paper fixes a 256-byte chunk and a 2.56 ns clock in the testbed; at higher line rates the zero-queuing guarantee would require either faster matching or larger chunks, and the latency cost of larger chunks for tiny 8-byte read requests is not quantified.
- Because one-sided writes pay a notification-and-grant round trip before sending data, a workload dominated by small one-sided writes would stress whether that RTT/2 overhead stays negligible in practice.
- A consequence the authors leave implicit is that a working EDM at scale would undercut the main argument for deploying CXL or InfiniBand purely for memory pooling, since Ethernet already exists in every rack and would also carry the disaggregated memory traffic.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes EDM, an Ethernet fabric for memory disaggregation that moves the remote-memory protocol stack into the Physical Coding Sublayer of the Ethernet PHY and adds a centralized, in-PHY scheduler at the switch. The scheduler builds a priority-based online maximal matching and creates PHY-level virtual circuits, with the goal of eliminating queuing, layer-2 forwarding, and transport-stack latency for memory traffic. The paper reports an unloaded remote read/write fabric latency of roughly 300 ns on an FPGA testbed, claims this is 3.7--12.7x lower than raw Ethernet, RoCEv2, and hardware TCP/IP, and uses a C network simulator to claim that under high load the average latency stays within 1.3x the unloaded value and that average message completion times are within 1.2--1.4x of ideal for several disaggregated workloads.
Significance. The unloaded-latency result is a substantial empirical contribution: it is measured on an FPGA testbed, backed by a detailed cycle-level breakdown in Table 1 and Figure 5, and accompanied by a public Verilog artifact. The idea of implementing remote memory access in the PHY, with custom 66-bit block types and intra-frame preemption, is novel and potentially impactful for memory disaggregation. The centralized in-PHY scheduler is also a well-motivated design point. The high-load claim, however, is the main reason the paper's central latency story generalizes beyond the unloaded case, and that claim currently rests on a simulator parameterization that is internally inconsistent with the scheduler timing model. The high-load result therefore needs to be re-established with a consistent set of parameters before the paper's conclusions are fully supported.
major comments (3)
- [§4.3, with §3.1.2 and §3.1.3] The simulator in §4.3 uses a 256 B chunk at 100 Gbps with N=144 nodes, but it never states the scheduler clock. If the simulator inherits the 2.56 ns testbed clock used in Figure 5, then forming a maximal matching takes 3·log2(144) ≈ 21.5 cycles, or about 55 ns, while a 256 B chunk on a 100 Gbps link transmits in 20.5 ns. This violates the line-rate condition derived in §3.1.3, which requires the chunk transmission time to be at least the matching latency to keep links busy. Under the 2.56 ns clock the per-port throughput would be capped near 256 B / 55 ns ≈ 4.7 Gbps, not 100 Gbps, so the 1.3x high-load latency figure is not supported by the stated parameters. If instead the scheduler runs at the 3 GHz ASIC clock of §4.1, the 256 B chunk is viable, but then the simulator does not match the testbed timing used elsewhere in §4.3. Please state the scheduler clock explicitly and rerun with a consistent chunk/clock combination, e.g., a 688 B chunk at 2.56 ns, or 256 B with the 3 GHz ASIC scheduler.
- [Figure 8 caption / inserted note] The manuscript itself admits a second simulator defect: read flows are initialized with 8 B RREQs, but the load accounting uses the 64 B RRES size, so the reported network load overstates the real injected traffic. The statement that this factor 'can be offsetted by real workloads' is not quantified and is not a substitute for correct load accounting. Because Figure 8a and the application MCT results in Figure 8b are the evidence for the high-load latency claim, the load definition must be corrected and the simulations repeated before the 1.3x claim can be accepted.
- [§4.2.1] The text states that 'testbed experiments showed that even under interference from IP traffic, EDM maintained a near-constant ~300 ns remote memory access latency,' but no such experiment, figure, or measurement procedure is presented in §4.2. Either the data should be reported, or the claim should be removed or explicitly labeled as a qualitative observation.
minor comments (3)
- [Figure 8] The inserted note about RREQ/RRES size accounting is written in informal, ungrammatical prose ('offsetted') and should be moved into the methodology section as a clearly stated assumption or limitation.
- [§4.2.2] The paper refers to a 'cycle-accurate FPGA hardware simulator' for the YCSB evaluation; please clarify whether this simulator is the same implementation as the testbed, and how its timing parameters were derived.
- [References] Reference [49] (Scale-out NUMA) appears to be duplicated as [50]; the duplicate should be removed and subsequent citations renumbered.
Circularity Check
No significant circularity: unloaded latency is measured, high-load behavior is simulated from the design, and self-citations are non-load-bearing building blocks.
full rationale
The paper's central latency result (~300 ns) is a direct FPGA measurement from an implemented testbed (Table 1, Figure 5), not a fitted or self-referential prediction. The high-load result (within 1.3x unloaded) comes from a C network simulator that implements the scheduler's grant algorithm; the zero-queuing property is a design invariant of maximal matching, not an empirical claim derived from itself. Empirical parameter choices (X=3, chunk size 256B) are simulator settings, not parameters fitted to the reported target metric. Citations to prior PHY idle-block work [36,37,39,60,91] and to constant-time ordered-list hardware structures [57-59,63] include the authors' own papers, but the EDM implementation and ASIC synthesis provide independent support; these citations are building blocks, not unverified premises that force the conclusion. The Figure 8 note admitting that read flows use 8B RREQs while load is accounted at 64B RRES, and the possible mismatch between scheduler matching latency and chunk transmission time, are correctness/consistency risks, not circular derivations. No step in the derivation chain equates an output to an input by construction.
Assumptions & free parameters
free parameters (3)
- chunk_size_c =
256 B in simulations; 128 B for line-rate 512-port estimate
- max_active_notifications_X =
3
- scheduler_clock_rate =
3 GHz (ASIC estimate)
assumptions (4)
- domain assumption The Ethernet PCS interface between encoder and scrambler permits insertion of custom 66-bit block types and reuse of idle blocks without breaking clock recovery or scrambling.
- domain assumption Memory traffic demand is known in advance: read sizes are available from RREQ and write sizes from explicit notifications.
- ad hoc to paper Link data corruption is non-transient and can be handled by disabling the link rather than retransmission.
- domain assumption A single switch with hundreds of ports is the target topology; multi-hop fabrics are out of scope.
invented entities (2)
-
/M*/, /N/, and /G/ 66-bit PHY block types
-
In-PHY virtual circuits
Cite this review
Pith. "Pith review of EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation." pith.science (2026). https://pith.science/paper/HWICA5OG
@misc{pith2026241108300,
author = {Pith},
title = {Pith review of: EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation},
year = {2026},
howpublished = {\url{https://pith.science/paper/HWICA5OG}},
note = {Machine review of arXiv:2411.08300}
}
abstract
Achieving low remote memory access latency remains the primary challenge in realizing memory disaggregation over Ethernet within the datacenters. We present EDM that attempts to overcome this challenge using two key ideas. First, while existing network protocols for remote memory access over the Ethernet, such as TCP/IP and RDMA, are implemented on top of the MAC layer, EDM takes a radical approach by implementing the entire network protocol stack for remote memory access within the Physical layer (PHY) of the Ethernet. This overcomes fundamental latency and bandwidth overheads imposed by the MAC layer, especially for small memory messages. Second, EDM implements a centralized, fast, in-network scheduler for memory traffic within the PHY of the Ethernet switch. Inspired by the classic Parallel Iterative Matching (PIM) algorithm, the scheduler dynamically reserves bandwidth between compute and memory nodes by creating virtual circuits in the PHY, thus eliminating queuing delay and layer 2 packet processing delay at the switch for memory traffic, while maintaining high bandwidth utilization. Our FPGA testbed demonstrates that EDM's network fabric incurs a latency of only $\sim$300 ns for remote memory access in an unloaded network, which is an order of magnitude lower than state-of-the-art Ethernet-based solutions such as RoCEv2 and comparable to emerging PCIe-based solutions such as CXL. Larger-scale network simulations indicate that even at high network loads, EDM's average latency remains within 1.3$\times$ its unloaded latency.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan
Mohammad Alizadeh, Albert Greenberg, David A. Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan. Data Center TCP (DCTCP) . SIGCOMM, 2010
2010
-
[2]
Less Is More: Trading a Little Band- width for Ultra-Low Latency in the Data Center
Mohammad Alizadeh, Abdul Kabbani, Tom Edsall, Balaji Prabhakar, Amin Vahdat, and Masato Yasuda. Less Is More: Trading a Little Band- width for Ultra-Low Latency in the Data Center . NSDI, 2012
2012
-
[3]
pFabric: Minimal Near-optimal Datacenter Transport
Mohammad Alizadeh, Shuang Yang, Milad Sharif, Sachin Katti, Nick McKeown, Balaji Prabhakar, and Scott Shenker. pFabric: Minimal Near-optimal Datacenter Transport. SIGCOMM, 2013
2013
-
[4]
Can far memory improve job throughput? EuroSys, 2020
Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout, Marcos K Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker. Can far memory improve job throughput? EuroSys, 2020
2020
-
[5]
Optimal Oblivious Reconfig- urable Networks
Daniel Amir, Tegan Wilson, Vishal Shrivastav, Hakim Weatherspoon, Robert Kleinberg, and Rachit Agarwal. Optimal Oblivious Reconfig- urable Networks. STOC, 2022
2022
-
[6]
High Speed Switch Scheduling for Local Area Networks
Thomas Anderson, Susan Owicki, James Saxe, and Charles Thacker. High Speed Switch Scheduling for Local Area Networks . TOCS, 1993
1993
-
[7]
Sirius: A Flat Datacenter Network with Nanosecond Optical Switching
Hitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe, Istvan Haller, Krzysztof Jozwik, Fotini Karinou, Sophie Lange, Kai Shi, Benn Thomsen, and Hugh Williams. Sirius: A Flat Datacenter Network with Nanosecond Optical Switching. SIGCOMM, 2020
2020
-
[8]
Disaggregating Stateful Net- work Functions
Deepak Bansal, Gerald DeGrace, Rishabh Tewari, Michal Zygmunt, James Grantham, Silvano Gai, Mario Baldi, Krishna Doddapaneni, Arun Selvarajan, Arunkumar Arumugam, Balakrishnan Raman, Avijit Gupta, Sachin Jain, Deven Jagasia, Evan Langlais, Pranjal Srivastava, Rishiraj Hazarika, Neeraj Motwani, Soumya Tiwari, Stewart Grant, Ranveer Chandra, and Srikanth Ka...
2023
Show all 95 references
-
[9]
Miller, Vishal Shrivastav, Pankaj Mehra, Matthew Boisvert, Avi Silberschatz, and Peter Alvaro
Daniel Bittman, Robert Soulé, Ethan L. Miller, Vishal Shrivastav, Pankaj Mehra, Matthew Boisvert, Avi Silberschatz, and Peter Alvaro. Don’t Let RPCs Constrain Your API . HotNets, 2021
2021
-
[10]
SIGCOMM, 2013
Pat Bosshart, Glen Gibb, Hun-Seok Kim, George Varghese, Nick McKe- own, Martin Izzard, Fernando Mujica, and Mark Horowitz.Forwarding Metamorphosis: Fast Programmable Match-Action Processing in Hard- ware for SDN. SIGCOMM, 2013
2013
-
[11]
Online Algorithms for Maximum Cardinality Matching with Edge Arrivals
Niv Buchbinder, Danny Segev, and Yevgeny Tkach. Online Algorithms for Maximum Cardinality Matching with Edge Arrivals . ESA, 2017
2017
-
[12]
dcPIM: Near-optimal Proactive Datacenter Transport
Qizhe Cai, Mina Tahmasbi Arashloo, and Rachit Agarwal. dcPIM: Near-optimal Proactive Datacenter Transport. SIGCOMM, 2022
2022
-
[13]
Rethinking Software Runtimes for Disaggregated Memory
Irina Calciu, M Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli. Rethinking Software Runtimes for Disaggregated Memory. ASPLOS, 2021
2021
-
[14]
Project Pberry: FPGA Acceleration for Remote Memory
Irina Calciu, Ivan Puddu, Aasheesh Kolli, Andreas Nowatzyk, Jayneel Gandhi, Onur Mutlu, and Pratap Subrahmanyam. Project Pberry: FPGA Acceleration for Remote Memory . HotOS, 2019
2019
-
[15]
Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of Memory
Xinyi Chen, Liangcheng Yu, Vincent Liu, and Qizhen Zhang. Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of Memory . SIGCOMM, 2023
2023
-
[16]
Credit-Scheduled Delay- Bounded Congestion Control for Datacenters
Inho Cho, Keon Jang, and Dongsu Han. Credit-Scheduled Delay- Bounded Congestion Control for Datacenters . SIGCOMM, 2017
2017
-
[17]
dRMT: Disaggre- gated Programmable Switching
Sharad Chole, Andy Fingerhut, Sha Ma, Anirudh Sivaraman, Shay Var- gaftik, Alon Berger, Gal Mendelson, Mohammad Alizadeh, Shang-Tse Chuang, Isaac Keslassy, Ariel Orda, and Tom Edsall. dRMT: Disaggre- gated Programmable Switching. SIGCOMM, 2017
2017
-
[18]
Benchmarking cloud serving systems with YCSB
Brian F Cooper, Adam Silberstein, Erwin Tam, Raghu Ramakrishnan, and Russell Sears. Benchmarking cloud serving systems with YCSB . 2010
2010
-
[19]
NSDI, 2014
Aleksandar Dragojević, Dushyanth Narayanan, Miguel Castro, and Orion Hodson.{FaRM}: Fast Remote Memory . NSDI, 2014
2014
-
[20]
Helios: a hybrid electrical/optical switch architecture for modular data centers
Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, and Amin Vahdat. Helios: a hybrid electrical/optical switch architecture for modular data centers . SIGCOMM, 2010
2010
-
[21]
Maltz, and Albert Greenberg
Daniel Firestone, Andrew Putnam, Sambhrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian Caulfield, Eric Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitu P...
2018
-
[22]
Network Requirements for Resource Disaggregation
Peter X Gao, Akshay Narayan, Sagar Karandikar, Joao Carreira, Sangjin Han, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker. Network Requirements for Resource Disaggregation . OSDI, 2016
2016
-
[23]
Gao, Akshay Narayan, Gautam Kumar, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker
Peter X. Gao, Akshay Narayan, Gautam Kumar, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker. PHost: Distributed near-Optimal Data- center Transport over Commodity Network Fabric . CoNEXT, 2015
2015
-
[24]
Dan Gibson, Hema Hariharan, Eric Lance, Moray McLaren, Behnam Montazeri, Arjun Singh, Stephen Wang, Hassan M. G. Wassel, Zhehua Wu, Sunghwan Yoo, Raghuraman Balasubramanian, Prashant Chandra, Michael Cutforth, Peter Cuy, David Decotigny, Rakesh Gautam, Alex Iriza, Milo M. K. M...
2022
-
[25]
Memory Pooling With CXL
Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Sangwon Lee, and Myoungsoo Jung. Memory Pooling With CXL . MICRO, 2023
2023
-
[26]
Direct Access, High-Performance Memory Disaggregation with {DirectCXL}
Donghyun Gouk, Sangwon Lee, Miryeong Kwon, and Myoungsoo Jung. Direct Access, High-Performance Memory Disaggregation with {DirectCXL}. USENIX ATC 22, 2022
2022
-
[27]
Efficient Memory Disaggregation with Infiniswap
Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, and Kang G Shin. Efficient Memory Disaggregation with Infiniswap . NSDI, 2017
2017
-
[28]
Re-architecting datacenter networks and stacks for low latency and high performance
Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew Moore, Gianni Antichi, and Marcin Wojcik. Re-architecting datacenter networks and stacks for low latency and high performance . SIGCOMM, 2017
2017
-
[29]
On the Off-chip Memory Latency of Real-Time Systems: Is DDR DRAM Really the Best Option? https://arxiv.org/pdf/ 1810.07059.pdf, 2018
Mohamed Hassan. On the Off-chip Memory Latency of Real-Time Systems: Is DDR DRAM Really the Best Option? https://arxiv.org/pdf/ 1810.07059.pdf, 2018
2018 arXiv
-
[30]
The nanoPU: A Nanosecond Network Stack for Datacenters
Stephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen, Muham- mad Shahbaz, Changhoon Kim, and Nick McKeown. The nanoPU: A Nanosecond Network Stack for Datacenters . OSDI, 2021
2021
-
[31]
Congestion avoidance and control
Van Jacobson. Congestion avoidance and control . SIGCOMM, 1988
1988
-
[32]
FireSim: FPGA- Accelerated Cycle-Exact Scale-Out System Simulation in the Public Cloud
Sagar Karandikar, Howard Mao, Donggyu Kim, David Biancolin, Alon Amid, Dayeol Lee, Nathan Pemberton, Emmanuel Amaro, Colin Schmidt, Aditya Chopra, Qijing Huang, Kyle Kovacs, Borivoje Nikolic, Randy Katz, Jonathan Bachrach, and Krste Asanovic. FireSim: FPGA- Accelerated Cycle-E...
2018
-
[33]
The New Intel ® Xeon® Processor Scalable Family (Formerly Skylake-SP)
Akhilesh Kumar. The New Intel ® Xeon® Processor Scalable Family (Formerly Skylake-SP). HotChips, 2017
2017
-
[34]
The part-time parliament
Leslie Lamport. The part-time parliament . ACM Transactions on Computer Systems, 1998
1998
-
[35]
Yanfang Le, Radhika Niranjan Mysore, Lalith Suresh, Gerd Zellweger, Sujata Banerjee, Aditya Akella, and Michael M. Swift. PL2: Towards Predictable Low Latency in Rack-Scale Networks . https://arxiv.org/abs/ 2101.06537, 2021
2021 arXiv
-
[36]
Globally Synchronized Time via Datacenter Networks
Ki Suh Lee, Han Wang, Vishal Shrivastav, and Hakim Weatherspoon. Globally Synchronized Time via Datacenter Networks . SIGCOMM, 2016
2016
-
[37]
PHY Covert Chan- nels: Can you see the Idles? NSDI, 2014
Ki Suh Lee, Han Wang, and Hakim Weatherspoon. PHY Covert Chan- nels: Can you see the Idles? NSDI, 2014
2014
-
[38]
Mind: In-Network Memory Man- agement for Disaggregated Data Centers
Seung-seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong, and Abhishek Bhattacharjee. Mind: In-Network Memory Man- agement for Disaggregated Data Centers . SOSP, 2021
2021
-
[39]
Seer: Enabling Future-A ware Online Caching in Networked Systems
Jason Lei and Vishal Shrivastav. Seer: Enabling Future-A ware Online Caching in Networked Systems . NSDI, 2024. EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation ASPLOS ’25, March 30–April 3, 2025, Rotterdam, Netherlands
2024
-
[40]
A Case Against CXL Memory Pooling
Philip Levis, Kun Lin, and Amy Tai. A Case Against CXL Memory Pooling. HotNets, 2023
2023
-
[41]
Pond: CXL-based Memory Pooling Systems for Cloud Platforms
Huaicheng Li, Daniel S Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, et al. Pond: CXL-based Memory Pooling Systems for Cloud Platforms. ASPLOS, 2023
2023
-
[42]
HPCC: High Precision Congestion Control
Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang, Zheng Cao, Ming Zhang, Frank Kelly, Mohammad Al- izadeh, and Minlan Yu. HPCC: High Precision Congestion Control . SIG- COMM, 2019
2019
-
[43]
Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness, Alex C
William M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness, Alex C. Snoeren, and George Porter.Expanding across time to deliver bandwidth efficiency and low latency . NSDI, 2020
2020
-
[44]
Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C
William M. Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C. Snoeren, and George Porter. RotorNet: A Scal- able, Low-complexity, Optical Datacenter Network. SIGCOMM, 2017
2017
-
[45]
SilkRoad: Making Stateful Layer-4 Load Balancing Fast and Cheap Using Switching ASICs
Rui Miao, Hongyi Zeng, Changhoon Kim, Jeongkeun Lee, and Minlan Yu. SilkRoad: Making Stateful Layer-4 Load Balancing Fast and Cheap Using Switching ASICs. SIGCOMM, 2017
2017
-
[46]
TIMELY: RTT-based Congestion Control for the Datacenter
Radhika Mittal, Terry Lam, Nandita Dukkipati, Emily Blem, Hassan Wassel, Monia Ghobadi, Amin Vahdat, Yaogong Wang, David Wether- all, and David Zats. TIMELY: RTT-based Congestion Control for the Datacenter. SIGCOMM, 2015
2015
-
[47]
Homa: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities
Behnam Montazeri, Yilong Li, Mohammad Alizadeh, and John Ouster- hout. Homa: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities. SIGCOMM, 2018
2018
-
[48]
Rolf Neugebauer, Gianni Antichi, José Fernando Zazo, Yury Audzevich, Sergio López-Buedo, and Andrew W. Moore. Understanding PCIe performance for end host networking . SIGCOMM, 2018
2018
-
[50]
Scale-out NUMA
Stanko Novakovic, Alexandros Daglis, Edouard Bugnion, Babak Falsafi, and Boris Grot. Scale-out NUMA. ASPLOS, 2014
2014
-
[51]
Zero-queue
Jonathan Perry, Amy Ousterhout, Hari Balakrishnan, Devavrat Shah, and Hans Fugal. Fastpass: A Centralized "Zero-queue" Datacenter Net- work. SIGCOMM, 2014
2014
-
[52]
Caulfield, Eric S
Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fow- ers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim, Sitaram Lanka, James Larus, Eric Pet...
2014
-
[53]
OSDI, 2020
Zhenyuan Ruan, Malte Schwarzkopf, Marcos K Aguilera, and Adam Belay.{AIFM}:High-Performance, Application-Integrated Far Memory. OSDI, 2020
2020
-
[54]
Implementing Fault-tolerant Services using the State Machine Approach: A Tutorial
Fred Schneider. Implementing Fault-tolerant Services using the State Machine Approach: A Tutorial. ACM Computing Surveys, 1990
1990
-
[55]
Schuh, Arvind Krishnamurthy, David Culler, Henry M
Henry N. Schuh, Arvind Krishnamurthy, David Culler, Henry M. Levy, Luigi Rizzo, Samira Khan, and Brent E. Stephens. CC-NIC: a Cache- Coherent Interface to the NIC . ASPLOS, 2024
2024
-
[56]
LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation
Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation . OSDI, 2018
2018
-
[57]
Fast, Scalable, and Programmable Packet Scheduler in Hardware
Vishal Shrivastav. Fast, Scalable, and Programmable Packet Scheduler in Hardware. SIGCOMM, 2019
2019
-
[58]
Programmable Multi-Dimensional Table Filters for Line Rate Network Functions
Vishal Shrivastav. Programmable Multi-Dimensional Table Filters for Line Rate Network Functions . SIGCOMM, 2022
2022
-
[59]
Stateful Multi-Pipelined Programmable Switches
Vishal Shrivastav. Stateful Multi-Pipelined Programmable Switches . SIGCOMM, 2022
2022
-
[60]
Globally Synchronized Time via Datacenter Networks
Vishal Shrivastav, Ki Suh Lee, Han Wang, and Hakim Weatherspoon. Globally Synchronized Time via Datacenter Networks . Transactions on Networking, 2019
2019
-
[61]
Shoal: A Network Architecture for Disaggregated Racks
Vishal Shrivastav, Asaf Valadarsky, Hitesh Ballani, Paolo Costa, Ki Suh Lee, Han Wang, Rachit Agarwal, and Hakim Weatherspoon. Shoal: A Network Architecture for Disaggregated Racks . NSDI, 2019
2019
-
[62]
StRoM: Smart Remote Memory
David Sidler, Zeke Wang, Monica Chiosa, Amit Kulkarni, and Gustavo Alonso. StRoM: Smart Remote Memory . EuroSys, 2020
2020
-
[63]
Programmable Packet Scheduling at Line Rate
Anirudh Sivaraman, Suvinay Subramanian, Mohammad Alizadeh, Sharad Chole, Shang-Tse Chuang, Anurag Agrawal, Hari Balakrishnan, Tom Edsall, Sachin Katti, and Nick McKeown. Programmable Packet Scheduling at Line Rate . SIGCOMM, 2016
2016
-
[64]
IEEE Standard for Ethernet
10.1109/IEEESTD.2022.9844436. IEEE Standard for Ethernet . IEEE Std 802.3-2022 (Revision of IEEE Std 802.3-2018), 2022
2022
-
[65]
Berkeley Big Data Bench- mark
https://amplab.cs.berkeley.edu/benchmark/. Berkeley Big Data Bench- mark. AMP Lab, UC Berkeley, 2014
2014
-
[66]
Compare-and-swap
https://en.wikipedia.org/wiki/Compare-and-swap. Compare-and-swap. Wikipedia
-
[67]
Infiniband
https://en.wikipedia.org/wiki/InfiniBand. Infiniband. Wikipedia
-
[68]
Priority Encoder
https://en.wikipedia.org/wiki/Priority_encoder. Priority Encoder. Wikipedia
-
[69]
RDMA over converged Ethernet
https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet. RDMA over converged Ethernet . Wikipedia
-
[70]
Shortest Re- maining Time First
https://en.wikipedia.org/wiki/Shortest_remaining_time. Shortest Re- maining Time First. Wikipedia
-
[71]
The Machine
https://en.wikipedia.org/wiki/The_Machine_(computer_ architecture). The Machine. Wikipedia
-
[72]
Corundum
https://github.com/corundum/corundum. Corundum. GitHub
-
[73]
NVIDIA NVLink and NVSwitch Technical Overview
https://images.nvidia.com/content/pdf/nvswitch-technical- overview.pdf. NVIDIA NVLink and NVSwitch Technical Overview . NVIDIA Corporation
-
[74]
Broadcom Delivers Industry’s First 51.2-Tbps Co-Packaged Optics Ethernet Switch Platform for Scalable AI Systems
https://investors.broadcom.com/news-releases/news-release- details/broadcom-delivers-industrys-first-512-tbps-co-packaged- optics. Broadcom Delivers Industry’s First 51.2-Tbps Co-Packaged Optics Ethernet Switch Platform for Scalable AI Systems . Broadcom
-
[75]
OpenCAPI Specifica- tions
https://opencapi.org/technical/specifications/. OpenCAPI Specifica- tions. OpenCAPI Consortium
-
[76]
The New World of 400 Gbps Ethernet
https://www.accton.com/Technology-Brief/the-new-world-of-400- gbps-ethernet/. The New World of 400 Gbps Ethernet . Accton
-
[77]
Tofino Switch
https://www.barefootnetworks.com. Tofino Switch. Intel
-
[78]
CCIX Base Specification 1.0
https://www.ccixconsortium.com/library/specification/. CCIX Base Specification 1.0. CCIX Consortium Inc
-
[79]
CXL 3.0 Specification
https://www.computeexpresslink.org/download-the-specification . CXL 3.0 Specification. Compute Express Link Consortium Inc
-
[80]
Intel Rack Scale Design: Just what is it? Intel
https://www.datacenterdynamics.com/en/opinions/intel-rack-scale- design-just-what-is-it/ . Intel Rack Scale Design: Just what is it? Intel
-
[81]
Stratix 10 FPGA
https://www.intel.com/content/dam/www/programmable/us/en/ pdfs/literature/hb/stratix-10/s10-overview .pdf. Stratix 10 FPGA. Intel
-
[82]
Stratix V FPGA
https://www.intel.com/content/dam/www/programmable/us/en/ pdfs/literature/hb/stratix-v/stx5_51001.pdf. Stratix V FPGA. Intel
-
[83]
Utra Path Interconnect
https://www.intel.com/content/www/us/en/products/details/fpga/ agilex.html. Utra Path Interconnect. Intel
-
[84]
CXL Is Dead In The AI Era
https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era . CXL Is Dead In The AI Era . SemiAnalysis
-
[85]
DC Ultra RTL Synthesis
https://www.synopsys.com/implementation-and-signoff/rtl- synthesis-test/dc-ultra .html. DC Ultra RTL Synthesis . Synopsys
-
[86]
UCIe 1.0 Specification
https://www.uciexpress.org/specification. UCIe 1.0 Specification. Uni- versal Chiplet Interconnect Express
-
[87]
XC50256 CXL2.0/PCle5.0 switch
https://www.xconn-tech.com/product. XC50256 CXL2.0/PCle5.0 switch. XconnTech
-
[88]
Alveo U200 Data Center Accelerator Card
https://www.xilinx.com/products/boards-and-kits/alveo/u200 .html. Alveo U200 Data Center Accelerator Card . AMD Xilinx
-
[89]
Priority-based Flow Control
http://www.ieee802.org/1/pages/802.1bb.html. Priority-based Flow Control. IEEE DCB. 802.1Qbb, 2011
2011
-
[90]
Semeru: A Memory-Disaggregated Managed Runtime
Chenxi Wang, Haoran Ma, Shi Liu, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen, Michael D Bond, Ravi Netravali, Miryung Kim, and Guo- qing Harry Xu. Semeru: A Memory-Disaggregated Managed Runtime . OSDI, 2020. ASPLOS ’25, March 30–April 3, 2025, Rotterdam, Netherlands Weigao Su and V...
2020
-
[91]
Timing is Everything: Accurate, Minimum Overhead, A vailable Bandwidth Estimation in High-Speed Wired Networks
Han Wang, Ki Suh Lee, Erluo Li, Chiun Lin Lim, Ao Tang, and Hakim Weatherspoon. Timing is Everything: Accurate, Minimum Overhead, A vailable Bandwidth Estimation in High-Speed Wired Networks. IMC, 2014
2014
-
[92]
Aurelia: CXL Fabric with Tentacle
Shu-Ting Wang and Weitao Wang. Aurelia: CXL Fabric with Tentacle. WORDS, 2023
2023
-
[93]
Is Tail-Optimal Scheduling Possible? Operations Research, INFORMS, 2012
Adam Wierman and Bert Zwart. Is Tail-Optimal Scheduling Possible? Operations Research, INFORMS, 2012
2012
-
[94]
Redy: Remote Dynamic Memory Cache
Qizhen Zhang, Philip A Bernstein, Daniel S Berger, and Badrish Chan- dramouli. Redy: Remote Dynamic Memory Cache . https://arxiv.org/ abs/2112.12946, 2021
2021 arXiv
-
[95]
Congestion Control for Large-Scale RDMA Deployments
Yibo Zhu, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mo- hamad Haj Yahia, and Ming Zhang. Congestion Control for Large-Scale RDMA Deployments. SIGCOMM, 2015
2015
-
[96]
Understanding and Mitigating Packet Corruption in Data Center Networks
Danyang Zhuo, Monia Ghobadi, Ratul Mahajan, Klaus-Tycho Förster, Arvind Krishnamurthy, and Thomas Anderson. Understanding and Mitigating Packet Corruption in Data Center Networks . SIGCOMM, 2017
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.