Pith. sign in

REVIEW 3 major objections 3 minor 95 references

EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation

T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read EDM claims that moving the remote-memory protocol stack into the Ethernet PHY and adding a centralized in-switch scheduler cuts remote read and write latency to roughly 300 ns, beating RoCEv2 and TCP/IP by an order of magnitude and…

desk verdict Solid FPGA demonstration of a sub-300ns in-PHY memory fabric, but the high-load latency claims rest on a simulator with inconsistent parameters. read the letter →

arxiv 2411.08300 v4 pith:HWICA5OG submitted 2024-11-13 cs.OS cs.NI

classification cs.OScs.NI
keywords memorydisaggregationEthernetPHYremotelatencyin-networkschedulerparalleliterativematchingvirtualcircuitsFPGAprototypesmall-messagetransport
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

EDM sets out to make Ethernet a practical fabric for memory disaggregation by attacking the two biggest sources of remote-access latency: the host protocol stack and the switch. Its first move is radical: instead of layering RDMA or TCP/IP on top of the Ethernet MAC, EDM puts the entire remote-memory protocol inside the Physical Coding Sublayer (PCS), where data moves in 66-bit blocks and idle gaps can carry memory traffic. Its second move is to place a centralized scheduler inside the switch PHY that grants bandwidth in the form of short virtual circuits, so memory messages never queue or wait for layer-2 processing. On an FPGA testbed, the paper reports fabric latency of roughly 300 ns for remote reads and writes, which is 3.7 to 12.7 times lower than raw Ethernet, RoCEv2, and hardware TCP/IP, and comparable to CXL. Simulations at 144 nodes keep average latency within 1.3 times the unloaded value even at high load, which matters because it suggests remote memory over commodity Ethernet could escape the congestion collapse that plagues small-message traffic.

What carries the argument

The load-bearing machinery is the Ethernet Physical Coding Sublayer (PCS), the part of the PHY that encodes data into 66-bit blocks, repurposed as the data path for memory messages via new /M*/, /N/, and /G/ block types. On the switch, the key mechanism is a centralized scheduler that implements priority-based Parallel Iterative Matching (PIM), the classic switch-scheduling algorithm that matches inputs to outputs iteratively, extended so each iteration takes a constant three clock cycles through ordered-list hardware and a priority encoder. The scheduler's grants create virtual circuits in the PHY between source and destination ports, which is what eliminates queuing, drops, and layer-2 forwarding at the switch for memory traffic.

What would settle it

Build or simulate a 512-port EDM switch at the claimed 3 GHz clock and measure whether a maximal matching is actually formed in about 27 cycles; if the matching latency exceeds the transmission time of the 128-byte or 256-byte chunk, the zero-queuing promise fails on the first incast, and average latency under 80 percent all-to-all load will exceed the claimed 1.3 times unloaded.

Watch

Extended reading notes

Core claim

The central claim is that the Ethernet MAC layer is not an unavoidable price of using Ethernet for remote memory access. By operating directly in the Physical Coding Sublayer, EDM carries memory messages as custom 66-bit PHY block types, reuses inter-frame gap bits, and achieves intra-frame preemption at 66-bit granularity, eliminating the 64-byte minimum frame, the IFG overhead, and head-of-line blocking caused by large non-memory frames. The companion claim is that a centralized scheduler implemented in the switch PHY can form a priority-based online maximal matching fast enough to reserve bandwidth in advance, so memory traffic is forwarded through virtual circuits with zero queuing and zero layer-2 processing delay. On the testbed, remote reads and writes complete through the fabric in about 300 ns, and in simulation the average completion time stays within 1.2 to 1.4 times the ideal for several disaggregated application workloads, outperforming reactive transports and CXL-style credit flow control under load.

Load-bearing premise

The scheduler's zero-queuing and high-utilization guarantees depend on it forming a maximal matching in about 3 times log(N) clock cycles at a clock rate fast enough that the chosen chunk size keeps every link busy; if computing the matching takes longer than sending a chunk, links go idle and queues build.

Editorial extensions

If this is right

  • If EDM's numbers hold, remote memory reads and writes over Ethernet can run at latencies previously associated with PCIe-attached CXL, so memory disaggregation does not require a separate fabric.
  • Memory traffic can share links with IP and storage traffic without sacrificing small-message latency, because 66-bit-granularity preemption bounds interference from large frames.
  • Removing the TCP/IP or RDMA transport stack from the memory data path eliminates per-packet encapsulation and congestion-control delays, simplifying the host NIC for memory traffic.
  • With a scheduler guaranteeing zero queuing, applications see predictable memory-access latency even under incast-like many-to-one loads.
  • The ASIC synthesis suggests the scheduler fits within a modern switch chip: roughly 10 mm2 of area and about 1 MB of SRAM for a 512-port switch, which is a concrete scaling path beyond the two-port FPGA prototype.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper does not explore is whether the same PHY-level multiplexing can carry other latency-critical traffic, such as coherence or synchronization primitives, alongside memory and IP traffic; only memory messages and grants are defined.
  • The paper fixes a 256-byte chunk and a 2.56 ns clock in the testbed; at higher line rates the zero-queuing guarantee would require either faster matching or larger chunks, and the latency cost of larger chunks for tiny 8-byte read requests is not quantified.
  • Because one-sided writes pay a notification-and-grant round trip before sending data, a workload dominated by small one-sided writes would stress whether that RTT/2 overhead stays negligible in practice.
  • A consequence the authors leave implicit is that a working EDM at scale would undercut the main argument for deploying CXL or InfiniBand purely for memory pooling, since Ethernet already exists in every rack and would also carry the disaggregated memory traffic.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper proposes EDM, an Ethernet fabric for memory disaggregation that moves the remote-memory protocol stack into the Physical Coding Sublayer of the Ethernet PHY and adds a centralized, in-PHY scheduler at the switch. The scheduler builds a priority-based online maximal matching and creates PHY-level virtual circuits, with the goal of eliminating queuing, layer-2 forwarding, and transport-stack latency for memory traffic. The paper reports an unloaded remote read/write fabric latency of roughly 300 ns on an FPGA testbed, claims this is 3.7--12.7x lower than raw Ethernet, RoCEv2, and hardware TCP/IP, and uses a C network simulator to claim that under high load the average latency stays within 1.3x the unloaded value and that average message completion times are within 1.2--1.4x of ideal for several disaggregated workloads.

Significance. The unloaded-latency result is a substantial empirical contribution: it is measured on an FPGA testbed, backed by a detailed cycle-level breakdown in Table 1 and Figure 5, and accompanied by a public Verilog artifact. The idea of implementing remote memory access in the PHY, with custom 66-bit block types and intra-frame preemption, is novel and potentially impactful for memory disaggregation. The centralized in-PHY scheduler is also a well-motivated design point. The high-load claim, however, is the main reason the paper's central latency story generalizes beyond the unloaded case, and that claim currently rests on a simulator parameterization that is internally inconsistent with the scheduler timing model. The high-load result therefore needs to be re-established with a consistent set of parameters before the paper's conclusions are fully supported.

major comments (3)
  1. [§4.3, with §3.1.2 and §3.1.3] The simulator in §4.3 uses a 256 B chunk at 100 Gbps with N=144 nodes, but it never states the scheduler clock. If the simulator inherits the 2.56 ns testbed clock used in Figure 5, then forming a maximal matching takes 3·log2(144) ≈ 21.5 cycles, or about 55 ns, while a 256 B chunk on a 100 Gbps link transmits in 20.5 ns. This violates the line-rate condition derived in §3.1.3, which requires the chunk transmission time to be at least the matching latency to keep links busy. Under the 2.56 ns clock the per-port throughput would be capped near 256 B / 55 ns ≈ 4.7 Gbps, not 100 Gbps, so the 1.3x high-load latency figure is not supported by the stated parameters. If instead the scheduler runs at the 3 GHz ASIC clock of §4.1, the 256 B chunk is viable, but then the simulator does not match the testbed timing used elsewhere in §4.3. Please state the scheduler clock explicitly and rerun with a consistent chunk/clock combination, e.g., a 688 B chunk at 2.56 ns, or 256 B with the 3 GHz ASIC scheduler.
  2. [Figure 8 caption / inserted note] The manuscript itself admits a second simulator defect: read flows are initialized with 8 B RREQs, but the load accounting uses the 64 B RRES size, so the reported network load overstates the real injected traffic. The statement that this factor 'can be offsetted by real workloads' is not quantified and is not a substitute for correct load accounting. Because Figure 8a and the application MCT results in Figure 8b are the evidence for the high-load latency claim, the load definition must be corrected and the simulations repeated before the 1.3x claim can be accepted.
  3. [§4.2.1] The text states that 'testbed experiments showed that even under interference from IP traffic, EDM maintained a near-constant ~300 ns remote memory access latency,' but no such experiment, figure, or measurement procedure is presented in §4.2. Either the data should be reported, or the claim should be removed or explicitly labeled as a qualitative observation.
minor comments (3)
  1. [Figure 8] The inserted note about RREQ/RRES size accounting is written in informal, ungrammatical prose ('offsetted') and should be moved into the methodology section as a clearly stated assumption or limitation.
  2. [§4.2.2] The paper refers to a 'cycle-accurate FPGA hardware simulator' for the YCSB evaluation; please clarify whether this simulator is the same implementation as the testbed, and how its timing parameters were derived.
  3. [References] Reference [49] (Scale-out NUMA) appears to be duplicated as [50]; the duplicate should be removed and subsequent citations renumbered.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: unloaded latency is measured, high-load behavior is simulated from the design, and self-citations are non-load-bearing building blocks.

full rationale

The paper's central latency result (~300 ns) is a direct FPGA measurement from an implemented testbed (Table 1, Figure 5), not a fitted or self-referential prediction. The high-load result (within 1.3x unloaded) comes from a C network simulator that implements the scheduler's grant algorithm; the zero-queuing property is a design invariant of maximal matching, not an empirical claim derived from itself. Empirical parameter choices (X=3, chunk size 256B) are simulator settings, not parameters fitted to the reported target metric. Citations to prior PHY idle-block work [36,37,39,60,91] and to constant-time ordered-list hardware structures [57-59,63] include the authors' own papers, but the EDM implementation and ASIC synthesis provide independent support; these citations are building blocks, not unverified premises that force the conclusion. The Figure 8 note admitting that read flows use 8B RREQs while load is accounted at 64B RRES, and the possible mismatch between scheduler matching latency and chunk transmission time, are correctness/consistency risks, not circular derivations. No step in the derivation chain equates an output to an input by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 2 invented entities

The paper's performance claims rest on three fitted or assumed quantities: the chunk size, the maximum active notifications per source-destination pair, and the scheduler clock rate. The PHY-injection mechanism assumes prior PHY-idle manipulation results, and the reliability model assumes corruption is permanent. The invented block types and virtual circuits are internal to the design and have no external evidence beyond the prototype.

free parameters (3)
  • chunk_size_c = 256 B in simulations; 128 B for line-rate 512-port estimate
    Chunk size is a scheduler parameter that must exceed the matching latency times the link rate; 256 B is used in section 4.3 without demonstrating that the line-rate condition holds at the testbed clock.
  • max_active_notifications_X = 3
    Set to 3 as empirically found best in section 3.1.2; it bounds the per-destination notification queue and serializes outstanding messages per source-destination pair.
  • scheduler_clock_rate = 3 GHz (ASIC estimate)
    Used to claim 9 ns matching latency for a 512-port switch; the FPGA prototype runs at a 2.56 ns clock, so the ASIC clock is an unvalidated estimate.
assumptions (4)
  • domain assumption The Ethernet PCS interface between encoder and scrambler permits insertion of custom 66-bit block types and reuse of idle blocks without breaking clock recovery or scrambling.
    Invoked in section 3.2 to place EDM logic in the PCS; relies on prior PHY-channel work [36,37,60,91].
  • domain assumption Memory traffic demand is known in advance: read sizes are available from RREQ and write sizes from explicit notifications.
    Used in section 3.1.1 to build the demand notification queue; depends on the memory controller interface exposing request sizes.
  • ad hoc to paper Link data corruption is non-transient and can be handled by disabling the link rather than retransmission.
    Stated in section 3.3; EDM provides no end-to-end reliability for memory messages, which is a significant departure from transport expectations.
  • domain assumption A single switch with hundreds of ports is the target topology; multi-hop fabrics are out of scope.
    Stated in section 2.1; the scheduler creates virtual circuits within one switch and fault tolerance assumes a backup top-of-rack switch.
invented entities (2)
  • /M*/, /N/, and /G/ 66-bit PHY block types
    purpose: Carry memory messages, demand notifications, and grants inside the PCS; distinguish memory traffic from standard Ethernet blocks.
    These block types are defined by the paper and demonstrated only in the paper's own FPGA prototype; no external implementation or standard supports them.
  • In-PHY virtual circuits
    purpose: Forward memory data between matched source and destination ports without layer 2 processing.
    A logical construct of the scheduler; its properties are argued from the matching algorithm rather than independently measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation." pith.science (2026). https://pith.science/paper/HWICA5OG

@misc{pith2026241108300,
  author       = {Pith},
  title        = {Pith review of: EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HWICA5OG}},
  note         = {Machine review of arXiv:2411.08300}
}
abstract

Achieving low remote memory access latency remains the primary challenge in realizing memory disaggregation over Ethernet within the datacenters. We present EDM that attempts to overcome this challenge using two key ideas. First, while existing network protocols for remote memory access over the Ethernet, such as TCP/IP and RDMA, are implemented on top of the MAC layer, EDM takes a radical approach by implementing the entire network protocol stack for remote memory access within the Physical layer (PHY) of the Ethernet. This overcomes fundamental latency and bandwidth overheads imposed by the MAC layer, especially for small memory messages. Second, EDM implements a centralized, fast, in-network scheduler for memory traffic within the PHY of the Ethernet switch. Inspired by the classic Parallel Iterative Matching (PIM) algorithm, the scheduler dynamically reserves bandwidth between compute and memory nodes by creating virtual circuits in the PHY, thus eliminating queuing delay and layer 2 packet processing delay at the switch for memory traffic, while maintaining high bandwidth utilization. Our FPGA testbed demonstrates that EDM's network fabric incurs a latency of only $\sim$300 ns for remote memory access in an unloaded network, which is an order of magnitude lower than state-of-the-art Ethernet-based solutions such as RoCEv2 and comparable to emerging PCIe-based solutions such as CXL. Larger-scale network simulations indicate that even at high network loads, EDM's average latency remains within 1.3$\times$ its unloaded latency.

Figures

Figures reproduced from arXiv: 2411.08300 by the authors.

Figure 1
Figure 1. Memory disaggregation over Ethernet. PCIe, which dramatically reduce the interconnect latency to the NIC [73, 75, 78, 79, 86]. The demand for lower la￾tency for cloud services has also prompted tighter integration of the processor and memory with the network controller, which promises to reduce the processor/memory to NIC la￾tency to sub-100 ns. Such designs exist both in academia (e.g., nanoPU [30], soNUMA [49], FA… view at source ↗
Figure 2
Figure 2. Data path for memory traffic in EDM vs. in existing Ethernet fabrics for memory disaggregation. • Limitation 6: Queuing delay and drops at the switch. Most datacenter congestion control protocols are reactive in nature [1–3, 31, 42], i.e., they rely on some form of congestion feedback to handle congestion. Thus they can￾not eliminate queuing, especially for many-to-one (incast) traffic pattern. Memory traffic is ext… view at source ↗
Figure 3
Figure 3. EDM network stack. frame starts with an /S/ control block, followed by several /D/ data blocks, and ends with a /T/ control block. Ethernet also has a special /E/ control block that is typically used to make up the IFG. Each control block has a sync header value of "01" and an 8-bit block type, followed by 56 bits of payload (which could be frame data). The data block has sync header value of "10" and 64 bits of pay… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Testbed setup. 4 Evaluation We evaluated the performance of EDM using an FPGA-based hardware testbed and large-scale network simulations. 4.1 Hardware Prototype We implemented EDM’s traffic scheduler (§3.1) and the host and switch stacks (§3.2) in Verilog by modifying …
Figure 5
Figure 5. Figure 5: Breakdown of latency for EDM’s network fabric for 64 B read and write. TD+PD = transmission+propagation delay. A clock cycle is 2.56 ns. For details on the cycle num￾bers refer to §3.2.1 and §3.2.2. scheduler ensures no congestion and packet drops, thus ob￾viating the …
Figure 7
Figure 7. Figure 7: End-to-end latency for YCSB workload. 4.2.2 Real application performance. We implemented a remote key-value store on EDM’s testbed and used a cycle￾accurate FPGA hardware simulator to evaluate EDM against the baselines using the YCSB workload [18]. Bandwidth utilizatio…
Figure 8
Figure 8. Figure 8: Network simulation. would consequently block or slow down all other ingress ports that have traffic destined to the victim. This is similar to head-of-line blocking in PFC [95]. In addition, we also conducted an experiment that in￾cluded a mixture of RREQ and WREQ in d…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 78 canonical work pages

  1. [1]

    Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan

    Mohammad Alizadeh, Albert Greenberg, David A. Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan. Data Center TCP (DCTCP) . SIGCOMM, 2010

  2. [2]

    Less Is More: Trading a Little Band- width for Ultra-Low Latency in the Data Center

    Mohammad Alizadeh, Abdul Kabbani, Tom Edsall, Balaji Prabhakar, Amin Vahdat, and Masato Yasuda. Less Is More: Trading a Little Band- width for Ultra-Low Latency in the Data Center . NSDI, 2012

  3. [3]

    pFabric: Minimal Near-optimal Datacenter Transport

    Mohammad Alizadeh, Shuang Yang, Milad Sharif, Sachin Katti, Nick McKeown, Balaji Prabhakar, and Scott Shenker. pFabric: Minimal Near-optimal Datacenter Transport. SIGCOMM, 2013

  4. [4]

    Can far memory improve job throughput? EuroSys, 2020

    Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout, Marcos K Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker. Can far memory improve job throughput? EuroSys, 2020

  5. [5]

    Optimal Oblivious Reconfig- urable Networks

    Daniel Amir, Tegan Wilson, Vishal Shrivastav, Hakim Weatherspoon, Robert Kleinberg, and Rachit Agarwal. Optimal Oblivious Reconfig- urable Networks. STOC, 2022

  6. [6]

    High Speed Switch Scheduling for Local Area Networks

    Thomas Anderson, Susan Owicki, James Saxe, and Charles Thacker. High Speed Switch Scheduling for Local Area Networks . TOCS, 1993

  7. [7]

    Sirius: A Flat Datacenter Network with Nanosecond Optical Switching

    Hitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe, Istvan Haller, Krzysztof Jozwik, Fotini Karinou, Sophie Lange, Kai Shi, Benn Thomsen, and Hugh Williams. Sirius: A Flat Datacenter Network with Nanosecond Optical Switching. SIGCOMM, 2020

  8. [8]

    Disaggregating Stateful Net- work Functions

    Deepak Bansal, Gerald DeGrace, Rishabh Tewari, Michal Zygmunt, James Grantham, Silvano Gai, Mario Baldi, Krishna Doddapaneni, Arun Selvarajan, Arunkumar Arumugam, Balakrishnan Raman, Avijit Gupta, Sachin Jain, Deven Jagasia, Evan Langlais, Pranjal Srivastava, Rishiraj Hazarika, Neeraj Motwani, Soumya Tiwari, Stewart Grant, Ranveer Chandra, and Srikanth Ka...

Show all 95 references
  1. [9]

    Miller, Vishal Shrivastav, Pankaj Mehra, Matthew Boisvert, Avi Silberschatz, and Peter Alvaro

    Daniel Bittman, Robert Soulé, Ethan L. Miller, Vishal Shrivastav, Pankaj Mehra, Matthew Boisvert, Avi Silberschatz, and Peter Alvaro. Don’t Let RPCs Constrain Your API . HotNets, 2021

  2. [10]

    SIGCOMM, 2013

    Pat Bosshart, Glen Gibb, Hun-Seok Kim, George Varghese, Nick McKe- own, Martin Izzard, Fernando Mujica, and Mark Horowitz.Forwarding Metamorphosis: Fast Programmable Match-Action Processing in Hard- ware for SDN. SIGCOMM, 2013

  3. [11]

    Online Algorithms for Maximum Cardinality Matching with Edge Arrivals

    Niv Buchbinder, Danny Segev, and Yevgeny Tkach. Online Algorithms for Maximum Cardinality Matching with Edge Arrivals . ESA, 2017

  4. [12]

    dcPIM: Near-optimal Proactive Datacenter Transport

    Qizhe Cai, Mina Tahmasbi Arashloo, and Rachit Agarwal. dcPIM: Near-optimal Proactive Datacenter Transport. SIGCOMM, 2022

  5. [13]

    Rethinking Software Runtimes for Disaggregated Memory

    Irina Calciu, M Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli. Rethinking Software Runtimes for Disaggregated Memory. ASPLOS, 2021

  6. [14]

    Project Pberry: FPGA Acceleration for Remote Memory

    Irina Calciu, Ivan Puddu, Aasheesh Kolli, Andreas Nowatzyk, Jayneel Gandhi, Onur Mutlu, and Pratap Subrahmanyam. Project Pberry: FPGA Acceleration for Remote Memory . HotOS, 2019

  7. [15]

    Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of Memory

    Xinyi Chen, Liangcheng Yu, Vincent Liu, and Qizhen Zhang. Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of Memory . SIGCOMM, 2023

  8. [16]

    Credit-Scheduled Delay- Bounded Congestion Control for Datacenters

    Inho Cho, Keon Jang, and Dongsu Han. Credit-Scheduled Delay- Bounded Congestion Control for Datacenters . SIGCOMM, 2017

  9. [17]

    dRMT: Disaggre- gated Programmable Switching

    Sharad Chole, Andy Fingerhut, Sha Ma, Anirudh Sivaraman, Shay Var- gaftik, Alon Berger, Gal Mendelson, Mohammad Alizadeh, Shang-Tse Chuang, Isaac Keslassy, Ariel Orda, and Tom Edsall. dRMT: Disaggre- gated Programmable Switching. SIGCOMM, 2017

  10. [18]

    Benchmarking cloud serving systems with YCSB

    Brian F Cooper, Adam Silberstein, Erwin Tam, Raghu Ramakrishnan, and Russell Sears. Benchmarking cloud serving systems with YCSB . 2010

  11. [19]

    NSDI, 2014

    Aleksandar Dragojević, Dushyanth Narayanan, Miguel Castro, and Orion Hodson.{FaRM}: Fast Remote Memory . NSDI, 2014

  12. [20]

    Helios: a hybrid electrical/optical switch architecture for modular data centers

    Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, and Amin Vahdat. Helios: a hybrid electrical/optical switch architecture for modular data centers . SIGCOMM, 2010

  13. [21]

    Maltz, and Albert Greenberg

    Daniel Firestone, Andrew Putnam, Sambhrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian Caulfield, Eric Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitu P...

  14. [22]

    Network Requirements for Resource Disaggregation

    Peter X Gao, Akshay Narayan, Sagar Karandikar, Joao Carreira, Sangjin Han, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker. Network Requirements for Resource Disaggregation . OSDI, 2016

  15. [23]

    Gao, Akshay Narayan, Gautam Kumar, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker

    Peter X. Gao, Akshay Narayan, Gautam Kumar, Rachit Agarwal, Sylvia Ratnasamy, and Scott Shenker. PHost: Distributed near-Optimal Data- center Transport over Commodity Network Fabric . CoNEXT, 2015

  16. [24]

    Dan Gibson, Hema Hariharan, Eric Lance, Moray McLaren, Behnam Montazeri, Arjun Singh, Stephen Wang, Hassan M. G. Wassel, Zhehua Wu, Sunghwan Yoo, Raghuraman Balasubramanian, Prashant Chandra, Michael Cutforth, Peter Cuy, David Decotigny, Rakesh Gautam, Alex Iriza, Milo M. K. M...

  17. [25]

    Memory Pooling With CXL

    Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Sangwon Lee, and Myoungsoo Jung. Memory Pooling With CXL . MICRO, 2023

  18. [26]

    Direct Access, High-Performance Memory Disaggregation with {DirectCXL}

    Donghyun Gouk, Sangwon Lee, Miryeong Kwon, and Myoungsoo Jung. Direct Access, High-Performance Memory Disaggregation with {DirectCXL}. USENIX ATC 22, 2022

  19. [27]

    Efficient Memory Disaggregation with Infiniswap

    Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, and Kang G Shin. Efficient Memory Disaggregation with Infiniswap . NSDI, 2017

  20. [28]

    Re-architecting datacenter networks and stacks for low latency and high performance

    Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew Moore, Gianni Antichi, and Marcin Wojcik. Re-architecting datacenter networks and stacks for low latency and high performance . SIGCOMM, 2017

  21. [29]

    On the Off-chip Memory Latency of Real-Time Systems: Is DDR DRAM Really the Best Option? https://arxiv.org/pdf/ 1810.07059.pdf, 2018

    Mohamed Hassan. On the Off-chip Memory Latency of Real-Time Systems: Is DDR DRAM Really the Best Option? https://arxiv.org/pdf/ 1810.07059.pdf, 2018

  22. [30]

    The nanoPU: A Nanosecond Network Stack for Datacenters

    Stephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen, Muham- mad Shahbaz, Changhoon Kim, and Nick McKeown. The nanoPU: A Nanosecond Network Stack for Datacenters . OSDI, 2021

  23. [31]

    Congestion avoidance and control

    Van Jacobson. Congestion avoidance and control . SIGCOMM, 1988

  24. [32]

    FireSim: FPGA- Accelerated Cycle-Exact Scale-Out System Simulation in the Public Cloud

    Sagar Karandikar, Howard Mao, Donggyu Kim, David Biancolin, Alon Amid, Dayeol Lee, Nathan Pemberton, Emmanuel Amaro, Colin Schmidt, Aditya Chopra, Qijing Huang, Kyle Kovacs, Borivoje Nikolic, Randy Katz, Jonathan Bachrach, and Krste Asanovic. FireSim: FPGA- Accelerated Cycle-E...

  25. [33]

    The New Intel ® Xeon® Processor Scalable Family (Formerly Skylake-SP)

    Akhilesh Kumar. The New Intel ® Xeon® Processor Scalable Family (Formerly Skylake-SP). HotChips, 2017

  26. [34]

    The part-time parliament

    Leslie Lamport. The part-time parliament . ACM Transactions on Computer Systems, 1998

  27. [35]

    Yanfang Le, Radhika Niranjan Mysore, Lalith Suresh, Gerd Zellweger, Sujata Banerjee, Aditya Akella, and Michael M. Swift. PL2: Towards Predictable Low Latency in Rack-Scale Networks . https://arxiv.org/abs/ 2101.06537, 2021

  28. [36]

    Globally Synchronized Time via Datacenter Networks

    Ki Suh Lee, Han Wang, Vishal Shrivastav, and Hakim Weatherspoon. Globally Synchronized Time via Datacenter Networks . SIGCOMM, 2016

  29. [37]

    PHY Covert Chan- nels: Can you see the Idles? NSDI, 2014

    Ki Suh Lee, Han Wang, and Hakim Weatherspoon. PHY Covert Chan- nels: Can you see the Idles? NSDI, 2014

  30. [38]

    Mind: In-Network Memory Man- agement for Disaggregated Data Centers

    Seung-seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong, and Abhishek Bhattacharjee. Mind: In-Network Memory Man- agement for Disaggregated Data Centers . SOSP, 2021

  31. [39]

    Seer: Enabling Future-A ware Online Caching in Networked Systems

    Jason Lei and Vishal Shrivastav. Seer: Enabling Future-A ware Online Caching in Networked Systems . NSDI, 2024. EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation ASPLOS ’25, March 30–April 3, 2025, Rotterdam, Netherlands

  32. [40]

    A Case Against CXL Memory Pooling

    Philip Levis, Kun Lin, and Amy Tai. A Case Against CXL Memory Pooling. HotNets, 2023

  33. [41]

    Pond: CXL-based Memory Pooling Systems for Cloud Platforms

    Huaicheng Li, Daniel S Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, et al. Pond: CXL-based Memory Pooling Systems for Cloud Platforms. ASPLOS, 2023

  34. [42]

    HPCC: High Precision Congestion Control

    Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang, Zheng Cao, Ming Zhang, Frank Kelly, Mohammad Al- izadeh, and Minlan Yu. HPCC: High Precision Congestion Control . SIG- COMM, 2019

  35. [43]

    Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness, Alex C

    William M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness, Alex C. Snoeren, and George Porter.Expanding across time to deliver bandwidth efficiency and low latency . NSDI, 2020

  36. [44]

    Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C

    William M. Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C. Snoeren, and George Porter. RotorNet: A Scal- able, Low-complexity, Optical Datacenter Network. SIGCOMM, 2017

  37. [45]

    SilkRoad: Making Stateful Layer-4 Load Balancing Fast and Cheap Using Switching ASICs

    Rui Miao, Hongyi Zeng, Changhoon Kim, Jeongkeun Lee, and Minlan Yu. SilkRoad: Making Stateful Layer-4 Load Balancing Fast and Cheap Using Switching ASICs. SIGCOMM, 2017

  38. [46]

    TIMELY: RTT-based Congestion Control for the Datacenter

    Radhika Mittal, Terry Lam, Nandita Dukkipati, Emily Blem, Hassan Wassel, Monia Ghobadi, Amin Vahdat, Yaogong Wang, David Wether- all, and David Zats. TIMELY: RTT-based Congestion Control for the Datacenter. SIGCOMM, 2015

  39. [47]

    Homa: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities

    Behnam Montazeri, Yilong Li, Mohammad Alizadeh, and John Ouster- hout. Homa: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities. SIGCOMM, 2018

  40. [48]

    Rolf Neugebauer, Gianni Antichi, José Fernando Zazo, Yury Audzevich, Sergio López-Buedo, and Andrew W. Moore. Understanding PCIe performance for end host networking . SIGCOMM, 2018

  41. [50]

    Scale-out NUMA

    Stanko Novakovic, Alexandros Daglis, Edouard Bugnion, Babak Falsafi, and Boris Grot. Scale-out NUMA. ASPLOS, 2014

  42. [51]

    Zero-queue

    Jonathan Perry, Amy Ousterhout, Hari Balakrishnan, Devavrat Shah, and Hans Fugal. Fastpass: A Centralized "Zero-queue" Datacenter Net- work. SIGCOMM, 2014

  43. [52]

    Caulfield, Eric S

    Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fow- ers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim, Sitaram Lanka, James Larus, Eric Pet...

  44. [53]

    OSDI, 2020

    Zhenyuan Ruan, Malte Schwarzkopf, Marcos K Aguilera, and Adam Belay.{AIFM}:High-Performance, Application-Integrated Far Memory. OSDI, 2020

  45. [54]

    Implementing Fault-tolerant Services using the State Machine Approach: A Tutorial

    Fred Schneider. Implementing Fault-tolerant Services using the State Machine Approach: A Tutorial. ACM Computing Surveys, 1990

  46. [55]

    Schuh, Arvind Krishnamurthy, David Culler, Henry M

    Henry N. Schuh, Arvind Krishnamurthy, David Culler, Henry M. Levy, Luigi Rizzo, Samira Khan, and Brent E. Stephens. CC-NIC: a Cache- Coherent Interface to the NIC . ASPLOS, 2024

  47. [56]

    LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation

    Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation . OSDI, 2018

  48. [57]

    Fast, Scalable, and Programmable Packet Scheduler in Hardware

    Vishal Shrivastav. Fast, Scalable, and Programmable Packet Scheduler in Hardware. SIGCOMM, 2019

  49. [58]

    Programmable Multi-Dimensional Table Filters for Line Rate Network Functions

    Vishal Shrivastav. Programmable Multi-Dimensional Table Filters for Line Rate Network Functions . SIGCOMM, 2022

  50. [59]

    Stateful Multi-Pipelined Programmable Switches

    Vishal Shrivastav. Stateful Multi-Pipelined Programmable Switches . SIGCOMM, 2022

  51. [60]

    Globally Synchronized Time via Datacenter Networks

    Vishal Shrivastav, Ki Suh Lee, Han Wang, and Hakim Weatherspoon. Globally Synchronized Time via Datacenter Networks . Transactions on Networking, 2019

  52. [61]

    Shoal: A Network Architecture for Disaggregated Racks

    Vishal Shrivastav, Asaf Valadarsky, Hitesh Ballani, Paolo Costa, Ki Suh Lee, Han Wang, Rachit Agarwal, and Hakim Weatherspoon. Shoal: A Network Architecture for Disaggregated Racks . NSDI, 2019

  53. [62]

    StRoM: Smart Remote Memory

    David Sidler, Zeke Wang, Monica Chiosa, Amit Kulkarni, and Gustavo Alonso. StRoM: Smart Remote Memory . EuroSys, 2020

  54. [63]

    Programmable Packet Scheduling at Line Rate

    Anirudh Sivaraman, Suvinay Subramanian, Mohammad Alizadeh, Sharad Chole, Shang-Tse Chuang, Anurag Agrawal, Hari Balakrishnan, Tom Edsall, Sachin Katti, and Nick McKeown. Programmable Packet Scheduling at Line Rate . SIGCOMM, 2016

  55. [64]

    IEEE Standard for Ethernet

    10.1109/IEEESTD.2022.9844436. IEEE Standard for Ethernet . IEEE Std 802.3-2022 (Revision of IEEE Std 802.3-2018), 2022

  56. [65]

    Berkeley Big Data Bench- mark

    https://amplab.cs.berkeley.edu/benchmark/. Berkeley Big Data Bench- mark. AMP Lab, UC Berkeley, 2014

  57. [66]

    Compare-and-swap

    https://en.wikipedia.org/wiki/Compare-and-swap. Compare-and-swap. Wikipedia

  58. [67]

    Infiniband

    https://en.wikipedia.org/wiki/InfiniBand. Infiniband. Wikipedia

  59. [68]

    Priority Encoder

    https://en.wikipedia.org/wiki/Priority_encoder. Priority Encoder. Wikipedia

  60. [69]

    RDMA over converged Ethernet

    https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet. RDMA over converged Ethernet . Wikipedia

  61. [70]

    Shortest Re- maining Time First

    https://en.wikipedia.org/wiki/Shortest_remaining_time. Shortest Re- maining Time First. Wikipedia

  62. [71]

    The Machine

    https://en.wikipedia.org/wiki/The_Machine_(computer_ architecture). The Machine. Wikipedia

  63. [72]

    Corundum

    https://github.com/corundum/corundum. Corundum. GitHub

  64. [73]

    NVIDIA NVLink and NVSwitch Technical Overview

    https://images.nvidia.com/content/pdf/nvswitch-technical- overview.pdf. NVIDIA NVLink and NVSwitch Technical Overview . NVIDIA Corporation

  65. [74]

    Broadcom Delivers Industry’s First 51.2-Tbps Co-Packaged Optics Ethernet Switch Platform for Scalable AI Systems

    https://investors.broadcom.com/news-releases/news-release- details/broadcom-delivers-industrys-first-512-tbps-co-packaged- optics. Broadcom Delivers Industry’s First 51.2-Tbps Co-Packaged Optics Ethernet Switch Platform for Scalable AI Systems . Broadcom

  66. [75]

    OpenCAPI Specifica- tions

    https://opencapi.org/technical/specifications/. OpenCAPI Specifica- tions. OpenCAPI Consortium

  67. [76]

    The New World of 400 Gbps Ethernet

    https://www.accton.com/Technology-Brief/the-new-world-of-400- gbps-ethernet/. The New World of 400 Gbps Ethernet . Accton

  68. [77]

    Tofino Switch

    https://www.barefootnetworks.com. Tofino Switch. Intel

  69. [78]

    CCIX Base Specification 1.0

    https://www.ccixconsortium.com/library/specification/. CCIX Base Specification 1.0. CCIX Consortium Inc

  70. [79]

    CXL 3.0 Specification

    https://www.computeexpresslink.org/download-the-specification . CXL 3.0 Specification. Compute Express Link Consortium Inc

  71. [80]

    Intel Rack Scale Design: Just what is it? Intel

    https://www.datacenterdynamics.com/en/opinions/intel-rack-scale- design-just-what-is-it/ . Intel Rack Scale Design: Just what is it? Intel

  72. [81]

    Stratix 10 FPGA

    https://www.intel.com/content/dam/www/programmable/us/en/ pdfs/literature/hb/stratix-10/s10-overview .pdf. Stratix 10 FPGA. Intel

  73. [82]

    Stratix V FPGA

    https://www.intel.com/content/dam/www/programmable/us/en/ pdfs/literature/hb/stratix-v/stx5_51001.pdf. Stratix V FPGA. Intel

  74. [83]

    Utra Path Interconnect

    https://www.intel.com/content/www/us/en/products/details/fpga/ agilex.html. Utra Path Interconnect. Intel

  75. [84]

    CXL Is Dead In The AI Era

    https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era . CXL Is Dead In The AI Era . SemiAnalysis

  76. [85]

    DC Ultra RTL Synthesis

    https://www.synopsys.com/implementation-and-signoff/rtl- synthesis-test/dc-ultra .html. DC Ultra RTL Synthesis . Synopsys

  77. [86]

    UCIe 1.0 Specification

    https://www.uciexpress.org/specification. UCIe 1.0 Specification. Uni- versal Chiplet Interconnect Express

  78. [87]

    XC50256 CXL2.0/PCle5.0 switch

    https://www.xconn-tech.com/product. XC50256 CXL2.0/PCle5.0 switch. XconnTech

  79. [88]

    Alveo U200 Data Center Accelerator Card

    https://www.xilinx.com/products/boards-and-kits/alveo/u200 .html. Alveo U200 Data Center Accelerator Card . AMD Xilinx

  80. [89]

    Priority-based Flow Control

    http://www.ieee802.org/1/pages/802.1bb.html. Priority-based Flow Control. IEEE DCB. 802.1Qbb, 2011

  81. [90]

    Semeru: A Memory-Disaggregated Managed Runtime

    Chenxi Wang, Haoran Ma, Shi Liu, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen, Michael D Bond, Ravi Netravali, Miryung Kim, and Guo- qing Harry Xu. Semeru: A Memory-Disaggregated Managed Runtime . OSDI, 2020. ASPLOS ’25, March 30–April 3, 2025, Rotterdam, Netherlands Weigao Su and V...

  82. [91]

    Timing is Everything: Accurate, Minimum Overhead, A vailable Bandwidth Estimation in High-Speed Wired Networks

    Han Wang, Ki Suh Lee, Erluo Li, Chiun Lin Lim, Ao Tang, and Hakim Weatherspoon. Timing is Everything: Accurate, Minimum Overhead, A vailable Bandwidth Estimation in High-Speed Wired Networks. IMC, 2014

  83. [92]

    Aurelia: CXL Fabric with Tentacle

    Shu-Ting Wang and Weitao Wang. Aurelia: CXL Fabric with Tentacle. WORDS, 2023

  84. [93]

    Is Tail-Optimal Scheduling Possible? Operations Research, INFORMS, 2012

    Adam Wierman and Bert Zwart. Is Tail-Optimal Scheduling Possible? Operations Research, INFORMS, 2012

  85. [94]

    Redy: Remote Dynamic Memory Cache

    Qizhen Zhang, Philip A Bernstein, Daniel S Berger, and Badrish Chan- dramouli. Redy: Remote Dynamic Memory Cache . https://arxiv.org/ abs/2112.12946, 2021

  86. [95]

    Congestion Control for Large-Scale RDMA Deployments

    Yibo Zhu, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mo- hamad Haj Yahia, and Ming Zhang. Congestion Control for Large-Scale RDMA Deployments. SIGCOMM, 2015

  87. [96]

    Understanding and Mitigating Packet Corruption in Data Center Networks

    Danyang Zhuo, Monia Ghobadi, Ratul Mahajan, Klaus-Tycho Förster, Arvind Krishnamurthy, and Thomas Anderson. Understanding and Mitigating Packet Corruption in Data Center Networks . SIGCOMM, 2017

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.