REVIEW 2 cited by
UDON: A case for offloading to general purpose compute on CXL memory
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Upcoming CXL-based disaggregated memory devices feature special purpose units to offload compute to near-memory. In this paper, we explore opportunities for offloading compute to general purpose cores on CXL memory devices, thereby enabling a greater utility and diversity of offload. We study two classes of popular memory intensive applications: ML inference and vector database as candidates for computational offload. The study uses Arm AArch64-based dual-socket NUMA systems to emulate CXL type-2 devices. Our study shows promising results. With our ML inference model partitioning strategy for compute offload, we can place up to 90% data in remote memory with just 20% performance trade-off. Offloading Hierarchical Navigable Small World (HNSW) kernels in vector databases can provide upto 6.87$\times$ performance improvement with under 10% offload overhead.
Forward citations
Cited by 2 Pith papers
-
From Block to Byte: Transforming PCIe SSDs with CXL Memory Protocol and Instruction Annotation
A CXL-attached SSD using Determinism and Bufferability annotations is simulated to be 10.9x faster than PCIe memory expansion and to approach DRAM-like latency under high locality.
-
Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
A structured survey of CXL-based computing system research, organized by memory expansion, unified memory, and distributed memory pooling and sharing.
Discussion (0). Continue with ORCID to comment.