Pith. sign in

REVIEW 2 cited by

UDON: A case for offloading to general purpose compute on CXL memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02868 v1 pith:FLPOA22D submitted 2024-04-03 cs.ET

classification cs.ET
keywords memoryoffloadcomputedevicesoffloadingpurposegeneralinference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Upcoming CXL-based disaggregated memory devices feature special purpose units to offload compute to near-memory. In this paper, we explore opportunities for offloading compute to general purpose cores on CXL memory devices, thereby enabling a greater utility and diversity of offload. We study two classes of popular memory intensive applications: ML inference and vector database as candidates for computational offload. The study uses Arm AArch64-based dual-socket NUMA systems to emulate CXL type-2 devices. Our study shows promising results. With our ML inference model partitioning strategy for compute offload, we can place up to 90% data in remote memory with just 20% performance trade-off. Offloading Hierarchical Navigable Small World (HNSW) kernels in vector databases can provide upto 6.87$\times$ performance improvement with under 10% offload overhead.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Block to Byte: Transforming PCIe SSDs with CXL Memory Protocol and Instruction Annotation

    cs.AR 2025-06 conditional novelty 6.0 of 10

    A CXL-attached SSD using Determinism and Bufferability annotations is simulated to be 10.9x faster than PCIe memory expansion and to approach DRAM-like latency under high locality.

  2. Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey

    cs.DC 2024-12 conditional novelty 4.0 of 10

    A structured survey of CXL-based computing system research, organized by memory expansion, unified memory, and distributed memory pooling and sharing.

Pith tools