Pith. sign in

Paper Citation Record · LEDGER

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference

As of 9 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2502.07578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07578 v3

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:19:43.247215Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 121 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07bac523-f739-4068-87a5-c1bcda05eaac · outbound

This paper cites URL: https://azure.microsoft.com/en-us/ pricing/calculator/.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference URL: https://azure.microsoft.com/en-us/ pricing/calculator/

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.861786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.861786Z digest=sha256:c5bd7a8f73c8cfbdaa5894dc91ca9858e0268194add59074eaeb9dba997536f9

Observation f50db636-43f6-481c-b096-45b2a2d8b8c8 · outbound

This paper cites Enabling cxl memory expansion for in-memory database management systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Enabling cxl memory expansion for in-memory database management systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.866044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.866044Z digest=sha256:19306bc8e53d47fae54c711d0c3096a29a22d01acd0dbeac2299121dda0ea169

Observation 5089bf48-f3a9-4e3d-91cb-1453f5c22089 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.869539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.869539Z digest=sha256:8b88a3f792ff73dc1fe5bdd7400c1a08a41b992d772ebcdd1eab3215b3f37cc5

Observation a3d5d758-2180-4b31-b598-d9b0e46d418d · outbound

This paper cites Introducing the next generation of claude.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Introducing the next generation of claude

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.873549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.873549Z digest=sha256:ef8e58a846dc0eae91e5b299ed9ed642aea957fe9289161f3fcaebf38984fe66

Observation ea5c3df5-c2e5-413d-ab17-fb19b56d3d23 · outbound

This paper cites Exploiting CXL-based memory for distributed deep learn- ing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Exploiting CXL-based memory for distributed deep learn- ing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.877646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.877646Z digest=sha256:09d82777bfa2454e2271e556382b90347d28b1f593fb3577994c190a976e474b

Observation 5a0cd7bb-5661-46a5-b1b2-fb361d8d7820 · outbound

This paper cites Llama 2 70b: An mlperf inference benchmark for large language models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Llama 2 70b: An mlperf inference benchmark for large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.881241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.881241Z digest=sha256:a42a9db1465c0c4cebf94226ad1d3a5647b64d365f09a4fb95a5983d58850500

Observation a3965f40-faf1-48a8-8170-6704fd7edf22 · outbound

This paper cites Improving image generation with better captions.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Improving image generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.884793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.884793Z digest=sha256:17d32d515477fc2a2d5280e2d1304bf2f633cb0984625e2674d530c1f8788a46

Observation 24294e30-8d96-4e91-bc5c-60cdcf8f0080 · outbound

This paper cites 144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.888201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.888201Z digest=sha256:a5b180d5b1191197a2a9f81e36ba93f5c1b2ca9ec425c3acbb9dbd31c06d314c

Observation 7d94f444-8173-4e1f-9ad4-c70dc261920b · outbound

This paper cites The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.891765Z digest=sha256:e85e8e01e8aec4aebe51fd2132c8264cc0825140cbadd0d8ed4cfcf89dba36a9

Observation 5bbabc2f-30e6-4979-aaf5-7321ed902eef · outbound

This paper cites Patterson, and Krste Asanović.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Patterson, and Krste Asanović

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.896429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.896429Z digest=sha256:b96e4e58ae9ee49ef7c243b22f5735065a613a48fc9f8f13d63c7f2c5f7de284

Observation bbaeb890-7147-48e2-9b19-32d4dddaccfe · outbound

This paper cites Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.899955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.899955Z digest=sha256:d625983a90ec271a0c97ba0f733d91911979ac9539f9d8cb4fa4c0d759ec14f9

Observation 5a3e52fd-7d4b-4141-b10d-809ffeee38ff · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.903454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.903454Z digest=sha256:c1f3eef1aae14155895c983e8642498270c0fa966c8828922e46b66e10916b1b

Observation 2469b9de-ad81-44cc-9436-8113f685dad5 · outbound

This paper cites The true processing in memory accelerator.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The true processing in memory accelerator

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.907410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.907410Z digest=sha256:ed2435682633b3b4350e421aeea90fa9d860d3e38d217dc4e71ccacf0597c6dc

Observation 004a9ef1-3dcb-4734-b9bd-d136be30c219 · outbound

This paper cites To pim or not for emerging general purpose processing in ddr memory systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference To pim or not for emerging general purpose processing in ddr memory systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.910828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.910828Z digest=sha256:24389e7103405a27a311df7997eacaab7ac712d30d964cf7431d871ed776f420

Observation 338a6725-26d1-48f7-b93b-c2cf52df2949 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.917902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.917902Z digest=sha256:62b89367d4ffce22c8a058d05a20f2e0a97cc974ef6fd7ba5b5cf913045ab691

Observation cd1e20ff-666e-4342-a52f-208a3b487efe · outbound

This paper cites Dram spot price.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dram spot price

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.921715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.921715Z digest=sha256:bbff5bc5c6768fbcb6135b5be6320d68334e7a523206b84bb369b94b10390e88

Observation 49a53810-bc51-4a1c-9f6b-e73e391f1e6e · outbound

This paper cites The Llama 3 Herd of Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.926165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.926165Z digest=sha256:644697aeed57bc751935df77930f61e305ead775838fb56a41afa38571add13f

Observation 67955ca7-fbde-4b9f-8eb0-223e80525b45 · outbound

This paper cites The inference cost of search disruption – large language model cost analysis.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The inference cost of search disruption – large language model cost analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.930080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.930080Z digest=sha256:f1914dbf77f467157adb5abffd666129b7cdef5526add083b930c622de2e2868

Observation 12ba10d4-ae69-4261-a872-0cc372b2d438 · outbound

This paper cites Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.934251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.934251Z digest=sha256:2cac75ff15183dcfb8216405748abc8c81df0efa591e6a1bad4abc5cfc8620b2

Observation c3f77fda-f121-4c79-a0de-2993503f6b9d · outbound

This paper cites Pci interface ic.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Pci interface ic

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.938373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.938373Z digest=sha256:ada380d64ee0d5bd48cecb51b87206b957511d4c100b57b35655d9e10d7f8f9b

Observation 2f0894fd-ced2-40d0-91ef-bbda10643519 · outbound

This paper cites Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.942471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.942471Z digest=sha256:38dd77d879f573982b5c49b05da90a9ae1fcdec7b4351ec0f82b3487e929121a

Observation fde56906-0e9f-4ce9-a19b-78ad37b29a88 · outbound

This paper cites Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.946555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.946555Z digest=sha256:650d2a0833a84dad58e381eb69c776de7b0aa888ccacbf67b7213b844f4416a7

Observation 40162986-cd5d-489c-9cb5-40d60ace121d · outbound

This paper cites NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.950787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.950787Z digest=sha256:00dd51b747300ba127155da349ce4b071cb80d076121aa27d02e2c8c75e4a846

Observation 5c1f2059-3400-449b-b31b-ff9222b4fa4d · outbound

This paper cites Reinhardt, Adrian M.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Reinhardt, Adrian M

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.954152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.954152Z digest=sha256:4fbf872a6237825985933bfe82508f7cdce1164822dfbe67de5c8fa8fd688574

Observation 6456e43a-4f47-40ce-ba0a-2873b1e4bee8 · outbound

This paper cites Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.961275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.961275Z digest=sha256:a733bc8f5b8d967e8d62f17d62d5bc49fde75dd6cdbf6a84dd91635ddc12513e

Observation 77b8ec34-0861-4634-9898-0ae60995e6eb · outbound

This paper cites Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.964664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.964664Z digest=sha256:11a1a1d90730338a24982de3c06dcbeac497655ba58bd627ef1a69a75fd7aa5f

Observation 931f5695-aaf7-4f1c-9352-0a50f9f63405 · outbound

This paper cites Evaluating machine learningworkloads on memory-centric comput- ing systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Evaluating machine learningworkloads on memory-centric comput- ing systems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.968712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.968712Z digest=sha256:4b75d71350123bd01afc947b6b4759e3b502131603d1cfc0c5eb76f40adf392b

Observation c6cc619e-5bb2-46e2-a359-a533ca66c343 · outbound

This paper cites Our next-generation model: Gemini 1.5.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Our next-generation model: Gemini 1.5

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.972030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.972030Z digest=sha256:a9e3d9304687d5301d399a39080824daf9a6ec802b6b5e73f9352cf3a852da31

Observation 834a3a4f-a801-4228-9c53-1d452bad5f9f · outbound

This paper cites Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.975332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.975332Z digest=sha256:1408ef20ed9eb16f7c3263a4b949b646cea18cf072d3f98430423d076056b324

Observation 49574b5b-b973-451a-8eed-66df7d31407a · outbound

This paper cites Direct access, High-Performance memory disaggregation with DirectCXL.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Direct access, High-Performance memory disaggregation with DirectCXL

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.978839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.978839Z digest=sha256:519756fe74516b71001f579c59adfd48e9794eab8e5ce30a3dae68b81871013c

Observation f5295c49-69b5-4e67-8994-0e33a8f01a05 · outbound

This paper cites OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.982214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.982214Z digest=sha256:e729e43d683da10e33b91ffc909953799d7b016d59e3fb853dbc13eabf45a3ab

Observation 61eaf1de-bd24-4108-903c-c4c1a42c4044 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.985722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.985722Z digest=sha256:e28df85ad5a276a718d0ff68cc63dd60348d5dfae4f91d6178d32fa71c412bf9

Observation 9b69936e-63a4-47c0-955d-64219e5c4490 · outbound

This paper cites ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.989518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.989518Z digest=sha256:db7ae30893fe123e5ecb7b4a07c4ae9a820e8f11c486857ca2db0d4693f81cbe

Observation 5e296e87-2f2c-4d9a-b59c-ac1b3e695651 · outbound

This paper cites Deep residual learning for image recognition.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Deep residual learning for image recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.992996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.992996Z digest=sha256:613dbc5f78228eb5ceedf00dce971545606bb5e53fdee95dbf864e435580a408

Observation 55134ec4-639a-426d-8520-975c003e047c · outbound

This paper cites Gaussian Error Linear Units (GELUs).

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Gaussian Error Linear Units (GELUs)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.996285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.996285Z digest=sha256:1c78beca8799c523a427331de9efad13ac91b192212d020b7245dfc8a6bcb3ed

Observation cb395ee5-4e1b-416a-ae1e-4824d597c81c · outbound

This paper cites Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.999831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.999831Z digest=sha256:8d8c19bb10cd13aa4ed2d9b2ee4539c35c67b06e3705b982ed390205b8e303ee

Observation 63c96630-b1d1-448b-b1fc-ed12aa84ec4d · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Le, Yonghui Wu, and Zhifeng Chen

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.007401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.007401Z digest=sha256:e343fa3a0983bf19b0939d35cb9c2e2ae4551e359dbc89380dcf399eb38abcc6

Observation 6ffc5613-87bf-401f-8af7-621eb91e81f2 · outbound

This paper cites BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.011573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.011573Z digest=sha256:1e22b0dac320eb90df2d9a1a5a4384fe8bdf40934a82cdecfe33ba6156b1ff05

Observation ba3a06f7-092e-4df1-a37c-510a3da0fa64 · outbound

This paper cites Floatpim: In-memory acceleration of deep neural network training with high precision.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Floatpim: In-memory acceleration of deep neural network training with high precision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.015085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.015085Z digest=sha256:d1be9d7d424231d690b77d5d85d0cba994aab8b17f1c6b7551f4b50235710985

Observation 9ba074ee-9d04-4044-a038-a29e28cb3de0 · outbound

This paper cites Ultra-efficient processing in-memory for data intensive applications.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Ultra-efficient processing in-memory for data intensive applications

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.018239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.018239Z digest=sha256:c5fb812d80006d26ad4d8215ae483c933fcc550875a7f2247c4ebd7da1c2d65a

Observation fed0dce8-a0f9-46db-bfbe-c59958960241 · outbound

This paper cites Intel xeon gold 6430 processor, 60m cache, 2.10 ghz.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Intel xeon gold 6430 processor, 60m cache, 2.10 ghz

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.021491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.021491Z digest=sha256:724fdc393b13d6456c5c00f77b2345159879abfd22c1449d184558c069150f6c

Observation bfa1399a-05e0-4d73-a5a6-e7ba13adfdf0 · outbound

This paper cites Intel ® xeon® gold 6430 processor.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Intel ® xeon® gold 6430 processor

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.025589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.025589Z digest=sha256:f87684736db28c23807a3de5c90ca270ce8b92440764364287b25dbf0268c758

Observation 3d974faa-ce32-4c60-9c7f-19365e6705cd · outbound

This paper cites CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.029008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.029008Z digest=sha256:b814ae1b9be2fb1b1b3633640e9132334fcec35c5406813df4819621002a089c

Observation bf1c202c-5700-43bd-b906-d8d3953b23e6 · outbound

This paper cites Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.032314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.032314Z digest=sha256:d988fbfac08f7ff449edabce0c7bc747370368b2adef0e129df23d09f8e2484e

Observation f3183397-8d8e-4962-b7fd-85ba7ed408bb · outbound

This paper cites Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.035525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.035525Z digest=sha256:0c47530a23580cc1cd92e6ea42885179b19804d6e761a63ad1c3b56143d3cbee

Observation 551d1872-b606-4835-927f-e7255e0853e1 · outbound

This paper cites ChatGPT for good? On opportunities and challenges of large language models for education.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ChatGPT for good? On opportunities and challenges of large language models for education

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.039355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.039355Z digest=sha256:3116d8f2d1ca25ff5fa4958eaf9c5da7eb08e0df853668f27cd0d09a2dddd86a

Observation fc0b1791-a46a-4a6c-bef5-518b86cd102e · outbound

This paper cites Rec- nmp: Accelerating personalized recommendation with near-memory processing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Rec- nmp: Accelerating personalized recommendation with near-memory processing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.043715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.043715Z digest=sha256:2b3aa63ce0421048f4b21734f1ab9b2cb998cdb6485a9703001f6900853c1c01

Observation 40984173-3fef-493d-8435-905a989496c9 · outbound

This paper cites Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.047402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.047402Z digest=sha256:b614c723cafb0def559523e61190d8edb5b0c2b0d58273ffe89198706adc02e3

Observation 21964e30-ba5f-42c4-93fe-95844210fa51 · outbound

This paper cites Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.051215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.051215Z digest=sha256:a3b511a8224411c0eb436214b13a9fc26cb96e7898008aee8d572cad86bf4b9c

Observation 815bf70f-471f-4ee4-9d71-5072dcb57c42 · outbound

This paper cites Samsung PIM/PNM for Transfmer Based AI: Energy Efficiency on PIM/PNM Cluster.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Samsung PIM/PNM for Transfmer Based AI: Energy Efficiency on PIM/PNM Cluster

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.424550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.055256Z digest=sha256:49f006397c5d6f3572a9fa0e430dcb89291460ed9bd465ccad5f9add38122813

Observation 24dbc041-3b61-4467-884a-3dfc5ade2e11 · outbound

This paper cites A 1ynm 1.25v 8gb 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep learning application.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A 1ynm 1.25v 8gb 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep learning application

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.059497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.059497Z digest=sha256:6ffdd8303e48118a2202efa352bd7bcd0f7f0c67ce77205fd8c60cf95b2be300

Observation b6cef06e-7a03-471a-b491-2f78bf7b2609 · outbound

This paper cites Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.411377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.063790Z digest=sha256:159b130631bc0980fff23d3cd26251d3fe36317d5d00ba24c3bed2d86633e9a2

Observation 6588b9e2-114f-418f-a03a-f1a17dc44c75 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Efficient memory management for large language model serving with pagedattention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.067968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.067968Z digest=sha256:3b2ea82689d1d85817d07f1c234f4041d82fe5b31815546054ae9c1f481e5367

Observation 0be80827-8f78-46d0-aab6-14544a8d6650 · outbound

This paper cites Memory-centric computing with sk hynix’s domain-specific memory.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Memory-centric computing with sk hynix’s domain-specific memory

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.072246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.072246Z digest=sha256:bdbc88aad31e2e4243444ae36d29b178d98f55a0124bd9bfac3376fb29d3b8dc

Observation c7b25c34-da8e-45b6-a33f-722a70341f7f · outbound

This paper cites System architecture and software stack for GDDR6-AiM.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference System architecture and software stack for GDDR6-AiM

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.391031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.076737Z digest=sha256:b0fd9c1b1c21ec22254e8c3fb994575ff65396a00f6ed721e40a1b6b75f8f080

Observation 7b24437a-9736-4834-b18c-7bb0c0f435ba · outbound

This paper cites 25.4 a 20nm 6gb function-in-memory DRAM, based on HBM2 with a 1.2 tflops pro- grammable computing unit using bank-level parallelism, for machine learning applications.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 25.4 a 20nm 6gb function-in-memory DRAM, based on HBM2 with a 1.2 tflops pro- grammable computing unit using bank-level parallelism, for machine learning applications

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.380426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.080673Z digest=sha256:06df70bc2b7de627c28a563a68a174fb8336e9a996be32256a7900d4033c7ee2

Observation 66bdc8f1-1e91-4704-bee7-d804f371433f · outbound

This paper cites Improving in-memory database operations with acceleration DIMM (AxDIMM).

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Improving in-memory database operations with acceleration DIMM (AxDIMM)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.369207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.084956Z digest=sha256:e1f5e46e8664eba517dcc7456cba3e91551ceb97b0feda37d0174fbb536eb646

Observation bc0b94e1-27db-470f-a9be-167adc88e9e6 · outbound

This paper cites Using machine learning to increase yield and lower packaging costs.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Using machine learning to increase yield and lower packaging costs

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.357587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.088957Z digest=sha256:ffc85c0e4152b4edb2d2659c4860fb208c98f6dae2985002b8860a12e4564e64

Observation 5d1065ab-8c94-415c-9c2c-93facde77733 · outbound

This paper cites A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.347503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.093213Z digest=sha256:8276647ac09f09648258d19eef8bcbb760deacf7756508a6636b568e036c6349

Observation 06c69c93-84af-4188-b7d6-b4ebf002e139 · outbound

This paper cites Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott Lee, Ishwar Agarwal, Mark D.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott Lee, Ishwar Agarwal, Mark D

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.336520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.097580Z digest=sha256:c7f6181d4646e535900cb3bf3364c89e7549f653de2efb5c9257d75e7ba510e0

Observation 4cdfb794-01a7-46ab-a526-cb5ae6f9a145 · outbound

This paper cites Accelerating distributed reinforcement learning with in-switch computing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Accelerating distributed reinforcement learning with in-switch computing

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.326047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.101725Z digest=sha256:2fd0c151333bd3f71ecf745eef4badbd6651b00ba933b69bd50184d72a8e4687

Observation 8ed1efdb-7d91-4d21-8099-e68bb427a939 · outbound

This paper cites Specification.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Specification

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.315839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.105664Z digest=sha256:8011b98cc0d5b5f525cf54dbcda14290daa143abf9678532ec50d3a86e085dbc

Observation 7f709682-02b2-49d6-829d-af86e7374b5a · outbound

This paper cites DeepSeek-V3 Technical Report.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference DeepSeek-V3 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.109950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.109950Z digest=sha256:874e547f764a8965c69d3287cc5731071a7bd6c9b8403dc8e211da74d337ddfa

Observation 510e802c-84e8-47b2-af67-c7f69a80ffe1 · outbound

This paper cites Enmc: Extreme near-memory classification via approximate screening.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Enmc: Extreme near-memory classification via approximate screening

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.305602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.113850Z digest=sha256:0364a969e84dc12eff39db9de635bb9f9f71b4f28f4e21d5c40c275c356d4bfa

Observation 9ccfcc69-ef8b-44b3-a38d-05032c7029c6 · outbound

This paper cites Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.294964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.118085Z digest=sha256:ae4a19222f32cd5f3c7b385afc96295c2a0832f33bf7bc552f1372ad1944fe13

Observation 0ae30fee-06ff-434c-8ca7-5fb36673ebfc · outbound

This paper cites Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.122136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.122136Z digest=sha256:bf8128acea3290b839f4d1fce466e8295d706ad152744665b1ac0bc5d3cf112e

Observation bac7644a-2243-4259-a206-b048edc8780a · outbound

This paper cites A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.283985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.126343Z digest=sha256:0a37409eee7ce87cc06bc5b855a3d86dbaa4b237c652f55209285df7c05d6a9f

Observation 91daf12f-993f-4c47-a980-e9d882c50483 · outbound

This paper cites Dram power calculator.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dram power calculator

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.273545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.129668Z digest=sha256:50f64667358dd220dce7a39b8cb7db8e16b725d0f70041e766ee0592354e37d4

Observation b421ff6f-4c59-477b-b23b-21d96fadea73 · outbound

This paper cites Nvidia shipped 3.76m data center gpus in 2023.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia shipped 3.76m data center gpus in 2023

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.263356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.133028Z digest=sha256:abdc5ad94706d379289a02eb2deb0d87a17de4d8401cd0164e971cb6c3c54ae1

Observation ab8cb036-44dc-4216-915c-da2638c00132 · outbound

This paper cites Supply chain aware computer architecture.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Supply chain aware computer architecture

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.253293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.136219Z digest=sha256:5c1ac5f9e8ed91f0bf1e9439a2c905f6bbd401c531b80f4d695e5142bb5bd58b

Observation 7d85537c-4844-4d37-9329-0f38af56f576 · outbound

This paper cites 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for recommendation system.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for recommendation system

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.243312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.139776Z digest=sha256:7779c9f49795a65ccf3117c7eee64ccedb930b52b03b262a7a2c0bda1ccb0645

Observation ff55d8ee-85ff-4168-86b3-cdfc2e145b2e · outbound

This paper cites Introduction to the nvidia dgx a100 system.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Introduction to the nvidia dgx a100 system

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.232936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.143328Z digest=sha256:1a9bdc6954f82ca0b3c389f3ceb11283e240edee823d54db50b96faea8e1e247

Observation 87ffbc70-e3a9-4f62-a1c7-08606df975ee · outbound

This paper cites Nvidia a100 tensor core gpu.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia a100 tensor core gpu

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.222631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.146694Z digest=sha256:1e866831a01e8c0bdb80a54783add6712cddddf2ef44ad64192ce341f24f3a4f

Observation b9db1fba-f9ff-41f1-839c-6a6dfc0ed566 · outbound

This paper cites Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.212585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.150077Z digest=sha256:453f8ee5b35eaf6b3dd918229f29cb7b00699078b0243042fdf74cdabe3bf37f

Observation 4f6571af-9e27-4194-9641-a9036d6ecbba · outbound

This paper cites Nvlink and nvlink switch.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvlink and nvlink switch

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.202214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.153989Z digest=sha256:a6ce20092c6e6ec2759e543d27c3ea21df2b9f254a6936a422034163210fbc63

Observation 3cb0cb57-9174-4a9d-8d7c-0057f438580d · outbound

This paper cites an unresolved cited work.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:19:44.191897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.157526Z digest=sha256:52ad180ba85497e3e0864359cf55f5b001e734c54f3c0355cf7b5399f8a52287

Observation 48f7f33b-48ef-4d04-9e6f-567b0376ba51 · outbound

This paper cites Fine- grained dram: Energy-efficient dram for extreme bandwidth systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Fine- grained dram: Energy-efficient dram for extreme bandwidth systems

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.181932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.160924Z digest=sha256:ed2ce476be3902d05e8152c6956cbc2483223249e3a5cc619790147d0ac8e64f

Observation 001a95ce-8a7c-4781-a9a6-d1e1842e31df · outbound

This paper cites Bureau of Labor Statistics.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Bureau of Labor Statistics

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.171164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.164318Z digest=sha256:5d453b3a5bb59b6d29049a215b189d2c77720fd2d6c0bf92e5527b8cc5ab5b47

Observation 93dc2701-7455-4c73-a6e6-33f54ddaa771 · outbound

This paper cites Accelerating neural network inference with processing-in-dram: From the edge to the cloud.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Accelerating neural network inference with processing-in-dram: From the edge to the cloud

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.161590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.168553Z digest=sha256:844172476cb9575f940529d8e4ff5e0089f230ecd8897acb011e8e35f9faddf1

Observation 53c05752-51aa-45d7-a70a-58eac2fb0983 · outbound

This paper cites Gpt-4 turbo and gpt-4.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Gpt-4 turbo and gpt-4

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.151325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.172611Z digest=sha256:6166810d0b9c1ea3defb9a6ca06bbec04b96266d692ff59bfb7065b5e635efcd

Observation 89e01c35-16d1-4ced-bc9e-a20aeb8f2d32 · outbound

This paper cites Learning to reason with llms.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Learning to reason with llms

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.140848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.176883Z digest=sha256:b3255d913202b8136aaa973d1355ec4c200878494580531b5425a57a6a21f069

Observation 72d9eac6-99da-4454-9711-51cc0a1679c7 · outbound

This paper cites Video generation models as world simulators.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Video generation models as world simulators

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.130360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.180173Z digest=sha256:0e39e59cea80c7b45d0b9ab48b861ca58aa3e6354cd9b7670f4a30ed6a9b149a

Observation 05f5e22d-b5ed-44e4-91e8-f04aa921dd0d · outbound

This paper cites GPT-4 Technical Report.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GPT-4 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.183524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.183524Z digest=sha256:a93b12d703dbc34d67c0b800a6678df20eee3ddfda53872126f6fc9526cb41e5

Observation b83c9968-c212-40fa-9f43-80d845986aac · outbound

This paper cites Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.116094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.186916Z digest=sha256:dd53e06555ff9a67159f002e6304ede786f48a34464f74b27af1839d8c32ec83

Observation 1d76f3a5-652f-412f-a3fa-fa8c10785fb6 · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer-based generative model inference.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Attacc! unleashing the power of pim for batched transformer-based generative model inference

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.104257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.190248Z digest=sha256:5dad34caf11b1fb709e4c67dd88356328e3022f9cf0efa18c398ec582a2c1831

Observation 49de6f65-d730-407d-8c97-2d2b9afe42ae · outbound

This paper cites Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.092592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.197770Z digest=sha256:a6b8a5b8051c5777a3670c2bfcf75983ec01279f7d2825ab0f8ec4fa0c78cba2

Observation a1d117e2-4a53-43a7-8f1d-f5aca6135e4e · outbound

This paper cites An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.081069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.201368Z digest=sha256:68d19275ceed4574ebf34b1c4811f8e60ea6937e80ea34a1b3a473bdab7b151d

Observation c2b6d119-1ac8-4b0c-8f60-3eb64c6beb38 · outbound

This paper cites an unresolved cited work.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.193719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.193719Z digest=sha256:59bc1beab45b93210087bb0bafa0e8446709fad966ecea5315701263911534e3

Observation 5b2912e3-6b3b-4e5b-ba80-6981b4c294a7 · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Splitwise: Efficient gen- erative llm inference using phase splitting

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.059017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.208695Z digest=sha256:a013aea2df9b6b02e3e121790d7ef0d9e8c6fca7b094df89787fe8ec5b88e3b4

Observation b30cd212-45ff-4832-aeff-ec2c75b2361f · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Movie Gen: A Cast of Media Foundation Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.212984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.212984Z digest=sha256:3cc2f25fcb8ddc533c765487fa450fd63165222fb5b7a979bcb0b325f8d3787d

Observation cb0aafde-3aa6-426a-ac18-5890ceeaef73 · outbound

This paper cites Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.070811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.204959Z digest=sha256:95630e7b91f8d8b3881dd7a5e432c521be1d661104e920109a0f31a2dcad7f37

Observation 50869f67-984a-48bc-a6f2-53098521ce11 · outbound

This paper cites FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.046764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.221118Z digest=sha256:77f651b70d26210279380247ec317d8ff56655268b78141f54640a98f21f84ce

Observation 7770edb4-b003-40ce-8c5b-e8e21091736f · outbound

This paper cites Dota: detect and omit weak attentions for scalable transformer acceleration.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dota: detect and omit weak attentions for scalable transformer acceleration

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.036444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.225224Z digest=sha256:107e97f8d51396d9d5d51d89296d093d27da0775095b3432bfaf1568d61f9c25

Observation 31adadca-0d46-4a13-9f5d-f7a7364b8f78 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.216975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.216975Z digest=sha256:e5d222bbfc28b072e3eb09aa754ab6dc1f8ea3a2f110f2835490fb24a6e42801

Observation a8ad555e-ed1a-4fa2-9fc5-6b722b428f21 · outbound

This paper cites Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.025911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.233212Z digest=sha256:ccca99789fbc116bb6c5164df95a89097b0967a57c2c9f60c2a53aa15c36c00b

Observation 52e7404c-5666-46c4-bf2d-99a4b3108973 · outbound

This paper cites 8gb gddr6 sgram c-die.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 8gb gddr6 sgram c-die

Reference 97

Resolution
verified exact
raw_fallback, observed 2026-08-08T12:19:43.535628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.236583Z digest=sha256:bf9016de43313c6f5a51d3aa751773cabaeb5e47975001876e4735502e226e8b

Observation 347e29e6-146d-4f10-8c3c-f9e18db03e1d · outbound

This paper cites Searching for Activation Functions.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Searching for Activation Functions

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.229683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.229683Z digest=sha256:89439c96191b22774a6f70b53c54f9874a8dcdc7026d77f427b26e40225e3fdf

Observation 36b826bd-1753-4dde-8de6-b8fb57a89e45 · outbound

This paper cites GLU Variants Improve Transformer.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GLU Variants Improve Transformer

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.243256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.243256Z digest=sha256:af7c60a893de1ca2ea4c696357fba0d2cbf42190051cd968b37df9b803d8f7f9

Observation 1106256b-df8e-4e41-a921-c63c4186a17a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.247215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.247215Z digest=sha256:a374f60bc919370b9374f54c45ef8830d0b8b9226f87b5211e8e003c298efa4c

Observation 7992e68b-9baf-43c4-87c3-91a2b39d4c87 · outbound

This paper cites Generative ai winds in memory semiconduc- tors, total demand for server drams is declining.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Generative ai winds in memory semiconduc- tors, total demand for server drams is declining

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.015080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:43.239755Z digest=sha256:de95280a74f25083beeeb2cb0886f3a6762479985c0496c32af8e44ea475378f

Pith citing papers

No inbound Pith citation observations are available.