Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:19:43.247215Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2502.07578.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:19:43.247215Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 121 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 07bac523-f739-4068-87a5-c1bcda05eaac · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference URL: https://azure.microsoft.com/en-us/ pricing/calculator/
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50db636-43f6-481c-b096-45b2a2d8b8c8 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Enabling cxl memory expansion for in-memory database management systems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5089bf48-f3a9-4e3d-91cb-1453f5c22089 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d5d758-2180-4b31-b598-d9b0e46d418d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Introducing the next generation of claude
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5c3df5-c2e5-413d-ab17-fb19b56d3d23 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Exploiting CXL-based memory for distributed deep learn- ing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0cd7bb-5661-46a5-b1b2-fb361d8d7820 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Llama 2 70b: An mlperf inference benchmark for large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3965f40-faf1-48a8-8170-6704fd7edf22 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Improving image generation with better captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24294e30-8d96-4e91-bc5c-60cdcf8f0080 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d94f444-8173-4e1f-9ad4-c70dc261920b · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbabc2f-30e6-4979-aaf5-7321ed902eef · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Patterson, and Krste Asanović
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbaeb890-7147-48e2-9b19-32d4dddaccfe · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a3e52fd-7d4b-4141-b10d-809ffeee38ff · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2469b9de-ad81-44cc-9436-8113f685dad5 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The true processing in memory accelerator
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 004a9ef1-3dcb-4734-b9bd-d136be30c219 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference To pim or not for emerging general purpose processing in ddr memory systems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338a6725-26d1-48f7-b93b-c2cf52df2949 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd1e20ff-666e-4342-a52f-208a3b487efe · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dram spot price
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a53810-bc51-4a1c-9f6b-e73e391f1e6e · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67955ca7-fbde-4b9f-8eb0-223e80525b45 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The inference cost of search disruption – large language model cost analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12ba10d4-ae69-4261-a872-0cc372b2d438 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f77fda-f121-4c79-a0de-2993503f6b9d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Pci interface ic
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0894fd-ced2-40d0-91ef-bbda10643519 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde56906-0e9f-4ce9-a19b-78ad37b29a88 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40162986-cd5d-489c-9cb5-40d60ace121d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1f2059-3400-449b-b31b-ff9222b4fa4d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Reinhardt, Adrian M
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6456e43a-4f47-40ce-ba0a-2873b1e4bee8 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b8ec34-0861-4634-9898-0ae60995e6eb · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931f5695-aaf7-4f1c-9352-0a50f9f63405 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Evaluating machine learningworkloads on memory-centric comput- ing systems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6cc619e-5bb2-46e2-a359-a533ca66c343 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Our next-generation model: Gemini 1.5
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834a3a4f-a801-4228-9c53-1d452bad5f9f · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49574b5b-b973-451a-8eed-66df7d31407a · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Direct access, High-Performance memory disaggregation with DirectCXL
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5295c49-69b5-4e67-8994-0e33a8f01a05 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61eaf1de-bd24-4108-903c-c4c1a42c4044 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b69936e-63a4-47c0-955d-64219e5c4490 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e296e87-2f2c-4d9a-b59c-ac1b3e695651 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Deep residual learning for image recognition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55134ec4-639a-426d-8520-975c003e047c · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Gaussian Error Linear Units (GELUs)
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb395ee5-4e1b-416a-ae1e-4824d597c81c · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c96630-b1d1-448b-b1fc-ed12aa84ec4d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Le, Yonghui Wu, and Zhifeng Chen
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ffc5613-87bf-401f-8af7-621eb91e81f2 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3a06f7-092e-4df1-a37c-510a3da0fa64 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Floatpim: In-memory acceleration of deep neural network training with high precision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba074ee-9d04-4044-a038-a29e28cb3de0 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Ultra-efficient processing in-memory for data intensive applications
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed0dce8-a0f9-46db-bfbe-c59958960241 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Intel xeon gold 6430 processor, 60m cache, 2.10 ghz
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa1399a-05e0-4d73-a5a6-e7ba13adfdf0 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Intel ® xeon® gold 6430 processor
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d974faa-ce32-4c60-9c7f-19365e6705cd · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf1c202c-5700-43bd-b906-d8d3953b23e6 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3183397-8d8e-4962-b7fd-85ba7ed408bb · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 551d1872-b606-4835-927f-e7255e0853e1 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ChatGPT for good? On opportunities and challenges of large language models for education
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0b1791-a46a-4a6c-bef5-518b86cd102e · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Rec- nmp: Accelerating personalized recommendation with near-memory processing
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40984173-3fef-493d-8435-905a989496c9 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21964e30-ba5f-42c4-93fe-95844210fa51 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 815bf70f-471f-4ee4-9d71-5072dcb57c42 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Samsung PIM/PNM for Transfmer Based AI: Energy Efficiency on PIM/PNM Cluster
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 24dbc041-3b61-4467-884a-3dfc5ade2e11 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A 1ynm 1.25v 8gb 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep learning application
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6cef06e-7a03-471a-b491-2f78bf7b2609 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6588b9e2-114f-418f-a03a-f1a17dc44c75 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Efficient memory management for large language model serving with pagedattention
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be80827-8f78-46d0-aab6-14544a8d6650 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Memory-centric computing with sk hynix’s domain-specific memory
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b25c34-da8e-45b6-a33f-722a70341f7f · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference System architecture and software stack for GDDR6-AiM
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b24437a-9736-4834-b18c-7bb0c0f435ba · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 25.4 a 20nm 6gb function-in-memory DRAM, based on HBM2 with a 1.2 tflops pro- grammable computing unit using bank-level parallelism, for machine learning applications
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66bdc8f1-1e91-4704-bee7-d804f371433f · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Improving in-memory database operations with acceleration DIMM (AxDIMM)
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc0b94e1-27db-470f-a9be-167adc88e9e6 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Using machine learning to increase yield and lower packaging costs
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d1065ab-8c94-415c-9c2c-93facde77733 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 06c69c93-84af-4188-b7d6-b4ebf002e139 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott Lee, Ishwar Agarwal, Mark D
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4cdfb794-01a7-46ab-a526-cb5ae6f9a145 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Accelerating distributed reinforcement learning with in-switch computing
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ed1efdb-7d91-4d21-8099-e68bb427a939 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Specification
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f709682-02b2-49d6-829d-af86e7374b5a · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference DeepSeek-V3 Technical Report
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510e802c-84e8-47b2-af67-c7f69a80ffe1 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Enmc: Extreme near-memory classification via approximate screening
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ccfcc69-ef8b-44b3-a38d-05032c7029c6 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0ae30fee-06ff-434c-8ca7-5fb36673ebfc · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac7644a-2243-4259-a206-b048edc8780a · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91daf12f-993f-4c47-a980-e9d882c50483 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dram power calculator
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b421ff6f-4c59-477b-b23b-21d96fadea73 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia shipped 3.76m data center gpus in 2023
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ab8cb036-44dc-4216-915c-da2638c00132 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Supply chain aware computer architecture
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d85537c-4844-4d37-9329-0f38af56f576 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for recommendation system
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff55d8ee-85ff-4168-86b3-cdfc2e145b2e · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Introduction to the nvidia dgx a100 system
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 87ffbc70-e3a9-4f62-a1c7-08606df975ee · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia a100 tensor core gpu
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b9db1fba-f9ff-41f1-839c-6a6dfc0ed566 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f6571af-9e27-4194-9641-a9036d6ecbba · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvlink and nvlink switch
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3cb0cb57-9174-4a9d-8d7c-0057f438580d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48f7f33b-48ef-4d04-9e6f-567b0376ba51 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Fine- grained dram: Energy-efficient dram for extreme bandwidth systems
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 001a95ce-8a7c-4781-a9a6-d1e1842e31df · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Bureau of Labor Statistics
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93dc2701-7455-4c73-a6e6-33f54ddaa771 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Accelerating neural network inference with processing-in-dram: From the edge to the cloud
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 53c05752-51aa-45d7-a70a-58eac2fb0983 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Gpt-4 turbo and gpt-4
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 89e01c35-16d1-4ced-bc9e-a20aeb8f2d32 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Learning to reason with llms
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72d9eac6-99da-4454-9711-51cc0a1679c7 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Video generation models as world simulators
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05f5e22d-b5ed-44e4-91e8-f04aa921dd0d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GPT-4 Technical Report
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b83c9968-c212-40fa-9f43-80d845986aac · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1d76f3a5-652f-412f-a3fa-fa8c10785fb6 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Attacc! unleashing the power of pim for batched transformer-based generative model inference
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 49de6f65-d730-407d-8c97-2d2b9afe42ae · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1d117e2-4a53-43a7-8f1d-f5aca6135e4e · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2b6d119-1ac8-4b0c-8f60-3eb64c6beb38 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Unresolved cited work
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2912e3-6b3b-4e5b-ba80-6981b4c294a7 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Splitwise: Efficient gen- erative llm inference using phase splitting
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b30cd212-45ff-4832-aeff-ec2c75b2361f · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Movie Gen: A Cast of Media Foundation Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0aafde-3aa6-426a-ac18-5890ceeaef73 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50869f67-984a-48bc-a6f2-53098521ce11 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7770edb4-b003-40ce-8c5b-e8e21091736f · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dota: detect and omit weak attentions for scalable transformer acceleration
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 31adadca-0d46-4a13-9f5d-f7a7364b8f78 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ad555e-ed1a-4fa2-9fc5-6b722b428f21 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 52e7404c-5666-46c4-bf2d-99a4b3108973 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 8gb gddr6 sgram c-die
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 347e29e6-146d-4f10-8c3c-f9e18db03e1d · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Searching for Activation Functions
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b826bd-1753-4dde-8de6-b8fb57a89e45 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GLU Variants Improve Transformer
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1106256b-df8e-4e41-a921-c63c4186a17a · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7992e68b-9baf-43c4-87c3-91a2b39d4c87 · outbound
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Generative ai winds in memory semiconduc- tors, total demand for server drams is declining
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.