Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:22:04.793768Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 5 inbound Pith citation observations for arXiv:2411.17309.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:22:04.793768Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T06:28:07.670564Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
62 of 62 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation ec2ed864-d673-4722-90fa-f9d636df04ea · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A comprehensive overvi ew of large language models, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cdb41c4-29a2-4d0a-bb09-fb74eb051919 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A survey of large language models, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01cbf9b5-fbac-4419-9edb-16d701342905 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Saddam Hossain Mukta, K aniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Marufatul Jannat Mim, Jubaer Ahmad, Mohammed Eu nus Ali, and Sami Azam
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f821116-6339-4df7-8422-a983f7b47dfa · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2cd94a8f-5589-4d3b-b045-3b16be50ebe0 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Recurrent neural net- work based language model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8ee1d6f7-0fb5-492a-8dd8-bcd999fdf59f · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Attention is all you need
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4b7f7a7-198b-4b15-a930-52a562033cd6 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad03c6a8-d628-49db-b77c-d0f15f98cd23 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Improving language understanding by generative pre-training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 58c67ba7-272b-4db8-b0fe-8db904a66f5c · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Exploring the limits of transfer learnin g with a unified text-to-text transformer
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9959e50a-0876-45ea-9afa-8c1200c8b7a2 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Bart: Denoising sequence-to-s equence pre-training for natural language generation, translation, and comprehension, 2019
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0f0da0d2-f1d6-4ecb-83af-0c0b12c8d03c · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Language models are few-shot learners
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b89764bf-6c00-43c6-b345-043c1236a1f7 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37350396-3204-490a-92c7-4e74ae6fd3ea · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Webgpt: Browser-assisted question-answering with human feedback, 2022
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2c879a45-0831-4332-a1ad-2d923f862514 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Llama: Open and efficient foundation l anguage models, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b55f7ed-02fc-467e-8a67-e6b43b1510c8 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3bf8726d-6506-4670-97b8-6ed10b8da332 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e857cee-a669-4b93-8d92-1db8ee03356c · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shak- eri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, J onathan H
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 78efecaa-c42f-4b7b-a3a2-4f47430e5c50 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 021b9e3e-d59e-4875-bd44-9983b2eccc1d · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference GPT-4 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84aa659-acc1-4d04-9059-ad75dc245da2 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a5e1773e-8849-439b-970b-a5786de139a3 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Webster and Chunyu Kit
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1ba81683-63e0-42ce-bff3-563afefe0d86 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Distributed representations of words and phrases and their compositionality, 2013
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ef9f4a-03dd-4c0a-a8ab-118a1c5a9c46 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Glove: Global vectors for word representation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e0578b4-fec3-4b4d-a9fb-881ff3a4c7e6 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Root mean square layer normalization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 05071be5-f98e-4a91-bfe1-29cb4fef8345 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f75e728-95aa-43aa-b199-e68151eae8b8 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference vllm: Easy, fast, and cheap llm serving with p agedattention
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f9e9dda-47e7-45e1-b74f-f1fcfa31fe9d · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Dissecting batching effects in gpt inferen ce, 2023
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9e108453-2dcb-464a-82ca-e1f06128b168 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mobilellm: Optimizing sub-billion parameter language models for on-device use ca ses, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a88d004-292a-4952-a924-db19515621dd · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Octopus v2: On-device language model for super agent, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbc00a9-b313-482b-a44c-988c663e74ff · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A survey on hardware accelerato rs for large language models, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ae04b50f-f2f1-40cf-8f13-3ccb7e590a56 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Ene rgy and policy considerations for deep learning in nlp, 2019
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70ebaa6c-7349-4f9d-bca3-3b020dc7d2b4 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 68c9198f-9cb6-4c65-9356-69a2e4d1b3e5 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mahoney, and Kurt Keutzer
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 084126a7-03c0-4906-8ff2-ca5aca5bfdbd · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference From wor ds to watts: Benchmarking the energy costs of large language model inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 401e5b96-f03b-4ed6-8b58-eed264dae6e5 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Me asuring and improving the energy efficiency of large language models inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 34093139-1c1b-46e0-bba1-579680996b3e · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Risks and benefits of large language models for the environment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 135cf3a2-c618-4053-b74b-3318178b39b5 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A short survey of viewing large languag e models in legal aspect, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 391a95d7-9599-480a-9a9d-db49df27cf13 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference What does it mean for a language model to preserve privacy? In Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, pages 2280–2292, 2022
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0b75a610-c378-4972-9867-61fa8a3c0023 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Y ou are what you write: Preserving privacy in the era of large language models, 2022
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2326507b-179d-42af-aa61-2a9bcafe802e · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Deli ver high performance ml inference with aws inferentia
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9c83a7b9-1261-451f-8a98-389e0c3f9c36 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mm1: Methods, analysis & insights from multimo dal llm pre-training, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7bd2a334-5f1d-4d49-884e-ba05c4f45c0e · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language mo dels
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 98d44f72-f20a-41c3-9f8d-53f922b4c328 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Compression of generative pre-trained language models via quantizatio n, 2022
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7c857d5b-d3ca-47ea-a961-6322d3ad5fe3 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mahoney, and Kurt Keutzer
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cd873ac-af68-43eb-bda9-d5ae4b9fc840 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Onebit: Towards extremely low-bit large language models, 2 024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 97442a05-4323-4615-a962-a9e8855852e8 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference The era of 1-bit llms: All large lang uage models are in 1.58 bits, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 454ca70c-e3ee-440b-9391-a7fcc22f822c · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Llm-pruner: On the structural pruning of large language mod- els
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0f579aff-c75b-4d47-8acd-1ef769fda8a1 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference From dense to sparse: Contrastive pruning for better pre-trained lang uage model compression
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7971c7d1-dcb8-4bdf-bf50-f4eff84e164d · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A s urvey on model compression for large language models, 2023
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 03ab731c-db83-4481-b2d8-fb16f09e2daf · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Distill ing the knowledge in a neural network, 2015
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 439f6cc7-c116-4059-abbb-87227f76885f · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Knowledge distillation: A survey
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 71280841-43d0-4ece-9a94-f49cf061451d · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference De Lima, Hamid Farzaneh, and Jeronimo Castrillon
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 63e1812a-a159-4e0d-ba55-3193a165e25e · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A Modern Primer on Processing in Memory, pages 171–243
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c63990ee-a8b6-4a53-888b-d1ac89d0b452 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference High-speed emerging memories for ai hardwar e accelerators
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a240b566-1cf7-4912-8c5d-5e765c8d0464 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference The breakthrough memory solutions for improved perfo rmance on llm inference
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5bdf3035-69fa-413a-864f-38ad766bca71 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Oliveira, and Onur Mutlu
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 434583d4-6302-40fe-ad3c-80b4121bf811 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Energy efficiency impa ct of processing in memory: A comprehensive review of workloads on the upmem architecture
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f01ef35c-cdd6-42ab-a09b-e5ee369c8e45 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Technical report, Qualcomm, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4625f2ac-3218-4515-a7ba-cf0c113b9fd8 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c551978f-5160-48cf-bd94-d259895def6c · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Accessed: 2024-07-11
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 210ed5ae-eaa8-421c-924d-1efdbeb6eff2 · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Introducing Apple’s On-Device and Server Found ation Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e52482d3-145e-4d4d-ac38-58cfb605fccd · outbound
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Achieving High Mixtral 8x7B Performance with N VIDIA H100 Tensor Core GPUs and TensorRT- LLM
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation af81e374-6ef4-4fa4-8a5f-54a701a8f4c5 · inbound
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33084609-477b-4384-9bcc-18af6a435940 · inbound
Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d94236d-d27e-4b61-9891-1facd1175a8e · inbound
Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c35b4897-1a8d-4cf0-9f5f-617a6e7d4710 · inbound
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a0c12c67-db0c-433e-a46c-8e6237efc2eb · inbound
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.