Pith. sign in

Paper Citation Record · LEDGER

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 5 inbound Pith citation observations for arXiv:2411.17309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17309 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:22:04.793768Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:28:07.670564Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ec2ed864-d673-4722-90fa-f9d636df04ea · outbound

This paper cites A comprehensive overvi ew of large language models, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A comprehensive overvi ew of large language models, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.974735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.478481Z digest=sha256:966652978a7f2142c17e7bfe1ac3ddc8573c7480827bf4871a4c4c63f38e6859

Observation 7cdb41c4-29a2-4d0a-bb09-fb74eb051919 · outbound

This paper cites A survey of large language models, 2023.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A survey of large language models, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.956356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.484696Z digest=sha256:6e4a4da435eb79ec907f489c555a37b7e34231af84d310b5d95cd0be60005cfd

Observation 01cbf9b5-fbac-4419-9edb-16d701342905 · outbound

This paper cites Saddam Hossain Mukta, K aniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Marufatul Jannat Mim, Jubaer Ahmad, Mohammed Eu nus Ali, and Sami Azam.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Saddam Hossain Mukta, K aniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Marufatul Jannat Mim, Jubaer Ahmad, Mohammed Eu nus Ali, and Sami Azam

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.938270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.490020Z digest=sha256:801e6a850595cb64097cb681de4a102fbae5395cd031d574111d5393a319c3d9

Observation 2f821116-6339-4df7-8422-a983f7b47dfa · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.919848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.496427Z digest=sha256:5b9ea2dd4d8bb6682ddcaf968855bc9053e1e21ceb38a4c3e78cdf6ec0735ef6

Observation 2cd94a8f-5589-4d3b-b045-3b16be50ebe0 · outbound

This paper cites Recurrent neural net- work based language model.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Recurrent neural net- work based language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.903307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.501734Z digest=sha256:475bac237217306250e849a0dc3f411abc509d1f974ba9db63c957804781758b

Observation 8ee1d6f7-0fb5-492a-8dd8-bcd999fdf59f · outbound

This paper cites Attention is all you need.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Attention is all you need

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.876612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.507598Z digest=sha256:2300c83f8e1fb16d8fad4022da1330b392c7637ecf3004b7849bf9dbcf1fe741

Observation b4b7f7a7-198b-4b15-a930-52a562033cd6 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding, 2019.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.859662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.514225Z digest=sha256:b1f68c674818acc1dfb9487d22ed39db298d17d165c7e4c7156156aa7927ba1b

Observation ad03c6a8-d628-49db-b77c-d0f15f98cd23 · outbound

This paper cites Improving language understanding by generative pre-training.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Improving language understanding by generative pre-training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.842184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.519951Z digest=sha256:9bf59b1046d8496378e86de5b0746f6a3107dc0b9e418d14cfe5ccd0f35bec69

Observation 58c67ba7-272b-4db8-b0fe-8db904a66f5c · outbound

This paper cites Exploring the limits of transfer learnin g with a unified text-to-text transformer.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Exploring the limits of transfer learnin g with a unified text-to-text transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.825475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.524881Z digest=sha256:957596593dd5aa868a8ce67c76b77965efad8058924ba7e6b647ad10aa23bafa

Observation 9959e50a-0876-45ea-9afa-8c1200c8b7a2 · outbound

This paper cites Bart: Denoising sequence-to-s equence pre-training for natural language generation, translation, and comprehension, 2019.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Bart: Denoising sequence-to-s equence pre-training for natural language generation, translation, and comprehension, 2019

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.808688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.530215Z digest=sha256:a671a516322333f0fe0aedcb43eea9f08e6f31d7df2edc29d96dc1d2b71ffa96

Observation 0f0da0d2-f1d6-4ecb-83af-0c0b12c8d03c · outbound

This paper cites Language models are few-shot learners.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Language models are few-shot learners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.791731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.535107Z digest=sha256:cdadd1721e6e8b10bae01658258723ab23e8c072c994ba5c7a626fd9de4455ba

Observation b89764bf-6c00-43c6-b345-043c1236a1f7 · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.771576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.540449Z digest=sha256:455f6894b2a4a3fc36d93170de969790c028844130dba074e6f84318f41e2051

Observation 37350396-3204-490a-92c7-4e74ae6fd3ea · outbound

This paper cites Webgpt: Browser-assisted question-answering with human feedback, 2022.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Webgpt: Browser-assisted question-answering with human feedback, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.753744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.546122Z digest=sha256:0a96cca8f878acab9175a0ba769eb62219788e32efde86f5e56ad72a33fb0957

Observation 2c879a45-0831-4332-a1ad-2d923f862514 · outbound

This paper cites Llama: Open and efficient foundation l anguage models, 2023.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Llama: Open and efficient foundation l anguage models, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.736753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.550920Z digest=sha256:c366e1a69857b370a169025dd892c0ede9a80b7ff5b0da0d186b4177e72b5a55

Observation 6b55f7ed-02fc-467e-8a67-e6b43b1510c8 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.718782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.556118Z digest=sha256:35e4e17fa2f70c5572092ee6993d7fa1a511796423f7ccbbb32bf748a0f3496c

Observation 3bf8726d-6506-4670-97b8-6ed10b8da332 · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.699520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.561303Z digest=sha256:3938f06dc3869f32a34b2cd26a7ed9076d722b9014153e8e343ce2645dfb2802

Observation 2e857cee-a669-4b93-8d92-1db8ee03356c · outbound

This paper cites Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shak- eri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, J onathan H.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shak- eri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, J onathan H

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.683093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.566424Z digest=sha256:54753a2f7e25a7eceb9e025d5794a3252aeda9c3a8a20dc97e969ff85f701888

Observation 78efecaa-c42f-4b7b-a3a2-4f47430e5c50 · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.666431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.571547Z digest=sha256:76a390ba54b4fcfd4cfeccfab04ce0021655dff8fbeeccd3dec096c3ef461f55

Observation 021b9e3e-d59e-4875-bd44-9983b2eccc1d · outbound

This paper cites GPT-4 Technical Report.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:22:04.576637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:22:04.576637Z digest=sha256:a1c280bc9a6608de41ed4a34c0b5e09a505f03cc28cb143e2a423252aab71938

Observation b84aa659-acc1-4d04-9059-ad75dc245da2 · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.649320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.582025Z digest=sha256:ceac9af74e3840e9cac063c534f2d2ca8d6c385a9e9655d2471b7aa76b39b847

Observation a5e1773e-8849-439b-970b-a5786de139a3 · outbound

This paper cites Webster and Chunyu Kit.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Webster and Chunyu Kit

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.632979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.586930Z digest=sha256:a880cc26b5dca4034112ca1adc6d35714cb8e57943a10779a92f800d5d809228

Observation 1ba81683-63e0-42ce-bff3-563afefe0d86 · outbound

This paper cites Distributed representations of words and phrases and their compositionality, 2013.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Distributed representations of words and phrases and their compositionality, 2013

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:22:04.591779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:22:04.591779Z digest=sha256:f5a090d226731a25df1d86aa5647cb0216c675afa3571e4aba914cd1ff8e69b8

Observation 34ef9f4a-03dd-4c0a-a8ab-118a1c5a9c46 · outbound

This paper cites Glove: Global vectors for word representation.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Glove: Global vectors for word representation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.601861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.596644Z digest=sha256:662060ea1d0ca7c91696f75156975ac2120e5c8c4117fb9e0d8ffe62ea0d1549

Observation 2e0578b4-fec3-4b4d-a9fb-881ff3a4c7e6 · outbound

This paper cites Root mean square layer normalization.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Root mean square layer normalization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.581655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.601566Z digest=sha256:c7cc979e5ed6b392a46ee647f717e820d1ee8fea498e27ac9e2cf730a4f57678

Observation 05071be5-f98e-4a91-bfe1-29cb4fef8345 · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.558736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.606451Z digest=sha256:c971669b0382b7f6ddd98c7bbeda47e97091c2c470f5ce6579ee0f189971768f

Observation 6f75e728-95aa-43aa-b199-e68151eae8b8 · outbound

This paper cites vllm: Easy, fast, and cheap llm serving with p agedattention.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference vllm: Easy, fast, and cheap llm serving with p agedattention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.537133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.611654Z digest=sha256:e2324f53e714623cb39587da6b7477b35eb87e27e1f6c0fa6786d944e57559ec

Observation 9f9e9dda-47e7-45e1-b74f-f1fcfa31fe9d · outbound

This paper cites Dissecting batching effects in gpt inferen ce, 2023.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Dissecting batching effects in gpt inferen ce, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.513341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.616494Z digest=sha256:7bb10b877594274cdc761cc012c67cf1fc9afe381a4d23e5d67817c2ce3a2e54

Observation 9e108453-2dcb-464a-82ca-e1f06128b168 · outbound

This paper cites Mobilellm: Optimizing sub-billion parameter language models for on-device use ca ses, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mobilellm: Optimizing sub-billion parameter language models for on-device use ca ses, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.495627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.621945Z digest=sha256:382c1b98e20cedc5bae8e701d2a17dc8923ce7fd957679a44d9eac0a0f53b47d

Observation 2a88d004-292a-4952-a924-db19515621dd · outbound

This paper cites Octopus v2: On-device language model for super agent, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Octopus v2: On-device language model for super agent, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:22:04.626809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:22:04.626809Z digest=sha256:6f58d27b292f4477ee58b62eebb034a019f24bc92c430525fef32174d7e79d4b

Observation ddbc00a9-b313-482b-a44c-988c663e74ff · outbound

This paper cites A survey on hardware accelerato rs for large language models, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A survey on hardware accelerato rs for large language models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.467549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.631779Z digest=sha256:3a03dddc637739221167731a66cce64b3ebc8e08e6add7f6dc69c241d89aefe6

Observation ae04b50f-f2f1-40cf-8f13-3ccb7e590a56 · outbound

This paper cites Ene rgy and policy considerations for deep learning in nlp, 2019.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Ene rgy and policy considerations for deep learning in nlp, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.449579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.636403Z digest=sha256:23edcdffad95265495a892ecc654613bf1508cc12cbdc08998e075d85ea21fa9

Observation 70ebaa6c-7349-4f9d-bca3-3b020dc7d2b4 · outbound

This paper cites an unresolved cited work.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:22:05.429206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.641330Z digest=sha256:8b56ad816ab6b709fa7b036af8b0fb2aa29987842c6c90589171ba3daa726be6

Observation 68c9198f-9cb6-4c65-9356-69a2e4d1b3e5 · outbound

This paper cites Mahoney, and Kurt Keutzer.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mahoney, and Kurt Keutzer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.410165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.646001Z digest=sha256:106d7650e4f2315b0f67d2ecba096ba937e933fa30edbd969c4b1ded5f6bc215

Observation 084126a7-03c0-4906-8ff2-ca5aca5bfdbd · outbound

This paper cites From wor ds to watts: Benchmarking the energy costs of large language model inference.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference From wor ds to watts: Benchmarking the energy costs of large language model inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.392666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.650726Z digest=sha256:b4b0a62b24cff62f36afdbfa9484c57a8912aeba16a435b9c249629af8d9fc8a

Observation 401e5b96-f03b-4ed6-8b58-eed264dae6e5 · outbound

This paper cites Me asuring and improving the energy efficiency of large language models inference.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Me asuring and improving the energy efficiency of large language models inference

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.375090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.656916Z digest=sha256:cfa48fbedddd5d9650c6a42b1504a0f9b4e52dd3c5f3d96f40349f18009e3a6d

Observation 34093139-1c1b-46e0-bba1-579680996b3e · outbound

This paper cites Risks and benefits of large language models for the environment.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Risks and benefits of large language models for the environment

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.357353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.662417Z digest=sha256:6b22131969653c97957be83a005093b641677c17fb55bde295c1921c5189bea6

Observation 135cf3a2-c618-4053-b74b-3318178b39b5 · outbound

This paper cites A short survey of viewing large languag e models in legal aspect, 2023.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A short survey of viewing large languag e models in legal aspect, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.339485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.667481Z digest=sha256:191c6b2996fb62887b29a7c7d651277dcf77929bdb3896c69052028c04811003

Observation 391a95d7-9599-480a-9a9d-db49df27cf13 · outbound

This paper cites What does it mean for a language model to preserve privacy? In Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, pages 2280–2292, 2022.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference What does it mean for a language model to preserve privacy? In Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, pages 2280–2292, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.322171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.672762Z digest=sha256:fbfb8c619b06eece05fefe163eddf97026f774faf278a062c90725c2884fd0b7

Observation 0b75a610-c378-4972-9867-61fa8a3c0023 · outbound

This paper cites Y ou are what you write: Preserving privacy in the era of large language models, 2022.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Y ou are what you write: Preserving privacy in the era of large language models, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.304921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.677567Z digest=sha256:a948f2fe7b7c9d4b5af4c8a0640389a7a394c5c668a08fb4a3a6d6c8b66a0ff0

Observation 2326507b-179d-42af-aa61-2a9bcafe802e · outbound

This paper cites Deli ver high performance ml inference with aws inferentia.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Deli ver high performance ml inference with aws inferentia

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.285540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.682320Z digest=sha256:c64a021fcaae2d98ebb5c152e57b0b93faa7f73119ef138489c78180c1505c35

Observation 9c83a7b9-1261-451f-8a98-389e0c3f9c36 · outbound

This paper cites Mm1: Methods, analysis & insights from multimo dal llm pre-training, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mm1: Methods, analysis & insights from multimo dal llm pre-training, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.266413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.687110Z digest=sha256:2c82fc8fe2ab4b33db96f86adaba6cb74dce053be4dbb8d1c4a1a83689796153

Observation 7bd2a334-5f1d-4d49-884e-ba05c4f45c0e · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language mo dels.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language mo dels

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.249324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.693099Z digest=sha256:e677002582a01d29daf6d3e7add99b9b700b847caf392fa88acb13324a8af75b

Observation 98d44f72-f20a-41c3-9f8d-53f922b4c328 · outbound

This paper cites Compression of generative pre-trained language models via quantizatio n, 2022.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Compression of generative pre-trained language models via quantizatio n, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.232371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.698101Z digest=sha256:b47770027bd7a8cb8e374532a8efc9f0cbf2bc1a08f97e7c163b910ff208e0e6

Observation 7c857d5b-d3ca-47ea-a961-6322d3ad5fe3 · outbound

This paper cites Mahoney, and Kurt Keutzer.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Mahoney, and Kurt Keutzer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.214957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.703891Z digest=sha256:16ae774bc5d3e48e95ee23310d3072b894c5dd0cbb27bc06dcbf11ee859a1f5b

Observation 7cd873ac-af68-43eb-bda9-d5ae4b9fc840 · outbound

This paper cites Onebit: Towards extremely low-bit large language models, 2 024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Onebit: Towards extremely low-bit large language models, 2 024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.192944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.709035Z digest=sha256:2f2c4f61bc5e6cda3241e07ab736387760b373f77f5c1cfb1940c9051437201d

Observation 97442a05-4323-4615-a962-a9e8855852e8 · outbound

This paper cites The era of 1-bit llms: All large lang uage models are in 1.58 bits, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference The era of 1-bit llms: All large lang uage models are in 1.58 bits, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.174903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.713994Z digest=sha256:3263cd6b3505fa36dd25a6510765d3b3b8969506b01f321275d390ea80380ec1

Observation 454ca70c-e3ee-440b-9391-a7fcc22f822c · outbound

This paper cites Llm-pruner: On the structural pruning of large language mod- els.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Llm-pruner: On the structural pruning of large language mod- els

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.156196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.719045Z digest=sha256:42a24231338f55571d23f1e1c0d543b89e70e6b8132fa554af63fd99378a66cb

Observation 0f579aff-c75b-4d47-8acd-1ef769fda8a1 · outbound

This paper cites From dense to sparse: Contrastive pruning for better pre-trained lang uage model compression.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference From dense to sparse: Contrastive pruning for better pre-trained lang uage model compression

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.137016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.723777Z digest=sha256:ba85a7108bb7355f98efaa7ae82deab770060e73a19c05ee7e0fc598ab8394a5

Observation 7971c7d1-dcb8-4bdf-bf50-f4eff84e164d · outbound

This paper cites A s urvey on model compression for large language models, 2023.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A s urvey on model compression for large language models, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.116672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.728663Z digest=sha256:58b1f76d3d8475574e498d18b339acdae710a40d1b07ad72a4230e2defc78e3c

Observation 03ab731c-db83-4481-b2d8-fb16f09e2daf · outbound

This paper cites Distill ing the knowledge in a neural network, 2015.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Distill ing the knowledge in a neural network, 2015

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.099988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.733590Z digest=sha256:d6f9e07de894dd217e4f90e5ae89756878ec0bfb5a3cbf070b7a67ea901bf4ab

Observation 439f6cc7-c116-4059-abbb-87227f76885f · outbound

This paper cites Knowledge distillation: A survey.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Knowledge distillation: A survey

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.082243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.738393Z digest=sha256:62b21f46dfa9c90eb4a2f33dde52e92c697093067f46a57dd5cbf7b75e6687b9

Observation 71280841-43d0-4ece-9a94-f49cf061451d · outbound

This paper cites De Lima, Hamid Farzaneh, and Jeronimo Castrillon.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference De Lima, Hamid Farzaneh, and Jeronimo Castrillon

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.065206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.743633Z digest=sha256:b1c76b5015012e40fe0c2d4d46e27d4fbc22bd60fa6320a7a1a0b17874e3ece9

Observation 63e1812a-a159-4e0d-ba55-3193a165e25e · outbound

This paper cites A Modern Primer on Processing in Memory, pages 171–243.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference A Modern Primer on Processing in Memory, pages 171–243

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.043911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.748254Z digest=sha256:56d6ec65f20881372abe871c31a51d26570d7e75a8dd77f3fcc177e3c1557f33

Observation c63990ee-a8b6-4a53-888b-d1ac89d0b452 · outbound

This paper cites High-speed emerging memories for ai hardwar e accelerators.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference High-speed emerging memories for ai hardwar e accelerators

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.024572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.753366Z digest=sha256:494b26fba313c86bb9d064edd779982827a777a541ca3b448b6b362fd0ade1ea

Observation a240b566-1cf7-4912-8c5d-5e765c8d0464 · outbound

This paper cites The breakthrough memory solutions for improved perfo rmance on llm inference.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference The breakthrough memory solutions for improved perfo rmance on llm inference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:05.005213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.758120Z digest=sha256:328711a93834ea0d38e00a4afa8902cf5830f357e4ecf07f0b1ec3a5c5042fe1

Observation 5bdf3035-69fa-413a-864f-38ad766bca71 · outbound

This paper cites Oliveira, and Onur Mutlu.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Oliveira, and Onur Mutlu

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:04.981952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.763341Z digest=sha256:7da886ce43fc432a93fee4dbd05ebe0fd6f934190d7851b7baf6baddeaecb711

Observation 434583d4-6302-40fe-ad3c-80b4121bf811 · outbound

This paper cites Energy efficiency impa ct of processing in memory: A comprehensive review of workloads on the upmem architecture.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Energy efficiency impa ct of processing in memory: A comprehensive review of workloads on the upmem architecture

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:04.963238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.767996Z digest=sha256:d21baf2a548cf2c632f3d0e7a7e3261a9576eb061a7f2a5fa321af24fead1dbc

Observation f01ef35c-cdd6-42ab-a09b-e5ee369c8e45 · outbound

This paper cites Technical report, Qualcomm, 2024.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Technical report, Qualcomm, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:04.945984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.772859Z digest=sha256:d0d547eecd0106e561be6821afb036a625a674a0e5bce897f4c78e4a85c4e528

Observation 4625f2ac-3218-4515-a7ba-cf0c113b9fd8 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:22:04.777569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:22:04.777569Z digest=sha256:e5152cf3bac785b4194883c9283b1c8bebb43345660f6aaf07cbb0abaabbd366

Observation c551978f-5160-48cf-bd94-d259895def6c · outbound

This paper cites Accessed: 2024-07-11.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Accessed: 2024-07-11

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:04.928678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.783813Z digest=sha256:dff953b467fe04ae6eaa2b1e4f5d642f85464c9a1e885116da9d9b8394a6f70e

Observation 210ed5ae-eaa8-421c-924d-1efdbeb6eff2 · outbound

This paper cites Introducing Apple’s On-Device and Server Found ation Models.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Introducing Apple’s On-Device and Server Found ation Models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:04.911018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.788924Z digest=sha256:2d23173695ba7a1fef0cef5f4381ee31ad9db8252f70e936f6f9ab7c3cb70967

Observation e52482d3-145e-4d4d-ac38-58cfb605fccd · outbound

This paper cites Achieving High Mixtral 8x7B Performance with N VIDIA H100 Tensor Core GPUs and TensorRT- LLM.

PIM-AI: A Novel Architecture for High-Efficiency LLM Inference Achieving High Mixtral 8x7B Performance with N VIDIA H100 Tensor Core GPUs and TensorRT- LLM

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:22:04.889189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:22:04.793768Z digest=sha256:68ec7b5b9113c77d2180f3f1a1779e1cbe71d45e8c702a9c71758dc79f7464d0

Pith citing papers

Observation af81e374-6ef4-4fa4-8a5f-54a701a8f4c5 · inbound

The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency cites this paper.

The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency PIM-AI: A Novel Architecture for High-Efficiency LLM Inference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:00.979649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:46:11.859634Z digest=sha256:66cf17d162aa39e0f866b227fee7b8ca8728ddb6c8131da8105cfd24f88f033b

Observation 33084609-477b-4384-9bcc-18af6a435940 · inbound

Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory cites this paper.

Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory PIM-AI: A Novel Architecture for High-Efficiency LLM Inference

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.467021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T19:51:10.637303Z digest=sha256:95810b7119ff5b0834371b6c75f74cbaf15f4d131264f35b27116f81965e2321

Observation 3d94236d-d27e-4b61-9891-1facd1175a8e · inbound

Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory cites this paper.

Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory PIM-AI: A Novel Architecture for High-Efficiency LLM Inference

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:54.941479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T19:51:10.637303Z digest=sha256:fdaffad8121bd294e8a64fcf6a17731d51f267905aa29401cb84a58f8e9020e2

Observation c35b4897-1a8d-4cf0-9f5f-617a6e7d4710 · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices PIM-AI: A Novel Architecture for High-Efficiency LLM Inference

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:14:14.929917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T03:09:57.069086Z digest=sha256:e2ffeff70e052a398672be41aec27b2ca859063cc47e58f64d58a471e7ce5614

Observation a0c12c67-db0c-433e-a46c-8e6237efc2eb · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices PIM-AI: A Novel Architecture for High-Efficiency LLM Inference

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.897941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T06:28:07.670564Z digest=sha256:5eb23c5485b31a1e6901812f026ba8e937399f0505ac80a566febadc62db107d