Pith. sign in

Paper Citation Record · LEDGER

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

As of 4 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 2 inbound Pith citation observations for arXiv:2511.06838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.06838 v4

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:16:27.497363Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:52:39.415504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact10
  • verified fuzzy66
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91c0cbda-410f-4f64-9607-3dcedc025466 · outbound

This paper cites AMD INSTINCT™ MI350X GPU.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AMD INSTINCT™ MI350X GPU

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.381545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:6591cd66ce71020ee6bf96d3682e4fa747de6e35b9562b93d1d4a7da377849da

Observation 0466a4a0-b073-41c4-90d2-aa0ad4611b80 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.385570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5c6f53a2a733b80123ca5966df2079df08fc4f4b98acd4a83ee6d389d264ece6

Observation 5ffe8591-3f1b-422f-be20-b97d8420901e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.193780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:40a661ae4d88821a79c07388c2fd5e7691571a0049cca5db94a77e3d0edd9819

Observation c9058645-61af-400b-9d8b-72c2ae33df5f · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.388138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7cef491c1f2b7ff9b442b7e919fd900e70e3fbdfd5a4392403b61ddd03c3a7c4

Observation 8ca165b9-7a3c-442c-ad7f-7a252ee8e115 · outbound

This paper cites CACTI 7: New tools for interconnect exploration in innovative off-chip memories.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats CACTI 7: New tools for interconnect exploration in innovative off-chip memories

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.473674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:9df0c5b01e2fbe9d6c288c0487020cde48facc919003382780b7937f29a22f56

Observation 86fd75e1-3aa8-4d7a-be30-4c2a9b166c8f · outbound

This paper cites Language Models are Few-Shot Learners.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Language Models are Few-Shot Learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.358350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:1804dee978053d32d62bfed2590297ad6eff9c000a9403508c5a92292f449ba7

Observation 94742906-3c9b-4726-92fd-5aafa83d2ef5 · outbound

This paper cites BitMoD: Bit-serial Mixture-of- Datatype LLM Acceleration.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats BitMoD: Bit-serial Mixture-of- Datatype LLM Acceleration

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.501615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7cce9e4cac40817f26fce480b4d89cbdfd48070d65ce34ea9f4fc7908f17b089

Observation 3bfbabd8-5e52-45fe-956a-2cf4d8fd5442 · outbound

This paper cites Ecco: Improving Memory Band- width and Capacity for LLMs via Entropy-Aware Cache Compression.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ecco: Improving Memory Band- width and Capacity for LLMs via Entropy-Aware Cache Compression

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.476680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:eea58f9ec554aef91285bc255ed59b036a54c94e1e38e8da06e99f1a5ddf4963

Observation 5b80588a-b746-4f55-9cb6-9fa6eb4302f7 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.226910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b2ba3c66a1e330d3bdc061c3c410a4c25033a28db55f07e64bb8aff1ebfb94f6

Observation 0a46fa4d-337c-4a22-a73d-0f95d19bc213 · outbound

This paper cites DeepSeek R1.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats DeepSeek R1

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.351770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:4c401118ddac04b9cfa6235e22703cf84b7d4dd7a27ad06cd7823b5e8ab3b8ef

Observation 7ce4bcba-230c-4174-a0e2-6326a867e548 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.221843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:848fc88dbea63a5013330b2956d858427aba80addfcce32fb52b634ac73071db

Observation d89f3c01-9891-409b-ae6c-ab27d4be13f7 · outbound

This paper cites The true Processing In Memory accelerator.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The true Processing In Memory accelerator

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.344619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:660f9242566f3515ec0354662561f469f0ce3183af028434e99af7bc6c2bd1c2

Observation 0aa2b922-6b06-47ce-a15c-99b451485171 · outbound

This paper cites Documenting large webtext corpora: A case study on the colossal clean crawled corpus.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Documenting large webtext corpora: A case study on the colossal clean crawled corpus

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.349072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:6f2c7009f4c55227f285bbb601bc1285f6e9184483f20f4a4a9adf2f40bef2a4

Observation 4d617caa-5ed4-4368-a3c5-f0e4732d6285 · outbound

This paper cites Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.342113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:1c56e7f7c38183c873c0c5d7107c557b7c854c452a1e1ea9464ec5068925c5dd

Observation 92c72bcc-f4ed-4cf7-82c9-048c642bef3b · outbound

This paper cites Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.433912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:3f3f57c8b7549eec1e408ec99b86b96ac0b7345fefeb8615578e34a3ae84c279

Observation fd0a485e-3ccd-4422-9514-ede7d9e78e35 · outbound

This paper cites GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.338936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:761090bcf47e8cb3a8176aedf6d4216719f732517bca4d7b2b2a86a6e536ca01

Observation fe1e1bfc-e012-415b-a232-62872a374461 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.232926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:80cab9f245849bf8f6576909e23951622eb5aab972c2a60b5160eaed9f921be1

Observation a7d3720c-d2c1-4338-80dc-2e0477ee3806 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:20:32.211612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:0f8bda8c7085bcb37cb2b77f13d0ac19e28edf945bef123cde3f223f9ffe7987

Observation 730c8e44-014a-4199-bf83-7497416de775 · outbound

This paper cites Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.391034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d1289cc8736c710217e4d524e34edf29e5499a0f783bcd61dd6491c62fee5d15

Observation b3e9ca14-f00a-4855-871a-65d16eeecdcd · outbound

This paper cites OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.403649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:6996a4d1b44f592e1499054782bcd5bea3afd99250cbc96bdb0ea82fcf53351c

Observation a301124c-f97e-48c7-bce5-84946469ebf0 · outbound

This paper cites ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.555058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:4d3172b82d0e2c73bd6f754525f176d9b9aa1018fbe1b1c22fd536f5fa9ac305

Observation f4ae97eb-9a0a-448a-a701-865a7cfd32ab · outbound

This paper cites Newton: A DRAM-maker’s Accelerator-in- Memory (AiM) Architecture for Machine Learning.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Newton: A DRAM-maker’s Accelerator-in- Memory (AiM) Architecture for Machine Learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.557822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:0aef4ba7474bce71eba970a3e9b78dd833a1390928e3dd2aae8df884d1dd5595

Observation 4258aa8e-5cc0-4a2e-89a0-78261b683fac · outbound

This paper cites LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture- Dataflow Co-Optimization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture- Dataflow Co-Optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.562943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:c199e1da23b0b7b0ccf7560098470fc00c3ef82cc6b035d63de405ee0397c5f5

Observation 7636e3b1-8978-4041-a7ac-96c90721699b · outbound

This paper cites How Would the Viewer Feel? Estimating Wellbeing from Video Scenarios.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats How Would the Viewer Feel? Estimating Wellbeing from Video Scenarios

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-05-18T00:20:33.547004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b0c4dfba77117a4edb31f930082fff14e789351fa5ff0d59be578c656b210965

Observation 223b3dd1-693a-4b74-bb2a-7b2810876f34 · outbound

This paper cites NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.549665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:829451c4ab9e822b96afe223dc508fbfe66f3161b8879ddcff30ea8370572539

Observation ec386510-1f74-4678-8581-679a8b1da606 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.565565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:eea2645352531507bfc4650d7b1dba2339ff9157ee81a9ea44e4ddfce29ca97f

Observation b18f5743-db24-4491-acf6-16b565ecb9f0 · outbound

This paper cites M-ANT: Efficient Low-bit Group Quantization 13 for LLMs via Mathematically Adaptive Numerical Type.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats M-ANT: Efficient Low-bit Group Quantization 13 for LLMs via Mathematically Adaptive Numerical Type

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.537830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:e85d7d4043a6e5f21674e6f318a0cf32ed41de7181174ed85522f57d2c0d8c92

Observation 80654548-de0e-471d-8be6-49fe29b618bc · outbound

This paper cites PLAIN: Leveraging High Internal Bandwidth in PIM for Accelerating Large Language Model Inference via Mixed-Precision Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats PLAIN: Leveraging High Internal Bandwidth in PIM for Accelerating Large Language Model Inference via Mixed-Precision Quantization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.543646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7e38d99ccd30c44da0156ebc64fbb3e1563222ef0b685da66d57ca7285407eef

Observation 51813f2e-16d3-4639-ac90-41c9a433f04b · outbound

This paper cites FIGNA: Integer unit-based accelerator design for fp-int gemm preserving numerical accuracy.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FIGNA: Integer unit-based accelerator design for fp-int gemm preserving numerical accuracy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.531034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:6e9ed5d248faf1073b26402b8b76005afb8a63e6b40ba227bea976f841942f29

Observation 4d924579-6091-423b-b186-3eb49ad1d2e8 · outbound

This paper cites BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.533641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:8b49da5032d7460440a3addf995dde0ec560980821a2d0c29bdd80cfcff601cb

Observation da94f300-dd5a-4261-a94a-4a3b8d313b0b · outbound

This paper cites High Bandwidth Memory DRAM.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory DRAM

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.525525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:297b994c94ba6cd2901e92ce1a0ebe3ddaf0de9ae880cd0dce2e35aea0f950ed

Observation 2529d643-ad30-41e9-8c70-a908eea353ea · outbound

This paper cites High Bandwidth Memory (HBM3) DRAM.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory (HBM3) DRAM

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.516901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:96e5ef7815dca6299508c921a3df2ec3ccd5cbe379e9e9deac46233baffed796

Observation b29bf160-1749-4b53-921e-7d66abece790 · outbound

This paper cites High Bandwidth Memory (HBM4) DRAM.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory (HBM4) DRAM

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.519297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:a1572cec3e4e819e1f94a2a4ed7e45fd0226fbf0d008e8c7024ef10bc863dabc

Observation 3200a273-e128-414a-9922-878b8eddc9a0 · outbound

This paper cites Mistral 7B.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Mistral 7B

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.203892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:0388282d2ddc5787fe6ddd9d62f723cba1af6c8da512230ed0718fcd293ffd28

Observation 57c2c5a6-3de8-43ed-80b8-51209280a230 · outbound

This paper cites Ten Lessons From Three Generations Shaped Google’s TPUv4i: Industrial Product.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ten Lessons From Three Generations Shaped Google’s TPUv4i: Industrial Product

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.510404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ab3b5004f660570e727d9a11cae548d500850c12895108a033feda33b905c3d7

Observation 18815ced-57ba-4c65-9637-f4cad0e07032 · outbound

This paper cites SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.512917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:11a54573530792a7b70a9f503eaf7b41ad9eefb6ddbaf76fd7beb8e12d6868c3

Observation 98a5828b-bcab-47f5-b38c-8744fa088213 · outbound

This paper cites Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.521864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:76dee7f48c8adde71d56a0074f5944b40494f07f48b84575e4f53cb95fb19f08

Observation 4119634a-03db-4c84-8e7c-bf3219ed4ba0 · outbound

This paper cites Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.540602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:051df41c9910fd592f16324314cfb7a11815a5738b92b41aa443236ed6fe2833

Observation 13268a75-31ce-41c5-afb3-3b5db08816e7 · outbound

This paper cites Pimba: A Processing- in-Memory Acceleration for Post-Transformer Large Language Model Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Pimba: A Processing- in-Memory Acceleration for Post-Transformer Large Language Model Serving

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.560520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:23f206faf7bbe87900797a0ad6d27c455a68076092c7b2ebc776d1b635bda743

Observation 4bed79c1-8e12-4a1a-9097-fb40f4493735 · outbound

This paper cites Tender: Accelerating Large Language Mod- els via Tensor Decomposition and Runtime Requantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Tender: Accelerating Large Language Mod- els via Tensor Decomposition and Runtime Requantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.495865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ddd807e663a0ab3c23c2feafc83992206d59b5233f89616c5fe3ee35c2fdeb19

Observation 2d268b78-fc49-4542-b8da-6d5404da4f2d · outbound

This paper cites MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.505237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:478cb2f8046a74e12f09bf722dd30453a529793d2a470ea1716738c17dd45a21

Observation cfe3a37d-6b99-4a4b-9a51-624b39404277 · outbound

This paper cites A 1ynm 1.25V 8Gb 16Gb/s/Pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep Learning Application.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats A 1ynm 1.25V 8Gb 16Gb/s/Pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep Learning Application

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.491616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:0854efea2bcd153b08bd856ad74a8564ecdf20192d00e5df0d18f05ac8068c2c

Observation 86221d5b-e162-42e8-ab32-738c4119ca63 · outbound

This paper cites Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.486287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d85b450d0e4379c3fe451b1b2af251b13d84a2995b288deac6ae6615f7a84109

Observation 910caeb7-a98b-4182-a0ef-9451d363122b · outbound

This paper cites H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.488924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:3421cfd5d5a19c2dbc897c3b80f7cee0ed78a05ab3d6d5c15bc0775783c8da42

Observation f029eb9a-5252-4338-bc54-5c292db6e0fe · outbound

This paper cites ORCHES: Orchestrated Test-Time- Compute-based LLM Reasoning on Collaborative GPU-PIM HEteroge- neous System.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ORCHES: Orchestrated Test-Time- Compute-based LLM Reasoning on Collaborative GPU-PIM HEteroge- neous System

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.507758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:89c7825fa0cc81cd62a8be37416ea0815021cf46b8a1805b3cb7de51953bd334

Observation a7c7dd20-0c88-4026-adb0-4a3e17011c33 · outbound

This paper cites AWQ: Activation-aware Weight Quan- tization for LLM Compression and Acceleration.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AWQ: Activation-aware Weight Quan- tization for LLM Compression and Acceleration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.483421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:062928b9f6db47ce2726ad22dc36ef536cd8f9bdf293ea726077236890522946

Observation 7eb3f5a8-fed2-4360-b683-6b1fbe435314 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.463754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7494eed26f66499d1901a570f10938fdfa2e97e03348b6396645dbe250107275

Observation d00ef8dd-72a6-4f6e-aa24-477d713a4d0a · outbound

This paper cites SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Ef- ficient Encoding.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Ef- ficient Encoding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.467168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7c01cc9a48c2332550fc4e53b344b1d17f26ee98e2a0762838397b8ee0556719

Observation 4b9b8b7c-91c6-44dc-a0c2-0e6b37189d8c · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.470279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:f7938b6dcddf187498812668955608a306d88f6d3146c5a1f70a3a1aba5a55ae

Observation 370a2073-acda-4e41-a9a1-f4bbd43360f2 · outbound

This paper cites Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.479541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b32bc61267c0d9827577e21f60a9233adfa48d957a0ab7b261c08c117d6b1c17

Observation f560dd81-1e11-4663-be9e-d10ed09d5862 · outbound

This paper cites Pointer sentinel mixture models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Pointer sentinel mixture models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.498845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:0463309a7c499eb66c7185d734b04bd1f64ca9ac78581eec1e414c627ca8c921

Observation 04b1a893-0387-4d43-9ed3-4cbe65cf78bd · outbound

This paper cites Introducing Llama 3.1: Our most capable models to date.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Introducing Llama 3.1: Our most capable models to date

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.452529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d8ab760c977cf7886f6411191eadef5084425938708e32fbf5b0767b2b27fc36

Observation 5b87c798-2369-44d1-8b9a-d1c7a15dd751 · outbound

This paper cites Llama-3.2-90B-Vision-Instruct.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Llama-3.2-90B-Vision-Instruct

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.458550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:82069e75fb189a51a292884b2d3b8a86a6c2529a6ea602da36a2e003290f1ed0

Observation b2b80751-7dd5-49d2-8f89-c2600d8ab179 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.448130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d0e2bfd474db82a7d10ad6b83abc4885caf2db0c24a1868c4de08ae5f34bdd1c

Observation e16d8a53-e957-4f1f-8ae8-234a62a67c34 · outbound

This paper cites Meta Llama 2.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Meta Llama 2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.455706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ec9a001d37be7da2d4fb2fe51065dc232f39211679d4370536c1ec52663afa73

Observation 474e1aa6-4ec8-40a1-b4fe-a129476d8109 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.528443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:1a0e9f11652d00cdb85a598269369f52b3d9779c9484cf512dda23018a028689

Observation bb3509a2-068d-4a9b-b2d5-f5c029e1ec85 · outbound

This paper cites FP8 Formats for Deep Learning.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FP8 Formats for Deep Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.216570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d60e7bb23b833f7faf4e7f7e2fda60cc7093a171fa7b101e01bc41ee5a978941

Observation 5093e4e9-adfc-4cd1-828d-101a3284c8da · outbound

This paper cites Introducing NVFP4 for Efficient and Accurate Low-Precision Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Introducing NVFP4 for Efficient and Accurate Low-Precision Inference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.437286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:421ca5d56a6cb993d73b994306d92bd94c31bec894b2647a83fe4ae1c69b099b

Observation c1b5315e-0c17-41a2-9933-d09ccb181804 · outbound

This paper cites NVIDIA Blackwell GPU Architecture.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats NVIDIA Blackwell GPU Architecture

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.441718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:921453ff76cf01acc2833f154858fd3531b085892bbf79e20cde4ea77897370d

Observation ac45eb04-401d-49ec-99e3-160476b0c899 · outbound

This paper cites Openai o3-mini.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Openai o3-mini

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.429274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b1ccefb331cb36179c40e18517828478108b30ca2b9fd1626c17c7565c968c07

Observation 93764bbe-b56a-4fb8-91c1-507055548e46 · outbound

This paper cites Gsm8k dataset.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Gsm8k dataset

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.426603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:171c46c6eb40123116274cefae3d0ba54f2071e19dc34630d957cda15894c322

Observation ec768789-ed4b-46d9-b381-57899850a37f · outbound

This paper cites FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.393649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:bb7ad63c1b3dbb9db84e62feb161e9835854b17055a64248e61e7f5cd6d9e187

Observation 93f03422-2c83-4024-b65d-1f3bfbff073f · outbound

This paper cites AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.444850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:2f024ffb69e0bde506fc203d99954ede3e6c02cf7322ea509d017a96cf8fc375

Observation 913f3582-5df7-48b6-83e0-a79860ba73bc · outbound

This paper cites MicroScopiQ: Ac- celerating Foundational Models through Outlier-Aware Microscaling Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats MicroScopiQ: Ac- celerating Foundational Models through Outlier-Aware Microscaling Quantization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.354666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:df35693d629d4d8a31990153091aa6e6cf1e83ab72fa2200d7b696dade8f71ee

Observation 82a7c380-c6ea-4179-8f22-dd0fa2eceba7 · outbound

This paper cites With Shared Microexponents, A Little Shifting Goes a Long Way.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats With Shared Microexponents, A Little Shifting Goes a Long Way

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.421658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:3d90571adc69158941f9a444352c8fe1cdffd3b0e01ee4ee23b97ea930d939fc

Observation bf252da6-c9db-4ab1-a5d5-5bf9851bd74e · outbound

This paper cites IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.424042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:a80e7babd6f582d742a5776fe4fa3faabeea41a8f72ccc9491386bc26b281e80

Observation bbb88c1a-9395-482d-8aad-0a3394ea7e9c · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.418893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b0cedc153a39cc79f8ee50a25cc35d8453f4464474833ca425767b71016f6688

Observation dfc59e7d-6eae-497f-9578-b4905e295fd1 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.361607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:c1b2da9f9ee9a966c91e9c2873c5d110d0a872a27b102e676a566c088150f245

Observation 461f9047-e068-43af-98fe-113a9a93a2bf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LLaMA: Open and Efficient Foundation Language Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.170271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:132a54a889cab83b1260b643d9507b6480fa42739649cb84232d921990ad0c3a

Observation a0879e73-781c-4dcd-a536-8001cd9782b6 · outbound

This paper cites FP8 versus INT8 for efficient deep learning inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FP8 versus INT8 for efficient deep learning inference

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:20:32.176875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:864bd57649849d798dea179d3394691867bd59efc1a6e6835992603eab999322

Observation 3c4815e1-f3c6-4828-9b6f-ad776c5e7580 · outbound

This paper cites ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:20:32.183663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:041e3108d92065648df1da9e9b154c245eff21c7b71846b76974eefbf94cfc77

Observation 49cc89dc-9e41-440a-8596-e2db780c35d4 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.413716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ae5c09491d24a3ed666bf7076b10058d4470551fa7fd3170326ec42df496f2b5

Observation 2a03912f-72ba-44c6-9c7c-0db39722b8e6 · outbound

This paper cites Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.416502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:561c7b475149ed921a607961c94d7eac1aedc5cb2ea20c5c4cf0a8f63b6f67a6

Observation 5eb34fdd-7deb-4ff6-aee9-ffd31afc069a · outbound

This paper cites Qwen2.5 Technical Report.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Qwen2.5 Technical Report

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.188473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:aa65e2cbfe14da439c1c79579d414e5026d74af30a723bcf9bf1b42e493cc087

Observation 6aca5194-9473-4521-8077-568ed4212ed6 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.411176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:705a3cba253b4ef666bf5166882feba9e97d94ced4c77368147dd35e3bccbda9

Observation a23971a4-014b-412e-b08c-da7d65ce5ceb · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batch- ing.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batch- ing

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.400872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:73888c5f337997734223248f4b29e955e07b6597a73f9c49b751fd56817ec078

Observation 7cce09f4-7b68-494d-bc38-8972158ae4e1 · outbound

This paper cites SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.396613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d1bcc54573b58131e6f3dc72c9ac5a3d620eccd0de82b86303c7f3d2c0a84926

Observation 0e81c5fb-20f9-4e49-a0a4-3ed5e59f177f · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput- optimized Large Language Model Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats DistServe: Disaggregating Prefill and Decoding for Goodput- optimized Large Language Model Serving

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.406530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:f04d3ba9b5dae602008ca39d8c9a8808e7557d74d6f02e050697f4edec5c6165

Observation f067f8df-4765-48fc-aab2-61a31f82e5ad · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats A Survey on Efficient Inference for Large Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.198469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:1cbae9684790fa5ca14f5dee1b1ab4d0f3440389d224cbd4cba036d69c3ade38

Pith citing papers

Observation 6e5d9942-adc4-496a-9eb2-3d3afd6144d7 · inbound

CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device cites this paper.

CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:52:39.415504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:52:39.415504Z digest=sha256:a118f1e33808e23989fcecca4709ca8202b3ef33811e63183cf271092a65c8cd

Observation 3c97f8a5-465f-478d-86b0-258194387a0e · inbound

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference cites this paper.

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T22:36:29.218635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:36:29.218635Z digest=sha256:1810c220a77f33f06d1f3dcd11f5ff6626135d98a1c73d4add03ba03d139e6b3