Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:16:27.497363Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 2 inbound Pith citation observations for arXiv:2511.06838.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:16:27.497363Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T09:52:39.415504Z
A source-named dated measurement, never combined with another source.
Source: cited_works
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 91c0cbda-410f-4f64-9607-3dcedc025466 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AMD INSTINCT™ MI350X GPU
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0466a4a0-b073-41c4-90d2-aa0ad4611b80 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ffe8591-3f1b-422f-be20-b97d8420901e · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9058645-61af-400b-9d8b-72c2ae33df5f · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ca165b9-7a3c-442c-ad7f-7a252ee8e115 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats CACTI 7: New tools for interconnect exploration in innovative off-chip memories
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 86fd75e1-3aa8-4d7a-be30-4c2a9b166c8f · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Language Models are Few-Shot Learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94742906-3c9b-4726-92fd-5aafa83d2ef5 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats BitMoD: Bit-serial Mixture-of- Datatype LLM Acceleration
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3bfbabd8-5e52-45fe-956a-2cf4d8fd5442 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ecco: Improving Memory Band- width and Capacity for LLMs via Entropy-Aware Cache Compression
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b80588a-b746-4f55-9cb6-9fa6eb4302f7 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a46fa4d-337c-4a22-a73d-0f95d19bc213 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats DeepSeek R1
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ce4bcba-230c-4174-a0e2-6326a867e548 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d89f3c01-9891-409b-ae6c-ab27d4be13f7 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The true Processing In Memory accelerator
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0aa2b922-6b06-47ce-a15c-99b451485171 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4d617caa-5ed4-4368-a3c5-f0e4732d6285 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92c72bcc-f4ed-4cf7-82c9-048c642bef3b · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd0a485e-3ccd-4422-9514-ede7d9e78e35 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe1e1bfc-e012-415b-a232-62872a374461 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7d3720c-d2c1-4338-80dc-2e0477ee3806 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 730c8e44-014a-4199-bf83-7497416de775 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3e9ca14-f00a-4855-871a-65d16eeecdcd · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a301124c-f97e-48c7-bce5-84946469ebf0 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f4ae97eb-9a0a-448a-a701-865a7cfd32ab · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Newton: A DRAM-maker’s Accelerator-in- Memory (AiM) Architecture for Machine Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4258aa8e-5cc0-4a2e-89a0-78261b683fac · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture- Dataflow Co-Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7636e3b1-8978-4041-a7ac-96c90721699b · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats How Would the Viewer Feel? Estimating Wellbeing from Video Scenarios
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 223b3dd1-693a-4b74-bb2a-7b2810876f34 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec386510-1f74-4678-8581-679a8b1da606 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b18f5743-db24-4491-acf6-16b565ecb9f0 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats M-ANT: Efficient Low-bit Group Quantization 13 for LLMs via Mathematically Adaptive Numerical Type
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 80654548-de0e-471d-8be6-49fe29b618bc · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats PLAIN: Leveraging High Internal Bandwidth in PIM for Accelerating Large Language Model Inference via Mixed-Precision Quantization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51813f2e-16d3-4639-ac90-41c9a433f04b · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FIGNA: Integer unit-based accelerator design for fp-int gemm preserving numerical accuracy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4d924579-6091-423b-b186-3eb49ad1d2e8 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da94f300-dd5a-4261-a94a-4a3b8d313b0b · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory DRAM
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2529d643-ad30-41e9-8c70-a908eea353ea · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory (HBM3) DRAM
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b29bf160-1749-4b53-921e-7d66abece790 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory (HBM4) DRAM
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3200a273-e128-414a-9922-878b8eddc9a0 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Mistral 7B
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 57c2c5a6-3de8-43ed-80b8-51209280a230 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ten Lessons From Three Generations Shaped Google’s TPUv4i: Industrial Product
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 18815ced-57ba-4c65-9637-f4cad0e07032 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98a5828b-bcab-47f5-b38c-8744fa088213 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4119634a-03db-4c84-8e7c-bf3219ed4ba0 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13268a75-31ce-41c5-afb3-3b5db08816e7 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Pimba: A Processing- in-Memory Acceleration for Post-Transformer Large Language Model Serving
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4bed79c1-8e12-4a1a-9097-fb40f4493735 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Tender: Accelerating Large Language Mod- els via Tensor Decomposition and Runtime Requantization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d268b78-fc49-4542-b8da-6d5404da4f2d · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cfe3a37d-6b99-4a4b-9a51-624b39404277 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats A 1ynm 1.25V 8Gb 16Gb/s/Pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep Learning Application
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 86221d5b-e162-42e8-ab32-738c4119ca63 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 910caeb7-a98b-4182-a0ef-9451d363122b · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f029eb9a-5252-4338-bc54-5c292db6e0fe · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ORCHES: Orchestrated Test-Time- Compute-based LLM Reasoning on Collaborative GPU-PIM HEteroge- neous System
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7c7dd20-0c88-4026-adb0-4a3e17011c33 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AWQ: Activation-aware Weight Quan- tization for LLM Compression and Acceleration
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7eb3f5a8-fed2-4360-b683-6b1fbe435314 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d00ef8dd-72a6-4f6e-aa24-477d713a4d0a · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Ef- ficient Encoding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b9b8b7c-91c6-44dc-a0c2-0e6b37189d8c · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 370a2073-acda-4e41-a9a1-f4bbd43360f2 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f560dd81-1e11-4663-be9e-d10ed09d5862 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Pointer sentinel mixture models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04b1a893-0387-4d43-9ed3-4cbe65cf78bd · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Introducing Llama 3.1: Our most capable models to date
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b87c798-2369-44d1-8b9a-d1c7a15dd751 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Llama-3.2-90B-Vision-Instruct
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2b80751-7dd5-49d2-8f89-c2600d8ab179 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e16d8a53-e957-4f1f-8ae8-234a62a67c34 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Meta Llama 2
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 474e1aa6-4ec8-40a1-b4fe-a129476d8109 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bb3509a2-068d-4a9b-b2d5-f5c029e1ec85 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FP8 Formats for Deep Learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5093e4e9-adfc-4cd1-828d-101a3284c8da · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Introducing NVFP4 for Efficient and Accurate Low-Precision Inference
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c1b5315e-0c17-41a2-9933-d09ccb181804 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats NVIDIA Blackwell GPU Architecture
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ac45eb04-401d-49ec-99e3-160476b0c899 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Openai o3-mini
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93764bbe-b56a-4fb8-91c1-507055548e46 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Gsm8k dataset
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec768789-ed4b-46d9-b381-57899850a37f · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93f03422-2c83-4024-b65d-1f3bfbff073f · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 913f3582-5df7-48b6-83e0-a79860ba73bc · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats MicroScopiQ: Ac- celerating Foundational Models through Outlier-Aware Microscaling Quantization
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 82a7c380-c6ea-4179-8f22-dd0fa2eceba7 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats With Shared Microexponents, A Little Shifting Goes a Long Way
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bf252da6-c9db-4ab1-a5d5-5bf9851bd74e · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bbb88c1a-9395-482d-8aad-0a3394ea7e9c · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dfc59e7d-6eae-497f-9578-b4905e295fd1 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 461f9047-e068-43af-98fe-113a9a93a2bf · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LLaMA: Open and Efficient Foundation Language Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0879e73-781c-4dcd-a536-8001cd9782b6 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FP8 versus INT8 for efficient deep learning inference
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c4815e1-f3c6-4828-9b6f-ad776c5e7580 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49cc89dc-9e41-440a-8596-e2db780c35d4 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a03912f-72ba-44c6-9c7c-0db39722b8e6 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5eb34fdd-7deb-4ff6-aee9-ffd31afc069a · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Qwen2.5 Technical Report
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6aca5194-9473-4521-8077-568ed4212ed6 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a23971a4-014b-412e-b08c-da7d65ce5ceb · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batch- ing
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7cce09f4-7b68-494d-bc38-8972158ae4e1 · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0e81c5fb-20f9-4e49-a0a4-3ed5e59f177f · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats DistServe: Disaggregating Prefill and Decoding for Goodput- optimized Large Language Model Serving
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f067f8df-4765-48fc-aab2-61a31f82e5ad · outbound
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats A Survey on Efficient Inference for Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e5d9942-adc4-496a-9eb2-3d3afd6144d7 · inbound
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c97f8a5-465f-478d-86b0-258194387a0e · inbound
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.