Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:55:39.939185Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2504.14365.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:55:39.939185Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation feff1650-5512-4333-bc7f-d613e2abc38b · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb08ee07-d818-4ab1-b3c2-7509d7755767 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f01a62a1-c212-41c3-b459-786b31024d0b · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Piqa: Reasoning about physical commonsense in natural language,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ba9cf5-9096-4800-9b53-67fa94a1226b · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f049cc-0091-4420-bcbf-285874938aea · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 15.3 a 65nm 3t dynamic analog ram- based computing-in-memory macro and cnn accelerator with retention enhancement, adaptive analog sparsity and 44tops/w system energy efficiency,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8409119a-f975-465b-92e1-2547f0b4589c · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be94c62b-0908-4d96-8646-dc9d7216ef70 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7c3575eb-77e7-44bf-9365-960f9e3a1da8 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Integrating nvidia deep learning accelerator (nvdla) with risc-v soc on firesim,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 816f2d6a-c339-4941-aac0-aa102802f77c · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator SparseGPT: Massive language models can be accurately pruned in one-shot,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b50e1cca-8265-4adf-8ac7-9c6fd2389271 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A 5-nm 254-tops/w 221-tops/mm 2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3ae0549c-5f48-418f-a0ee-4e3340d4174c · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a8ebf6-d390-4981-8c03-809a8c1ac92f · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8dfdcf7-c4c0-4172-b270-c6554b6ac0a7 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Sparsity-aware and re-configurable npu architecture for samsung flagship mobile soc,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 22bda09e-5628-4bf7-b578-9f922554373a · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40dc742b-5d3b-4b12-b8e4-8ba803f9102a · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Mixtral of Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b20204-95bd-4824-bfdb-ebdf3ac01d1f · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Colonnade: A reconfigurable sram-based digital bit- serial compute-in-memory macro for processing neural networks,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a7f4da27-88c5-4cdb-8141-86dfb17bd647 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Laxor: A bit-accurate bnn accelerator with latch-xor logic for local computing,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c664d95e-6800-4dfa-be1a-f0650edc5944 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef034e36-283e-4d01-8115-93165073cefc · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 33.2 a fully integrated analog reram based 78.4 tops/w compute-in-memory chip with fully parallel mac computing,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fec4ce2b-e162-4afd-a23a-e209b548fb74 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebbceff-bbea-4cc7-86c4-0f8533207d63 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator An Analysis of Neural Language Modeling at Multiple Scales
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7740e6a5-7649-41cc-85b7-70e94a6de7eb · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Introducing meta llama 3: The most capable openly available llm to date,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741c9a68-bb6f-446f-894f-8d427b1382f5 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Accelerating Sparse Deep Neural Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9a6263-ec33-4302-81e1-78e97098e5bf · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Scnn: An accelerator for compressed-sparse convolutional neural networks,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3d9673d9-00f4-47d6-8dc6-4566fa98def2 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2250370-b668-487f-bbe1-dd0c694b9875 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b1bd7b-ebd7-4b34-b300-c6fb91a465ae · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27eab695-f837-4151-9afc-9d683e024110 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b805c6e2-3702-4750-a9e2-3bda45d77c94 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Dnnweaver: From high-level deep network models to fpga acceleration,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01979481-af26-4f28-a4b9-bc31527625f4 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Towards vqa models that can read,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922dc18b-f23d-47dc-9312-f5e1d34ba640 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Sp-imc: A sparsity aware in-memory-computing macro in 28nm cmos with configurable sparse representation for highly sparse dnn workloads,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2c59a99b-55cc-4004-81b3-792439f849a2 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A fully-digital and row-pipelined compute-in-memory neural network accelerator with soc-level benchmarking for ar/vr applications,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f9baaa9-8a58-42b3-b106-8c9f5798265a · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A simple and effective pruning approach for large language models,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a266c4-d203-4ace-891a-bf43674befb2 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b6dd9d-a4d4-471b-ba1c-a5fa3e2f6be5 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Sdp: Co-designing algorithm, dataflow, and architecture for in-sram sparse nn acceleration,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6e68ec14-afeb-4b99-89cf-52cb180c80c8 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Highlight: Efficient and flexible dnn acceleration with hierarchical structured sparsity,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1c7ab78a-537e-42c3-8244-0f2ba667a316 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A task-centric angle of llm pre-trained weights through sparsity,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 331331cc-b9bd-4f5b-ad0d-bbc9a4105c54 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Outlier weighed layerwise sparsity: A missing secret sauce for pruning llms to high sparsity,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a168e307-4eb4-429e-a4a5-50c6103df820 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Compute-in-memory chips for deep learning: Recent trends and prospects,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 080ca778-31e3-4487-8262-2b95ec81283d · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 15.2 a 2.75-to-75.9 tops/w computing-in-memory nn processor supporting set-associate block-wise zero skipping and ping- pong cim with simultaneous computation and weight updating,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b9bce9e-a36b-402e-98a8-026e8992f655 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96b563b-51a0-4dbc-a7bc-9693216e49b0 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A 28-nm 18.7 tops/mm2 89.4-to-234.6 tops/w 8b single-finger edram compute- in-memory macro with bit-wise sparsity aware and kernel-wise weight update/refresh,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation feb2bccb-3a17-4750-8c1c-da429d3217f3 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3139199b-6603-4da6-be46-fcbe77bdd443 · outbound
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A digital sram computing-in-memory design utilizing activation unstructured sparsity for high-efficient dnn inference,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.