Pith. sign in

Paper Citation Record · LEDGER

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2504.14365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14365 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:55:39.939185Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation feff1650-5512-4333-bc7f-d613e2abc38b · outbound

This paper cites Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.720185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.720185Z digest=sha256:dcd413290a5b0b2096261b6dc80bb86746736c75be894a555bacb479ccd95364

Observation fb08ee07-d818-4ab1-b3c2-7509d7755767 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.726020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.726020Z digest=sha256:246ae2af2a3d84bf48f5f7a05a765d06020c0fdc254990b14d043675ea58059d

Observation f01a62a1-c212-41c3-b459-786b31024d0b · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Piqa: Reasoning about physical commonsense in natural language,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.731716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.731716Z digest=sha256:44107b96f7371b92e0b3ecad86eea014855d90391210e74fa7e7020f5e79da03

Observation c7ba9cf5-9096-4800-9b53-67fa94a1226b · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.736357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.736357Z digest=sha256:e328666a736d1697bab3db8f428f94eef673a2dd9fc888b433a4212d049d687a

Observation 28f049cc-0091-4420-bcbf-285874938aea · outbound

This paper cites 15.3 a 65nm 3t dynamic analog ram- based computing-in-memory macro and cnn accelerator with retention enhancement, adaptive analog sparsity and 44tops/w system energy efficiency,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 15.3 a 65nm 3t dynamic analog ram- based computing-in-memory macro and cnn accelerator with retention enhancement, adaptive analog sparsity and 44tops/w system energy efficiency,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.510595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.741431Z digest=sha256:8b2591f1dd6719576fdc29e4736dde52ebe9ebeab413de12269376554185ad83

Observation 8409119a-f975-465b-92e1-2547f0b4589c · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.745841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.745841Z digest=sha256:85be27d5dea0080bfbedd45eb7ea183ce1a543ceac441341fc7f56499e03d8af

Observation be94c62b-0908-4d96-8646-dc9d7216ef70 · outbound

This paper cites Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:55:40.122741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.751705Z digest=sha256:624f568b5688dfb190c6ccac85dab2c9c8282341ccb10fbfb3e28b2fd659f315

Observation 7c3575eb-77e7-44bf-9365-960f9e3a1da8 · outbound

This paper cites Integrating nvidia deep learning accelerator (nvdla) with risc-v soc on firesim,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Integrating nvidia deep learning accelerator (nvdla) with risc-v soc on firesim,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.496947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.756137Z digest=sha256:07382149cd4722dcdb3992d231e04688fdca329230f91745fd1781562bba92ee

Observation 816f2d6a-c339-4941-aac0-aa102802f77c · outbound

This paper cites SparseGPT: Massive language models can be accurately pruned in one-shot,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator SparseGPT: Massive language models can be accurately pruned in one-shot,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.479748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.760266Z digest=sha256:0b71849e43ac58346882cb7e907f4972089ac4fa1119f416fc62fbdecfabbd43

Observation b50e1cca-8265-4adf-8ac7-9c6fd2389271 · outbound

This paper cites A 5-nm 254-tops/w 221-tops/mm 2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A 5-nm 254-tops/w 221-tops/mm 2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.463964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.764168Z digest=sha256:fd8cf67aa975ad3c2c6f99dd52a4b4afd63d76adb7a16bc8c757bc2ddd9ca61d

Observation 3ae0549c-5f48-418f-a0ee-4e3340d4174c · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.769632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.769632Z digest=sha256:5a726732599dcf955d11537e20ed4c7d8ac2a66bba50024ca0a710b6e13ccfab

Observation e8a8ebf6-d390-4981-8c03-809a8c1ac92f · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.776256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.776256Z digest=sha256:6ac02a41df43927d60da5d7cf851faf7bf92ec18d2745dbfe53305294b7235b6

Observation f8dfdcf7-c4c0-4172-b270-c6554b6ac0a7 · outbound

This paper cites Sparsity-aware and re-configurable npu architecture for samsung flagship mobile soc,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Sparsity-aware and re-configurable npu architecture for samsung flagship mobile soc,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.445608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.782007Z digest=sha256:18cebd408303d97d774119dc2c76d117708d40c645d38c0926fed20f5ecb3772

Observation 22bda09e-5628-4bf7-b578-9f922554373a · outbound

This paper cites Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.429527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.786250Z digest=sha256:4040b18893ba4e9b6c30b51f7264813de21dd927b2722e78f2d6ab36f94cb715

Observation 40dc742b-5d3b-4b12-b8e4-8ba803f9102a · outbound

This paper cites Mixtral of Experts.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.790833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.790833Z digest=sha256:691e654494120fac6b3f736e8a8b31fde39ef6c53dd040eadfa1b5c44542dc94

Observation 29b20204-95bd-4824-bfdb-ebdf3ac01d1f · outbound

This paper cites Colonnade: A reconfigurable sram-based digital bit- serial compute-in-memory macro for processing neural networks,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Colonnade: A reconfigurable sram-based digital bit- serial compute-in-memory macro for processing neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.416050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.797376Z digest=sha256:18cdebf59e287ce645845c96bcc952cca956a534a81c009282f07f500bbd3d03

Observation a7f4da27-88c5-4cdb-8141-86dfb17bd647 · outbound

This paper cites Laxor: A bit-accurate bnn accelerator with latch-xor logic for local computing,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Laxor: A bit-accurate bnn accelerator with latch-xor logic for local computing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.394018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.801646Z digest=sha256:d1ffb093b2e768e85c3b64d7063cb5ff97948c688f5ae8702d2ddbd1a5b81652

Observation c664d95e-6800-4dfa-be1a-f0650edc5944 · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.807957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.807957Z digest=sha256:6c7f1f9e6de996d21446fc12a6c8404bca20a831c9cf1c5fec4cca179cb0f1ea

Observation ef034e36-283e-4d01-8115-93165073cefc · outbound

This paper cites 33.2 a fully integrated analog reram based 78.4 tops/w compute-in-memory chip with fully parallel mac computing,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 33.2 a fully integrated analog reram based 78.4 tops/w compute-in-memory chip with fully parallel mac computing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.368998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.814513Z digest=sha256:2174b24c6c2b70ecb1ecc299297eff24408e25c44141531938db8bea9830d3ea

Observation fec4ce2b-e162-4afd-a23a-e209b548fb74 · outbound

This paper cites S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.825661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.825661Z digest=sha256:ab4ec9902796541312bac5685ab12f7d3cc25fc15fcdfb03b878ad906317ac15

Observation cebbceff-bbea-4cc7-86c4-0f8533207d63 · outbound

This paper cites An Analysis of Neural Language Modeling at Multiple Scales.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator An Analysis of Neural Language Modeling at Multiple Scales

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.830657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.830657Z digest=sha256:3ca867f5f9b17a135a3d958db90d4a5b1b0974b482f5744638ce270744ef3f2d

Observation 7740e6a5-7649-41cc-85b7-70e94a6de7eb · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Introducing meta llama 3: The most capable openly available llm to date,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.835343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.835343Z digest=sha256:7ecd3eb4beec58fb8c91485b9859b8695fd0be5836f308c0cbeceb4bcc7763f5

Observation 741c9a68-bb6f-446f-894f-8d427b1382f5 · outbound

This paper cites Accelerating Sparse Deep Neural Networks.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Accelerating Sparse Deep Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.839688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.839688Z digest=sha256:646f3cabf6bcef89567fa38f257859cc1e446fbf35fa60f31a8db7b97323af02

Observation ea9a6263-ec33-4302-81e1-78e97098e5bf · outbound

This paper cites Scnn: An accelerator for compressed-sparse convolutional neural networks,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Scnn: An accelerator for compressed-sparse convolutional neural networks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.338531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.844740Z digest=sha256:f341f0f44de35997553f0cdd412cbb2ac20d80de0990bc799965d6448e36f51d

Observation 3d9673d9-00f4-47d6-8dc6-4566fa98def2 · outbound

This paper cites FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.849507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.849507Z digest=sha256:bfa510cceff654b846f2a8ab394a60a02296cc0afa55de80098eacd34bff06e1

Observation e2250370-b668-487f-bbe1-dd0c694b9875 · outbound

This paper cites MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.854178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.854178Z digest=sha256:5d1d313300c230566463ecc622f27f7d8bc269eeb23eee280609d5aeed53735f

Observation 19b1bd7b-ebd7-4b34-b300-c6fb91a465ae · outbound

This paper cites Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.859331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.859331Z digest=sha256:64ac6d55ca6b7bcdf8a0f43636b917eef4328ab1c4ddb1cf2f26181c29f2fbe0

Observation 27eab695-f837-4151-9afc-9d683e024110 · outbound

This paper cites Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.324944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.863524Z digest=sha256:1c915e16a2eb814366b207c8fb56a793e3ccb67dbe7c99b42baf95724d85e080

Observation b805c6e2-3702-4750-a9e2-3bda45d77c94 · outbound

This paper cites Dnnweaver: From high-level deep network models to fpga acceleration,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Dnnweaver: From high-level deep network models to fpga acceleration,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.868416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.868416Z digest=sha256:4250dc7bcc30562092be11886cff7adbf2c8f1445e66d52694e3bea6ab613a05

Observation 01979481-af26-4f28-a4b9-bc31527625f4 · outbound

This paper cites Towards vqa models that can read,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Towards vqa models that can read,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.874234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.874234Z digest=sha256:743c4f5722635db855454e5d6357c5efd8aefd276566b08259b751906534871d

Observation 922dc18b-f23d-47dc-9312-f5e1d34ba640 · outbound

This paper cites Sp-imc: A sparsity aware in-memory-computing macro in 28nm cmos with configurable sparse representation for highly sparse dnn workloads,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Sp-imc: A sparsity aware in-memory-computing macro in 28nm cmos with configurable sparse representation for highly sparse dnn workloads,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.298516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.879190Z digest=sha256:efa1f0f4eac521027ba15399eae4365cff1642a2ed11ee6e632b94f5ffa2c4ec

Observation 2c59a99b-55cc-4004-81b3-792439f849a2 · outbound

This paper cites A fully-digital and row-pipelined compute-in-memory neural network accelerator with soc-level benchmarking for ar/vr applications,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A fully-digital and row-pipelined compute-in-memory neural network accelerator with soc-level benchmarking for ar/vr applications,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.883551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.883551Z digest=sha256:2343487f860ed904d4ec9ec4d810584f021c22abac944ffa0525fcbbb3e0708a

Observation 0f9baaa9-8a58-42b3-b106-8c9f5798265a · outbound

This paper cites A simple and effective pruning approach for large language models,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A simple and effective pruning approach for large language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.887458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.887458Z digest=sha256:d206463ffd08c6fcc3ad4b35beb9749790a05c4f45490b22ad5f7ca481e7d0f2

Observation 90a266c4-d203-4ace-891a-bf43674befb2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.891688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.891688Z digest=sha256:df7253a3dbc8cb57daee6cc40f2c264053d9876ee7269e97b367824e1c258f46

Observation 79b6dd9d-a4d4-471b-ba1c-a5fa3e2f6be5 · outbound

This paper cites Sdp: Co-designing algorithm, dataflow, and architecture for in-sram sparse nn acceleration,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Sdp: Co-designing algorithm, dataflow, and architecture for in-sram sparse nn acceleration,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.268520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.896547Z digest=sha256:d3400c292a940039cf743edc99559441b785aee8fbbc9ea2ef5ec940d9671d33

Observation 6e68ec14-afeb-4b99-89cf-52cb180c80c8 · outbound

This paper cites Highlight: Efficient and flexible dnn acceleration with hierarchical structured sparsity,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Highlight: Efficient and flexible dnn acceleration with hierarchical structured sparsity,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.254592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.900849Z digest=sha256:33adfce9ef6509658e3552c65bf2dd59e0f908501a6ca1aa04286ab7a494f898

Observation 1c7ab78a-537e-42c3-8244-0f2ba667a316 · outbound

This paper cites A task-centric angle of llm pre-trained weights through sparsity,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A task-centric angle of llm pre-trained weights through sparsity,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.241471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.904752Z digest=sha256:457f7e04d5e7bc6ad8478cb9dbdfdc591b5aa4a6d71a0fc09f1d05651a1ceaca

Observation 331331cc-b9bd-4f5b-ad0d-bbc9a4105c54 · outbound

This paper cites Outlier weighed layerwise sparsity: A missing secret sauce for pruning llms to high sparsity,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Outlier weighed layerwise sparsity: A missing secret sauce for pruning llms to high sparsity,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.227953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.908929Z digest=sha256:8b84ec3d9d4f3c6b47a5ce8e7bef5534bacccc8dbc823383f74c15dcf08d8e76

Observation a168e307-4eb4-429e-a4a5-50c6103df820 · outbound

This paper cites Compute-in-memory chips for deep learning: Recent trends and prospects,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Compute-in-memory chips for deep learning: Recent trends and prospects,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.913686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.913686Z digest=sha256:ef5ba820e9d07e0ca3b1862a049aa218048a288646811866bc94aad8174b993b

Observation 080ca778-31e3-4487-8262-2b95ec81283d · outbound

This paper cites 15.2 a 2.75-to-75.9 tops/w computing-in-memory nn processor supporting set-associate block-wise zero skipping and ping- pong cim with simultaneous computation and weight updating,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 15.2 a 2.75-to-75.9 tops/w computing-in-memory nn processor supporting set-associate block-wise zero skipping and ping- pong cim with simultaneous computation and weight updating,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.208706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.920416Z digest=sha256:a3287c1a4b7e84aeb440c7dd85f9c9fb4f3e89724539e1ae6e5774a5c1d882e8

Observation 9b9bce9e-a36b-402e-98a8-026e8992f655 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.925019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.925019Z digest=sha256:d9e49fcbc47419dba69526d81f42b25a48e8917ded75d411b1911d7857c40c23

Observation d96b563b-51a0-4dbc-a7bc-9693216e49b0 · outbound

This paper cites A 28-nm 18.7 tops/mm2 89.4-to-234.6 tops/w 8b single-finger edram compute- in-memory macro with bit-wise sparsity aware and kernel-wise weight update/refresh,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A 28-nm 18.7 tops/mm2 89.4-to-234.6 tops/w 8b single-finger edram compute- in-memory macro with bit-wise sparsity aware and kernel-wise weight update/refresh,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.196376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.929262Z digest=sha256:4fef57c1abcccab207d14155f048d034388b9e7d0249e052a67ad69a4143334c

Observation feb2bccb-3a17-4750-8c1c-da429d3217f3 · outbound

This paper cites Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:39.934771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:39.934771Z digest=sha256:198db1542f8da73256bbef29050d9b2e72fa98b7820ea5f533a59b507b8e0d87

Observation 3139199b-6603-4da6-be46-fcbe77bdd443 · outbound

This paper cites A digital sram computing-in-memory design utilizing activation unstructured sparsity for high-efficient dnn inference,.

Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator A digital sram computing-in-memory design utilizing activation unstructured sparsity for high-efficient dnn inference,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:55:40.183762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:55:39.939185Z digest=sha256:0500b9588f3194967375dd167f8b1e08c0d970db19041005885a02da74e817d8

Pith citing papers

No inbound Pith citation observations are available.