Pith. sign in

Paper Citation Record · LEDGER

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.09291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09291 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:14.200548Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact5
  • verified fuzzy30
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e32df82e-0b82-4bde-b96e-a6fc77a698a9 · outbound

This paper cites Attention is all you need,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.880373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.880373Z digest=sha256:20acfb276af4fca9d9a1713d0a5383ed1dea673c4ede7ff40d0c4268248dd2ba

Observation 36b613a3-f314-42a5-aeaf-e9b2e1b0b291 · outbound

This paper cites GPT-4 Technical Report.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.893367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.893367Z digest=sha256:c4f6c8d872445e01879ac2295b77ea840665103645a645c8ec7b18c1b5294045

Observation a455d192-1b8b-4c51-a974-5444b2ff48f5 · outbound

This paper cites Edgellm: Fast on-device llm inference with speculative decoding,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Edgellm: Fast on-device llm inference with speculative decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.322895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.898827Z digest=sha256:3ee9139ea51066c1eeefbbb6bd085a59bc4683e48945b3eefca03c55c258707a

Observation 0a85edc2-af46-4ffc-8b5d-7f707be0105f · outbound

This paper cites Tenet: An efficient sparsity-aware lut-centric architec- ture for ternary llm inference on edge,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Tenet: An efficient sparsity-aware lut-centric architec- ture for ternary llm inference on edge,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.904026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.904026Z digest=sha256:a9c9583b3f9079be649c39ee0826398da8d3be3f8f7c32f3e7a54446f7c61b5c

Observation 1fddd498-6867-41f5-8f01-7f2b81bebb31 · outbound

This paper cites Efficient inference for edge large language models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Efficient inference for edge large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.282071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.909424Z digest=sha256:36b28622ee1ade67dfa74108683d7f027b07d943e8258b6f56dbf15730399779

Observation 3573c1ed-d213-43fe-a8a0-7e6dd3caa118 · outbound

This paper cites Hqp: Sensitivity-aware hybrid quantization and pruning for ultra-low-latency edge ai inference,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Hqp: Sensitivity-aware hybrid quantization and pruning for ultra-low-latency edge ai inference,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:11:15.277640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.917125Z digest=sha256:4b9af706c6c5e55e5d90e939abc64bac1dfa6f6f7e6b4313ade4a8011f1c51bf

Observation 7256c3b9-6579-4967-b4c3-72bcde7efa2e · outbound

This paper cites Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.922408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.922408Z digest=sha256:001f55540f7ae36d816efd6a7afcfdb85bb7b5411827dcb7b0121c3099430101

Observation b6dab895-3ee1-4ebe-bd9e-cebfc9fc7376 · outbound

This paper cites A unified and resource-aware framework for adaptive inference optimization on edge devices,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge A unified and resource-aware framework for adaptive inference optimization on edge devices,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.250404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.928976Z digest=sha256:fb6392afaf7396f87e1cfa7c73edfb1092f71110f062e14d760ae57967dc2004

Observation 57d5f87a-7660-4883-bf2d-aff7ae0f1811 · outbound

This paper cites Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:11:15.103327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.934956Z digest=sha256:ac82e8f5de37d9f48908c2aefb7cd8d0cae435b91504149d44a8dced2caa4bbc

Observation 48e8ba41-e9f3-4d7f-ad30-6138c2f9806e · outbound

This paper cites H2eal: Hybrid-bonding architecture with hybrid sparse attention for efficient long-context llm inference,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge H2eal: Hybrid-bonding architecture with hybrid sparse attention for efficient long-context llm inference,

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:11:15.063507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.939638Z digest=sha256:152402a5693addb3b75dd6747375b13e7aabe2d7b35e40743a78872bee49a5c6

Observation d735169a-fb09-4d58-b1c0-f61e8b9e3d60 · outbound

This paper cites Gptq: Accurate post-training quantization for generative pre-trained transformers,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Gptq: Accurate post-training quantization for generative pre-trained transformers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.218896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.945458Z digest=sha256:9f3de1e370769c8b3fb3c4a8c818afa8e9d71802cb889d68dc052c107e2fe249

Observation 778cd9c9-cea3-4f58-906c-e3619c3399f9 · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device LLM compression and acceleration,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Awq: Activation-aware weight quanti- zation for on-device LLM compression and acceleration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.188325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.950811Z digest=sha256:86c8e71c78f797d4865942d8772bbd1dc1c3182ea211e29743eb20d071c09d59

Observation 53f80e26-2ee7-465b-b94e-33f9963f40e2 · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.955411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.955411Z digest=sha256:0ea62fa3c9d41b59983af4406b9863c927000ef9aec08c93d563fdaab63392be

Observation d368b82d-868b-453d-8ba0-de167cbda201 · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge FlatQuant: Flatness Matters for LLM Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.960472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.960472Z digest=sha256:82c54b2df85f8a6334fb4ca616680fa9b10d3001f4dfd00e3085f21d469467c0

Observation f1c7cdd9-0473-491b-aedb-ad8728347b6c · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.966910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.966910Z digest=sha256:42712984ef2bb09b72da05bd4dc432c50ef2db61ffdbb971ad577682bad0c0f3

Observation 6cdfe297-8f34-48e2-950f-3814763d6db7 · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.144878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.973125Z digest=sha256:936bef9205532b78340ea2ba5a9942aeb58a91b2042d5abd02ee0f7d42063280

Observation 046f9a38-d0fb-48d5-ae75-62d1045143ce · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sparsegpt: Massive language models can be accurately pruned in one-shot,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.122255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.978271Z digest=sha256:f50728c81d938cb70432da0585763fd787d28cf223a0a8180ca0e316b3214c39

Observation bb98d5e5-50cf-4d10-a4c9-069f951d2b3c · outbound

This paper cites Unleashing network/accelerator co-exploration potential on fpgas: A deeper joint search,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Unleashing network/accelerator co-exploration potential on fpgas: A deeper joint search,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.098421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.982786Z digest=sha256:1dd5c5d36febd36ebefa9a223ae29ec63516781921becb02ed3d582204b4bf17

Observation cf7b5df2-95e6-4630-af16-49bcb67304b7 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.988363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.988363Z digest=sha256:d7e6ade7cc02f7c4fc3e0b951da913849be57906bbc73e20ea992322e7ffe40e

Observation 8bb0eb05-3851-4281-a247-d7730da7223b · outbound

This paper cites Foundation models in autonomous driving: A survey on scenario generation and scenario analysis,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Foundation models in autonomous driving: A survey on scenario generation and scenario analysis,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.993348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.993348Z digest=sha256:a5b07e2708f887d08d55661e3f9a1c82b18e9a4584c1cb75eb2543a9715b992b

Observation 44a86edf-1e17-4d32-9cb3-94e35d5ba7b7 · outbound

This paper cites Intelligent Assistant Language Understanding On Device.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Intelligent Assistant Language Understanding On Device

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:11:14.819211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:13.998090Z digest=sha256:3b2171ee4e8ac9ddc4adec47feaffbac5a3f4297cfbb71c51932c36c0469e8f6

Observation 4f410ef4-782c-4eb8-9ae9-3e3918c0dd92 · outbound

This paper cites Private llm inference on consumer black- well gpus: A practical guide for cost-effective local deployment in smes,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Private llm inference on consumer black- well gpus: A practical guide for cost-effective local deployment in smes,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.004286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.004286Z digest=sha256:5d5a3f2a46c5785a24127588fc4d356e5d5dc7299de6b633a7780f69b5e6bcc4

Observation 5df39e1f-77ec-492d-913f-d281794a4cc8 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Efficient memory management for large language model serving with pagedattention,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.009653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.009653Z digest=sha256:049aade6ab71f94e4353fc9f0bc04f6f32c10c279519dcd7f9ef50e930347645

Observation 096f2b0e-e92e-48ec-9a70-4aef88b8f3e1 · outbound

This paper cites Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.036825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.014423Z digest=sha256:ac7acc931c5e34cba8d16466738db74a73134fd4233e97f634391bc22d5e85f7

Observation 64b24222-b831-4819-b475-270e1a2c6a8a · outbound

This paper cites Wrp: Weight recover prune for structured sparsity,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Wrp: Weight recover prune for structured sparsity,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.008394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.019429Z digest=sha256:9a08baedb7a48820bddfa344fccac1db2dbf7b9c8e027a0175f412eea9530a3c

Observation 73ef787b-cd4f-4dcb-a3d3-5df2d51b054a · outbound

This paper cites Structured pruning for large language models using coupled components elimination and minor fine-tuning,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Structured pruning for large language models using coupled components elimination and minor fine-tuning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.978334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.025339Z digest=sha256:a28e4de9d749e7e3343ae611151c31a4c60ae3b95f23359fb72f2150e6d51509

Observation c213d2d9-454f-4641-86e3-b607f2ef8cec · outbound

This paper cites Unisparta: A unified sparse tensor program tuning framework,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Unisparta: A unified sparse tensor program tuning framework,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.957569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.031314Z digest=sha256:5c5e2d74179a5e5eb4ac47e9512fb449d902d46784ed85ceeb95d43a65b24fc4

Observation 9f400c61-4807-41f8-8509-6baea235b01b · outbound

This paper cites Closertome: A unified framework for accurate and transferable latency prediction across heterogeneous devices,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Closertome: A unified framework for accurate and transferable latency prediction across heterogeneous devices,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.934916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.037833Z digest=sha256:3e8ec0f78992efa8f425e01dd1fe02ba930e8e642f995349d07a06c0912f26ee

Observation 87a07782-4901-43b4-b6c3-1f6bdbf0b776 · outbound

This paper cites Sparse gpu kernels for deep learning,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sparse gpu kernels for deep learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.042758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.042758Z digest=sha256:8faf881f1656171c93c4a863f8de0a9333d4c08bcd587b5b6781b938a4a67545

Observation 476a3036-6e4d-4ff3-8694-2888c44c5441 · outbound

This paper cites Dtc-spmm: Bridging the gap in accelerating general sparse matrix multiplication with tensor cores,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Dtc-spmm: Bridging the gap in accelerating general sparse matrix multiplication with tensor cores,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.899565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.049399Z digest=sha256:a515c0cbae2e3a94ae979e2146d23634d0864a3907062a38a5569099e200c30f

Observation 9dc8fdac-502c-4518-a87a-36225c41839f · outbound

This paper cites High Performance Unstructured SpMM Computation Using Tensor Cores.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge High Performance Unstructured SpMM Computation Using Tensor Cores

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:11:14.386302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.057480Z digest=sha256:63e9627fc55efe0e7ae831576abfa3316e3fbdec3f8617ee88ac761f9f49d829

Observation d625940a-3aa5-4f92-93e4-2177901cead1 · outbound

This paper cites Acc-spmm: Accelerating general-purpose sparse matrix-matrix multiplication with gpu tensor cores,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Acc-spmm: Accelerating general-purpose sparse matrix-matrix multiplication with gpu tensor cores,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.879741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.063332Z digest=sha256:5965b5a99353028d09241c1a08b71cd9e58d79dfa7359de74bb25265c37970d6

Observation 9a51e21c-3b40-411b-b54c-891c67399322 · outbound

This paper cites V oltrix: Sparse {Matrix-Matrix}multiplication on tensor cores with asynchronous and balanced kernel optimization,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge V oltrix: Sparse {Matrix-Matrix}multiplication on tensor cores with asynchronous and balanced kernel optimization,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.849359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.069653Z digest=sha256:ff87fc8fef027ed2b16366d5767f32e65d7d5671764d840786ecf202fac38e77

Observation ca3a53d4-3eb1-4f08-b9c0-1b8529d384b0 · outbound

This paper cites {SparTA}:{Deep-Learning}model sparsity via{Tensor- with-Sparsity-Attribute},.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge {SparTA}:{Deep-Learning}model sparsity via{Tensor- with-Sparsity-Attribute},

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.819336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.075875Z digest=sha256:cbefcad56301e8a932d7c52f488d5534d8079b1cea1fb52b4887e9cae6910e6d

Observation e644f87f-7eff-4d7d-83f3-95f372ccef9c · outbound

This paper cites cusparse library user guide,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge cusparse library user guide,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.796539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.081652Z digest=sha256:bb7f992c140721f62acb591f2f36ac4d9388a38f4fc2138cf795174985199a15

Observation 685a4805-e82b-42ba-8f36-c622fb9910e5 · outbound

This paper cites cuSPARSELt,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge cuSPARSELt,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.768465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.086991Z digest=sha256:1c7f02a4445d0f82a1630449a69edb5b3e51ce8ffbbce1172fa444c95a9a0c29

Observation adbbfad2-d07e-4028-99f4-d68fe5a3c4eb · outbound

This paper cites Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.093933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.093933Z digest=sha256:2f52822f131dbfa1bc8d82f098f3fdfac0de7725604f75e7aba56ebfa2be7ab5

Observation 040defa0-ce08-4624-9015-0898cbafeab1 · outbound

This paper cites Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.739592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.101175Z digest=sha256:6eec1f095011873032229b48e76b93cb58d02f324036f62b59dfaf144c2f11e1

Observation b9119f88-8de1-4c92-bbe8-22b3a84e8f60 · outbound

This paper cites Sputnik: Sparse matrix multiplication on gpus,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sputnik: Sparse matrix multiplication on gpus,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.714739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.110607Z digest=sha256:b6a44c12181e823f3c98f33ed04df2a4c55a38a7a5c3d116b064151cff4f70ad

Observation 8fcc6f41-34f2-4be4-9336-81e88b79fe16 · outbound

This paper cites Nvidia jetson agx orin developer kit,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Nvidia jetson agx orin developer kit,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.684324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.116153Z digest=sha256:cfcf78262be0f4025772271681ba35395e407b0e228e5896a4ec3b1ed2917cb4

Observation 8156f3f9-4bd9-431b-86ac-7285f34dd375 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.123256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.123256Z digest=sha256:2be0175caa5fa7b461d7711a52506a69861516301176baab992f37a6c521c064

Observation 86034bed-f3e5-4c73-8d50-8c8053f20aab · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge OPT: Open Pre-trained Transformer Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.129460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.129460Z digest=sha256:96e08a1795ac53caaed7759a7d58e4a718f710ff3b6a04c5be65bf087e844833

Observation b6761266-fdb3-45b4-9124-e88596bb7f7f · outbound

This paper cites Qwen2 Technical Report.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.135742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.135742Z digest=sha256:057c21903c820200b1881aa45f9c1f00601b741c541a4b2c53f4db701b2ba98f

Observation a1e28cf9-2e28-41fb-99b9-4fa3ed02c662 · outbound

This paper cites Llama 3: Open foundation and instruction-tuned models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Llama 3: Open foundation and instruction-tuned models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.654539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.141376Z digest=sha256:26b548f5fa8a05078e43a1f33989f2f5075ece9fec87db6d4495e7163f4a2875

Observation d504adc6-2e79-4e11-bd1f-0cd281c722d6 · outbound

This paper cites Cutlass: Cuda templates for linear algebra subroutines and solvers,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Cutlass: Cuda templates for linear algebra subroutines and solvers,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.629743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.148396Z digest=sha256:47b019daa548eaa7bbb1e44f050614e03374b9020dcc8d826857d431b2b1849a

Observation d85dbe4e-a52c-4760-8aff-b0bb09672f72 · outbound

This paper cites cublas library user guide,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge cublas library user guide,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.604215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.159695Z digest=sha256:cb6a0566bb26d08f0c3962bd898bac5b57b5109b55481f66dc4f77846b4b2274

Observation d43fd8cd-d093-4e7e-819c-db8ce32c10cb · outbound

This paper cites Fastertransformer,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Fastertransformer,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.576894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.169034Z digest=sha256:404e8ffe46c4b76425c1b9c014326bf8b5f111714b58ca5fb12278babeaaa297

Observation ffc38f0b-64fa-4337-8039-4a66f96141f2 · outbound

This paper cites A simple and effective pruning approach for large language models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge A simple and effective pruning approach for large language models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.554020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.177358Z digest=sha256:b112a3fceb779d7bafc654f7d9f16b1b1443c53cdfb4798ab98141cfdfa74ff5

Observation 2e7f0333-732b-4d8f-8bc7-9e1ba16b237d · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.532365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.184734Z digest=sha256:ef8d80f4c8e3745a300c57bf7901b0ffb6203fd14bee83c0f1c166206d507fb2

Observation bd7c9003-2240-4437-a743-b7f1649931cb · outbound

This paper cites He examines various aspects of embedded systems, with a focus on performance, availability, flexibility, and energy efficiency.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge He examines various aspects of embedded systems, with a focus on performance, availability, flexibility, and energy efficiency

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.488007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.200548Z digest=sha256:f7161e1edbb296b217ec0269396373efd3b3d83a59bc9152b4ecd75831c79b4b

Observation 2f1cdba3-8dda-42f6-86bb-783c038b9f1a · outbound

This paper cites He is currently pursuing the Eng.D.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge He is currently pursuing the Eng.D

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.510650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T20:11:14.192355Z digest=sha256:9c11ba7294519d8905921710d2cb3bcbc47518ee49eb0101e45715e88ce3df29

Pith citing papers

No inbound Pith citation observations are available.