Pith. sign in

Paper Citation Record · LEDGER

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration

As of 7 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2607.08993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08993 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T01:10:03.032181Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0b5da88-3461-4e1b-ba7c-1c43d0b44aab · outbound

This paper cites AWQ GitHub repository.https://github.com/mit-han-lab/llm- awq/.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration AWQ GitHub repository.https://github.com/mit-han-lab/llm- awq/

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:43e9b7c956d199618c78139f6934417469fff28bd3f89d20bbec97e6295c8e54

Observation 75cb4bde-1478-4aa0-9a54-6541ff851bb5 · outbound

This paper cites NVIDIA H100 NVL GPU.https://www.nvidia.com/content/dam/ en-zz/Solutions/Data-Center/h100/PB-11773-001_v01.pdf.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration NVIDIA H100 NVL GPU.https://www.nvidia.com/content/dam/ en-zz/Solutions/Data-Center/h100/PB-11773-001_v01.pdf

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:5f734390682f87125cfe5d1151ebf10510492afc3057aa4a82b1cf930abac678

Observation e32f8ed7-e5e3-4ce6-a0c8-eb02cfa8b192 · outbound

This paper cites vLLM GitHub repository.https://github.com/vllm-project/vllm/.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration vLLM GitHub repository.https://github.com/vllm-project/vllm/

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:92669f5339f806c8ae18cee7bce686355679b4692084bbdc1fdc81520c8218b4

Observation 9fe71c5b-fa21-4327-adf6-2c68779d7cc7 · outbound

This paper cites FloTHERM.https://plm.sw.siemens.com/.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration FloTHERM.https://plm.sw.siemens.com/

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:b1884ab2b539abbac4dab17a3abf05b6b7544113e22ba61ec5c6d81c3ca981a3

Observation 0834a8ef-37d7-4dc4-8520-bba7e97f0cf6 · outbound

This paper cites NVIDIA Management Library (NVML).https://developer.nvidia.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration NVIDIA Management Library (NVML).https://developer.nvidia

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:b1cb9dd4a98ba0bd3f4a7ae264adeea547a2356e9049808b0862bf4f477db046

Observation 8a691a48-1048-4c3a-a072-e04c04d4baca · outbound

This paper cites NVIDIA NSight Compute.https://developer.nvidia.com/nsight- compute/.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration NVIDIA NSight Compute.https://developer.nvidia.com/nsight- compute/

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:70fd044f7c141d4d4aebf1f998bc66496454347d79835c4fe518a96b431ebc47

Observation 5732377d-1e4b-4b31-b50b-cfffbde9943c · outbound

This paper cites NVIDIA NSight Systems.https://developer.nvidia.com/nsight- systems/.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration NVIDIA NSight Systems.https://developer.nvidia.com/nsight- systems/

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:89335fe2e76bb655c70b5c46b02464e6347842d7cede9894727d5d533b61b280

Observation ad02d7b2-92de-47cd-953b-c9fb3f20ee56 · outbound

This paper cites Synopsys Design Compiler.https://www.synopsys.com/.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Synopsys Design Compiler.https://www.synopsys.com/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:08d34fbe1705b820db6bc6f61f0dfd9caa2d08390c14e07518d607d6631f4e1a

Observation 364ad72c-885f-4461-a02e-de699e9cd544 · outbound

This paper cites TensorRT-Weight-only-quantization.https://developer.nvidia.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration TensorRT-Weight-only-quantization.https://developer.nvidia

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:a023294685a82497e497f32db37f38b417a8dda390d053ca7b1dea6d2ef8a94d

Observation 8eaf4db0-5bc8-4285-b113-eeaed3c7ea1a · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:4bc7ef5b2ec55f1395c7b76c5b35a64590c5f589fea39baada573a6d8da30e3b

Observation 9e0fb216-1738-4ce7-ab87-26bbea732c5d · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:23c5e60b073b156c1b7edfe9e5f47924247400aebb3a8072e65166a4a848d919

Observation f4c52379-ea00-4419-92e4-8a20241bd9e4 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:12cf67983a41f51b21b53467866ea5eb03dbbc342a46baa5b6a9ee2f3370974a

Observation b20419c8-df38-4b1f-96c6-757744719bf3 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:f268e5e191b34367edffe55adf23f8c8943a4aa6ed846895443f27af8eaf63d0

Observation d73e7b3d-48ab-4e0b-8ac7-9010744c85f8 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:2026ddb003b0463995c1402113c8f491bb6b5c78c0c0b8a20562e8181022d7d8

Observation b4700904-5106-4134-923c-ade1f166edaf · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:75e432bfc3dd1b667bf30d2dc3f9006350f93b17a02f9e6bacd716e8d45c858b

Observation a615ed01-ab18-4a72-9243-85f86be8cb49 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:ef68e9b1b5d8591ea92d797459f61a32d75f34a70b756350b23dd28779b2159b

Observation 06d0d6f4-93c2-4097-bc99-994db77fcc43 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:17aeb4b448491fcb54530996656042ae7eb298d2add6ac4b3c38498456cb0af7

Observation aa7457c3-a334-4755-bcba-67cf9bc0e0bc · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:d19f77fdf31165367f30407b327ffcca90b031addf78292612f0ecfa5f54c336

Observation 62199408-a4ae-4a31-bc47-668e1b04ea85 · outbound

This paper cites int8 (): 8-bit matrix multiplication for transformers at scale.Advances in neural information processing systems35 (2022), 30318–30332.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration int8 (): 8-bit matrix multiplication for transformers at scale.Advances in neural information processing systems35 (2022), 30318–30332

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:05e38f099533db8073439c8cf5189a07041342065bca22b86b654ec196d46da5

Observation eae854f0-15a1-4525-89a3-866236a892c1 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:49ff79af1a8effb76500538bcd5178148ee770e13375ec15de18457bf99e86ac

Observation e3f26909-a5f9-4f0f-b0b8-05684a9315eb · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:3dc5482cad94867437003b617fdf11f5d227d037a372e9a7471956cde8793a9a

Observation b59f9414-ac4d-458b-a75a-2a4c0cd259ea · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:824ffbd627d96bfa88261e41caad101af5345d6fb88a0c32a7248a31bc16a8c7

Observation c5470fe1-aa42-468c-9014-771ae014542c · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:120633b10841118309d4ee61a16ce3bf0b7ca4fb3894881f433d0c5784c7d4df

Observation e8cad5e9-a754-4629-a443-ddae322c182e · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:000f4a42eb2115754167f5424d6e1be0c55a1f9196ea32ec026c261fcc9bee8a

Observation 39c9cb58-7ea9-4cb6-a102-43e231dbd6b2 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:d9bfde0ae3074743b4127b089cff97cd8f0aded5329419ec1bf6401773058593

Observation 2d05d036-7c23-4829-8711-39ea3c512b63 · outbound

This paper cites InProceedings of the IEEE conference on computer vision and pattern recognition.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration InProceedings of the IEEE conference on computer vision and pattern recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:1dd3fe4b84ca762ab9558c4f6fad09a56dd4f7a592f7caa4b652ee3f296b81fa

Observation 4d7daa7a-264a-4b8e-8ed7-4f36d87f3f69 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:13ca011d905d71950d2b7c4975ff5be09d036f21d223425d090f6c48a556b447

Observation 54c32933-5f26-4daf-add0-51d2c0ed6151 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:051270761538ce6eaec177237ac226bd8ee280e84dba99bb4b738fc8329f6ddc

Observation 39bb0f80-d69c-4e7a-99c3-571c7bd01733 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:3281dc6234b264c2fa7035102d31bcb575cf47f57e8b66a8e8f572eb6878e26f

Observation ad10a104-3c2b-42cf-b3ae-8e9a5289b9b0 · outbound

This paper cites Mistral 7B.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Mistral 7B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:75b9da3516a73182fa668258cf54b1bf5447818bd3819777f60988f83a3b64d6

Observation ab16eb30-2ff8-40ca-a949-64208d19e93a · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:5c4c6338cd9c1da355c8668375f73c91e889648b92ed1794b120b9e38c4e8d6d

Observation b60fbefa-2b07-40bd-94a2-baa13222deaa · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:ae3b08fd51939d3ed04369d441685fb8c0405326e2e39074c3d15733a2cea76b

Observation 30e63d77-7751-4ebf-8c4f-0973aa7fe04f · outbound

This paper cites Scaling Laws for Neural Language Models.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Scaling Laws for Neural Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:4348bc47bc68e868ed9af2f92b0d9da7e569d4a201b3ef36206f0db5c6857301

Observation da63b521-8e3f-4dfa-ac1e-9c2115f95ec7 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:85b9adb9b0293e8d4121fec90e384174abc300d09ee34b4478c2ed3a0ee6d049

Observation 99b8a849-93a2-4da8-8cea-84f48c3cd097 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:b1c6ebf6a75244b6e3cf49e4595092939418d5dfb76f3a0b3263bcf3e479bbe8

Observation cac02070-cbb8-444d-a001-f36861d197c5 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:6c94aee61a1dad05b23df827c259e98734b1d31c77ec067ec96f24b2deda5cf4

Observation 8f8ecf25-754e-4290-a1a0-249cdffe1e21 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:8e7e19f4f122c5ef20112e63984699dc67c8bc9a6b6134be3fbdc80173d3ccfe

Observation eda5d5ed-c6f7-473f-83d0-bae92b3b2c37 · outbound

This paper cites InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:d6c96e34af9eb3bed495465aa15c55a244bb331cc99bbeb930d19ffdce295f02

Observation d1a0d435-c214-4a66-9637-94f85c195e8f · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:69cd7ba4a1aa8020fb305d15a03bfe4d0f2d946a0785a09f99fb4285508a8b4d

Observation 91ab3514-5ab9-417b-bd6b-a58cc4aea220 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration SqueezeLLM: Dense-and-Sparse Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:b1808e1e35b2fe26e21f7eec65d0cd8a7fb0830fc81aff8f9a88e9d19e94c7f4

Observation f490cb86-8e3d-4172-bceb-09c3baccc34d · outbound

This paper cites QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:bf2aab4977dd7aa594f4ccd49cdea65209be6e28ba84ae54664ca7748c6a3040

Observation f7dec1dd-1d58-4a1b-9ef3-6ad2c3e1e001 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:4dac974c3197d2a73dbd95e0ff8182a1b13fda2711a206718b20c2b170cced6f

Observation e816ef1b-7fdb-4471-861d-9691d04c1a61 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:32e3df80c1be48e24f8ccd8b245b57fe4585de69faef2d0e85f801565646c58a

Observation 8b64c419-c8a0-42e5-8637-620c9d6d7203 · outbound

This paper cites InProceedings of the 29th symposium on operating systems principles.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration InProceedings of the 29th symposium on operating systems principles

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:f48f62652dc32baae71fa284bbfb9b45ce282a147776e719faf72efa00d1f588

Observation 41acf1b4-ec0f-42d2-b850-c3d67c808778 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:ae309ca7e4d237c19654b9a0f1b7d00b702172b5edba9a042f3c34a5c8634472

Observation 3ae34834-5a8e-4042-aa34-1384379d097a · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:aa98e6e986b6ef90e9e7d2d21502ff6b868265bb988913abab64cdb75913d2ac

Observation 7541a11a-411f-469e-9d96-8e7e278cc77f · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:470af24e1a2af915068a7c334d4fec74abace5d58e76aeae4d2675aae1afb295

Observation b1cd48d4-d862-4eae-9072-1bc8fd4d00ee · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:8d3321ddc2f01ff120f468f1249f8ead987b987622f6daab6709eeb7b35f5763

Observation 669e74c3-96c1-4e65-a054-3d87dc1fc759 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:cb4e1ea8a765574c2e877dcc090a31f3cc78101c7988312a5366f1d453bc70e7

Observation 1754fb07-c4dc-4794-a7a4-574e7ca1d8c2 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:102c0f2365dbcb744a614d44bce581ec55f4a6a56e1559f6cf6181b1afb56234

Observation 109cca2d-fb6e-4581-b0ea-04d9c188e1fe · outbound

This paper cites Custom 8-bit floating point value format for reducing shared memory bank conflict in approximate nearest neighbor search.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Custom 8-bit floating point value format for reducing shared memory bank conflict in approximate nearest neighbor search

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:d4c809b58a60fccb046b64cb9951287c2fd0f98483c629db3f3cf83a5644338c

Observation 136220bb-1fef-44b7-9aba-4e218fb10ad2 · outbound

This paper cites TorchAO: PyTorch-Native Training-to-Serving Model Optimization.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration TorchAO: PyTorch-Native Training-to-Serving Model Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:63d5986920290da993137c4afd0d31d62d2587351a13c9328322f0da4c980b21

Observation 20ce84db-bde6-4c76-8ab3-534611cca6c7 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:3c0d28ed9a937af3fa1cfc100b2a0b1e0ee5f3bb06339211c3373ed9e303f80d

Observation 08d576a2-5722-4cce-9b40-07d1df0ff7e7 · outbound

This paper cites LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:45b2244acd559756cb2d14e54af48fa4b5f9f0a9535dce34830ea7ce29bcd024

Observation 7c84cb65-d0f1-4084-9d7f-a044577e5025 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:01664e9f0c7e19d9db2bca63c9c8fd73a9537ad9c859b7ea0f323641f8ead970

Observation d31f3997-2717-4fba-9eb4-d29a815ff4bb · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:fcfa113967ee556b94832bd99aba46d44b7a4de968cc903bd968a2b5d830a1c0

Observation be9634d8-71ee-496e-95e2-4bc3d0cc10b6 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:a117a78a9faef2571ad7e2c3e27a7087125661be828490c066b332d7dc51a145

Observation ecb37a61-e90a-48eb-828e-db65c7bc8149 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:400d95edb75e60cd2824d563d7444c3acc4b414b6766320f348b49ce68010583

Observation a3852807-a9ec-4c5a-b593-97bf3d9065fc · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:63ae26d2abe94ff230e13dfaf4740ccd7ca31fea6f8d6cf798568eb664c8acdc

Observation d018821f-a67e-46a5-83fc-3b18162893b6 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:4edfea8c48e5f1685432377fa3d12bf10908958d53c6f5624b26ad08bbca1c2a

Observation b7c99f50-5133-4152-96cb-0d1c4fb7d7a2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:5d67b245bb2476951d55d2f5c28bff7a18bc8b3a8c98030ca401a0a72152a238

Observation 9a4b6702-558f-42fe-bc96-ad54b479efb5 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:6f12f3d4801fa74251552c521df4114bb3c9815a367eebb15ca44dddc40c9ce4

Observation be6d6557-35cd-46b2-9265-59c83c760d20 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:b7f7a295c9d33ba34d5dd1031e4738cf92dfc9ca6546e822bd4a6460f0bd54b7

Observation 9de50d70-9815-49b2-b999-dacac0297846 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:accbc68a24948ce869a74efec369b0c7d8b6744b9b75fe2bc996807d505f7c5f

Observation 3c86b8dc-6499-472a-9169-a1acafae269b · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:96b77dfe4912d35246923e7c37cfe9c09a48aa2b3ac5063bb4590181441324d9

Observation bd387cc3-d292-46f2-99a4-6da9b54f602f · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:47b786b8867a37b83bd135b7c7c55dd2905b5688055673d691f322c110f3246d

Observation 3cbed698-3d63-4a38-9eee-f8d4d27f6faa · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:3e6b597e8b149920f8a971b87796a0a24226fe0e2e0718cc3b7a3361acdce5a3

Observation 427d0969-7cae-4a4c-823b-c7d78a072d25 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:135e2d9c80d85a2d830ad4132d2e1872095caa236c1a7269b26643a0707a1a1a

Observation 92a29df8-67f2-4c9d-8ecf-74a21687d03e · outbound

This paper cites Qwen3 Technical Report.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Qwen3 Technical Report

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:be65bb64583c7c0b5026be51eca3f6b9bff25c54425da20f9fc62aedd19fe331

Observation e651f12e-b2dd-48c8-b6ce-42bf956a1d32 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:22ba5e481b5ed9f18f19333d3afe9e1395d01e6cae463c2bcd4af9a6a9c53426

Observation b486abc4-7a29-4502-8f1f-6183d49cd646 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration OPT: Open Pre-trained Transformer Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:069a6c1598bb2b96872b9fe234d6cacad903f88da4d4404d1c9e2e31d7ea39ea

Observation 176971e1-129e-44fe-9653-e42be1804a59 · outbound

This paper cites MixPE: Quantization and Hardware Co-design for Efficient LLM Inference.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration MixPE: Quantization and Hardware Co-design for Efficient LLM Inference

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:580ce91e4e658cc16adeda5dea29189782b0feb73b0a9d256c51ac36317edbc8

Observation 36fb5bac-05d9-43d0-a6a8-cc62a15caf02 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:bacead99557a7b9d4ace651f13001733b3ba530f0cb11f7e3cd1f43f5460a2e6

Observation a3946790-ddaa-4c53-8f9c-2b36cf80babe · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:98009255fbe76972aab270db147924faef2be4ffe8e11af51b0034d4cc772ec6

Observation 737934a7-5e78-4e67-84bd-4936196c692b · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:9153582e80ef0106f05794487d4f385f9c0dd078e8eb0c251f849baf4df8e213

Observation 765e83b0-8928-4bf3-a551-1317f99b9402 · outbound

This paper cites an unresolved cited work.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:9d2077fbfa7b475f6240d4abb69079f31e87b7a56e78c839e5c5c9980d30a9ac

Pith citing papers

No inbound Pith citation observations are available.