Pith. sign in

Paper Citation Record · LEDGER

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving

As of 15 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2509.01229.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01229 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:51:01.240091Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T12:38:31.783807Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T12:41:22.816856Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33d69fc8-5ca4-482c-8991-68d8b707075c · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:58.727339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:58.727339Z digest=sha256:8b044cc73d65bf0eb8706348e1f72e5e7c1d6da4b64a99f4aac989e776a7c561

Observation 36a89caf-b328-4462-9615-5b9d20f67e4c · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.804292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:50:58.918382Z digest=sha256:39dd64536c5e7e0fad913c64258872315ff63adf911cd084856444c63f9f0f63

Observation 866efa00-67b3-4e79-9ef7-4b994c45f1be · outbound

This paper cites Low-Rank Quantization-Aware Training for LLMs.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Low-Rank Quantization-Aware Training for LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.043745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.043745Z digest=sha256:2123ac03c34425611754e6c40bc045f9657e13d8221b19b0e3ee870afa9d8e46

Observation 4c4bb3cb-8bb4-4c7c-9d5e-ead0b2eef072 · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.236249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.236249Z digest=sha256:44dc252f5e2469d3171a60d48dd45c264aad82b31af3ce17368a8123a2b1c8d4

Observation cda9b768-4f76-48e3-9f42-ba2f16c6e77e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.329622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.329622Z digest=sha256:defee710e081647fa01d873f456bc0bf3ab934e97f08543951105b845b24737b

Observation f424dd9e-3948-4a54-9a70-5e333c78988d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.389003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.389003Z digest=sha256:fa403dd3e9f57376c09e5ed4794ff86204170f469c588d29978023bca9ffcd2e

Observation ab9adc55-311a-4df8-b4e0-92636de41a5c · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.790186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:50:59.491951Z digest=sha256:edda7001f8f6b2e2dcca8d7ab064f35f7035534eda08ec743bc4aeaab81d8ed9

Observation 2b4168fd-093a-4644-a754-43deece68509 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.565844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.565844Z digest=sha256:0a39216652ac422104c18c86bab233baeaf852aa42ec6cb25aadb2820876b610

Observation 26d02f9d-50c9-4cad-abeb-719682daf6ff · outbound

This paper cites The Llama 3 Herd of Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.639928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.639928Z digest=sha256:8f659ba51911747d06df0d78b99bd35effadd322ea775583673a0768d56795c9

Observation ab963050-78d5-4e73-b9b2-da5785f3f1da · outbound

This paper cites Mistral 7B.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.710091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.710091Z digest=sha256:d2eccaa519a6b30dfc12f1edabfd41b2a6cf8700d530ba58eae55f291a7374db

Observation 8961faf1-5dfd-496a-aa6c-8be7e5f45def · outbound

This paper cites Mixtral of Experts.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Mixtral of Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.789990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.789990Z digest=sha256:a8231d5f65a95ec2d02afa8e71a23de94df05260d2556c1c564021376851c7bd

Observation 2e5cd3d2-23aa-4f5c-98a8-c68ab9b78c3d · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:50:59.871156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:50:59.871156Z digest=sha256:856c8f708c5e693862d803948415aaf4bed31bfab7eb34568544db18c1331d2d

Observation cb58c9ee-55a7-4e39-b26f-664f097f45a3 · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.766163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:50:59.935473Z digest=sha256:f35dfd6949d0d0e19ed2f0af7cb99b2d349b81528e21280827110e5f48020731

Observation dd669daa-669b-4d2e-999d-a92959458e57 · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.752455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:00.008632Z digest=sha256:c658f7b058759828abc5851d3ff660e60755593cc707f5e750a50fc1a8719d5d

Observation 28a5a33a-5a88-4d97-acff-9206d9d920fd · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:00.115901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:00.115901Z digest=sha256:8048d467b151d940425e7bce4b1f8785106b575f332c245a2985be4ddf14ea90

Observation 16d86e1e-37b1-41af-baf9-37c7abe3fa62 · outbound

This paper cites COMET: Towards Partical W4A4KV4 LLMs Serving.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving COMET: Towards Partical W4A4KV4 LLMs Serving

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:51:01.499545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:00.225256Z digest=sha256:44391ac1f8936d74a7e79b5b828992372f109f2676c5ec5e842cad629abf76fb

Observation bbe18e3d-6411-4361-856c-30c03ce71e59 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:00.337520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:00.337520Z digest=sha256:742502370fd151c89669477358662eb41ec08e9f33d315d6d74142ca84e3925c

Observation f33c51b6-621f-4308-85da-6c2360e1cdaa · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving SpinQuant: LLM quantization with learned rotations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:00.448570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:00.448570Z digest=sha256:a23f9a255432e307993dbf4f7033444d1303132026977c26145e4727b66a8ac9

Observation a2f26be9-8d3d-4e3e-b6aa-fdcca2ba7596 · outbound

This paper cites Pointer Sentinel Mixture Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Pointer Sentinel Mixture Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:00.524851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:00.524851Z digest=sha256:459cd497b30e7684f91645ed95253c09bf9b3df2501ac8c1a42ed0165d744046

Observation 647173bf-de49-49d0-98b8-bdd9b4e5e82f · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.738910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:00.635341Z digest=sha256:f3dd58c8a48541e0ce245f6e8eaf720a36421448a7285579b6a2631d804eb5c2

Observation d30b226d-8ee0-4c29-991d-e5be8695c406 · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:00.710364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:00.710364Z digest=sha256:c4d5f3ea514548ecdb7a26ce3a849407fa5dc39c55367e33fc9f3a9a23ea50ed

Observation 722f42d6-ee3f-4cd1-b8bc-b777d2f32e5d · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.715591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:00.803648Z digest=sha256:e598753e10a23ec73602facefb15972443f4f9ef4b226eb24797350fd22c7eb0

Observation 23d72f5b-203a-437d-b972-e538b94b3b7c · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:00.895682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:00.895682Z digest=sha256:7cb7ca1db637be88d131e39ecd5e3f7be8bf982e95f0d5703f8e1c9c9a0dfd9f

Observation 821ad1a4-168c-4ef8-900d-89c007ed6c61 · outbound

This paper cites Squat: Quant Small Language Models on the Edge.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Squat: Quant Small Language Models on the Edge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:51:01.418595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:00.964875Z digest=sha256:75cde094a51a5327841482a9cead26e3134b596dc5477bcfcd60631291ca1de3

Observation 095fda24-251a-40d6-874d-3d58dca313f3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.077785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.077785Z digest=sha256:577e50a8351617683cc2c6d15960879622ebf9695719a44e850c101fdb96bf25

Observation 16b23d92-a37b-439a-b139-dd9610cfdf59 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.158011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.158011Z digest=sha256:ef81c2fa3c77fa48dc6e21f92ab0f9347d14d350331ea6f648fc36d1574fa4d0

Observation f68407fc-0970-4982-8aff-875ce0a25f00 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.198277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.198277Z digest=sha256:fcbe7b0e67b44142199cd17d08764fa130c088c10ce2a3b03e2af376cdbf8efd

Observation 057bd4f2-a8c0-4035-9284-998b1ad53b9a · outbound

This paper cites Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.202661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.202661Z digest=sha256:645330bd0bfe6189a0578f148eb4444e8401059c4185bf18904367ab5ef98ad8

Observation e22a4f09-6b85-47ab-8808-4fe552f61a7b · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.207097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.207097Z digest=sha256:45ee5e9c5ba40fbc6d13710598404ef08e6b2553bceb435bf1f007afe0796eb6

Observation ca793098-5702-4d56-9da1-abd372637612 · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving OneBit: Towards Extremely Low-bit Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.211009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.211009Z digest=sha256:99a3a4b2c2374a6f398c94776bd985c7b599aa254abd131b040455adf79ea96d

Observation 4bfdf7f1-8425-4bf5-bb83-577889c2d4be · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Yi: Open Foundation Models by 01.AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.215194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.215194Z digest=sha256:40c190c03e25b2caa3ba7c25e512fa5b93d6ba1781da68de95af5458e440d467

Observation 8a0e2c2e-003c-44b8-bea8-7160c5f5309f · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.691688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:01.219435Z digest=sha256:0d77c0cee6d226830e7c088bd21d6856dd53654ca98441c95ebc7e2162433236

Observation 916dd6c4-d250-4e2a-9174-e4c7c035e62c · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.223251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.223251Z digest=sha256:c10655b1904214d423b4d7a72fb68133577c6907620d71f7d60d0d7f3aa097c7

Observation 49c6fe9f-3712-4dab-becc-206455f21cbc · outbound

This paper cites QQQ: Quality Quattuor-Bit Quantization for Large Language Models.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.231907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.231907Z digest=sha256:5568c42b2e0bb6df54a2ad14b63a7e0701cd81c763e35ac4e24e92f7f98f1d52

Observation 74c042ce-ba89-427b-a4d8-66b5bd8a9d7a · outbound

This paper cites an unresolved cited work.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:51:01.667807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:51:01.236043Z digest=sha256:cf08e6981b297ef7b5a9b073bbb4ea0eb7b04e0e7bfe5049d974061359125f81

Observation 905c8c31-43ef-4744-a325-f3a148f02657 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.240091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.240091Z digest=sha256:bc7df17d53bd4c27a0c531fc9abb6a0adee96aa0ff6544b827269f1e6bd7029d

Observation 0e1da5af-0771-4817-955d-d6ae33c03305 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.227376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.227376Z digest=sha256:66c1ac9bee795d6b6a050966b62dd66b06ecf93eca65caf00129e009b3aea99c

Pith citing papers

Observation 70332ef3-f172-4b80-a889-757a0a393ff8 · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.819166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:461d41b23f91733e30a461a95c0bd36da8cb3bfb5b465feadaac9fff85cd2a78

Observation e57b1a6c-4eb1-4c46-aff3-a37ea38ccb25 · inbound

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference cites this paper.

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:58:34.214757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:57:18.944617Z digest=sha256:b29a9afd567a07a19f3638a43bf6c60c50e37c1854d6a810bd7e7034d76feef9