Pith. sign in

Paper Citation Record · LEDGER

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2508.10395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10395 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:37:15.720135Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c1eef06-70f8-4abe-a5f2-a32206256e35 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.313365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:11.718326Z digest=sha256:ce723b96859b4f6330131295d35e5118eec50f776739f80e0f229497b23527d5

Observation 7b5c91f9-60f9-4118-a663-eaaae8bdc347 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Longbench: A bilingual, multitask benchmark for long context understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.302365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:11.802675Z digest=sha256:835b365204bea8b644b7c96727499c01c9c19444163c0f7528e9140296c1d883

Observation 3f524b65-db83-47c7-aff0-08794c3652b9 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:11.866338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:11.866338Z digest=sha256:e74ebd08801a35f3d6f932c439f75141f8124ec1519892f163a8ec160783e0cb

Observation 2323e10e-2681-401c-858c-2587f9bb0dbb · outbound

This paper cites xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:11.923378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:11.923378Z digest=sha256:233c44b2e64f7bd20c1c6009623fce8abcf3cacfa89d7673f9646768750ac9e8

Observation d4076f9b-e933-4659-91f9-57fa21f5e753 · outbound

This paper cites Palu: Compressing kv-cache with low-rank projection.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Palu: Compressing kv-cache with low-rank projection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.283985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:12.013603Z digest=sha256:36460442e4a09b473c237efdd534a3e0119e256788813755cefd061275b68164

Observation afbd76c2-0382-4128-88f0-b7113a4bcdad · outbound

This paper cites Palm: Scaling language modeling with pathways.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Palm: Scaling language modeling with pathways

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.072745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.072745Z digest=sha256:caefd0cf8cbc2dbd98c37fd598d349768cba3490216e4550977df057a9c05cea

Observation 8f1ba5c9-83b4-4e4c-90bf-d2082ebd2070 · outbound

This paper cites Training verifiers to solve math word problems, 2021.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Training verifiers to solve math word problems, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.157932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.157932Z digest=sha256:6b2cee037fd5d55b36c4a30731e62db7a2bb1d0561ea0f3454ebc03c6f3178d4

Observation 15832a7a-6a4d-4b8e-ae9c-372c89b42146 · outbound

This paper cites The Llama 3 Herd of Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.244737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.244737Z digest=sha256:a3e5c5901d4a9baa14a812457fc8f787ad11731b90adb9605dfeaf0018173ad0

Observation 0aa150db-c316-4ac7-98f4-70d155a558a1 · outbound

This paper cites The language model evaluation harness, 07 2024.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization The language model evaluation harness, 07 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.334012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.334012Z digest=sha256:6ee0b0206297b9be1ddd4776d496287d3fef5c8d4fffdec0485edf23ace95361

Observation c15ce080-4ef0-490b-a6b5-09817e47bc31 · outbound

This paper cites Fast state restoration in llm serving with hcache.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Fast state restoration in llm serving with hcache

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.250581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:12.399294Z digest=sha256:f71930e72ad5be204e53b4809c60ecc613e99ce9472d2d08ead9e79e3c0b3901

Observation 7e72e24a-621d-4c84-b7fc-4b430db5f4b6 · outbound

This paper cites Ai and memory wall.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Ai and memory wall

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.239833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:12.480539Z digest=sha256:a1079bb26719ccf12f00fa073a619711b7e71dd096a2783c4c241644705f8ea5

Observation e5e2d711-c226-46b9-899b-1bb84e4f5ea4 · outbound

This paper cites Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.575486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.575486Z digest=sha256:490952c807923e85899508ccbac8e6397b8640a18f15e1f337969815b1d4d7e2

Observation 9855f0ae-3af8-485a-be6a-c2bcf512839b · outbound

This paper cites PolarQuant: Quantizing KV Caches with Polar Transformation.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization PolarQuant: Quantizing KV Caches with Polar Transformation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.660388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.660388Z digest=sha256:8847197003404b5d0656478a90092a56fb603dd8a957aaa9d785089f44eb6518

Observation 3968a25b-49c7-4b2c-a48d-fec801aa1f4c · outbound

This paper cites Zipcache: Accurate and efficient kv cache quantization with salient token identification.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Zipcache: Accurate and efficient kv cache quantization with salient token identification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.228837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:12.754333Z digest=sha256:5214bc97125debf44a65b7d3e3a60f3a8078088b1c915e77cc14b7baebb171d7

Observation 631dec7d-3102-4409-a021-98f27c8b7cb3 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.874359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.874359Z digest=sha256:04b2c39a022b9e9988c503d8aa75d1fe98808ca5774505cbcc5d87e053f1fc52

Observation a2d231dd-b670-40fb-aa4a-bff9f51e1fe8 · outbound

This paper cites Checkmate: Breaking the memory wall with optimal tensor rematerialization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Checkmate: Breaking the memory wall with optimal tensor rematerialization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.989497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.989497Z digest=sha256:d9cead6a0c9ba70af6a084ee9a1feb85820269d05021f64aa93960ade05dab27

Observation ce27969a-f689-4ccb-9730-79aa815ebc1d · outbound

This paper cites Residual connections encourage iterative inference.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Residual connections encourage iterative inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.205712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.070387Z digest=sha256:2ade1fbaadb87888dee8868961e6a7d9b753c101ab97ff1ceb89a58984dd5e8d

Observation 4b9171b1-b443-4307-925a-a6c3eb5e615d · outbound

This paper cites Mistral 7B.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Mistral 7B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.135459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.135459Z digest=sha256:db428787e8aa7b5aefcdadd17f71d0cd860dbd8e4912707be4683d997a201b02

Observation 6c9e2511-d8cd-4035-97bf-314a49baa76e · outbound

This paper cites Mahoney, Sophia Shao, and Amir Gholami.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Mahoney, Sophia Shao, and Amir Gholami

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.195749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.221672Z digest=sha256:d7a99de45190b5b558719ecaa56b044e23ba028bb5ad34f0ffe1f5bec7a0c713

Observation 16de2b6d-3f68-4667-ae85-a82b0265a614 · outbound

This paper cites Squeezellm: Dense-and-sparse quantization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Squeezellm: Dense-and-sparse quantization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.186133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.297646Z digest=sha256:a5c3102e88b855e694c4eacddc3e4b51acefc1f3580d54bbdc44a24769c158bb

Observation 66d614fa-0a90-4ecb-9a66-70505ce68cc1 · outbound

This paper cites Efficient llm inference with activation checkpointing and hybrid caching.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Efficient llm inference with activation checkpointing and hybrid caching

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.347682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.347682Z digest=sha256:81b538952319114cb20f47aede7b781cde88f016a9e64a3e2efd1a2d26eeedc5

Observation aac11420-51eb-4b86-9915-ce5839478833 · outbound

This paper cites Intactkv: Improving large language model quantization by keeping pivot tokens intact.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Intactkv: Improving large language model quantization by keeping pivot tokens intact

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.176321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.424917Z digest=sha256:32d5d9070642ce495a21d5b945ce8ff813ffa57dba093f5505e572c5942f281e

Observation 6ec37ac1-38a1-4b1f-993f-383e9afd8792 · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Kivi: A tuning-free asymmetric 2bit quantization for kv cache

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.510276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.510276Z digest=sha256:80052fddb3f88e9caa6640b34a5deb9109d63ffa02e3ae9323fb1ab880ec3a7a

Observation 5075b5bd-4990-41d5-9345-1bff31bedf0c · outbound

This paper cites Pointer sentinel mixture models, 2016.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Pointer sentinel mixture models, 2016

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.576293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.576293Z digest=sha256:1fbd7fa2762c5196d45955be74de440e099f281d03b8e0ce8a0382f29189c4ec

Observation 2a402916-1310-475a-9143-b456d4ba0e3c · outbound

This paper cites Llama 3.1: https://ai.meta.com/blog/meta-llama-3-1 , 2024.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Llama 3.1: https://ai.meta.com/blog/meta-llama-3-1 , 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.154060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.666970Z digest=sha256:7bf967004346941c8d6cc6a3367257bbbed999973de249f56dc79014c534f6f4

Observation 0a2e3a84-3e2c-44d5-bbdd-27c0cbbb5682 · outbound

This paper cites an unresolved cited work.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.746229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.746229Z digest=sha256:05d4a3de379e1cf56de17ae8aac3285e4a39dd1c422a43209eef543ffe452a3c

Observation 07d52c83-fe05-43e7-b10e-22b823fdc017 · outbound

This paper cites Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.137087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.829043Z digest=sha256:b0146c4854bbbf03338edf7fa7bcdde24e1ff3308d91976bde1dc3ae3f45f861

Observation 1c5359fa-7684-4b7f-928f-63b05c1448c7 · outbound

This paper cites Eigen attention: Attention in low-rank space for kv cache compression.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Eigen attention: Attention in low-rank space for kv cache compression

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.126325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.892816Z digest=sha256:c38027d6a961b1e66961af190a522f873539a1a2beb38d8d1b20012aac8f8c5c

Observation fdbb958f-2b72-49ba-abe2-c377589a0e98 · outbound

This paper cites Loki: Low-rank keys for efficient sparse attention.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Loki: Low-rank keys for efficient sparse attention

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.116849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:13.962201Z digest=sha256:e0e4f7b6e1d7d96890bbb6065acac9d0162f6582c3e6cd8f2b2467cb6ad46529

Observation 7d51c3f3-3efb-44ae-9b02-a22e32eeed36 · outbound

This paper cites Roformer: Enhanced trans- former with rotary position embedding.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Roformer: Enhanced trans- former with rotary position embedding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.107223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:14.040351Z digest=sha256:03065042cef39dc9cf05181677a6eedf3cb5f346ce632064a710017d210bcfa1

Observation 06966c43-271f-4c7e-8221-fa75e02016c8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Gemini: A Family of Highly Capable Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.106347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.106347Z digest=sha256:f58f64406f4a95a0fde41aad7473f2488b316d4481bae6033ecbccba7aaaf646

Observation 136be86f-4dfc-459c-addb-b7adc00bee00 · outbound

This paper cites an unresolved cited work.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:37:16.097187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:14.177445Z digest=sha256:13c4233cc5a5201ad77c103620be1645149fd1ae5033a8ead99db665d361df0e

Observation 7c385cdb-0eda-40c7-b86b-c8145ab0f465 · outbound

This paper cites Quantspec: Self-speculative decoding with hierarchical quantized kv cache.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Quantspec: Self-speculative decoding with hierarchical quantized kv cache

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.087682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:14.251524Z digest=sha256:d17ecedb142418e6531577536f1b6ae0829b0bf388ef761e6668a5801b099271

Observation 047f1126-b2f5-4fa9-9538-fa1d4a02bbd8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.314482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.314482Z digest=sha256:3a93a75fd115bb1394fe853cee45d20261af6865882edbd25a639c35d027c9e0

Observation 0ac099aa-ddd4-4725-bf1a-129e143d515e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.389766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.389766Z digest=sha256:b79b720873c0902b604cfa5df7067ad1c80ef007112f4dcbae377346663bbfd8

Observation 4a0741e7-31e9-4492-bafa-6d4b2456f1bf · outbound

This paper cites Attention is all you need.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Attention is all you need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.452375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.452375Z digest=sha256:14c47122037c30a119143bfdf73cb8397ebee98f2236da79cf091384ac29cd5c

Observation 4745bc3d-d048-4169-bed1-831077355f13 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Roofline: an insightful visual performance model for multicore architectures

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.495134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.495134Z digest=sha256:3c89837fe70d8213020e887d4e6b6cfff32fa6a1c9e416d5b56823fa11438d78

Observation 3720f747-8b2a-4343-9bae-c838fb1fd24c · outbound

This paper cites PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.599574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.599574Z digest=sha256:b6c56530689d46ef93fcecea79b178464bc8bee602ccaf8bb0b191bdac5ce819

Observation 15ecedfa-918c-4f51-bf87-f9124faf0059 · outbound

This paper cites Efficient streaming language models with attention sinks.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Efficient streaming language models with attention sinks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.064850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:14.615011Z digest=sha256:687885811c09221921edb7cf77738d88aece24e07bd84e0551c938e70d6f9228

Observation f5905c12-24a7-4e6a-b390-da8ece273cf7 · outbound

This paper cites On layer normalization in the transformer architecture.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization On layer normalization in the transformer architecture

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.637536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.637536Z digest=sha256:766b9de7e6fa316d958206e51243c52bd60f1085d47a75c0e52c525cbe81ad12

Observation cef110ff-40c1-4fcc-b6d4-7d18d26e9899 · outbound

This paper cites Recalkv: Low-rank kv cache compression via head reordering and offline calibration.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Recalkv: Low-rank kv cache compression via head reordering and offline calibration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.749127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.749127Z digest=sha256:e1913986344e49b3b05d5da535d5853773217f39fe3e941bdec05e46c0ecbdca

Observation 56482788-ff5a-4cc4-b718-e7969d227cea · outbound

This paper cites El- attention: Memory efficient lossless attention for generation.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization El- attention: Memory efficient lossless attention for generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.048265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:14.924191Z digest=sha256:8db5843c786dc39865e2986440c9c1093c873ab0c361e84f476bb1d2368c075e

Observation c7653971-b57a-42d5-8ea4-c3692c18ec05 · outbound

This paper cites Qwen3 technical report, 2025.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Qwen3 technical report, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.107048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.107048Z digest=sha256:df07aa94ac72010e2d62de082aba4191c02cf7c41400be229d26465536038de1

Observation 45cd5222-97fa-4974-a33c-7dbc48f255f2 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.322633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.322633Z digest=sha256:fa345f7c45eb76c495dd3b924955397ef3ee543e117241bec09311ed0bf9255f

Observation 84d11a4a-1084-4bf8-812f-086dfa615dee · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.475893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.475893Z digest=sha256:00d4890396574b79d5b62264866b8cf781cc19e4d2841cba50d1a83054b8fb2b

Observation 88e694ca-5ae1-49eb-acb9-af558ae55074 · outbound

This paper cites Root mean square layer normalization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Root mean square layer normalization

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.031243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:15.673100Z digest=sha256:662c165e16af485acf3b4c9313b41193413ef61702668c001e585d30d8ad3b1f

Observation 594bfca9-8323-4366-8d46-5c149d425ec0 · outbound

This paper cites LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.716757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.716757Z digest=sha256:8b3ccc9f96150f9ba9b3af0edc83a2732596674e86525d13b8dfdd2bb87c388b

Observation 94d290ba-be83-46d2-a099-fdd937a0a2c2 · outbound

This paper cites an unresolved cited work.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Unresolved cited work

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T20:37:16.020486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:37:15.720135Z digest=sha256:6d2f163e9d0d0f33de32f129d04c4fbb32502f70a8c1914205565d94a163faa6

Pith citing papers

No inbound Pith citation observations are available.