Pith. sign in

Paper Citation Record · LEDGER

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression

As of 17 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2505.06901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06901 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:36:42.436933Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ba44858-284c-495e-850b-baf81a56841b · outbound

This paper cites 2023.AMD Instinct™MI300X Accelerator: Technical Overview.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression 2023.AMD Instinct™MI300X Accelerator: Technical Overview

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:36:43.733598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.142378Z digest=sha256:c3ae14714be9b33675e47e004f8d0c2399d2339d65ebc32879e4728f06032e06

Observation ea690f7a-7b70-4fcf-a2ec-9f76244832d3 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.720182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.147134Z digest=sha256:84f62f6ed453211fa72d3782f59f8c0cb67843d7b5ac5bbaef23c37a0ab5403c

Observation 0eed6b22-e74f-4328-8951-48535e6c1226 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.151190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.151190Z digest=sha256:3c6da3cf8586c984dd6df452c3905a5fa8f89cae07e8bc9d933967a87f2d1793

Observation b7128f30-57b7-4429-8fac-d2d768147546 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.155706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.155706Z digest=sha256:753758b5fb5d9d4702fcf58e985f138297114eb907f523f3289a4b518e166a64

Observation 1b6e17f6-115b-48c5-aebf-a2a510d2fd6e · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.164075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.164075Z digest=sha256:1d9beacbf8c0bece4d9b97917dff22eda4b4f475ac8985d7cf3b7a6512c9ee49

Observation ae70cfc8-7e2f-4f06-944a-2f544117442e · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.692728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.168378Z digest=sha256:6a0dcddbdd18dec3be9dedd0f16fdb5ee9f41b5e26fdf1b1438b98e241dbcc1c

Observation fbbf2c78-8a1b-4e24-90b2-eed0231aac3c · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.173455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.173455Z digest=sha256:fb3f3c1149cbc8fab77de5143f7bbb73ab8857d0ffd77f48f201794480fad812

Observation 7c1d2d16-ceb5-4ce0-8c86-c496ac3df8fe · outbound

This paper cites Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUs.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.177685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.177685Z digest=sha256:60cebdbbcbab32d4614f35abeeca9280bb60ffb2b065e62d9759f3f5141502a8

Observation 01bedb1e-4807-42c1-bb62-9aad6cee91e3 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.182102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.182102Z digest=sha256:932c929ff6580d961216e58e7d7b86a0dfa38e508569ef3897b1713fbf927ed9

Observation bfbe14cd-d1b8-4b1e-8da5-b89d2741f42b · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression QLoRA: Efficient Finetuning of Quantized LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.186179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.186179Z digest=sha256:301c3e8db419a7594f52fccdf94f7c5dbea5c2a6eff6273a710d8344b1e95452

Observation 2d37c38d-2f87-42f5-8d8a-fc66dd6bc80d · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.194595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.194595Z digest=sha256:b6703b63db822875940982227f756b34a98e36b55310b5aad2ceca6d71889710

Observation f2bf3ee0-e093-4a24-8b77-744ac2ad06e7 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.198654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.198654Z digest=sha256:d0ea8ba27bb3a6855fa60d4ba30f9e9f4d5ff0ceea3222f15f3202efad6b8559

Observation f2bf3a08-1563-46c5-9e40-93449848118f · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.202764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.202764Z digest=sha256:118e33e5197f8c75815783dd3a592597e03441ad156f8f18e6e48a28b91788d1

Observation 18fb1da2-22d7-4652-ac62-c5f9d2f85d66 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.206873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.206873Z digest=sha256:9dd6866882d41ec57b99a243c75e129f7a5c9ffbdbecd17ac61f271e0d0f5ad6

Observation 93d97360-809d-4fc6-adf2-2ffe738e7b34 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.678971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.211016Z digest=sha256:e66c28e9f049e21d668cda6a71c7f4a67ee6c086247b3d52449509feeb48451b

Observation dd0b94f4-1226-47bb-98c5-b4e2392e10c8 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.214966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.214966Z digest=sha256:8a260cdcb0fdf509646c06ed36d7b0cd7d55b7d0825ce4e5c88ee2c009ab29b9

Observation d70f915d-8b2d-4619-b5af-620a56bd5df9 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.653270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.218980Z digest=sha256:c7e5ed0c08036c279632a9404426d891e7ad8f37e7f243b29a660e05b9e0d1a7

Observation 186bc34c-753a-4e13-91ff-94fc7b7db737 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.227106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.227106Z digest=sha256:b39c9cc414c1ab4b21ce5b70863991a99880809e056081d6de4c06bfb96e251b

Observation 1ae65579-1ac3-47c5-b7ac-dedb2e444839 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.634748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.231005Z digest=sha256:54e7e8fd75ff224bb40e80a839fdfea6afa23714e06350934859787dbdf34513

Observation b67f1e06-250e-4912-bb6d-192b6ccb2219 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.234680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.234680Z digest=sha256:94bc1dad5bfbc6c4132dc3d20a6a79f0819d8a708e06d755ab7736455ccba543

Observation 24195b5b-5e41-4a25-8dd4-f8e4a5affd25 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.238452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.238452Z digest=sha256:04aedfe76b2364b05e26b2fba195ab284f3117c6fed76c661199e33ef315aa7b

Observation d770ce9c-ac0e-46d1-b221-fd2f52a04d89 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.242177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.242177Z digest=sha256:239a00190c56d3c6b52c390f1e024c28d5bdfb782eb7897af4936b7233c2460f

Observation c8fb3e41-15a1-4341-885b-58b97e80ee96 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.608129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.246005Z digest=sha256:cda4d9482c3e81911dad7e10c24dba0a625e6b500a7cddef08d73044dc7eec48

Observation 01eee136-35ac-4030-ae3a-0f9c4bdf5391 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.249513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.249513Z digest=sha256:4e397089c88530a629a5df534c0bcc229dd7f56183d4bd43e90d5d64d15bfb11

Observation 6dea8987-df02-4fa3-b4a5-5dd2d129cbd7 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.595849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.253339Z digest=sha256:069663a6525acc9c9adcc00ddebc5181ea99d05089b790e38d35532e11bdf951

Observation 1bc47a00-7d92-4512-be9c-e6aea921184b · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.257055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.257055Z digest=sha256:e9a39c229b99f965d66a94affc9bc49f1b659aede71acc984231e45ea95c8b78

Observation d0af8b2f-7d8f-4b4b-b2a7-10822f65be99 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.261185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.261185Z digest=sha256:a336663fa396bf967d2cdff11c5bb51936d6160d13450af5170aab705857a8bf

Observation b10a2068-e407-4559-a329-69e4a0fc8533 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Towards Reasoning in Large Language Models: A Survey

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.264954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.264954Z digest=sha256:d7a14244182d2bbc92bf976e2afc698b21530a09237356a6a898d2b34986ad44

Observation 49833275-173d-4355-bd48-fe13ad47a197 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.576643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.269030Z digest=sha256:45e1d185c2c7b4b9ffac41fff75a8d2388383c7e0d57813f619722b5283fe31e

Observation 0f615f21-e179-4fc4-a69c-edada259c8f1 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.563348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.272739Z digest=sha256:f7e88dd9e849c327e302812eb4e1e33f8fd05d484f7acd3586ad50ebdaad1b1a

Observation d5f07562-c26f-4da9-ac35-d13c1ceda427 · outbound

This paper cites Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.276351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.276351Z digest=sha256:0851366fac08dc86c8cf417f7916040e5cb96e40c5efcdf0399b32f3001f28ab

Observation 5c6b69ca-48d9-4583-b5e2-4408869278bd · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.280103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.280103Z digest=sha256:8a55d5ae7f8e6a6a6db8943ba9ccd56111a3d7777f29b1edb45594bbf776873a

Observation eb5746a8-9e82-4021-9f07-d2c0d63d8665 · outbound

This paper cites Mixtral of Experts.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.283492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.283492Z digest=sha256:6a9ea39f6ac0849204b3ef84f94501af739f6e29b35ba75b75b5ad3ad87b047c

Observation 4707d447-edad-4fee-b2a6-86716e015dc7 · outbound

This paper cites Scaling Laws for Neural Language Models.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Scaling Laws for Neural Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.287406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.287406Z digest=sha256:96ad9b6bc678a0b85d5e128fb6ccc18e4366c9151f493a8b918f28dd5f3c4e81

Observation 4363da4d-df96-41ee-a730-3d5fdbc79550 · outbound

This paper cites Aamodt, and Timothy G.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Aamodt, and Timothy G

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.291153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.291153Z digest=sha256:0736c859f2e95e0bd4bcea63c9d337286e170bedbf2ede6bf2cd890914ebc195

Observation bfaae37a-2dbb-4bac-9dff-6cb3a04119de · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression SqueezeLLM: Dense-and-Sparse Quantization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.294573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.294573Z digest=sha256:dc3efbf8eb263bc2fb0e5a88424ccbe72e85a298da74bd1470a20ee7619d3bf3

Observation 2d0e2cc0-40ab-4dca-a2b6-45e6a1e09118 · outbound

This paper cites Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.298592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.298592Z digest=sha256:face93e1165cb2ad661a8ff12eba11bd5e9b107dea92ab603601566649a7d987

Observation 85c98218-db0e-44fb-84dc-48c290e92afc · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.302129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.302129Z digest=sha256:4e2d042715ce69c6f9fe5f85fab38ceea10fa66a8b08671d54ea4ee8a8e6dd8b

Observation 2407bd37-3565-4402-8567-0cb1374c2dc2 · outbound

This paper cites Analyzing Machine Learning Workloads Using a Detailed GPU Simulator.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Analyzing Machine Learning Workloads Using a Detailed GPU Simulator

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.305885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.305885Z digest=sha256:3ef6152544a2d9008975babc3da8d71125fc7d379a6b732c3ca8b3be2447cede

Observation 1e02f5d3-6361-4b28-a192-c7398f968843 · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.309597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.309597Z digest=sha256:da6941fb515fe0802786431329aaa662cd178a72b743e4df16f074415f6b5cb1

Observation d465a060-e80a-4919-bbf4-81f4f305f7e1 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.550666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.313451Z digest=sha256:2f0b99af3fd72a51a89b83451808cbd037916cdf2881f3d73b7cfa209f6ce9cc

Observation 8cb22b35-3b3d-4c3e-a0e3-0b8e0b465e69 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.317271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.317271Z digest=sha256:186eb7aa3703cb4eb12b7940f3106e96f015e00600a69f92b5878a8324200fc0

Observation 1e983ade-44c3-4dad-ab47-200baeeeffe8 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.538976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.321134Z digest=sha256:f09ec56ef8ae17b9970098c0907bc8f00eee5ea495f51d7c73e8fa8690836a45

Observation 749c5ed9-e825-4c75-bbe0-2a5a436704aa · outbound

This paper cites Pointer Sentinel Mixture Models.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Pointer Sentinel Mixture Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.324901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.324901Z digest=sha256:905d809cf4d132feeecec40a8fa46f73bc72f9e5cf155eed3ab6d71294d42c98

Observation 78ecc4e0-8103-4f4f-8ca6-e92f957a1b54 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.328921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.328921Z digest=sha256:49c97ae27cfb4c587d289c22d6ddbc8a2ebd10fb932ef92fd834055912dc2a6d

Observation 561cd770-6ed1-4ea5-ad76-3f3c5a6997c6 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.517798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.332424Z digest=sha256:89ae7d041532616b10c1e0a1cd63bd64045347f280191a9df9cf7d188d999de2

Observation 1959a8d6-55bb-4f2f-9ad0-b9f788f61cf5 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.505307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.336111Z digest=sha256:e7a5793cdf241767b47ff521c9227f02bf873797ba3c64b558ceff14d546b74e

Observation 09ba496e-b8e2-4b7f-8ffb-ce4a7b2b6052 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.494255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.339348Z digest=sha256:ec4ea13bc6f28cd5c1d39fd00c2174103d75b042034f99338cca738edd057b7c

Observation 2ded34c2-9efb-4be0-8fb9-53f8733a70f5 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.483004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.342687Z digest=sha256:d4920ef6faea7a5d84076b2ab0f4cd848e78f20ea289a60b5b044096bc7d1b5c

Observation 7797d6a1-0c47-4245-a646-81215e2102a0 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.470944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.346148Z digest=sha256:1b0c1e5e6cd77a10d36d0d05f3ec1457a19d8540df57f1ce075440330168563e

Observation 0a8891b9-3be5-449e-93c4-c10578a64d1a · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.459409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.349595Z digest=sha256:f871ec4443f1725bf648af6fcbdfbd82e3603bae59011464030fda42fa54de4b

Observation 0c89043f-eb67-43f1-883e-856c0c9ffe4c · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.447328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.353237Z digest=sha256:6c5dce8b3ab7ebe7e46b6031dcc3a6cf1c5b86d13342036179fc8444cf6022fe

Observation a3f0707f-5710-4a6e-8d8d-ab54cc6ae2ac · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.435447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.357141Z digest=sha256:9d164e5e86d201d98cac45cc53dc006ee02642f60b179088ba669653e0a3b109

Observation f639c17c-6756-4a77-9ea4-de74e1098382 · outbound

This paper cites GPT-4 Technical Report.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression GPT-4 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.360730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.360730Z digest=sha256:02706b4a8a5054f7591a92d4c78cf60a31d45d6a79698ac826ca6d0c097b6f3b

Observation 5fb55292-f727-4954-aaa7-ff7ba3cff7bd · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Splitwise: Efficient generative LLM inference using phase splitting

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.365015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.365015Z digest=sha256:7c9dd7953f032fa78016e6e93ec6fb9d46d47fe8580069f4184b6f8fc1d5cffa

Observation c8f90535-7c58-416b-9013-06d8fde66837 · outbound

This paper cites Gibbons, Michael A.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Gibbons, Michael A

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.369077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.369077Z digest=sha256:72eaa2e2c223c0bc5b07c2d4867635cc5f07214e8c20a000b65cfd84c1eb0dfb

Observation 78c62cb9-34d8-44db-af9c-746c0a9193d9 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.423588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.373488Z digest=sha256:fe8afd70dd8529020379458a23c5b7615344084bce58328bfaf72aa648a3dae5

Observation dd563f12-7bdd-4e50-9187-edd04ebc34eb · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.378095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.378095Z digest=sha256:fbfad628e434286c8c7231c603ee858d1b30888b7c88459897b6854a87fa8e19

Observation ff5ed6d3-412c-4f4c-993f-4ffb51d5628c · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.382309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.382309Z digest=sha256:0494a347717f4efe84116a1ce2ff84fa298968139dcbe9ddb5fe0f8d9828e82e

Observation ede45b41-043e-4c1d-a40a-bc47571200ff · outbound

This paper cites 2017.High perfor- mance computing: modern systems and practices.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression 2017.High perfor- mance computing: modern systems and practices

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:36:43.404917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.386187Z digest=sha256:f71cf8b17b0899a2966e9cdc57786b8b6b8917decc6ffb16c33255a36cb4c1b0

Observation 60f7d2ff-7b0e-47c1-8f07-cec00f912d6d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.389811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.389811Z digest=sha256:aa0a9732e8de160ea6d68ef2c83f1699c4fb266d1ea62d35b97e05ccbd9872b7

Observation 01b85d50-cf37-4668-9877-8166b9afb727 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.393553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.393553Z digest=sha256:2ec904c4be33c86a496acfdb981754b8fe6449c36195589490f86912f955c88f

Observation f28783a7-fb44-4199-9980-abdf5031f734 · outbound

This paper cites Attention Is All You Need.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Attention Is All You Need

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.397502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.397502Z digest=sha256:da0e3a7c339390796001e011a86a37a66cf2d778d27356a4689eaf1f3cb61a04

Observation 0f32195b-1722-4334-afd9-6948acc521df · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.401214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.401214Z digest=sha256:0fd3802a95550c75b8b44dad6504d7dfb8e5450d6c74a4fbdcbb264f05acfb5e

Observation 6c557f79-84ef-45c9-99e8-06f5dd0dcf0c · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.404990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.404990Z digest=sha256:ae8413bcf0b6ca16937709115f0933d49d86a1c426c5f3f60f1adf0513943099

Observation fa450781-395d-47fe-ae02-8d469f824514 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.408693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.408693Z digest=sha256:a8d298285c9646a4ba87d5448f8dc955f14838ce85f26f1c7b5aeaeb89ceaa70

Observation 1d1036fc-d1e5-4d1e-9812-3c6eb9964d61 · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.412472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.412472Z digest=sha256:18eeaf2a1da7c9f987b35df6f3f11a3f2d454a30179a8a4ae2443375cc264d21

Observation 9fb2d416-6127-4d2b-8aee-b54a14438a3a · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.420509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.420509Z digest=sha256:39baffe7d81d12ea455665c2b8ffa0190c403533f80cebe704eff5133d842c49

Observation 63c20992-d18b-46b9-b860-4613b516d4f0 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.424522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.424522Z digest=sha256:f71196ca9133ce2d88bb837aa8f23c341f69b20fc9625c024ef234577c9cbad9

Observation 212a593c-b1d5-4228-8c15-29aa3e597ce1 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.429090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.429090Z digest=sha256:9a975617567daa707cb964866f0034186409eabb336dad1ed5d86b3459620a27

Observation 58e2e9f4-a798-4e41-9146-566e93cbd31b · outbound

This paper cites an unresolved cited work.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:36:43.369704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:36:42.433152Z digest=sha256:74d179de69f5455b1e9cc98cdc626dfbb5260dff1ad183f8b5478194b46a24ab

Observation 0c60819d-faaa-4ca9-b939-0cad6dba670d · outbound

This paper cites Qplacer: Frequency-Aware Component Placement for Superconducting Quantum Computers.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression Qplacer: Frequency-Aware Component Placement for Superconducting Quantum Computers

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.436933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.436933Z digest=sha256:500d3adb654f4a6a38f24e5ec3c3f21a0ccd142e5f49d4c623716669effc2c27

Observation 55673ec5-571b-488b-9c6c-d0ab1bf27d3a · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.159428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.159428Z digest=sha256:bfc9e5182afa134c2ac8b3a0464323a19443e3ac93b710f638d0210c85fb697a

Observation 18894168-e18e-4da5-90f4-cf39fb7aec54 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.416370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.416370Z digest=sha256:cdf973298f5faf7ff64c8f2fa2e9f0b902cff9a481bb264efa6105a5b6fd65ab

Observation a243c6b4-75f8-42d5-a34e-428b35f5f34c · outbound

This paper cites https://doi.org/10.1109/MCAS.2024.3476008.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression https://doi.org/10.1109/MCAS.2024.3476008

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.222971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.222971Z digest=sha256:4714726f67695a647c9f87d7f6d486ac7a30f6bc0b1eaeff3679f571e165e1a7

Pith citing papers

No inbound Pith citation observations are available.