Pith. sign in

Paper Citation Record · LEDGER

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference

As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 3 inbound Pith citation observations for arXiv:2604.02985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.02985 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T18:16:30.369786Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:14:54.909730Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-12T06:31:28.986900Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact12
  • verified fuzzy4
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28adab7b-4bb4-46c1-b55a-f03ff7e762c0 · outbound

This paper cites LongBench.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference LongBench

Reference 1

Resolution
metadata mismatch
doi, observed 2026-05-13T18:18:05.621702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:564acef7f1ec0a71fde99817b8a3693108586f2c5bbe5db89de3e7fa0e778896

Observation 294b8e5b-28af-49a6-a275-001d2993b86a · outbound

This paper cites In: Pro- ceedings of the 38th International Conference on Neural Information Processing Systems.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Pro- ceedings of the 38th International Conference on Neural Information Processing Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:05.625102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:93dd440860dc351d708a71c097658e3cb630639265b1af5b3f287458bbfaeba5

Observation 6dfedbf5-2e3c-4b6c-9905-913c5cdc7f59 · outbound

This paper cites Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen

Reference 3

Resolution
metadata mismatch
doi, observed 2026-05-13T18:18:05.632931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:f2fe6bba150c4e23f3993e581a996fc4cd0aefb4d2c49f292b29cd6533cd9222

Observation 06e7ed54-3b81-4e9a-8b24-e2f82241ad59 · outbound

This paper cites The Llama 3 Herd of Models.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference The Llama 3 Herd of Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:06.145521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:290415eb407f2c20a94e5fa9e1cffe6fd25465af6b89fb1afe277c53355d917f

Observation fce97cce-7bf8-46bf-a2f0-20ed45f1ae59 · outbound

This paper cites Extending context window of large language models via semantic compression.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Extending context window of large language models via semantic compression

Reference 5

Resolution
metadata mismatch
doi, observed 2026-05-13T18:18:05.630303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:bd46428b18add7e14ee135acb73d4d4ca52d1456141ea692e624d6a74b33fd5c

Observation bb3a37f7-6184-4295-a236-88ec6b99c30e · outbound

This paper cites In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=uREj4ZuGJE 14 C.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=uREj4ZuGJE 14 C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T04:26:38.008610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:886b1c27789b63878d73bc7f0f558a184723c807d05714eb17eba96049d91f6d

Observation a709b686-b396-424a-81fd-a0bb5d898963 · outbound

This paper cites an unresolved cited work.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-14T04:26:38.012938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:8d267621e520fb374ef3ca7d90d59201981ec34e911ff8818982a5f32a0d70cd

Observation b024978c-9d65-47f0-97a2-45b4c624ed1b · outbound

This paper cites In: Workshop on Efficient Systems for Foundation Models II @ ICML2024 (2024),https://openreview.net/forum? id=vs6CCDuK7l.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Workshop on Efficient Systems for Foundation Models II @ ICML2024 (2024),https://openreview.net/forum? id=vs6CCDuK7l

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T04:26:38.018084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:48377a4c5681f6c346f813b0ba66524f70af64ce9b072e2af195c9b7af4112d8

Observation 2f50954e-3a7b-4f1b-82e5-be63799b5d37 · outbound

This paper cites Mistral 7B.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Mistral 7B

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:06.149281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:776f4ab77fb55d526ae0e238cfc73811ca6b1965ac9eac0c9960960c0e6b8794

Observation e092879b-010c-4e76-b4e6-2919daa85f4b · outbound

This paper cites LLMLingua: Com- pressing prompts for accelerated inference of large language models.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference LLMLingua: Com- pressing prompts for accelerated inference of large language models

Reference 10

Resolution
verified exact
doi, observed 2026-05-13T18:18:05.616939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:a012ddd21666492f951bf653b528bd12f902972ca8ed03c58db68b42eb9a3597

Observation fc270d1e-9054-46e8-9e7b-cd5ef2c721cd · outbound

This paper cites In: Zong, C., Xia, F., Li, W., Navigli, R.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Zong, C., Xia, F., Li, W., Navigli, R

Reference 11

Resolution
malformed identifier
doi_truncated, observed 2026-05-13T18:18:05.603082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:6629a89a361225b9e11d5bc13fb5aab22ebfea1dafe840dc2660da1d5fae1c0c

Observation ed8b035c-aa4f-4dc7-ad64-f7baf90776b5 · outbound

This paper cites Discrete prompt compression with reinforcement learning.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Discrete prompt compression with reinforcement learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:05.588591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:33a607938345323b0d6df1277c0b1866745e591be823f289fb674674360886cc

Observation 889c6aa0-9592-4ea1-adbc-058a886ac5f4 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:05.611089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:4f3ecb40d379ee717135eeb6f6f924c59e7de781d7a83911e3f4b31b39dce8d3

Observation e68f50bf-952f-414b-a0d8-abe9c5c5ca06 · outbound

This paper cites In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Dec 2023).https://doi.org/ 10.18653/v1/2023.emnlp-main.391.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Dec 2023).https://doi.org/ 10.18653/v1/2023.emnlp-main.391

Reference 14

Resolution
verified exact
doi, observed 2026-05-13T18:18:05.613504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:54e08bb5f59cb59f59c33fada36d2f3a693006ca02bbef5b83eecd5ad4d5e827

Observation a32d290c-6929-4d83-abb1-2a667db96e6b · outbound

This paper cites In: Findings of the Association for Computational Linguistics: EMNLP 2023 (Dec 2023).https: //doi.org/10.18653/v1/2023.findings-emnlp.655.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Findings of the Association for Computational Linguistics: EMNLP 2023 (Dec 2023).https: //doi.org/10.18653/v1/2023.findings-emnlp.655

Reference 15

Resolution
verified exact
doi, observed 2026-05-13T18:18:05.591527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:d1fd6041fc2237c9893003968341ea3a146e67d17410faee7bdd0c4e8ae04af8

Observation d3870341-3f76-4baf-989f-b65c4e757e1a · outbound

This paper cites In: Proceedings of the ACM on Web Conference 2025.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the ACM on Web Conference 2025

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:05.597918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:62ea849210ea9fa2b496d9fd872d038d645a750820b8a9e872e299fa6ee503ca

Observation 039af329-a585-4312-b9be-2376997d751f · outbound

This paper cites In: Proceedings of the 38th International Conference on Neural Information Processing Systems.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the 38th International Conference on Neural Information Processing Systems

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:05.606325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:c7b31dc2884cd2fbc481fd33fa92c1c2a87614d6748a8e7ae293cb6baaccf769

Observation 4967f847-cae6-4b96-8113-e5c5b2739600 · outbound

This paper cites In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=mqVgBbNCm9.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=mqVgBbNCm9

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T04:26:38.022648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:8de0380dd140e58b40268b27e60d31b4751c21c281d05d75ec32631330568bfa

Observation 8b7f5455-abe6-440d-8c4a-4fb04726474d · outbound

This paper cites In: Findings of the Association for Computational Linguistics: ACL 2024 (Aug 2024).

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Findings of the Association for Computational Linguistics: ACL 2024 (Aug 2024)

Reference 19

Resolution
verified exact
doi, observed 2026-05-13T18:18:05.600651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:77c3de8d7d540f6fabf4387a62d7e89b3564493327f080858b3801c811ee436f

Observation aaefbc62-f481-4607-98e9-0afe246ea94d · outbound

This paper cites In: Proceedings of the International Conference on Modeling, Natural Language Processing and Machine Learning.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the International Conference on Modeling, Natural Language Processing and Machine Learning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:05.594744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:72dcac0beb084a886755683d27cf5f47ad149515b6415eb2dd1e0d086b981533

Observation b2f4dcc1-5fdc-45aa-a1f3-fd409fbe5c17 · outbound

This paper cites Learning to Filter Context for Retrieval-Augmented Generation.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Learning to Filter Context for Retrieval-Augmented Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:06.153973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:fcd0d6c2005814cefd73852013f149212afceb25e78f0f267f44b771ec83d070

Observation 6030a2cc-fcb2-4875-98f0-03e9fe969c6b · outbound

This paper cites Transformers: State-of-the-Art Natural Language Processing.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Transformers: State-of-the-Art Natural Language Processing

Reference 22

Resolution
metadata mismatch
doi, observed 2026-05-13T18:18:05.627950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:f304a4783b16c0db6f4d4c9d074ce4604fc719dcfaa35e7f9ddfab0bba7ee5cf

Observation ad49da04-d1a2-4ba0-a544-f3635ebd9eb6 · outbound

This paper cites In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=mlJLVigNHp.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=mlJLVigNHp

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T04:26:38.027167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:bcb9b47c9e07402f1b0e1e1b16d5a6040a1bfed6850e8cca37b501a348451b74

Observation 847a7ab5-94ff-4be4-81a0-62f5544971ae · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference A Survey on Efficient Inference for Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.856027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:63adf5298fd19535289cb9b749a8ec5b2cb41fa975adfe8ec86c2c0a4ba51f50

Observation ce6a5940-5854-41ba-ba70-53a73b0aba63 · outbound

This paper cites ACM Trans.

Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference ACM Trans

Reference 25

Resolution
verified exact
doi, observed 2026-05-13T18:18:05.619283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T18:16:30.369786Z digest=sha256:c88c27a814e7b186f206bbf564bfac698e23cafe3306b972de884fb612041059

Pith citing papers

Observation 8675d93d-35d1-4324-a44f-73677f4eca8a · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:26:24.190105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:4f7e0eccaa6870cfc1f9577aa218b12002a22084f2e34c71879c89f780a1ada5

Observation 648fcaa2-331e-4c2c-8505-73091bc99a59 · inbound

Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference cites this paper.

Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:31:28.994381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:11:02.387702Z digest=sha256:fc506bb7d8dc23fd623e20013152a7282663da04b1573690e415127a8d97fd19

Observation f79c36e8-5ca0-49ed-ae22-5db0b9c5880d · inbound

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching cites this paper.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.909730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.909730Z digest=sha256:40adf223062937773d3a053771abc5b065f5ac7cb499cb8aff9a36012d111dab