Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T18:16:30.369786Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 3 inbound Pith citation observations for arXiv:2604.02985.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T18:16:30.369786Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:14:54.909730Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-12T06:31:28.986900Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 28adab7b-4bb4-46c1-b55a-f03ff7e762c0 · outbound
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 294b8e5b-28af-49a6-a275-001d2993b86a · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Pro- ceedings of the 38th International Conference on Neural Information Processing Systems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6dfedbf5-2e3c-4b6c-9905-913c5cdc7f59 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 06e7ed54-3b81-4e9a-8b24-e2f82241ad59 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fce97cce-7bf8-46bf-a2f0-20ed45f1ae59 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Extending context window of large language models via semantic compression
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bb3a37f7-6184-4295-a236-88ec6b99c30e · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=uREj4ZuGJE 14 C
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a709b686-b396-424a-81fd-a0bb5d898963 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b024978c-9d65-47f0-97a2-45b4c624ed1b · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Workshop on Efficient Systems for Foundation Models II @ ICML2024 (2024),https://openreview.net/forum? id=vs6CCDuK7l
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f50954e-3a7b-4f1b-82e5-be63799b5d37 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Mistral 7B
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e092879b-010c-4e76-b4e6-2919daa85f4b · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference LLMLingua: Com- pressing prompts for accelerated inference of large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc270d1e-9054-46e8-9e7b-cd5ef2c721cd · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Zong, C., Xia, F., Li, W., Navigli, R
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ed8b035c-aa4f-4dc7-ad64-f7baf90776b5 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Discrete prompt compression with reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 889c6aa0-9592-4ea1-adbc-058a886ac5f4 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e68f50bf-952f-414b-a0d8-abe9c5c5ca06 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Dec 2023).https://doi.org/ 10.18653/v1/2023.emnlp-main.391
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a32d290c-6929-4d83-abb1-2a667db96e6b · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Findings of the Association for Computational Linguistics: EMNLP 2023 (Dec 2023).https: //doi.org/10.18653/v1/2023.findings-emnlp.655
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3870341-3f76-4baf-989f-b65c4e757e1a · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the ACM on Web Conference 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 039af329-a585-4312-b9be-2376997d751f · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the 38th International Conference on Neural Information Processing Systems
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4967f847-cae6-4b96-8113-e5c5b2739600 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=mqVgBbNCm9
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8b7f5455-abe6-440d-8c4a-4fb04726474d · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Findings of the Association for Computational Linguistics: ACL 2024 (Aug 2024)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aaefbc62-f481-4607-98e9-0afe246ea94d · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: Proceedings of the International Conference on Modeling, Natural Language Processing and Machine Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2f4dcc1-5fdc-45aa-a1f3-fd409fbe5c17 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Learning to Filter Context for Retrieval-Augmented Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6030a2cc-fcb2-4875-98f0-03e9fe969c6b · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference Transformers: State-of-the-Art Natural Language Processing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad49da04-d1a2-4ba0-a544-f3635ebd9eb6 · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference In: The Twelfth International Conference on Learning Representations (2024),https://openreview.net/forum? id=mlJLVigNHp
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 847a7ab5-94ff-4be4-81a0-62f5544971ae · outbound
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference A Survey on Efficient Inference for Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ce6a5940-5854-41ba-ba70-53a73b0aba63 · outbound
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8675d93d-35d1-4324-a44f-73677f4eca8a · inbound
Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 648fcaa2-331e-4c2c-8505-73091bc99a59 · inbound
Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f79c36e8-5ca0-49ed-ae22-5db0b9c5880d · inbound
Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.