Pith. sign in

Paper Citation Record · LEDGER

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 73 inbound Pith citation observations for arXiv:2310.05736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:05:54.122394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.322096Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f72dd5e3-3850-44c9-8af0-63c3c00b9f94 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.647448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:01385a8fafcef7ab3f30deba521d85417d8c25686211d55cd70ee5bb5b9dee4a

Observation bec1097a-429a-4dc0-bea3-2c1bb52f6112 · inbound

AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models cites this paper.

AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:08:26.041444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T21:06:59.776841Z digest=sha256:a3c4a3342e726b502b604c8e8c09a50a503cc4d9a18427eca3590cca3e1f4963

Observation 038d4429-156d-4e4e-89ce-8cdbf1905052 · inbound

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning cites this paper.

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:38:24.967159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T20:36:55.159302Z digest=sha256:a4741bb418c9154b67c4c70cf18da005ca3f7ff21ff2e121d4e77cc591fc31e5

Observation 911414fe-e58b-4241-b153-8be80d211849 · inbound

JPPO: Joint Power and Prompt Optimization for Accelerated Large Language Model Services cites this paper.

JPPO: Joint Power and Prompt Optimization for Accelerated Large Language Model Services LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:40:21.563948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:40:21.563948Z digest=sha256:f40ce901687d9cffa9ada1faf45f0c28e9e1ce4a5d45e0fbdd460c2718dec86c

Observation 451bc9a2-8be9-46c1-a39b-5fa59b2ba18a · inbound

JPPO++: Joint Power and Denoising-inspired Prompt Optimization for Mobile LLM Services cites this paper.

JPPO++: Joint Power and Denoising-inspired Prompt Optimization for Mobile LLM Services LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:32:20.773340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:32:20.773340Z digest=sha256:a1e9d98bf09954cef08c701a08b74b7f0053dbfe0bb32122cec28532772cbf35

Observation 7794ffa5-db67-469b-b98c-b57fb734b362 · inbound

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens cites this paper.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.184668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.184668Z digest=sha256:ff38786d48143e344a2e94a5fb0b0f38e923d62d17996855323464d17441e143

Observation b1bc1a00-5a49-4a02-bf38-f9fe7df36884 · inbound

FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing cites this paper.

FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T14:57:27.548016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:57:27.548016Z digest=sha256:b06cce5beee1872480be51adbabb0e249e0f6fac7d5977b0bf5912dce45a2594

Observation 0deb27a7-78d4-4780-8ed7-2568d03ebad0 · inbound

C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness cites this paper.

C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:03.890113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:47:03.890113Z digest=sha256:f5d4a57d37df516d9f23960ee658737627e36b7500b057f1d739462fae90c1a4

Observation 6afb25e3-bfd5-40c4-a724-80806a27129d · inbound

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations cites this paper.

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:47:40.400393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:47:40.335477Z digest=sha256:1b68cd769dc49d764a7c83dbc612a231ac2a8c29a3853bd2b1e2497df92c732e

Observation ba44d3ee-c20e-4333-8ccd-0f54d631d154 · inbound

EvoWiki: Evaluating LLMs on Evolving Knowledge cites this paper.

EvoWiki: Evaluating LLMs on Evolving Knowledge LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:04:20.131090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:04:20.131090Z digest=sha256:75dcf1f4cc703a333398d967c260c84ee5cc34b9605a59feead0f2661d771252

Observation 6108fa53-a800-42be-9fb5-fa42c178d672 · inbound

EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent cites this paper.

EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:51.267386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:51.267386Z digest=sha256:9ee5dcdaa2e8d7dee2a1e29e211e3a9bea88cc248316a46393c78eb6522d45d5

Observation 8b42b6be-e6f7-43dc-bd46-68648dccc2f4 · inbound

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG cites this paper.

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:34.859209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:34.859209Z digest=sha256:bcb8d5aa564c40e5100847199ea003a64fb19397f331cd4476048cdfe37a635d

Observation d5a994b4-e873-4a42-8700-6d5efc7ef278 · inbound

LeMo: Enabling LEss Token Involvement for MOre Context Fine-tuning cites this paper.

LeMo: Enabling LEss Token Involvement for MOre Context Fine-tuning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:11.074196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:26:11.074196Z digest=sha256:a61d1f6206329302bbb2b5b0474796a3fe941612e6063629b6d1bf0d039c99a6

Observation c0805d18-4818-4e42-8e52-cf6f52f50a36 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.685933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.685933Z digest=sha256:3df1efc25e74c4e233ed68b82a195fb767d50cf7c870398fae5a908ea3dcd452

Observation e48c0a43-7e26-48da-9d2d-4c97cf3dd3fa · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:46:30.078718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:07fba49d0d975b04a3030b27e7075686077b40297e459c4522e10f8a8d117ad8

Observation 6e2ceb50-3c8b-41dc-9d28-4039d3bc7f44 · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.491572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.491572Z digest=sha256:b9cebc8d4130545a06e3538c21d375726b4f94d17ea5bfecb0347ab5b00a1588

Observation 030cb13d-2a63-4871-bc1b-00bf228d3047 · inbound

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression cites this paper.

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:05:54.122394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:05:54.122394Z digest=sha256:62cb023af397b67dd913aff024d8d617db809f64cfeea7c40194e068992ef1cb

Observation 2cbcf88d-b56b-401c-8ef1-a4aad3be0a73 · inbound

MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG cites this paper.

MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:43:21.595446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:43:21.595446Z digest=sha256:6a4cd5a532bff457dab8643a3e64656b77375ca61d3eab3ad951a8932b22aa61

Observation e4058a9f-2dcd-4658-b504-75b78e3cd4c5 · inbound

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models cites this paper.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.983020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.983020Z digest=sha256:e914b2e80a8533e94f682cd9189fea2ffa0d1013fd313e68d14cdb860619e019

Observation 445cf6ed-50af-4bb1-a32e-798ee654dd31 · inbound

QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization cites this paper.

QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:52.446651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:52.446651Z digest=sha256:1a7a254115e2d0a22bb63baf4c72053e8eb0596393a169e7f463414b1f7ac109

Observation afafabaa-3390-46ed-81b2-6029bc351209 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.213167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.213167Z digest=sha256:bc3a7107e8b897b6cefe13ab0f66e44e47cb9df20bc69fa3d0d90beac9a78b1c

Observation 7fb1ec33-d153-40f9-9093-18fd3a040d60 · inbound

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling cites this paper.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.729533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.729533Z digest=sha256:7f4a63a89c82d81c6802439f75444a596cf78ae0aa11b37397152dea22376d7d

Observation 0427e401-f35a-4281-b7c8-c7fb90d3b9e6 · inbound

Lossless Token Sequence Compression via Meta-Tokens cites this paper.

Lossless Token Sequence Compression via Meta-Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:26.584176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:26.584176Z digest=sha256:16eab0284d5509862fc546f651b2ab72091f1fce4dfd5a095eb1f91a3a6cac9e

Observation 9fb66aae-c29f-4639-b5d7-2df33cc99457 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.063414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.063414Z digest=sha256:93d80dbcb88b7e3773bb418d8e6ffc67810c266f317ea788204fb39afba9c888

Observation b22fdf2f-2736-4930-a9f5-d6438f68cbe7 · inbound

Brevity is the soul of sustainability: Characterizing LLM response lengths cites this paper.

Brevity is the soul of sustainability: Characterizing LLM response lengths LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:09.648283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:09.648283Z digest=sha256:fe71be135d871e7448da3b0a5a87a0b972828470dc266233b0fb45864a3bbf4d

Observation 12d84b6d-e533-4cd1-aaac-86815c522721 · inbound

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation cites this paper.

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:24.837387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:24.837387Z digest=sha256:29d6b14ffbbec2d31e1131a5060b2e0254b7c12e640ea7be2a3700178944cdcd

Observation ea12e6e0-b53f-4cc3-a8ce-11b12b42f102 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:17:24.592535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:fed33ba5eb250702c6124c7f26c3bcc10f804e35a72e8b2537962679599cee53

Observation d4cf843a-3838-4780-ac24-684bd888cb39 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:10.564649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:10.564649Z digest=sha256:eca4c937c0c271cb5190187d9f46a03dff9d704b07f7216889b60c606ea72c74

Observation 9c490a55-a641-41ed-979e-524003d7b96a · inbound

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression cites this paper.

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:01.641529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:52:01.641529Z digest=sha256:319063bfb2ff4ae94c9213d2a6b0a0e0a7e5f06a2da3a2da68b0d7a963f7596e

Observation cc7a1180-8b1d-4f7c-8378-d88f4f88a128 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:37.392346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:37.392346Z digest=sha256:77b027f9b543e670a818d66c29dd18e58ec04b5e13aa9b9ee13911aead810e26

Observation 9fd21ffa-571e-4863-abbb-a76960a04134 · inbound

ARC-Encoder: learning compressed text representations for large language models cites this paper.

ARC-Encoder: learning compressed text representations for large language models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:28:49.536406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:28:49.536406Z digest=sha256:3e6282c08f9f2ec09d949f9361a00ac6b904b167591eb2397cfca32cdd25a2a8

Observation 32c3b362-8b83-428a-a01d-689ab3ba003b · inbound

When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents cites this paper.

When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:06:11.078142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:06:11.078142Z digest=sha256:2885deee84e6b86f85baa0ea82923377365e03cdd434c90d32de2ef239143e16

Observation 1840f4d0-04f9-4b48-9270-0fabf218d965 · inbound

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression cites this paper.

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:44:11.449562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T13:43:51.127429Z digest=sha256:e3a97fed3780c19158017d2fa7946b8ca29287f9ca0acb086f35e64dc5cd2ae4

Observation 33b6f1fa-13fa-4b47-9045-1a7d96d29872 · inbound

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression cites this paper.

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:25:23.076743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:25:23.076743Z digest=sha256:db0a2cbac5a4362e8bfbedee84036732c9e868a10c07a9d19fb38e1be35de05c

Observation 24041952-da97-4dc5-9b99-fdc3f1637114 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:10:10.359193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T13:06:56.207692Z digest=sha256:0bbfec5641711c2dd5a57930d9dd07ec001e0bb1bfec6ff2b92bc31da212038c

Observation 10c5e7c7-b318-48a3-8ac3-99ae685b5df0 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T11:31:29.533983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T11:29:45.134060Z digest=sha256:8f01bce53a38435ad1298bee198489222302d45a727d1d993c43ad16820ab023

Observation b67fc6f0-ae67-4b8b-80ce-9d7915ae65f8 · inbound

LLM-assisted Agentic Edge Intelligence Framework cites this paper.

LLM-assisted Agentic Edge Intelligence Framework LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:00:03.440252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T13:56:03.688192Z digest=sha256:4bf760b0b404f0641b9055029dbea5aa9461b5ee9bf12504990db903c5fe38ac

Observation f698696a-d1b7-4447-b945-3c592ab986b2 · inbound

On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation cites this paper.

On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:27.553422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:00:56.512868Z digest=sha256:3c2e0f698aedd749ad0a429c27b858c02958c6ca7f8136dce55ddd6ad17d3491

Observation 3dc12810-508d-4da3-ab4e-f63eee353003 · inbound

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models cites this paper.

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:55:10.698328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T06:52:18.563146Z digest=sha256:519f44b987755029c1b564f0c3f81a4722f8374abd35ca9d04ad5cf3a2e32e0b

Observation d77242a0-5e09-4624-adee-921596b1bbc4 · inbound

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization cites this paper.

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.788133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T05:21:56.826488Z digest=sha256:b112448f28d4eba09946abe4cb7cc8a897638e387134a046545085c38eba083c

Observation ff9c3f20-8131-48f6-a877-a1e59cafc4cf · inbound

Supplement Generation Training for Enhancing Agentic Task Performance cites this paper.

Supplement Generation Training for Enhancing Agentic Task Performance LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.528114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T00:58:27.655909Z digest=sha256:8ce3112dc72cb59ee08905da3f379e8ec20ba7135d71ea619f3ddb4d1c012718

Observation 075c5251-e277-4d6d-aeec-c40261aeb228 · inbound

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference cites this paper.

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.779041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T14:09:30.821354Z digest=sha256:09c9d705f0bbb88032c9e189b2846639b4d5c9b156c0842dd98c189a7bd8d826

Observation a234ca39-dfe8-4eb4-b367-d6070dd1cf0f · inbound

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory cites this paper.

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:25.537117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-07T10:45:48.976501Z digest=sha256:917fadd9a5b2d07a3c0c654296cc995c4feea3e566dba6bdb363b4b8fa232119

Observation 3726b935-c200-4b81-b842-0f3460351fb0 · inbound

Budget-Aware Routing for Long Clinical Text cites this paper.

Budget-Aware Routing for Long Clinical Text LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:10.192361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T19:59:00.888734Z digest=sha256:75fc9d7782122fec134de443f1f1ca1f6ee944f2e2192e5c139222b7207fb880

Observation adc10e6b-b90c-4c19-b7d5-189f57fa8d6b · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:19.749858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T18:54:06.144968Z digest=sha256:09665f1233ba032cef2f5549041fe9aae86e3155adcc202769cb16bfabc97465

Observation c2fec4e9-8772-4d5f-9473-5a56d9782452 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:19:16.570350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T00:18:32.423103Z digest=sha256:ba197dbafd7b7d9f5f4ef1d653d7dcf45107daa1ec635b5e2eda9384ec7106b0

Observation 570d6085-1061-43d9-ba01-bf90ae8330e2 · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.196333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:d5a7d3a7da6b1280246026dd4d9dae0e1cd36ccd1a4e143fb9e4405a06784cca

Observation 0ce26496-205a-4f76-9f79-39760693bd75 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.513036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:55f42a4e4b40d1c092e4a105084618f5335ee043e05345903d22dbe58b7a4b23

Observation 7f1a60ae-bca1-40cc-b45f-89d97c3143b8 · inbound

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets cites this paper.

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.270362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T23:23:49.298204Z digest=sha256:d20ab08dd73975573ae2d07593920c4abfe9059a3acdc2729a3ac09df09add1f

Observation 94c48765-afd0-491c-8695-d7eb7ac9dc0e · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.695917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:56:13.058872Z digest=sha256:195dbbf9f6ad42591f1102c2c0fb69e436007b2495860925b81aada5d06fff2a

Observation 53d80952-6f94-4501-b0fa-8058eeabff75 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T07:44:25.325808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:44:25.325808Z digest=sha256:7b6edc966d3fadfa89348092b48263762e15c66db38c24f828a7fc1dda3e0e31

Observation ddddc700-5c17-4948-8aa4-b4d4e80ec9e8 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:57.478212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:f45a654d7cc8b680ff49ac723ad3af3fa0d78f1d2b20fcf5fa9274a7ce764159

Observation 2736a578-a94a-418a-9eeb-13cec35d42d4 · inbound

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents cites this paper.

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:25.260446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T19:47:08.844690Z digest=sha256:44007e58f20b083962a6c8dff6e5adb7de38965bb1f41cd9eddbd391c0c3adcc

Observation ca86e023-c913-4143-8476-247483c077a5 · inbound

End-to-End Context Compression at Scale cites this paper.

End-to-End Context Compression at Scale LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.616571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T16:36:54.699174Z digest=sha256:d0ae2becc0dcc186e7d35d9dfce7ff3d3bf64dc3b2c0d1e8c01f1c348d854283

Observation 38dda257-2d76-4c89-a68f-c7290b5a80a0 · inbound

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents cites this paper.

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.619433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T16:10:25.191306Z digest=sha256:875c2eb5bc2522b6300f8567e3239cdc78931b069e8626fe27c1dc57ac49579f

Observation 4555f900-c28b-4ff1-b6fd-9311a2aa28a7 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.905129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:f431b45bed809dedce30aa141a567826c0e80eac999c2161a51c32a5b74fe139

Observation e1434e72-7500-43ac-98d2-1752019b4273 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.936832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:d500985f45d09d6db183f7ef267a5ee835455038d4bcf1d6a556a2616b8f6fc2

Observation a3c17276-531a-447f-b190-2fb8246f4ece · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:13.256103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:13.256103Z digest=sha256:cf3d1a17e18cd695b4b757f6cd7ed9cd94c24e06e908441076f921f1cbf8af73

Observation 06e93a9e-23a9-4284-a100-c4a53829084e · inbound

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning cites this paper.

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:34:37.884491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T11:30:52.871762Z digest=sha256:e80314082265ba70e2b613d14b99d089c7683799fb63d1cfa71bd34cadf95eaf

Observation 4404e395-a8a2-428c-8e87-7a2eed6babba · inbound

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning cites this paper.

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:49:21.550972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-04T01:45:36.335508Z digest=sha256:62200f5006dee2ead37f392d83ce939a1aaaac6e949d16d9902cc4d3d427e697

Observation b4d2f544-20d4-495b-ae9b-7a7cbd752945 · inbound

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair cites this paper.

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.569952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T14:09:32.980488Z digest=sha256:8e40f6787ab4f697c6d43ebb4c6de4a46d1014a85174aa67929651c961dab3f1

Observation d4c13e6d-a18b-4ed4-bcea-7c1ee9e55a06 · inbound

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair cites this paper.

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T08:32:14.868469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:32:14.868469Z digest=sha256:75aef6de4fc80fe930c7e5bd33c17fa9bdc03c5c15036c2cacc5282ceb75fb56

Observation 051587d5-d924-4796-a7b1-950773355e3a · inbound

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation cites this paper.

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T05:14:41.315225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:14:41.315225Z digest=sha256:5b4953b2f8b6503248eaeeeada5bde8b0523034322f220c5473f65cd5cba6e0c

Observation c0bd2724-36b3-4170-98b8-ba9ce4697d40 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.323273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:47c1fae74679b5136156dcc00683648a48a4c8b07873cc64c83eb9627d7095b8

Observation 410403c7-29ae-4e64-a6a4-897774022f78 · inbound

Mach-Mind-4-Flash Technical Report cites this paper.

Mach-Mind-4-Flash Technical Report LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T03:29:34.486347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:29:34.486347Z digest=sha256:f962d34989e245df7de7bdcdf7127b58c74f399b14edca4332fea8909713551c

Observation 3a5f3d21-5cf8-4a9d-800c-d8a8c972387b · inbound

What Context Does a Coding Agent Actually Need to Act? cites this paper.

What Context Does a Coding Agent Actually Need to Act? LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T17:40:33.397259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:40:33.397259Z digest=sha256:ec7cfaa44e1b534a653ac47eed6eadc30c9f498cc5f75ae915ea8ffd9c931104

Observation fd748da0-4d6e-40ff-9f9d-b36300ca1cc9 · inbound

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching cites this paper.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.635607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.635607Z digest=sha256:207d9c9069e1878c8c77b7c6a0adf2e1f004916f17b46aff5e3f0824ffc2754a

Observation 0a17e1cd-b03b-4492-a933-1c39358fb63b · inbound

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning cites this paper.

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:14.315923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:14.315923Z digest=sha256:31bf90efdbe902a27db4ebf3eb5657c0200b314b2c0198903311d7a7fe1f54eb

Observation e27cbc51-3fc1-4a48-905d-fe59b0e9ded6 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:09.109669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:09.109669Z digest=sha256:4f7621a64403fca7b679cfbb9559cc94b6b8e5e83e7aa5260a48596f9c31246c

Observation b0406970-d688-4c57-a030-8dbe36cf166a · inbound

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing cites this paper.

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:37:06.990703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:37:06.990703Z digest=sha256:916c3d78a3cb5411a7a2143e1d2d12556bef29fa47ada329b8ebd44512bd633d

Observation cd698b93-d591-4e94-91ca-a68829c8f023 · inbound

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression cites this paper.

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:38.995166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:38.995166Z digest=sha256:5b6f32d50c9714dd2a8cea822614f97deb8ffdafb527b4e2ea76449f92d8b2a3

Observation f2c13c4a-5ffe-4121-8948-f0309e808f58 · inbound

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents cites this paper.

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:27:03.120485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:27:03.120485Z digest=sha256:3d1b45967d8dda24993fb814eee587b2ad0ca24fe5dc5bd57983c91241d0661d

Observation bd5a7c86-9ff8-40b0-b214-40f61224ee2d · inbound

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups cites this paper.

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:24:46.926491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:24:46.926491Z digest=sha256:00321966221f3e8d264563bcb8b34a81019e934ba0f508103ea5d24deb99a939