Pith. sign in

Paper Citation Record · LEDGER

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 73 inbound Pith citation observations for arXiv:2310.05736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:05:54.122394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.322096Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f72dd5e3-3850-44c9-8af0-63c3c00b9f94 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.647448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:da84f862ab9f7b6a9e5d9106e5417e36b09284415ece4f5d3efb12e6f065a88e

Observation bec1097a-429a-4dc0-bea3-2c1bb52f6112 · inbound

AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models cites this paper.

AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:08:26.041444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T21:06:59.776841Z digest=sha256:544fa5f4639e41ad21533ad7ddcdcd3888054f3fd5c0d933bb7517148fb7abad

Observation 038d4429-156d-4e4e-89ce-8cdbf1905052 · inbound

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning cites this paper.

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:38:24.967159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T20:36:55.159302Z digest=sha256:380ed45ed690673e205b14b73ae55b2e76c5d8c731df358ece4fdff34ea3f9b9

Observation 911414fe-e58b-4241-b153-8be80d211849 · inbound

JPPO: Joint Power and Prompt Optimization for Accelerated Large Language Model Services cites this paper.

JPPO: Joint Power and Prompt Optimization for Accelerated Large Language Model Services LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:40:21.563948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:40:21.563948Z digest=sha256:f40ce901687d9cffa9ada1faf45f0c28e9e1ce4a5d45e0fbdd460c2718dec86c

Observation 451bc9a2-8be9-46c1-a39b-5fa59b2ba18a · inbound

JPPO++: Joint Power and Denoising-inspired Prompt Optimization for Mobile LLM Services cites this paper.

JPPO++: Joint Power and Denoising-inspired Prompt Optimization for Mobile LLM Services LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:32:20.773340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:32:20.773340Z digest=sha256:a1e9d98bf09954cef08c701a08b74b7f0053dbfe0bb32122cec28532772cbf35

Observation 7794ffa5-db67-469b-b98c-b57fb734b362 · inbound

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens cites this paper.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.184668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.184668Z digest=sha256:ff38786d48143e344a2e94a5fb0b0f38e923d62d17996855323464d17441e143

Observation b1bc1a00-5a49-4a02-bf38-f9fe7df36884 · inbound

FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing cites this paper.

FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T14:57:27.548016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:57:27.548016Z digest=sha256:b06cce5beee1872480be51adbabb0e249e0f6fac7d5977b0bf5912dce45a2594

Observation 0deb27a7-78d4-4780-8ed7-2568d03ebad0 · inbound

C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness cites this paper.

C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:03.890113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:47:03.890113Z digest=sha256:f5d4a57d37df516d9f23960ee658737627e36b7500b057f1d739462fae90c1a4

Observation 6afb25e3-bfd5-40c4-a724-80806a27129d · inbound

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations cites this paper.

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:47:40.400393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T04:47:40.335477Z digest=sha256:b3ef082f21c80c087d38faa85e37498f3f459695405288b5903c9760e1ea91aa

Observation ba44d3ee-c20e-4333-8ccd-0f54d631d154 · inbound

EvoWiki: Evaluating LLMs on Evolving Knowledge cites this paper.

EvoWiki: Evaluating LLMs on Evolving Knowledge LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:04:20.131090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:04:20.131090Z digest=sha256:75dcf1f4cc703a333398d967c260c84ee5cc34b9605a59feead0f2661d771252

Observation 6108fa53-a800-42be-9fb5-fa42c178d672 · inbound

EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent cites this paper.

EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:51.267386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:51.267386Z digest=sha256:9ee5dcdaa2e8d7dee2a1e29e211e3a9bea88cc248316a46393c78eb6522d45d5

Observation 8b42b6be-e6f7-43dc-bd46-68648dccc2f4 · inbound

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG cites this paper.

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:34.859209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:34.859209Z digest=sha256:bcb8d5aa564c40e5100847199ea003a64fb19397f331cd4476048cdfe37a635d

Observation d5a994b4-e873-4a42-8700-6d5efc7ef278 · inbound

LeMo: Enabling LEss Token Involvement for MOre Context Fine-tuning cites this paper.

LeMo: Enabling LEss Token Involvement for MOre Context Fine-tuning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:11.074196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:26:11.074196Z digest=sha256:a61d1f6206329302bbb2b5b0474796a3fe941612e6063629b6d1bf0d039c99a6

Observation c0805d18-4818-4e42-8e52-cf6f52f50a36 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.685933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.685933Z digest=sha256:3df1efc25e74c4e233ed68b82a195fb767d50cf7c870398fae5a908ea3dcd452

Observation e48c0a43-7e26-48da-9d2d-4c97cf3dd3fa · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:46:30.078718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:90d7fc184c4abb61daa9a98307680f5baa7880bf566db822773deee5509d3114

Observation 6e2ceb50-3c8b-41dc-9d28-4039d3bc7f44 · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.491572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.491572Z digest=sha256:b9cebc8d4130545a06e3538c21d375726b4f94d17ea5bfecb0347ab5b00a1588

Observation 030cb13d-2a63-4871-bc1b-00bf228d3047 · inbound

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression cites this paper.

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:05:54.122394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:05:54.122394Z digest=sha256:62cb023af397b67dd913aff024d8d617db809f64cfeea7c40194e068992ef1cb

Observation 2cbcf88d-b56b-401c-8ef1-a4aad3be0a73 · inbound

MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG cites this paper.

MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:43:21.595446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:43:21.595446Z digest=sha256:6a4cd5a532bff457dab8643a3e64656b77375ca61d3eab3ad951a8932b22aa61

Observation e4058a9f-2dcd-4658-b504-75b78e3cd4c5 · inbound

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models cites this paper.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.983020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.983020Z digest=sha256:e914b2e80a8533e94f682cd9189fea2ffa0d1013fd313e68d14cdb860619e019

Observation 445cf6ed-50af-4bb1-a32e-798ee654dd31 · inbound

QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization cites this paper.

QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:52.446651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:52.446651Z digest=sha256:1a7a254115e2d0a22bb63baf4c72053e8eb0596393a169e7f463414b1f7ac109

Observation afafabaa-3390-46ed-81b2-6029bc351209 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.213167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.213167Z digest=sha256:bc3a7107e8b897b6cefe13ab0f66e44e47cb9df20bc69fa3d0d90beac9a78b1c

Observation 7fb1ec33-d153-40f9-9093-18fd3a040d60 · inbound

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling cites this paper.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.729533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.729533Z digest=sha256:7f4a63a89c82d81c6802439f75444a596cf78ae0aa11b37397152dea22376d7d

Observation 0427e401-f35a-4281-b7c8-c7fb90d3b9e6 · inbound

Lossless Token Sequence Compression via Meta-Tokens cites this paper.

Lossless Token Sequence Compression via Meta-Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:26.584176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:26.584176Z digest=sha256:9658ceb912351d3c18accb49ab3b47656c95b64c724c630086776a09ff6532d3

Observation 9fb66aae-c29f-4639-b5d7-2df33cc99457 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.063414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.063414Z digest=sha256:93d80dbcb88b7e3773bb418d8e6ffc67810c266f317ea788204fb39afba9c888

Observation b22fdf2f-2736-4930-a9f5-d6438f68cbe7 · inbound

Brevity is the soul of sustainability: Characterizing LLM response lengths cites this paper.

Brevity is the soul of sustainability: Characterizing LLM response lengths LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:09.648283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:09.648283Z digest=sha256:fe71be135d871e7448da3b0a5a87a0b972828470dc266233b0fb45864a3bbf4d

Observation 12d84b6d-e533-4cd1-aaac-86815c522721 · inbound

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation cites this paper.

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:24.837387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:24.837387Z digest=sha256:29d6b14ffbbec2d31e1131a5060b2e0254b7c12e640ea7be2a3700178944cdcd

Observation ea12e6e0-b53f-4cc3-a8ce-11b12b42f102 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:17:24.592535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:0f9e8f4b7e7855f2fa1336af408f2a7fc114edf61634c786ed9fae33e31d71f0

Observation d4cf843a-3838-4780-ac24-684bd888cb39 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:10.564649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:10.564649Z digest=sha256:eca4c937c0c271cb5190187d9f46a03dff9d704b07f7216889b60c606ea72c74

Observation 9c490a55-a641-41ed-979e-524003d7b96a · inbound

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression cites this paper.

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:01.641529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:52:01.641529Z digest=sha256:319063bfb2ff4ae94c9213d2a6b0a0e0a7e5f06a2da3a2da68b0d7a963f7596e

Observation cc7a1180-8b1d-4f7c-8378-d88f4f88a128 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:37.392346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:37.392346Z digest=sha256:77b027f9b543e670a818d66c29dd18e58ec04b5e13aa9b9ee13911aead810e26

Observation 9fd21ffa-571e-4863-abbb-a76960a04134 · inbound

ARC-Encoder: learning compressed text representations for large language models cites this paper.

ARC-Encoder: learning compressed text representations for large language models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:28:49.536406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:28:49.536406Z digest=sha256:3e6282c08f9f2ec09d949f9361a00ac6b904b167591eb2397cfca32cdd25a2a8

Observation 32c3b362-8b83-428a-a01d-689ab3ba003b · inbound

When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents cites this paper.

When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:06:11.078142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:06:11.078142Z digest=sha256:2885deee84e6b86f85baa0ea82923377365e03cdd434c90d32de2ef239143e16

Observation 1840f4d0-04f9-4b48-9270-0fabf218d965 · inbound

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression cites this paper.

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:44:11.449562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T13:43:51.127429Z digest=sha256:91971babcc925b6684b39dfd6d812dfd8f5d2a35c095fe7ddd5d65dad1196030

Observation 33b6f1fa-13fa-4b47-9045-1a7d96d29872 · inbound

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression cites this paper.

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:25:23.076743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:25:23.076743Z digest=sha256:db0a2cbac5a4362e8bfbedee84036732c9e868a10c07a9d19fb38e1be35de05c

Observation 24041952-da97-4dc5-9b99-fdc3f1637114 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:10:10.359193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T13:06:56.207692Z digest=sha256:08754484f7213d444677368c9fcb374fee4623d0c9da2760f91769997b21c0b8

Observation 10c5e7c7-b318-48a3-8ac3-99ae685b5df0 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T11:31:29.533983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T11:29:45.134060Z digest=sha256:560cfe00815231210b54a2d59b423fae4f88ceb8c1f31dc8765c1f44f687b4ac

Observation b67fc6f0-ae67-4b8b-80ce-9d7915ae65f8 · inbound

LLM-assisted Agentic Edge Intelligence Framework cites this paper.

LLM-assisted Agentic Edge Intelligence Framework LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:00:03.440252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T13:56:03.688192Z digest=sha256:874c657051e0112587b403655ddfc733cdb9d16babbb0d9697679c7e23fa7b80

Observation f698696a-d1b7-4447-b945-3c592ab986b2 · inbound

On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation cites this paper.

On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:27.553422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:00:56.512868Z digest=sha256:9e60497181f849747b5d1ab674f51934c834819c1bf76ce558bb159cfd0cdcab

Observation 3dc12810-508d-4da3-ab4e-f63eee353003 · inbound

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models cites this paper.

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:55:10.698328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T06:52:18.563146Z digest=sha256:5944cefebce3322ac6adae75132260fe37c94f34eadb0f0958b3b15aaeea6e38

Observation d77242a0-5e09-4624-adee-921596b1bbc4 · inbound

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization cites this paper.

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.788133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:21:56.826488Z digest=sha256:b8bb8917e5a1ff321ebecf34efd61922db70ae03726a26d3128e9d29ed205867

Observation ff9c3f20-8131-48f6-a877-a1e59cafc4cf · inbound

Supplement Generation Training for Enhancing Agentic Task Performance cites this paper.

Supplement Generation Training for Enhancing Agentic Task Performance LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.528114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T00:58:27.655909Z digest=sha256:33acdb8be574ccbcf83765df4637c7ffdea03cbee04c02943ca7496385294057

Observation 075c5251-e277-4d6d-aeec-c40261aeb228 · inbound

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference cites this paper.

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.779041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T14:09:30.821354Z digest=sha256:886a93ab581286c52c8a2806a8b08d7094b915191fa62a96c4cfcd8de0c15ec6

Observation a234ca39-dfe8-4eb4-b367-d6070dd1cf0f · inbound

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory cites this paper.

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:25.537117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-07T10:45:48.976501Z digest=sha256:920dbe4fe74b0277988eb5b453a2f966c1f7f45c56bd35424a1da56cedfcf2b2

Observation 3726b935-c200-4b81-b842-0f3460351fb0 · inbound

Budget-Aware Routing for Long Clinical Text cites this paper.

Budget-Aware Routing for Long Clinical Text LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:10.192361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T19:59:00.888734Z digest=sha256:fd43732dbca1896214ac207905a101fa69a90f487a07aa6ccc99b2acf383a410

Observation adc10e6b-b90c-4c19-b7d5-189f57fa8d6b · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:19.749858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T18:54:06.144968Z digest=sha256:ea6a29873803d3c458463b3e62fe5fbe3ba02874db167e2b33cb884f0fc359dd

Observation c2fec4e9-8772-4d5f-9473-5a56d9782452 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:19:16.570350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T00:18:32.423103Z digest=sha256:7ae74fdcb5bddb0e229d429733d450fb3f68a8c3c567e0c33684c0eee745d93d

Observation 570d6085-1061-43d9-ba01-bf90ae8330e2 · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.196333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:d032ae72dd8a1ffc7a9615851ddaee25b5d357c20a7c61e15340afdac244234e

Observation 0ce26496-205a-4f76-9f79-39760693bd75 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.513036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:37d70b0c6f93ece770936895089cc1dd2f986cae941b9904ee19c7c4e477c83c

Observation 7f1a60ae-bca1-40cc-b45f-89d97c3143b8 · inbound

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets cites this paper.

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.270362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-29T23:23:49.298204Z digest=sha256:2239bc783dc2453d8979c1e30a17dc505797cbc4a58933c5b4a1135f53b84604

Observation 94c48765-afd0-491c-8695-d7eb7ac9dc0e · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.695917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T10:56:13.058872Z digest=sha256:b3fb53a93812c5271f238cfa3c6a75005898fd94c75fbbca0ad3b78da33a29a9

Observation 53d80952-6f94-4501-b0fa-8058eeabff75 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T07:44:25.325808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:44:25.325808Z digest=sha256:7b6edc966d3fadfa89348092b48263762e15c66db38c24f828a7fc1dda3e0e31

Observation ddddc700-5c17-4948-8aa4-b4d4e80ec9e8 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:57.478212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:9605a7a5ea65db164189fec04d0cdcf9ee7d40b7969350b964bfffaad368d4cf

Observation 2736a578-a94a-418a-9eeb-13cec35d42d4 · inbound

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents cites this paper.

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:25.260446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T19:47:08.844690Z digest=sha256:3e365bc908f34d2fe34d8a28aa66aba1192c9647303fc1c718a816f909775232

Observation ca86e023-c913-4143-8476-247483c077a5 · inbound

End-to-End Context Compression at Scale cites this paper.

End-to-End Context Compression at Scale LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.616571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T16:36:54.699174Z digest=sha256:c2d1d598ba731df717fec6808cf3fb73630f26f27a77f4ae2a18a056353defc5

Observation 38dda257-2d76-4c89-a68f-c7290b5a80a0 · inbound

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents cites this paper.

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.619433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:10:25.191306Z digest=sha256:8b4a8515bb901dfb7b1cc01bb4c3def2dd0e96f22580442656e121ea2fe16361

Observation 4555f900-c28b-4ff1-b6fd-9311a2aa28a7 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.905129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:b72db59acca08b51f28140cf3f54ac1bd5c97b27da058b91395e237f48759c67

Observation e1434e72-7500-43ac-98d2-1752019b4273 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.936832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:5bf417d604b8cd601c44cb0b3f29e3739d9e4363e2e8089a4ebaabc5f9fae45a

Observation a3c17276-531a-447f-b190-2fb8246f4ece · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:13.256103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:13.256103Z digest=sha256:cf3d1a17e18cd695b4b757f6cd7ed9cd94c24e06e908441076f921f1cbf8af73

Observation 06e93a9e-23a9-4284-a100-c4a53829084e · inbound

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning cites this paper.

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:34:37.884491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T11:30:52.871762Z digest=sha256:82efb0b3cca41b1a2bce80241945ff0c879a6aca90689ba32996aca7efd1ce23

Observation 4404e395-a8a2-428c-8e87-7a2eed6babba · inbound

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning cites this paper.

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:49:21.550972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-04T01:45:36.335508Z digest=sha256:10770d3336ff52f0c5612cca0c1e9a44df227fad9c0b8af0bdd7217b03a2c034

Observation b4d2f544-20d4-495b-ae9b-7a7cbd752945 · inbound

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair cites this paper.

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.569952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T14:09:32.980488Z digest=sha256:519f76b108f0a90465a2a916ad70394f5b916880bbb30a8145c9e905e34f822a

Observation d4c13e6d-a18b-4ed4-bcea-7c1ee9e55a06 · inbound

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair cites this paper.

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T08:32:14.868469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:32:14.868469Z digest=sha256:190a6a4702e89e62e9f1e784c21b6fa906f6991b4b55f87a4f685738cc1a7ac0

Observation 051587d5-d924-4796-a7b1-950773355e3a · inbound

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation cites this paper.

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T05:14:41.315225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:14:41.315225Z digest=sha256:5b4953b2f8b6503248eaeeeada5bde8b0523034322f220c5473f65cd5cba6e0c

Observation c0bd2724-36b3-4170-98b8-ba9ce4697d40 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.323273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:31c62a79fdc8f71dac533ba3e1c783797f4901fdd0fabf299f7569a29ed498be

Observation 410403c7-29ae-4e64-a6a4-897774022f78 · inbound

Mach-Mind-4-Flash Technical Report cites this paper.

Mach-Mind-4-Flash Technical Report LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T03:29:34.486347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:29:34.486347Z digest=sha256:f962d34989e245df7de7bdcdf7127b58c74f399b14edca4332fea8909713551c

Observation 3a5f3d21-5cf8-4a9d-800c-d8a8c972387b · inbound

What Context Does a Coding Agent Actually Need to Act? cites this paper.

What Context Does a Coding Agent Actually Need to Act? LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T17:40:33.397259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:40:33.397259Z digest=sha256:37e7630fe0172558925e7930566457e7e2c068706d8ccd0fc72b9ffe339cebfe

Observation fd748da0-4d6e-40ff-9f9d-b36300ca1cc9 · inbound

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching cites this paper.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.635607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.635607Z digest=sha256:207d9c9069e1878c8c77b7c6a0adf2e1f004916f17b46aff5e3f0824ffc2754a

Observation 0a17e1cd-b03b-4492-a933-1c39358fb63b · inbound

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning cites this paper.

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:14.315923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:14.315923Z digest=sha256:31bf90efdbe902a27db4ebf3eb5657c0200b314b2c0198903311d7a7fe1f54eb

Observation e27cbc51-3fc1-4a48-905d-fe59b0e9ded6 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:09.109669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:09.109669Z digest=sha256:4f7621a64403fca7b679cfbb9559cc94b6b8e5e83e7aa5260a48596f9c31246c

Observation b0406970-d688-4c57-a030-8dbe36cf166a · inbound

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing cites this paper.

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:37:06.990703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:37:06.990703Z digest=sha256:916c3d78a3cb5411a7a2143e1d2d12556bef29fa47ada329b8ebd44512bd633d

Observation cd698b93-d591-4e94-91ca-a68829c8f023 · inbound

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression cites this paper.

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:38.995166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:38.995166Z digest=sha256:5b6f32d50c9714dd2a8cea822614f97deb8ffdafb527b4e2ea76449f92d8b2a3

Observation f2c13c4a-5ffe-4121-8948-f0309e808f58 · inbound

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents cites this paper.

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:27:03.120485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:27:03.120485Z digest=sha256:3d1b45967d8dda24993fb814eee587b2ad0ca24fe5dc5bd57983c91241d0661d

Observation bd5a7c86-9ff8-40b0-b214-40f61224ee2d · inbound

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups cites this paper.

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:24:46.926491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:24:46.926491Z digest=sha256:00321966221f3e8d264563bcb8b34a81019e934ba0f508103ea5d24deb99a939