Pith. sign in

Paper Citation Record · LEDGER

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

As of 6 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2605.11605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11605 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T02:00:02.786195Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact17
  • verified fuzzy44
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46427dae-2b9c-497e-a3fe-6b04efb5422f · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.647472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:a3e3c1180842269bb7e849d8dea5f13b92475c52fc097116f3e709d8ebf5b355

Observation 8db28d7e-ab8f-4a5f-bbd6-46f11ec9c95a · outbound

This paper cites Qwen2.5-VL Technical Report.arXiv.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Qwen2.5-VL Technical Report.arXiv

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.558683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:f267e75225c06f4c5297019e3a5318f623eb26566b37f038e02f20e4099fac85

Observation 4da67932-a6ff-47df-89b8-04d1729db22a · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs VGGSound: A Large-scale Audio-Visual Dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.536785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:20f1ff58dc058a1091c31378628f8688c2cff0c5582ccbc10a6d3568618930f0

Observation f6005e4a-e121-4901-b81e-a50ca4690868 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:02:06.377506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:4af107c43dd084da28b3edbde4eaae2c2b9a460c66088dc920e6fb2134f322d2

Observation 28bab313-2f09-48ce-a36f-85bae1915e0c · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.563568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:0a8a27d57789f49070191b6f77492e224fd9befe5b99488896a3b345f9b551b3

Observation 381010d3-a7e7-405f-b127-b642c472fc57 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.621269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:d883b495966454a7a180086077d1a942a5cbf18be35f58da946e0e504dc5fadd

Observation da416ff9-67f0-474a-a4c1-b6bc865a476f · outbound

This paper cites V AST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs V AST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.511176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:f07593665e50d98c056a974f5ddd79034f4aa9cebe997be7c565f23c2145c571

Observation ec5f479e-00d9-4dd0-a200-e539819b6a52 · outbound

This paper cites StreamingTOM: Streaming Token Compression for Efficient Video Understanding.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs StreamingTOM: Streaming Token Compression for Efficient Video Understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.575821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:a46e5b31bb6b411be37746dfd08422a55b5db1ee08f10d52357f57fd63132103

Observation 527c5493-22dc-4eff-981d-1df508d70c47 · outbound

This paper cites InternVL: Scaling Up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs InternVL: Scaling Up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.629594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:04bfb87566a792d9ab953419d660373ad06181356fb9a3b34e4fb0efa863f4d6

Observation a8b010d0-8bcd-4244-88d3-997299a7a457 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:02:06.362106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:b477ebf477b3bb20a51328eb73b88f54ac7eba9b888d9eaf85b12bfa62daeac9

Observation 5e5f7cbb-4923-4063-9cf1-cbffedce71ab · outbound

This paper cites Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.660046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:e9e2d11928037d601e305c2de384239dadb94d203e90f2cccb0070fa39dd9d75

Observation cd6ac54a-e8b6-44aa-866d-5d2882258ed2 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:02:06.342903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:6783f2fe0a9fbabbbabd4be39f9a898856af9290b876f8d4f6e300fa4131b83c

Observation da58de3f-3268-4dc4-a92d-3b6af83600dc · outbound

This paper cites OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:56:19.131480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:bf6c0c0d69fd4ff969b6c20490ded8c9ba4d55e89ebcb533f41eaed91fbbf0de

Observation c0cd9387-cf1c-4926-8411-c15b44abf422 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.519927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:6607ef7fce1ce691b238823270fcadf79ae0d848a8ba3f850a537d3d01b17317

Observation caa110bc-ebb6-4b78-919c-c55326dfc54f · outbound

This paper cites Video-MME: The First-Ever Compre- hensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Video-MME: The First-Ever Compre- hensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.524633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:e9229c3d61a46530aa125c5290e91aeba7426062cd49318c9504bb304ecf253b

Observation 9411a143-83ec-451f-b623-b634ac6ab116 · outbound

This paper cites Gemmeke, Daniel P.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Gemmeke, Daniel P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.495135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:861f35fdedf56e1f7be5769840ac5f6f32b4fd020ecb883d95bae215e0b3927f

Observation e53fa6ab-5e4a-4b4f-a2f2-98d15ec110bd · outbound

This paper cites EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-23T04:13:43.420881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:25bd9e2e48c11069759ff225e4fa9e5d17865ab20648bb376845d0b4e5d8bea6

Observation 8ae23397-ee69-450e-b7c1-5c6a6c4e45a3 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs AST: Audio Spectrogram Transformer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.498259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:18d54cd823356ee6d723a7e1ab73a2a403ec68df75ab4049a217abb034ba6de8

Observation 152b6d51-3fe6-47ea-a2b5-481acfc3cd9d · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs OneLLM: One Framework to Align All Modalities with Language

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.528855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:db20904a4cb82f9affd5f155bc317d6184b7b643826bd5571eda5f34f5e56d07

Observation bee92d2e-85e1-4b3b-be17-d9a63bdadcc3 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.492104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:6da4c62c3ef91e8a22ddb5f8c5fefc20d0cb7b346bc1956112bd140630c1e605

Observation 6ca3676d-8686-4580-99ba-1bf0f689ffd3 · outbound

This paper cites Language is Not All You Need: Aligning Perception with Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Language is Not All You Need: Aligning Perception with Language Models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.638217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:3bad4250ddc9a9aec4cb1e293cb82ff45e5e458f08626cbc1fb9929fa731b583

Observation 4a52e37a-bf83-473b-bf15-83d521e3d184 · outbound

This paper cites GPT-4o System Card.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs GPT-4o System Card

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:02:06.333048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:48c982ba069da9b6e1262211fa77f7346c871b2c13a6196b6d4d2518c25b88db

Observation 9ac3acee-bad6-4f77-95f5-92fd9e31f6fb · outbound

This paper cites Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.567710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:ab45cbbc0184a01bee77663d506aec2477d765cac78a8d8aa67b9883b5aa7a33

Observation 54f8b596-915b-44ee-b959-d5b8867ac574 · outbound

This paper cites STORM: Token-Efficient Long Video Understanding for Multimodal LLMs.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs STORM: Token-Efficient Long Video Understanding for Multimodal LLMs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.642525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:65060a477ca40f9762fcdf6a5a639221ebe509a8c6cc0bf5c77963bd94bdb43f

Observation 851988ae-deea-4a6e-8b7b-4d62c32f3d19 · outbound

This paper cites LLaV A-OneVision: Easy Visual Task Transfer.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs LLaV A-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.670390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:1a4f5a7fd1ce8dc47f22008618aec0bbed616f761564a820cac1eab7a5c66920

Observation 838e6d47-199d-4410-811c-b9dfff8a762f · outbound

This paper cites Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Omnivideobench: Towards audio-visual understanding evaluation for omni mllms

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.297828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:00eba951780e93177639c4ff6deb7e7238fa3ef21426d809f742d1b57dada69e

Observation 10487093-e774-4356-91b5-290f1ce7e967 · outbound

This paper cites BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.545948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:52151b77235bfa622930579cbcc8101262417c672f2a12b4ce785f858c466ed3

Observation 1d8fc051-5434-4c62-a4ca-94775cbbb317 · outbound

This paper cites VideoChat- Flash: Hierarchical Compression for Long-Context Video Modeling.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs VideoChat- Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.608600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:0b970924ae8f6260bf590c5a954df1bc16df3ac633e5d8038ca7319d5f802e80

Observation 9411bb27-a3a9-462a-9b9a-907b1a0b4bec · outbound

This paper cites Video-LLaV A: Learning United Visual Representation by Alignment Before Projection.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Video-LLaV A: Learning United Visual Representation by Alignment Before Projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.584502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:a4c4132eb27c957d56243ea45e6f4bbefa09f32ed6ceb36672215ba29b54638d

Observation 32dc153c-abe0-4935-a8aa-2f85ea2928bd · outbound

This paper cites Visual Instruction Tuning.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Visual Instruction Tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.587896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:21c8a0dfacf67389b564a36406f8349a0a4d9def799d255c0f93cca570fdb1ab

Observation 26cbbea0-d15b-44f0-984b-e058915d9b66 · outbound

This paper cites Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.600554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:6e9459fa59b3a5bf5858ff4f8be8ec51800ee20484126f9461f256522c4e070d

Observation 1dffdf43-1dd2-4240-b85f-02f3f24bb3b2 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.307495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:1c548db188a45694cd3aebcdef7f4900a06b46d3b135e72a28c7e810ede2919a

Observation 302028f4-4fd4-48aa-b57e-a31b9ad2f0af · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.633942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:63d56c0b0202f0fd099fef64c4a3a33a25dd917736db61cb78b85c32e5f8f7c9

Observation 1fdfb7a6-5a81-47c0-a436-b1e74a46124e · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.328981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:74d4cb32fb3b7052f31e4a4fa61dbc44611fe7d3f286bfe79c9270c60df60e74

Observation f268fcbe-6abb-4f73-b567-aaaa13494a6c · outbound

This paper cites Adapt- Token: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Adapt- Token: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.338129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:061a7137f8ca9be2a41ab55aea2fa26d66a411de395ae1f9d44a4f2005fc1c60

Observation 606d06ea-4a60-4297-a449-7b85aafe8c0a · outbound

This paper cites An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.302188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:acb9e3904ab7cd8e4846c94f872709c5723a39f09aa5e079961c6b06e57be0f7

Observation aa97041d-52d6-492a-b46b-d3942462140d · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.550241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:628fd8808b8cc9b012c1816cc783e8c503c9657a831c07a22ec7875e20cd5b64

Observation e06d4365-e617-4d24-935b-60af578bbd59 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Robust Speech Recognition via Large-Scale Weak Supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.651937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:cc0849858ae0507570f3a61dbbc002ecffe8dac802d5836d43132b0014d05a12

Observation 9396f416-d069-437d-a162-89689ff91f90 · outbound

This paper cites LLaV A-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs LLaV A-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.616654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:170204d1eaf2351ab1490dbb4ad2ceaa626423e19453f1666f8591a7528ea1d9

Observation 63155665-8e18-4e90-842a-cf4fe96392f5 · outbound

This paper cites an unresolved cited work.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-13T13:57:51.489430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:137da42dd7f7db09ef8d5ccfc60bd3d39a5706002a91e50c1fb9fd451ef3fa48

Observation 4ec4d13e-33cf-4bf4-9f1f-2c8a30f522d2 · outbound

This paper cites HoliTom: Holistic Token Merging for Fast Video Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs HoliTom: Holistic Token Merging for Fast Video Large Language Models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.541689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:800adf6ed51be9d53253bda3062920370999794c49a401016c0705898fcae438

Observation 2e737fa9-64ad-4e62-8cc2-09cc8f167e4e · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.625379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:6015631d5ef91138a73a1e3cfb795b80a408e92af6231286f05bcaf9438c0790

Observation aa610cd1-acad-4834-85b6-0e17081cc188 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.674737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:06eae4527104bb31b8a7a2fbe4987aa489dbda404a1267dba50bf5b96e675402

Observation 718c3c65-f48e-41d3-9823-54f62f526fec · outbound

This paper cites TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.367035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:76b7a48c8e6f9de68837b97a216754e1849e9c757996c266f15ee02e650092d0

Observation d84ee85f-ea73-4a96-a415-52cfd8bd78d7 · outbound

This paper cites video-SALMONN 2: Caption-enhanced audio-visual large language models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs video-SALMONN 2: Caption-enhanced audio-visual large language models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.372567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:c79872cea2a76c9e2b72e48b93f13f8900c15f569fa9520e7d88eddc2c388072

Observation f1ef4fce-7c5b-4824-9b4e-01de9400d0b7 · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.664257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:4c60c4e4b3e6472d099705e2918f5a9619707a704b9b3384862832e9c83770cf

Observation c216012e-e004-4c7c-bdff-4a4b597749f9 · outbound

This paper cites OmniZip: Audio- Guided Dynamic Token Compression for Fast Omnimodal Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs OmniZip: Audio- Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.595117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:3bd01a030ae2aefd892690dc5fca376aaec768caedd4e2f2618d57c8b9702957

Observation 97664f33-948d-41dd-a3c3-5e023b1d0c3a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:02:06.352132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:1162ec479f671ec7778bc8529f19b0a894d8ad39691744ca78143cd5ec84a43b

Observation 87bb9f0c-cbf7-4397-bc6b-92bff4b09512 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:02:06.357482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:00616bcd95e66cddccb8a9e164d7cd8d862be60074f2f187a403652f8128833a

Observation dbf3d468-c410-49d6-b64e-4cd66733f14c · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.604633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:e819fdc55e6252d23835a197de79846b25edc74f56788e1642f3eb983831d36a

Observation 7f846117-6b8e-4e01-8233-7e397344da6d · outbound

This paper cites Qwen2.5-Omni Technical Report.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Qwen2.5-Omni Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:02:06.287537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:5175c414b8b740e9c296f26091a3dd57ee778f934938e864f114064460881e27

Observation cd3c5254-0b18-4755-89c2-4793b69da1f5 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.580575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:ccf802b35b09350798ce75728088f1047d3a81fb0fc756b4905d6cf393e61ecc

Observation f38366e9-d8a8-4b47-8a40-c344eb1ddf9e · outbound

This paper cites A VQA: A Dataset for Audio-Visual Question Answering on Videos.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs A VQA: A Dataset for Audio-Visual Question Answering on Videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.612437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:d60fc4eef481512d08b1b1045d941bd674c88eb8628fef0cc011d0bfb7db24ec

Observation ba526c80-44e9-4a64-b708-13cb480b510b · outbound

This paper cites VisionZip: Longer is Better but Not Necessary in Vision Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs VisionZip: Longer is Better but Not Necessary in Vision Language Models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.515439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:9e58ec0040318630da0ad798246b8089edf22f8104397ab4d3bdf3a372d48dde

Observation 79822637-f29b-4ac7-8aba-e42ce82ba184 · outbound

This paper cites TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.656044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:1b0b2279740c1571d1f374de104a6b5c42f13aca1d5788ad1aaf76156b0bbca9

Observation 354ce40a-0724-4fc3-a2dd-47afb20ab06b · outbound

This paper cites CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.502615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:929c4280359fbc3942b4c522ed2677755bad4ffd9c6a6686c7aa0ab85c45c001

Observation 8af86676-dba3-40e7-a6d8-f67db91c3888 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.506746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:38e9d0ee9358722ecb87f926137430895f6d2ddb0888708d05bc2c981c357416

Observation d07c912e-6f1d-469c-9db0-8b7b692b9a32 · outbound

This paper cites RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.571838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:1189ca1f6ea28690d4aa582259a69ef4c48bec777f0bcc9684fb11be648af4b1

Observation 7262370a-9725-46c9-8304-e10518785462 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.312638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:fc91fe744ba26db37c8006d6530410660edee1fe2b04422f1d0a6c8f72b8c1a6

Observation 46ff8bec-d77a-4928-af29-debcd42e5ce9 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.591886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:768aab3aef7a2b6790a9511ded0f245481eb860cdf0a8403d6d1c6ea2d9832d4

Observation f7efcaf8-f075-4462-910f-86a11605f658 · outbound

This paper cites p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.532793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:3298382a7fd64868a5188a56af769fdc8f2b00660d817b503113e11c2c1be7e2

Observation fc9069bb-08d5-4a77-9a8a-a561fb4fb24d · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:57:51.554496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:0a97d785d09000be910e77295cd22c6155bcc549499fe16814d7d0f69c247cfc

Observation 9f50a6fc-3ad4-4b69-971b-5fa48ea23872 · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.318077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:29951bfd11e9823a10a3eb64d7c1e13678182fb09710c5fd7562b9fcc3f1e010

Observation 3810f0cd-55ee-4648-8a6d-a09e64088ba7 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Mlvu: Benchmarking multi-task long video understanding

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:06.323895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:963dccbbaa54482e9120f28c388e751605cbc10288cb57bac445de364d078425

Observation 6647d32c-8957-4093-802d-3ef894dd2295 · outbound

This paper cites plucked string instrument music.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs plucked string instrument music

Reference 65

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T02:02:06.292937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:49ec511264845a568f09113beaa1c9bb09f32eb0ed0d2515bc07990572c2d72f

Pith citing papers

No inbound Pith citation observations are available.