Pith. sign in

Paper Citation Record · LEDGER

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

As of 22 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2608.08794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08794 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:04.833294Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact3
  • verified fuzzy63
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 290bf7ea-f76f-42cf-9227-995e1e933d68 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Qwen2.5-Omni Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.323992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.323992Z digest=sha256:8674b30aacb9c59e93a969678ff16e5a0b92f036ce2f6b2f8960ee65b24e4f56

Observation 41c3e960-98cd-4cba-bd7e-e86c7778ccc3 · outbound

This paper cites 2026 , doi=.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs 2026 , doi=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.332170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.332170Z digest=sha256:71e60c2619db1068a16a65a7f7ee48dac0757e5f175180265fcae5c3d95d9a61

Observation 62f56076-ad9a-4f85-aad4-b4098d53d252 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.338349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.338349Z digest=sha256:09fac84526546b99070d2b2c07799fabde75bff20302d685731b03ed19004b0c

Observation 3af5e994-d826-4005-8944-7fc16d7c7306 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.343768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.343768Z digest=sha256:55b816cb397c01f5a5b865e0d414d9258830aa9d47d6018aebda5637cf91e2f6

Observation e4f1db77-f85a-4c11-ad23-b506e707f112 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.351569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.351569Z digest=sha256:de1d1ac04a4e6aabcaf38cc3fc10d4d027783bb256aa56f8ceb51c5f8f3d0b7a

Observation ccfd8d71-a591-4287-b8d5-55dec2339b09 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.358437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.358437Z digest=sha256:55f190300c8f425284b9711f2b34ef3ea30066834726490c78a49243db678db0

Observation 45a57932-d5ce-487a-86fd-7eeaf4f3842a · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.364008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.364008Z digest=sha256:e74d795bcf42719422cffbcfaf903c9f138b5045a079f2b980db8c92b9690cdc

Observation 906a29b3-ac21-44ff-a46c-00cde8e97541 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.370019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.370019Z digest=sha256:b82735e08b82d40264f943e18bddbc367b8ea9453ec026947c0fdfa8794ac80d

Observation 7129b256-d0af-42df-bc8f-08975a24980a · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.375441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.375441Z digest=sha256:2b2a56b6e702746337a01169d208021d82513030015dc18cd30dc31611cc6a27

Observation 1667be9d-b56b-4448-a0b2-48a09f29e18d · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.380504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.380504Z digest=sha256:c0a2bf695ba5dca6fc4b3647fe76356ae068574470fb7d28d0ebd30cb00d215e

Observation 56591d01-76df-48ab-b485-b3d663f24c4d · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.385833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.385833Z digest=sha256:ab9f01f7520d1fa6c2a19b2a696528376f8afca7470f67e4b6ab051c6ad3f568

Observation db313d49-0823-406e-ae18-6e7bfd001b55 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.391357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.391357Z digest=sha256:550eeb971bfc75bf597e7a11c4ad439c2749458ab7b33c50bbc69f6d1afe6388

Observation 02cd2c1c-c82c-457a-9021-8192c9482b9a · outbound

This paper cites LMM s-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LMM s-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.395879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.395879Z digest=sha256:56a86c6bfa81c8c70fbe3f32ae8755078e306b384ff669a7cc83f5fecbb074bb

Observation 94ddffae-aa2f-47ee-83fa-78e024b07ac4 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.401662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.401662Z digest=sha256:da18139c473128066526ce63b359643fbde13fb443c4af051374ec854ec49de6

Observation 9270c765-1ea8-472d-a095-d4f069c68473 · outbound

This paper cites Baichuan-Omni-1.5 Technical Report.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Baichuan-Omni-1.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.406829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.406829Z digest=sha256:7a00856916075ac085373e623cb95364d5086a62ab26e308e6f005f455749198

Observation 6682a2c7-af30-478b-a4c6-c44d318d0bc2 · outbound

This paper cites Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.415210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.411994Z digest=sha256:6edcd9d23865985f5fb969ba50fe557c572cd7edc0182f138213c216fcfc7f3f

Observation bc0daf9b-3ca3-4e6f-8379-677bd9ef0604 · outbound

This paper cites LLaVA-OneVision : Easy Visual Task Transfer.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LLaVA-OneVision : Easy Visual Task Transfer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.399988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.419681Z digest=sha256:65af67aff69d079bfb2093c216f764bd27ad5990f921cf4cf5ffd2ed8c4052b8

Observation d32a45fc-62fb-479d-9c08-12dee8d169ba · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.424942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.424942Z digest=sha256:7a69c2a803bf20ad83120d2aa931d6c8d23c924497b57d636cc2bb1321992673

Observation 13bd0a8e-d5a8-424c-9362-42bda027e546 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.430937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.430937Z digest=sha256:06c3d5878f87a671e6ebba9f2fa509728c1f0848e4d20b34e97c35792d91d567

Observation 012b1cc9-df6e-435f-aa48-8bbc157a53fd · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.437312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.437312Z digest=sha256:eb79f3e3f4f8bc566a098452c22bec43e7f560676676d573c64cbe1f92a4185b

Observation d1b2dc1d-f0bb-4b6f-b77c-f055451be77a · outbound

This paper cites LongVILA : Scaling Long-Context Visual Language Models for Long Videos.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LongVILA : Scaling Long-Context Visual Language Models for Long Videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.381587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.442884Z digest=sha256:870d91ff87a425bd7253b4f775d21c99aae23c459feb47de02f3ec1535e2af10

Observation 2efd3ad1-3fff-475a-b3bf-7a80ba2b9104 · outbound

This paper cites LongVU : Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LongVU : Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.362809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.449535Z digest=sha256:c96f339a9228e0eb6a7c5bde63d6d9d98b57a7951bdffb625cb55616f4dc0da0

Observation c29c0a63-d8a6-4103-a3e5-057db1a96a28 · outbound

This paper cites Video Instruction Tuning with Synthetic Data.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Video Instruction Tuning with Synthetic Data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.347526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.456081Z digest=sha256:7de282a6441daa21314bfaa8810118e9f130b075a6a4e49db8bf07c24b6572ec

Observation 0037e256-d9b4-4d41-becf-3bb0cec5eac6 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.331432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.461372Z digest=sha256:7cfc1ec97662d23af2ade97a38eda0735fbf5dc79cbea8e3a6103c417bba279c

Observation 046cf382-402c-452f-ae15-b74a854cb560 · outbound

This paper cites An Image Is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs An Image Is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.312580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.467654Z digest=sha256:69fafdfaeb725435b095a0f76ebf0bc4d864354b2e1b0d271cbfa943f2fa3cb3

Observation 20d38b10-bca2-4b62-92b1-39ae80471235 · outbound

This paper cites DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.294426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.472640Z digest=sha256:3c5870d7b51d5e04542e1ecfac5a276b1be20a5c55881600d123ad74ae7987d7

Observation 78fd3e83-f233-4299-aec1-21da43ac470d · outbound

This paper cites PruneVid : Visual Token Pruning for Efficient Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs PruneVid : Visual Token Pruning for Efficient Video Large Language Models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.274850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.478220Z digest=sha256:65e611ee001d444392612b82bdbecaaf3d9ee3b59c78291cc5bd482517e29d8b

Observation 5b8b6cc2-04d4-49c8-a708-386667c84984 · outbound

This paper cites FastVID : Dynamic Density Pruning for Fast Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs FastVID : Dynamic Density Pruning for Fast Video Large Language Models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.259224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.482966Z digest=sha256:b2ae711a0a4502a8cb0fa6c2d5c83016cf1378f5f1b72ca31a95f672122fbfe6

Observation 40ff97e5-5235-41da-9888-f1837cb7c238 · outbound

This paper cites HoliTom : Holistic Token Merging for Fast Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs HoliTom : Holistic Token Merging for Fast Video Large Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.241978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.487995Z digest=sha256:6076f3f2cc98edac94c8da67141834a504dd35d8de94fef83b4c096dce80cc47

Observation 21d7d245-73f0-40ff-ab58-ed4438551f09 · outbound

This paper cites OmniZip : Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniZip : Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.226536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.492968Z digest=sha256:14e21aeb1976824d011b1a889c3769c6116ed4322a53bc63f82389b9f62d952c

Observation a27be746-c0a6-49f5-a573-d0fc220d049d · outbound

This paper cites Multimodal Long Video Modeling Based on Temporal Dynamic Context.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Multimodal Long Video Modeling Based on Temporal Dynamic Context

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:34:05.242869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.496979Z digest=sha256:d9675568226d818df763bc383e94455ef4f99a7b3a93a895c0d825b7b847ddcc

Observation 620e0f27-9d62-4806-9aa2-c697a925b576 · outbound

This paper cites Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.211265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.501800Z digest=sha256:7c826d5a5bd50a883e09dd66b4c4fcf0d36dda79b717878fe91e1dc1c6b442b8

Observation 13fc8d40-fc58-4a20-b13a-31960f99fccd · outbound

This paper cites Aligned Better, Listen Better for Audio-Visual Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Aligned Better, Listen Better for Audio-Visual Large Language Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.195410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.508206Z digest=sha256:1873fd79667559f8a3ac247a4f34bdca7de200a746d83b9183d714fe56e80d88

Observation b0fc5cbd-8661-461c-b8f6-b9e5b68855d0 · outbound

This paper cites Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.178604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.513212Z digest=sha256:8cf239798689666d4dee4e9c5709f4e9f997c0aa53d9cf7220c157340b83566d

Observation 902cd03c-7cd2-4ea8-97cb-e73a02b6af51 · outbound

This paper cites AVHBench : A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AVHBench : A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.162128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.517575Z digest=sha256:44442714dd2ffb7a7f5caa511607e5110df7b24a34b9511c77ae5c1dc4eb8070

Observation d1693f6f-bee1-4d10-a7d9-7fe1406d7b6a · outbound

This paper cites AVCD : Mitigating Hallucinations in Audio-Visual Large Language Models Through Contrastive Decoding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AVCD : Mitigating Hallucinations in Audio-Visual Large Language Models Through Contrastive Decoding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.146078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.521649Z digest=sha256:d1f81026b62d97da9e18cde9f051566237ece860debdf7a8e5b88b1c091dc587

Observation d5cd0636-6da0-4fe1-99f2-9db246767965 · outbound

This paper cites AVQA : A Dataset for Audio-Visual Question Answering on Videos.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AVQA : A Dataset for Audio-Visual Question Answering on Videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.129647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.526838Z digest=sha256:ab814bacba533e16da596ad3bd034a6c92a63968cb2fa6fb20644898e5062930

Observation e22d1398-0f27-46ea-93c6-8dd4ff7cff4e · outbound

This paper cites Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:34:05.219231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.531789Z digest=sha256:fa80437a1588ec0ec39f438b92a2739fc75842e2576dc9e8382ba04b748520c1

Observation 3fe58d8c-1878-4b2f-9803-85a650505031 · outbound

This paper cites OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.537374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.537374Z digest=sha256:bbc2df28592fe185803bd7367e45db87af2c5295ffa6a58fe9a2a3d0a3de37de

Observation 91310692-7bc8-42b5-a9fa-1e60c267c475 · outbound

This paper cites The Platonic Representation Hypothesis.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs The Platonic Representation Hypothesis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.113358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.543084Z digest=sha256:8c84284391a214bdd5c947f30c252dd1ceb51a4aa2b37d4188915996ea2e848f

Observation 9ef7ba3c-f573-4ecf-be10-c6beb2478b62 · outbound

This paper cites Understanding the Emergence of Multimodal Representation Alignment.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Understanding the Emergence of Multimodal Representation Alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.099263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.548374Z digest=sha256:5b6748b729182bbc9b39e26a9306db3e9dd497654b2d79bc1587900a00c22b65

Observation 28e9c72b-d78b-4e6b-aa51-377fdc83cbf4 · outbound

This paper cites To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.084016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.553055Z digest=sha256:480d63d7fceccad0964079458e4d7df55f62f7d7edccc6213c309956df767270

Observation 15b7d1cb-f380-4f98-8243-c0785b5f547e · outbound

This paper cites Similarity of Neural Network Representations Revisited.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Similarity of Neural Network Representations Revisited

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.069222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.560219Z digest=sha256:fb21da66d566ed71da8eb3337870999ae145c284dac4d9955118dd44a3992919

Observation 9b0de87e-4077-4260-89a4-1fbb96097ea6 · outbound

This paper cites The Effective Rank: A Measure of Effective Dimensionality.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs The Effective Rank: A Measure of Effective Dimensionality

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.054139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.566143Z digest=sha256:de102964fa63ca1e6e8daca73689aa87dc2a231aaeee6cd7bc3eec32d6aa5b25

Observation 221ffa55-ae16-43d9-96de-a74f629bfd4a · outbound

This paper cites Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.039294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.572585Z digest=sha256:7d0d79e85ed322c1a029db51bfbcbff8d82e77a113d32d297635284043c8bb6c

Observation 26158e9b-58a6-4e1d-99cf-2eacee9f2c68 · outbound

This paper cites Sinkhorn Distances: Lightspeed Computation of Optimal Transport.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Sinkhorn Distances: Lightspeed Computation of Optimal Transport

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.022602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.580491Z digest=sha256:5153fdb34312ded1e122c1abe211f87727697b3cbb835fed4e6e587c87a99708

Observation 21d54485-0ab1-40e0-9e06-aa9ba9ab845e · outbound

This paper cites Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.004863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.585625Z digest=sha256:d429a89b74ef7869f524001ee5decb8f1a136ee17c256a577c2b48e8410b1c93

Observation 23000116-fa86-4888-9946-48e9440d1272 · outbound

This paper cites Adaptive Keyframe Sampling for Long Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Adaptive Keyframe Sampling for Long Video Understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.989671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.590459Z digest=sha256:b61ae029139a4328cc7e9e4862daeeda64f9bb16bc60b6d84f34b7ac1a104122

Observation 14738c7a-226a-43fa-a065-1b3d5d3f415e · outbound

This paper cites Clustering by Fast Search and Find of Density Peaks.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Clustering by Fast Search and Find of Density Peaks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.975420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.596465Z digest=sha256:2cc0b2fbfe7493b7b61b64152453ab7f97e4312607d6ed0c866b49058793b874

Observation 98afac82-1ff5-4321-90c8-ce5756a4307d · outbound

This paper cites and Yi, Li and Su, Hao and Guibas, Leonidas J.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Yi, Li and Su, Hao and Guibas, Leonidas J

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.960542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.601798Z digest=sha256:ff269604ccb46ac7abdf96cdcc5551e908313e587bd2df09246baf0a0103a7de

Observation eceea694-4d49-4366-bbd1-78cb6c2855a7 · outbound

This paper cites Attributing Response to Context: A Jensen--Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Attributing Response to Context: A Jensen--Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.944556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.606530Z digest=sha256:49cfe47a409bf68d1e77f745c53ee645df27972372dfa950a095e5f2fc31012d

Observation 99514c01-d188-4f6b-ac5e-d7c550e56d6a · outbound

This paper cites Quantifying the Plausibility of Context Reliance in Neural Machine Translation.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Quantifying the Plausibility of Context Reliance in Neural Machine Translation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.928618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.611419Z digest=sha256:1baf09e34772990c833021a0a7c44a711d9ef64dd84b35b6c2907cb5ae43ead2

Observation c0ae69a5-3a3f-4079-bbbd-8d8adb69f650 · outbound

This paper cites AgilePruner : An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AgilePruner : An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.914346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.616843Z digest=sha256:8083709783ccdde1a871d72b3c1d3560216dac14fbea8ed12d09f9014da5f1ef

Observation 2a750601-c2c6-4f01-b93e-ba0fbe99bf18 · outbound

This paper cites CoSeLECT : Adaptive Frame Selection for Video-Language Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs CoSeLECT : Adaptive Frame Selection for Video-Language Understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.899299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.621422Z digest=sha256:493e2f014e1e5834cc8423512b95d612713c032dbe4e380e25fdc450e466dc0b

Observation 7b694104-52af-45fd-a91d-f4678141e143 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.883881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.626942Z digest=sha256:68acb48d27e2970105827798b593242c902fbde55e98c920d58dbea214dd6354

Observation 926d55fb-cad7-4a8d-b1f0-7d9814db5e68 · outbound

This paper cites WorldSense : Evaluating Real-World Omnimodal Understanding for Multimodal LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs WorldSense : Evaluating Real-World Omnimodal Understanding for Multimodal LLMs

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.867057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.631808Z digest=sha256:f907837e9e7b1cdb239bbd89bc270e3e7969a3dce4ec120081d1eb395ac35c1e

Observation 64c8f548-9d0a-4641-9309-e91cefd6ff86 · outbound

This paper cites Daily-Omni : Towards Audio-Visual Reasoning with Temporal Alignment Across Modalities.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Daily-Omni : Towards Audio-Visual Reasoning with Temporal Alignment Across Modalities

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.636738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.636738Z digest=sha256:3851b2a655ab7ab6c6b0e61757c2598d7359fd79577a99772d95c3f05a154055

Observation d9ace57e-e148-4d8e-af3e-ce6cb554f901 · outbound

This paper cites Audio-Centric Video Understanding Benchmark without Text Shortcut.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Centric Video Understanding Benchmark without Text Shortcut

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.851676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.641834Z digest=sha256:204116ea0b5a5d6e4383e89025611d994f1a3eb98def05a3501ca5584ce5e291

Observation e3912377-8156-48c7-9e22-962c616d9b24 · outbound

This paper cites Diff-Foley : Synchronized Video-to-Audio Synthesis with Latent Diffusion Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Diff-Foley : Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.835798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.646684Z digest=sha256:882aa08468032a56fc41f1c77d47953b9fecdcbea88b039b8c05c5bd5560ed55

Observation 178ccf26-e96e-4c2c-a7a0-0e367b3799d9 · outbound

This paper cites Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.819206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.651487Z digest=sha256:4295995b77d89c380cbfacdb98f70e25843415bf7d0221e0cbb56575ebd63554

Observation 33098789-4471-45a3-a7c5-ca33a8f5a5f8 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.804628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.656263Z digest=sha256:fe332002648ecce35fbadee0dc623cf7be8c8d1e3cfd1a66ac7684281803c014

Observation ac9ede9c-465a-42dd-9f9a-382d03908a40 · outbound

This paper cites FlashVID : Efficient Video Large Language Models via Training-Free Tree-Based Spatiotemporal Token Merging.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs FlashVID : Efficient Video Large Language Models via Training-Free Tree-Based Spatiotemporal Token Merging

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.789612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.661286Z digest=sha256:1eabe3fef806ec3a3f30ccf97c80c5aa04485732a9224bf4cd66803729078f8f

Observation edad8f17-c128-48f6-890c-5c7e6e0e3ba8 · outbound

This paper cites TopV : Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs TopV : Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.773216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.666530Z digest=sha256:c863638e1ac5991f89d676753fb3b1e6440e3638d3a59ac7e1cfa4370b092290

Observation cbd5359d-9be7-44c6-8f2a-0f0c753fb91e · outbound

This paper cites PyramidDrop : Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs PyramidDrop : Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.757400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.672141Z digest=sha256:d832cac9eb99dea40d001e94cfc70cd25ce6e63cb5345d8e85872371315e9166

Observation 8719b179-d9f2-4c69-a8e0-2cd756b375cd · outbound

This paper cites VoCo-LLaMA : Towards Vision Compression with Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VoCo-LLaMA : Towards Vision Compression with Large Language Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.739265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.678298Z digest=sha256:957fc0470ce401c4699c4e62ec985a2609bed2954e2c7bafc5774b685a7f6c32

Observation 54a91276-250e-4db6-a8c7-3fe73e6ab9e1 · outbound

This paper cites TimeViper : A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs TimeViper : A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.722897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.684596Z digest=sha256:9ca1bb3524e3516da64c6cd574e6881fcf93a280b60d9f99d44e0ce7d781bc8f

Observation 1835568e-29b8-43c2-8073-3cba45d81be0 · outbound

This paper cites Token-Efficient Long Video Understanding for Multimodal LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Token-Efficient Long Video Understanding for Multimodal LLMs

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.706189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.692524Z digest=sha256:d4726dc5c31c4001624c2c518be92b04d611e4ea0ec15f11751926c72d95eb3c

Observation 73a0f927-b83c-46b7-8d13-6d73085bc84d · outbound

This paper cites BIMBA : Selective-Scan Compression for Long-Range Video Question Answering.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs BIMBA : Selective-Scan Compression for Long-Range Video Question Answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.689244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.701786Z digest=sha256:5f59f35db2650d79c10720981f045fb1a54320294a4daf073c1991ed429f6dbb

Observation 7a67e39e-9b33-4641-8368-58aee9764a30 · outbound

This paper cites AdaptInfer : Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AdaptInfer : Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.708133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.708133Z digest=sha256:80f9929fcc5e903ca368abfd9bbfcfb59c89c28bd99c6847cc2b3fbfe079b4a1

Observation e2a22333-4a4d-4a21-9f95-b7a7aa0a9d86 · outbound

This paper cites FastAV : Efficient Token Pruning for Audio-Visual Large Language Model Inference.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs FastAV : Efficient Token Pruning for Audio-Visual Large Language Model Inference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.713517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.713517Z digest=sha256:c3e9cff5c64f962c758065b6dbb04f821e2d6c635cb0edb8003a589ca39b322e

Observation 44a6249d-59a1-479a-8523-99c9fcc1da8a · outbound

This paper cites DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.718348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.718348Z digest=sha256:9351ee83b88049357c1fee3c6923c1b457025e4ea022f6016892e9dbd9373358

Observation c406ce49-8f1f-40a5-abe3-88b8e52210ba · outbound

This paper cites OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:34:04.944999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.723034Z digest=sha256:6dd87a6079d43b0907325144fb59a81b5dd2b6eb1e9fa5335abfcc8ba50e47f8

Observation 3157ce29-3cc7-4c34-9b3a-85498bb68e4a · outbound

This paper cites EchoingPixels : Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs EchoingPixels : Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.674878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.730466Z digest=sha256:e21f505125497a8048c359b55af6330c1c4e95705257eb6ab0cf695d37a181cc

Observation 65960827-8c85-46d3-9902-1a4db2967499 · outbound

This paper cites OmniSIFT : Modality-Asymmetric Token Compression for Efficient Omni-Modal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniSIFT : Modality-Asymmetric Token Compression for Efficient Omni-Modal Large Language Models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.660121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.737016Z digest=sha256:469be5e73b2ef7dac0e0f910792f83668b089897fc9f250a7150c3546a422d8d

Observation 725b0517-0e75-4d52-aaa9-4f046de9c8c8 · outbound

This paper cites Audio-Synchronized Visual Animation.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Synchronized Visual Animation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.644548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.743151Z digest=sha256:dfaf9432aa0b3364f957075c9c91b13107f6b4c4477539c2054b09505e4a6565

Observation a9ccf354-fc5b-42ce-8c72-2ce9b5bdff08 · outbound

This paper cites Objects that Sound.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Objects that Sound

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.627064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.749758Z digest=sha256:49dc6fce1b7d3d1119e28a342c25cd8a404253a5bf3abc186c131099f6fc23cb

Observation dd6af23b-38f1-4492-a931-706e676457f0 · outbound

This paper cites Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.610601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.755487Z digest=sha256:35c893cd1a4ca1385b412e89db29df77eb8a39c498a4b9ca27de0453af695218

Observation 95d0c50f-adb8-41f0-abff-99cc055106ae · outbound

This paper cites Anchor-Aware Deep Metric Learning for Audio-Visual Retrieval.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Anchor-Aware Deep Metric Learning for Audio-Visual Retrieval

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.593473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.762439Z digest=sha256:12d9c0125c586f23119a2b6748eacc0d6d3646185ed0ca1e34f73b8b6feeba10

Observation 27a66c54-1874-4f38-bc51-46fd99e6ce8f · outbound

This paper cites Audio-Visual LLM for Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Visual LLM for Video Understanding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.578614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.768260Z digest=sha256:72c2d59e51ddf2106898bbfb2db256cedff0b67c30d11b9c77c2bf800962806e

Observation 1ef972aa-acf3-41ab-8951-38a0fe816153 · outbound

This paper cites OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.773342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.773342Z digest=sha256:d0a2b212c24ed9c9bc513e2f1851c7727ef1f23f5edf86e3426b8277f75ffb16

Observation 6d2819aa-50b1-4714-a53d-552641e0bdea · outbound

This paper cites Stage-adaptive Token Selection for Efficient Omni-modal LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.779683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.779683Z digest=sha256:45577c38f60362d47c478f6598a21eba051c12ee580f0f7ab7639cea6633d113

Observation 9e82ba95-0465-4a0c-aa11-7e8ab7d3586d · outbound

This paper cites Temporal Auditory Acuity.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Temporal Auditory Acuity

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.563096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.785326Z digest=sha256:d343b2e471eeeb1f51fbb8fd9165d4ed5855bcea6dfd41266b34e246354a9000

Observation 56bc9745-d559-4a8c-9484-922bf615b84f · outbound

This paper cites and Warren, David H.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Warren, David H

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.548516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.791209Z digest=sha256:4154f7a16a62e8f299519da169ef3406387c56ed4d3c1a0363d18cbd13a10d5d

Observation a0b1ded9-fb54-41ec-8aa3-317b48451716 · outbound

This paper cites What You See Is What You Hear.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs What You See Is What You Hear

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.533231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.797499Z digest=sha256:ad76744e20515d65a2286c527d1d640ef1a4de25ba5c2c35f16544de38b1d778

Observation 6d880b55-d4cf-4bcf-a020-0ea538a89017 · outbound

This paper cites and Olshausen, Bruno A.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Olshausen, Bruno A

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.517596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.801575Z digest=sha256:41f232a5eddd9cac2b583b0ac3d20cf5cd1354e08a430bf7d9840541f554dcc0

Observation aca657c4-db95-4e83-bae4-dcefca7f6698 · outbound

This paper cites and Theunissen, Fr \'e d \'e ric E.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Theunissen, Fr \'e d \'e ric E

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.501640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.805958Z digest=sha256:d1375c9e775a5a1d222d6ceab2bc270f38a584fbee02b8a55c1c66a1cb7d62e3

Observation 1e95843f-0875-4405-958f-ac4d56a8a436 · outbound

This paper cites and Plomp, Reinier.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Plomp, Reinier

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.485468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.810078Z digest=sha256:dd26f0cf03a7274495271b5c77b7ad69446a7a768037094345b970f6b1ff16a6

Observation e42a8b96-ec5a-4b3f-8af7-b40c58a1d004 · outbound

This paper cites and Zeng, Fan-Gang and Kamath, Vivek and Wygonski, John and Ekelid, Michael.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Zeng, Fan-Gang and Kamath, Vivek and Wygonski, John and Ekelid, Michael

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.470311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.817120Z digest=sha256:896d528790a87abc130be89eb258b9488cce20ac9ff6517bda56f945fa368558

Observation 68185178-c98a-434c-b905-241a5b234e5c · outbound

This paper cites Different Languages, Similar Encoding Efficiency: Comparable Information Rates across the Human Communicative Niche.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Different Languages, Similar Encoding Efficiency: Comparable Information Rates across the Human Communicative Niche

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.454116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.822617Z digest=sha256:6f1e2f93c6cd22132be4d6b93b5b0727c206a517fc84b83f450d4882dddde5f2

Observation 76ff274f-e5be-43b2-bd9c-db1223759c91 · outbound

This paper cites and Poeppel, David.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Poeppel, David

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.436834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.827388Z digest=sha256:916e57d670a08c0a18cc3a2e5f7cf235606eebaff9d8c8057df8787fe025f9f2

Observation 5245dd19-7af5-4484-bfba-148ca37450c7 · outbound

This paper cites and Stevenson, Ryan A.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Stevenson, Ryan A

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.419165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.833294Z digest=sha256:29e308556dc40e2fa668352142f23732dd7a0a4c0b12b1d778af5f7db6629fa1

Pith citing papers

No inbound Pith citation observations are available.