Pith. sign in

Paper Citation Record · LEDGER

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

As of 23 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2608.08794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08794 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:04.833294Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact3
  • verified fuzzy63
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 290bf7ea-f76f-42cf-9227-995e1e933d68 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Qwen2.5-Omni Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.323992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.323992Z digest=sha256:8674b30aacb9c59e93a969678ff16e5a0b92f036ce2f6b2f8960ee65b24e4f56

Observation 41c3e960-98cd-4cba-bd7e-e86c7778ccc3 · outbound

This paper cites 2026 , doi=.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs 2026 , doi=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.332170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.332170Z digest=sha256:71e60c2619db1068a16a65a7f7ee48dac0757e5f175180265fcae5c3d95d9a61

Observation 62f56076-ad9a-4f85-aad4-b4098d53d252 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.338349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.338349Z digest=sha256:09fac84526546b99070d2b2c07799fabde75bff20302d685731b03ed19004b0c

Observation 3af5e994-d826-4005-8944-7fc16d7c7306 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.343768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.343768Z digest=sha256:55b816cb397c01f5a5b865e0d414d9258830aa9d47d6018aebda5637cf91e2f6

Observation e4f1db77-f85a-4c11-ad23-b506e707f112 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.351569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.351569Z digest=sha256:de1d1ac04a4e6aabcaf38cc3fc10d4d027783bb256aa56f8ceb51c5f8f3d0b7a

Observation ccfd8d71-a591-4287-b8d5-55dec2339b09 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.358437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.358437Z digest=sha256:55f190300c8f425284b9711f2b34ef3ea30066834726490c78a49243db678db0

Observation 45a57932-d5ce-487a-86fd-7eeaf4f3842a · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.364008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.364008Z digest=sha256:e74d795bcf42719422cffbcfaf903c9f138b5045a079f2b980db8c92b9690cdc

Observation 906a29b3-ac21-44ff-a46c-00cde8e97541 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.370019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.370019Z digest=sha256:b82735e08b82d40264f943e18bddbc367b8ea9453ec026947c0fdfa8794ac80d

Observation 7129b256-d0af-42df-bc8f-08975a24980a · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.375441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.375441Z digest=sha256:2b2a56b6e702746337a01169d208021d82513030015dc18cd30dc31611cc6a27

Observation 1667be9d-b56b-4448-a0b2-48a09f29e18d · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.380504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.380504Z digest=sha256:c0a2bf695ba5dca6fc4b3647fe76356ae068574470fb7d28d0ebd30cb00d215e

Observation 56591d01-76df-48ab-b485-b3d663f24c4d · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.385833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.385833Z digest=sha256:ab9f01f7520d1fa6c2a19b2a696528376f8afca7470f67e4b6ab051c6ad3f568

Observation db313d49-0823-406e-ae18-6e7bfd001b55 · outbound

This paper cites an unresolved cited work.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.391357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.391357Z digest=sha256:550eeb971bfc75bf597e7a11c4ad439c2749458ab7b33c50bbc69f6d1afe6388

Observation 02cd2c1c-c82c-457a-9021-8192c9482b9a · outbound

This paper cites LMM s-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LMM s-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.395879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.395879Z digest=sha256:56a86c6bfa81c8c70fbe3f32ae8755078e306b384ff669a7cc83f5fecbb074bb

Observation 94ddffae-aa2f-47ee-83fa-78e024b07ac4 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.401662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.401662Z digest=sha256:da18139c473128066526ce63b359643fbde13fb443c4af051374ec854ec49de6

Observation 9270c765-1ea8-472d-a095-d4f069c68473 · outbound

This paper cites Baichuan-Omni-1.5 Technical Report.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Baichuan-Omni-1.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.406829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.406829Z digest=sha256:7a00856916075ac085373e623cb95364d5086a62ab26e308e6f005f455749198

Observation 6682a2c7-af30-478b-a4c6-c44d318d0bc2 · outbound

This paper cites Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.415210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.411994Z digest=sha256:aec513eec408c3c036698e785b1ebe2ffebab9cfb0a33e4b7e414c1b2a98decd

Observation bc0daf9b-3ca3-4e6f-8379-677bd9ef0604 · outbound

This paper cites LLaVA-OneVision : Easy Visual Task Transfer.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LLaVA-OneVision : Easy Visual Task Transfer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.399988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.419681Z digest=sha256:6c25bc8a7493a7ccfb7e5affa8e4798db6461befd3f5d8dcb457bf1b80f71f3f

Observation d32a45fc-62fb-479d-9c08-12dee8d169ba · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.424942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.424942Z digest=sha256:7a69c2a803bf20ad83120d2aa931d6c8d23c924497b57d636cc2bb1321992673

Observation 13bd0a8e-d5a8-424c-9362-42bda027e546 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.430937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.430937Z digest=sha256:06c3d5878f87a671e6ebba9f2fa509728c1f0848e4d20b34e97c35792d91d567

Observation 012b1cc9-df6e-435f-aa48-8bbc157a53fd · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.437312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.437312Z digest=sha256:eb79f3e3f4f8bc566a098452c22bec43e7f560676676d573c64cbe1f92a4185b

Observation d1b2dc1d-f0bb-4b6f-b77c-f055451be77a · outbound

This paper cites LongVILA : Scaling Long-Context Visual Language Models for Long Videos.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LongVILA : Scaling Long-Context Visual Language Models for Long Videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.381587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.442884Z digest=sha256:14d7aea588feead79fa0977549e17e797e3c3be6e4b5efac1754efd447b9f424

Observation 2efd3ad1-3fff-475a-b3bf-7a80ba2b9104 · outbound

This paper cites LongVU : Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs LongVU : Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.362809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.449535Z digest=sha256:96777e9ffba7ac4e7103ae61a8e347941a1bcc7a02f8d5699407456dba860cfb

Observation c29c0a63-d8a6-4103-a3e5-057db1a96a28 · outbound

This paper cites Video Instruction Tuning with Synthetic Data.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Video Instruction Tuning with Synthetic Data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.347526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.456081Z digest=sha256:596e9f3810428852a43c2cb91ed1df488cbcb528c827a988766b643102ba3c26

Observation 0037e256-d9b4-4d41-becf-3bb0cec5eac6 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.331432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.461372Z digest=sha256:fb832c2d363e24dc1e3f4e0cbaa1082d7b7d756b2ea4887201d774cb6873e12f

Observation 046cf382-402c-452f-ae15-b74a854cb560 · outbound

This paper cites An Image Is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs An Image Is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.312580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.467654Z digest=sha256:67c3b646eb55e1589b5e152b9cf3e03ed7df2f6e03eac4ba7496067f9a59f93d

Observation 20d38b10-bca2-4b62-92b1-39ae80471235 · outbound

This paper cites DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.294426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.472640Z digest=sha256:59b5cd1723879612f482a562e45625cf9183076a98956c5cb1682a30c75a888f

Observation 78fd3e83-f233-4299-aec1-21da43ac470d · outbound

This paper cites PruneVid : Visual Token Pruning for Efficient Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs PruneVid : Visual Token Pruning for Efficient Video Large Language Models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.274850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.478220Z digest=sha256:491907ca12dc7e5f5fd00dd2bc8ae6a35b6ba88c3f339e2189f2dc96950e8296

Observation 5b8b6cc2-04d4-49c8-a708-386667c84984 · outbound

This paper cites FastVID : Dynamic Density Pruning for Fast Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs FastVID : Dynamic Density Pruning for Fast Video Large Language Models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.259224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.482966Z digest=sha256:3d64c070acec0f8fbc4be1a54ef61bd900dcdf65de22c7f9ff07f89fc7bddee3

Observation 40ff97e5-5235-41da-9888-f1837cb7c238 · outbound

This paper cites HoliTom : Holistic Token Merging for Fast Video Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs HoliTom : Holistic Token Merging for Fast Video Large Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.241978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.487995Z digest=sha256:6592be3921d771272f513cf4d0642cb2e1850bd02e94116a4aad62c37d739425

Observation 21d7d245-73f0-40ff-ab58-ed4438551f09 · outbound

This paper cites OmniZip : Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniZip : Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.226536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.492968Z digest=sha256:b6cf2e1f3cb157d61ab902a2bcfca5a5aa18056b11fa6cdb3e29b1a60f571a25

Observation a27be746-c0a6-49f5-a573-d0fc220d049d · outbound

This paper cites Multimodal Long Video Modeling Based on Temporal Dynamic Context.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Multimodal Long Video Modeling Based on Temporal Dynamic Context

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:34:05.242869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.496979Z digest=sha256:dfd400c5b9503137384b8261656ad685ce81e47d8e37ce4a7f756873c913a06a

Observation 620e0f27-9d62-4806-9aa2-c697a925b576 · outbound

This paper cites Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.211265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.501800Z digest=sha256:bdca53512296084d61245efbd21a45d4d68250d70ff30a768653340d7f1b775a

Observation 13fc8d40-fc58-4a20-b13a-31960f99fccd · outbound

This paper cites Aligned Better, Listen Better for Audio-Visual Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Aligned Better, Listen Better for Audio-Visual Large Language Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.195410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.508206Z digest=sha256:e163f5757d3d7c158eec7c6d6e354cad27c791db82c7f1975fad9d0953cc7667

Observation b0fc5cbd-8661-461c-b8f6-b9e5b68855d0 · outbound

This paper cites Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.178604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.513212Z digest=sha256:29e316ae46005fb828d38be8017fbc29a007c0d7ec65b9a451461918c4df5079

Observation 902cd03c-7cd2-4ea8-97cb-e73a02b6af51 · outbound

This paper cites AVHBench : A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AVHBench : A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.162128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.517575Z digest=sha256:64d9324ae08f54124aec644fca1080f425119f723749552e57b928e971b0f1fc

Observation d1693f6f-bee1-4d10-a7d9-7fe1406d7b6a · outbound

This paper cites AVCD : Mitigating Hallucinations in Audio-Visual Large Language Models Through Contrastive Decoding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AVCD : Mitigating Hallucinations in Audio-Visual Large Language Models Through Contrastive Decoding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.146078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.521649Z digest=sha256:c61d0fd914f9fbc08ebd94db7aae95d03ef2682d64d3f09df6acca7a49f764ef

Observation d5cd0636-6da0-4fe1-99f2-9db246767965 · outbound

This paper cites AVQA : A Dataset for Audio-Visual Question Answering on Videos.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AVQA : A Dataset for Audio-Visual Question Answering on Videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.129647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.526838Z digest=sha256:348f9abb607480a6bc5b3387b13cc1517b782872616d3c36236622fcde2ceb52

Observation e22d1398-0f27-46ea-93c6-8dd4ff7cff4e · outbound

This paper cites Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:34:05.219231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.531789Z digest=sha256:6c0986908a8169386cd87cb8ce59af313a698b456bf517c1e7d311691a0716e6

Observation 3fe58d8c-1878-4b2f-9803-85a650505031 · outbound

This paper cites OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.537374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.537374Z digest=sha256:bbc2df28592fe185803bd7367e45db87af2c5295ffa6a58fe9a2a3d0a3de37de

Observation 91310692-7bc8-42b5-a9fa-1e60c267c475 · outbound

This paper cites The Platonic Representation Hypothesis.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs The Platonic Representation Hypothesis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.113358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.543084Z digest=sha256:d84f1b8e0558b4ab588dd50ff8c756831d6226e8d80fdc5ef49027803f70d70f

Observation 9ef7ba3c-f573-4ecf-be10-c6beb2478b62 · outbound

This paper cites Understanding the Emergence of Multimodal Representation Alignment.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Understanding the Emergence of Multimodal Representation Alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.099263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.548374Z digest=sha256:2f94d15436f78d66706bbeffc5e4c70028e8d00951868177879ac5607cedf94e

Observation 28e9c72b-d78b-4e6b-aa51-377fdc83cbf4 · outbound

This paper cites To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.084016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.553055Z digest=sha256:2a6da00db48dc77a68033d88b7ff9a4eca27e361df8afbca2a35b976436e2327

Observation 15b7d1cb-f380-4f98-8243-c0785b5f547e · outbound

This paper cites Similarity of Neural Network Representations Revisited.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Similarity of Neural Network Representations Revisited

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.069222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.560219Z digest=sha256:ffc62c85532a797bace60191f6b5a027587357d317fc1fb15dcceb62fe80c5f1

Observation 9b0de87e-4077-4260-89a4-1fbb96097ea6 · outbound

This paper cites The Effective Rank: A Measure of Effective Dimensionality.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs The Effective Rank: A Measure of Effective Dimensionality

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.054139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.566143Z digest=sha256:e9ad4319f23e673f05f5bcc8ab8e50d67920309abde80f876155886aa775f027

Observation 221ffa55-ae16-43d9-96de-a74f629bfd4a · outbound

This paper cites Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.039294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.572585Z digest=sha256:560c9b7c7f48af85516d40afc4aedb35e592162ac4ed2e22d2ec531854ba5b81

Observation 26158e9b-58a6-4e1d-99cf-2eacee9f2c68 · outbound

This paper cites Sinkhorn Distances: Lightspeed Computation of Optimal Transport.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Sinkhorn Distances: Lightspeed Computation of Optimal Transport

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.022602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.580491Z digest=sha256:bca18b8f3cf681649e1637f0b0391d5cc9f2b3bca25ce525b1311dec0f0c3943

Observation 21d54485-0ab1-40e0-9e06-aa9ba9ab845e · outbound

This paper cites Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Attention-Weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:06.004863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.585625Z digest=sha256:8c0f21a3c4c29577347d608aa23008375084065f11723eaacf1446b30627dc47

Observation 23000116-fa86-4888-9946-48e9440d1272 · outbound

This paper cites Adaptive Keyframe Sampling for Long Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Adaptive Keyframe Sampling for Long Video Understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.989671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.590459Z digest=sha256:d90573e9355670bd1ed23897b05ff5967e1c11eb21587b8c523988dbefbed932

Observation 14738c7a-226a-43fa-a065-1b3d5d3f415e · outbound

This paper cites Clustering by Fast Search and Find of Density Peaks.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Clustering by Fast Search and Find of Density Peaks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.975420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.596465Z digest=sha256:4d772defece5e06115cb90d08f103ce5b24f1d10591233c436f29c391d4c03cb

Observation 98afac82-1ff5-4321-90c8-ce5756a4307d · outbound

This paper cites and Yi, Li and Su, Hao and Guibas, Leonidas J.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Yi, Li and Su, Hao and Guibas, Leonidas J

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.960542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.601798Z digest=sha256:6f99ce2754562bb8bb723c9729206d3eef05de300247053e0ea55da3216d2d7d

Observation eceea694-4d49-4366-bbd1-78cb6c2855a7 · outbound

This paper cites Attributing Response to Context: A Jensen--Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Attributing Response to Context: A Jensen--Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.944556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.606530Z digest=sha256:334efb182277b4f0e17eb28caff8569a60e403417d89eff6a1df6bbdd328a0ab

Observation 99514c01-d188-4f6b-ac5e-d7c550e56d6a · outbound

This paper cites Quantifying the Plausibility of Context Reliance in Neural Machine Translation.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Quantifying the Plausibility of Context Reliance in Neural Machine Translation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.928618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.611419Z digest=sha256:d90367fcebf72997cffecb6a17321a31f9b3fc3bc2ddb97fbe399bfc538d34e2

Observation c0ae69a5-3a3f-4079-bbbd-8d8adb69f650 · outbound

This paper cites AgilePruner : An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AgilePruner : An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.914346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.616843Z digest=sha256:3253fc12874e027b541da804fec25674b04ec76e31b4ef0c5abb33bcd20c1256

Observation 2a750601-c2c6-4f01-b93e-ba0fbe99bf18 · outbound

This paper cites CoSeLECT : Adaptive Frame Selection for Video-Language Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs CoSeLECT : Adaptive Frame Selection for Video-Language Understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.899299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.621422Z digest=sha256:1dd82fc25c15657d8c11f933dbd4e860f3266797cb677f5dc4e3b8061ffd1c9a

Observation 7b694104-52af-45fd-a91d-f4678141e143 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.883881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.626942Z digest=sha256:5b81d5b270e1c3f73b7f41705e9e291566ed809d650e1675fddb422c9639ffac

Observation 926d55fb-cad7-4a8d-b1f0-7d9814db5e68 · outbound

This paper cites WorldSense : Evaluating Real-World Omnimodal Understanding for Multimodal LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs WorldSense : Evaluating Real-World Omnimodal Understanding for Multimodal LLMs

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.867057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.631808Z digest=sha256:9ab91a5e3e85847fe035ae38baaafe645968dc89f5a669e78a8982ebb6930117

Observation 64c8f548-9d0a-4641-9309-e91cefd6ff86 · outbound

This paper cites Daily-Omni : Towards Audio-Visual Reasoning with Temporal Alignment Across Modalities.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Daily-Omni : Towards Audio-Visual Reasoning with Temporal Alignment Across Modalities

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.636738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.636738Z digest=sha256:3851b2a655ab7ab6c6b0e61757c2598d7359fd79577a99772d95c3f05a154055

Observation d9ace57e-e148-4d8e-af3e-ce6cb554f901 · outbound

This paper cites Audio-Centric Video Understanding Benchmark without Text Shortcut.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Centric Video Understanding Benchmark without Text Shortcut

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.851676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.641834Z digest=sha256:7629e5b1611c3487a4ec4717845fc38cbef32baebc5ed3d057dea6e0e5efe2b8

Observation e3912377-8156-48c7-9e22-962c616d9b24 · outbound

This paper cites Diff-Foley : Synchronized Video-to-Audio Synthesis with Latent Diffusion Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Diff-Foley : Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.835798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.646684Z digest=sha256:f592cb93e337117fee3dd96cb85b10402e78400df7dd85f7211a4a17151f5b48

Observation 178ccf26-e96e-4c2c-a7a0-0e367b3799d9 · outbound

This paper cites Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.819206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.651487Z digest=sha256:cb560f26c06d1680972e6177fac40857001900a99dca58bb6c0e0279255d0a76

Observation 33098789-4471-45a3-a7c5-ca33a8f5a5f8 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.804628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.656263Z digest=sha256:e4189dbcf60ae1b9942dea8895e5dd8970a759ecf298c4c32e981e29b12e4d42

Observation ac9ede9c-465a-42dd-9f9a-382d03908a40 · outbound

This paper cites FlashVID : Efficient Video Large Language Models via Training-Free Tree-Based Spatiotemporal Token Merging.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs FlashVID : Efficient Video Large Language Models via Training-Free Tree-Based Spatiotemporal Token Merging

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.789612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.661286Z digest=sha256:ddec59773af0db6cd657468cf50018e5b9faccabbb9e80cec54ade249ba2f184

Observation edad8f17-c128-48f6-890c-5c7e6e0e3ba8 · outbound

This paper cites TopV : Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs TopV : Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.773216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.666530Z digest=sha256:5db90341e7e8e69dc60208d83eef7d224a3ecd0df977b94865285b55f06e0d64

Observation cbd5359d-9be7-44c6-8f2a-0f0c753fb91e · outbound

This paper cites PyramidDrop : Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs PyramidDrop : Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.757400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.672141Z digest=sha256:997472071e5d30a9427ea9db82d939b388ffc6be6227c300f104e823cdefe08c

Observation 8719b179-d9f2-4c69-a8e0-2cd756b375cd · outbound

This paper cites VoCo-LLaMA : Towards Vision Compression with Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs VoCo-LLaMA : Towards Vision Compression with Large Language Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.739265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.678298Z digest=sha256:0726d137a76c716dce822f6d93918e7b4145058a554c501e71c4ada6e1b87361

Observation 54a91276-250e-4db6-a8c7-3fe73e6ab9e1 · outbound

This paper cites TimeViper : A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs TimeViper : A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.722897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.684596Z digest=sha256:8b4b1aea95f22050bb5e10633d3b515a4e7b3541a6ae96c760efac82ff186d31

Observation 1835568e-29b8-43c2-8073-3cba45d81be0 · outbound

This paper cites Token-Efficient Long Video Understanding for Multimodal LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Token-Efficient Long Video Understanding for Multimodal LLMs

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.706189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.692524Z digest=sha256:aa041cf6b2a520e0e5eb511faedd953c68951c059d09fa28cbb75b2187c7477f

Observation 73a0f927-b83c-46b7-8d13-6d73085bc84d · outbound

This paper cites BIMBA : Selective-Scan Compression for Long-Range Video Question Answering.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs BIMBA : Selective-Scan Compression for Long-Range Video Question Answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.689244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.701786Z digest=sha256:b90aa43dbdfb582755819f4efe084cc265697eee0a81331734a63898fe04062a

Observation 7a67e39e-9b33-4641-8368-58aee9764a30 · outbound

This paper cites AdaptInfer : Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs AdaptInfer : Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.708133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.708133Z digest=sha256:80f9929fcc5e903ca368abfd9bbfcfb59c89c28bd99c6847cc2b3fbfe079b4a1

Observation e2a22333-4a4d-4a21-9f95-b7a7aa0a9d86 · outbound

This paper cites FastAV : Efficient Token Pruning for Audio-Visual Large Language Model Inference.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs FastAV : Efficient Token Pruning for Audio-Visual Large Language Model Inference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.713517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.713517Z digest=sha256:c3e9cff5c64f962c758065b6dbb04f821e2d6c635cb0edb8003a589ca39b322e

Observation 44a6249d-59a1-479a-8523-99c9fcc1da8a · outbound

This paper cites DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.718348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.718348Z digest=sha256:9351ee83b88049357c1fee3c6923c1b457025e4ea022f6016892e9dbd9373358

Observation c406ce49-8f1f-40a5-abe3-88b8e52210ba · outbound

This paper cites OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:34:04.944999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.723034Z digest=sha256:b778f789fc0fd3b44a2e42b9e04476461000dd04e687618211163bfc26160b45

Observation 3157ce29-3cc7-4c34-9b3a-85498bb68e4a · outbound

This paper cites EchoingPixels : Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs EchoingPixels : Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.674878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.730466Z digest=sha256:b096427c2b91d50670121212a027947e5466bfc9f5ded548857b39f04aa388f7

Observation 65960827-8c85-46d3-9902-1a4db2967499 · outbound

This paper cites OmniSIFT : Modality-Asymmetric Token Compression for Efficient Omni-Modal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniSIFT : Modality-Asymmetric Token Compression for Efficient Omni-Modal Large Language Models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.660121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.737016Z digest=sha256:a6f4a2d8b8f81d57d17f4efa80da5aea761da560be1862cc22a9da179daf709d

Observation 725b0517-0e75-4d52-aaa9-4f046de9c8c8 · outbound

This paper cites Audio-Synchronized Visual Animation.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Synchronized Visual Animation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.644548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.743151Z digest=sha256:fcdaa329b93138fbf27bdbcf01adbaa1feec79c70a7e57e07b80653de95f0d4e

Observation a9ccf354-fc5b-42ce-8c72-2ce9b5bdff08 · outbound

This paper cites Objects that Sound.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Objects that Sound

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.627064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.749758Z digest=sha256:96be2c26b2d37427d10fccb0614d131660b9c375e0feb26adcf2cb89e7c8580c

Observation dd6af23b-38f1-4492-a931-706e676457f0 · outbound

This paper cites Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.610601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.755487Z digest=sha256:fb6d5ab52eb1cbe99d0cf2f2632c690b59b4adb219e1801aebc4936d95c99879

Observation 95d0c50f-adb8-41f0-abff-99cc055106ae · outbound

This paper cites Anchor-Aware Deep Metric Learning for Audio-Visual Retrieval.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Anchor-Aware Deep Metric Learning for Audio-Visual Retrieval

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.593473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.762439Z digest=sha256:d40b3fbebf4cff4e047ce1ae6a6b9cf20bed31961440c5d479eb69f9daa05759

Observation 27a66c54-1874-4f38-bc51-46fd99e6ce8f · outbound

This paper cites Audio-Visual LLM for Video Understanding.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Audio-Visual LLM for Video Understanding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.578614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.768260Z digest=sha256:14c8f74eade4238bc3204d4464efec08547027d01f134e4e139b2226637d1db3

Observation 1ef972aa-acf3-41ab-8951-38a0fe816153 · outbound

This paper cites OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.773342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.773342Z digest=sha256:d0a2b212c24ed9c9bc513e2f1851c7727ef1f23f5edf86e3426b8277f75ffb16

Observation 6d2819aa-50b1-4714-a53d-552641e0bdea · outbound

This paper cites Stage-adaptive Token Selection for Efficient Omni-modal LLMs.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.779683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.779683Z digest=sha256:45577c38f60362d47c478f6598a21eba051c12ee580f0f7ab7639cea6633d113

Observation 9e82ba95-0465-4a0c-aa11-7e8ab7d3586d · outbound

This paper cites Temporal Auditory Acuity.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Temporal Auditory Acuity

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.563096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.785326Z digest=sha256:b77910832f3e2b77f8d95bae06dc3c290f446b0cf4f452401d9fd036db9ac949

Observation 56bc9745-d559-4a8c-9484-922bf615b84f · outbound

This paper cites and Warren, David H.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Warren, David H

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.548516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.791209Z digest=sha256:938ebed7eb9c7b1f57522d4827bb921ad704a11f1f6e38660156803cba51d82f

Observation a0b1ded9-fb54-41ec-8aa3-317b48451716 · outbound

This paper cites What You See Is What You Hear.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs What You See Is What You Hear

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.533231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.797499Z digest=sha256:ba193849dd11140922e91964d20fd3e49c8bff514f9a829e1dc4eb33e2463668

Observation 6d880b55-d4cf-4bcf-a020-0ea538a89017 · outbound

This paper cites and Olshausen, Bruno A.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Olshausen, Bruno A

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.517596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.801575Z digest=sha256:8a903df6a3cdf63aafcf7185d1d982664721807585990f663249d23e7a4d4704

Observation aca657c4-db95-4e83-bae4-dcefca7f6698 · outbound

This paper cites and Theunissen, Fr \'e d \'e ric E.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Theunissen, Fr \'e d \'e ric E

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.501640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.805958Z digest=sha256:63b3036741a679bd2d3276f6a3b5b404894e4bbbf251a8322d6195ea028520c2

Observation 1e95843f-0875-4405-958f-ac4d56a8a436 · outbound

This paper cites and Plomp, Reinier.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Plomp, Reinier

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.485468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.810078Z digest=sha256:25949b37290c360e66a57b8683c68c2d56305334b13636643fb74739d68751df

Observation e42a8b96-ec5a-4b3f-8af7-b40c58a1d004 · outbound

This paper cites and Zeng, Fan-Gang and Kamath, Vivek and Wygonski, John and Ekelid, Michael.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Zeng, Fan-Gang and Kamath, Vivek and Wygonski, John and Ekelid, Michael

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.470311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.817120Z digest=sha256:fa6573c57b96cb3ccb3ec00faa45c2cb8c21667157cde6a630d11deb76ee702f

Observation 68185178-c98a-434c-b905-241a5b234e5c · outbound

This paper cites Different Languages, Similar Encoding Efficiency: Comparable Information Rates across the Human Communicative Niche.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Different Languages, Similar Encoding Efficiency: Comparable Information Rates across the Human Communicative Niche

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.454116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.822617Z digest=sha256:02169e8dfc2c9effcb445efc6a8dfd1c812d8fbe5d1855eeae6d74002f88a91e

Observation 76ff274f-e5be-43b2-bd9c-db1223759c91 · outbound

This paper cites and Poeppel, David.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Poeppel, David

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.436834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.827388Z digest=sha256:d51b6223e3de7b089ff92d3cc1990796a404edff49129659a888b35e90f765bd

Observation 5245dd19-7af5-4484-bfba-148ca37450c7 · outbound

This paper cites and Stevenson, Ryan A.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs and Stevenson, Ryan A

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:34:05.419165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-14T04:34:04.833294Z digest=sha256:598398fd8ad6a85beac946f0cf7e012bf94252e11806d2760869ea05c1bc2fd5

Pith citing papers

No inbound Pith citation observations are available.