Pith. sign in

Paper Citation Record · LEDGER

Probing Cross-modal Information Hubs in Audio-Visual LLMs

As of 6 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2605.10815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10815 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T03:04:56.762735Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact9
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c43446c-8217-4f62-9c1a-c5e70c1583dd · outbound

This paper cites Some modalities are more equal than others: Decoding and architecting multi- modal integration in mllms.arXiv preprint arXiv:2511.22826.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Some modalities are more equal than others: Decoding and architecting multi- modal integration in mllms.arXiv preprint arXiv:2511.22826

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T03:07:08.707153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:4a81aa5a84971cc91cd41b202a64757682456a53d56f845d5202d80e7bad9600

Observation b76aaba4-84db-42bb-824c-ca06d81be28b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Probing Cross-modal Information Hubs in Audio-Visual LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:07:08.704483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:5d66c9dc76849979a9ea259d77f39127c99338d72a3af68ee3213cfec7bd2d00

Observation 94f2382e-1661-4d72-a753-a3c1da1aa451 · outbound

This paper cites Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.686182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:4453cf74d0917d7cb60c5ba6ceb3825d479a70f3738ef0cdb542272c61d3df72

Observation bb5e01b9-98bb-4aa9-9a7e-d71a8b752076 · outbound

This paper cites Fork-merge decoding: Enhancing multimodal understanding in audio-visual large language models.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Fork-merge decoding: Enhancing multimodal understanding in audio-visual large language models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.695271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:546c9d466c7fae496fe6e980194e8ba4fe13a248a4025fbf793ae84d977572e3

Observation 26ee8227-5eea-45f8-9223-3cd80b427a22 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.701574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:02cfaec6a39f3ed740583f1a08d4148fddca03d065bd3fbdd50c37cf90e4d49b

Observation d5d3ea50-aaf8-4e5b-91ae-a3545a8b3d34 · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

Probing Cross-modal Information Hubs in Audio-Visual LLMs On the Audio Hallucinations in Large Audio-Video Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.692174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:ab7dc2fa16d6419d2df0d39eac06353de7ba84ae4b4c29e1afd2e2bc7fce8d57

Observation c7bcbc8a-7830-4d1f-81f0-311bc2a3803e · outbound

This paper cites video-SALMONN 2: Caption-enhanced audio-visual large language models.

Probing Cross-modal Information Hubs in Audio-Visual LLMs video-SALMONN 2: Caption-enhanced audio-visual large language models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.689231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:36128b1f39f9d8b24dcb85e12e7478b1ea7e47bc6f80e08bf88ad04b7f256afb

Observation 65dc4318-994a-43c6-bbf6-46277cfa1d9a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Probing Cross-modal Information Hubs in Audio-Visual LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:07:08.709905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:1692a911a46f5b44cb248e94be93ff3a52955a48e2c45c9171c3861a3fed61b8

Observation ea141ae5-e609-48b4-a234-162fe52b04bb · outbound

This paper cites Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.682622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:59e240bf61e10222a08d6f0f9e26f9d2cd7a7d1bdaf497f2c7eda62a972f85c9

Observation 4d1b3571-0695-433d-b690-0d883765f688 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Qwen2.5-Omni Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:07:08.698135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:a418048d46759e2db84f7157e1b9060eae8191a0a7962c97db0e112e65c1c556

Observation da79c865-359a-4ff5-9541-9a7530e157e5 · outbound

This paper cites The structure is organized as follows: A.

Probing Cross-modal Information Hubs in Audio-Visual LLMs The structure is organized as follows: A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:32:45.737554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:54cd6a44fffeb669213ddc9823aebb22a701517ee3706dbbd183d6936547bf97

Observation ae2f560c-04d6-4a1d-abc5-71ebd7422e4b · outbound

This paper cites Specifically, visual sink tokens in VLMs exhibit massive activation along the same dimensions as the BOS token in the base LLM.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Specifically, visual sink tokens in VLMs exhibit massive activation along the same dimensions as the BOS token in the base LLM

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:32:45.739514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:ba9b7cf2ef12098ac4a0575ba440cd1b2e45890c84ca884fe38bbbd9ee525c04

Observation 860189c5-28c2-4f94-975f-f46bec40e12d · outbound

This paper cites While the original VCD applies noise solely to the image modality, we extend this approach to the audio-visual domain.

Probing Cross-modal Information Hubs in Audio-Visual LLMs While the original VCD applies noise solely to the image modality, we extend this approach to the audio-visual domain

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:32:45.741872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:d89b08d8210ce4cb2d357cbf83a3ac769ad2fe91f3eaaf3889185aea5de24ca1

Observation f2a7b017-2a83-4388-8cf8-5716597debcf · outbound

This paper cites an unresolved cited work.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Unresolved cited work

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T12:32:45.743525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:9510e89a0c270181f1956264e089adfbb611967132b0266ec81d91acdae68a11

Observation 61663980-3933-4c73-8b59-681b57d5a65a · outbound

This paper cites an unresolved cited work.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Unresolved cited work

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T12:32:45.746978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:6571fb82989528356d17ba22b9009b6d7923fa8e3742d588b35be81748f73b63

Observation 8004c53a-042b-4592-8c59-4bf149a724f9 · outbound

This paper cites As shown in Tab.

Probing Cross-modal Information Hubs in Audio-Visual LLMs As shown in Tab

Reference 16

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T12:32:45.745197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:bf04e1e44e31f0cafcbb514cd391b7353f4a1c4a0c14529e47e0f809cb1d8bb6

Pith citing papers

No inbound Pith citation observations are available.