Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T03:04:56.762735Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2605.10815.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T03:04:56.762735Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8c43446c-8217-4f62-9c1a-c5e70c1583dd · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Some modalities are more equal than others: Decoding and architecting multi- modal integration in mllms.arXiv preprint arXiv:2511.22826
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b76aaba4-84db-42bb-824c-ca06d81be28b · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94f2382e-1661-4d72-a753-a3c1da1aa451 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb5e01b9-98bb-4aa9-9a7e-d71a8b752076 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Fork-merge decoding: Enhancing multimodal understanding in audio-visual large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 26ee8227-5eea-45f8-9223-3cd80b427a22 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5d3ea50-aaf8-4e5b-91ae-a3545a8b3d34 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs On the Audio Hallucinations in Large Audio-Video Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7bcbc8a-7830-4d1f-81f0-311bc2a3803e · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs video-SALMONN 2: Caption-enhanced audio-visual large language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65dc4318-994a-43c6-bbf6-46277cfa1d9a · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea141ae5-e609-48b4-a234-162fe52b04bb · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4d1b3571-0695-433d-b690-0d883765f688 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Qwen2.5-Omni Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da79c865-359a-4ff5-9541-9a7530e157e5 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs The structure is organized as follows: A
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae2f560c-04d6-4a1d-abc5-71ebd7422e4b · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Specifically, visual sink tokens in VLMs exhibit massive activation along the same dimensions as the BOS token in the base LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 860189c5-28c2-4f94-975f-f46bec40e12d · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs While the original VCD applies noise solely to the image modality, we extend this approach to the audio-visual domain
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2a7b017-2a83-4388-8cf8-5716597debcf · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61663980-3933-4c73-8b59-681b57d5a65a · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8004c53a-042b-4592-8c59-4bf149a724f9 · outbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs As shown in Tab
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.