Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.09093.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:42:55.863310Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:27:36.858275Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 823baf9b-99e4-45d8-b4a8-ebcf608cccb8 · inbound
A Comprehensive Overview of Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ec90a18-15d0-4e14-9d2d-7b78253eb0e9 · inbound
SALMONN: Towards Generic Hearing Abilities for Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e5f935e-91f7-42dd-90d9-b091d4afbb92 · inbound
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7b24515a-5acf-4e75-a8b7-b3aa0fb5913f · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad33c919-b2e0-490e-b34f-1e843972964e · inbound
A Survey on Knowledge Distillation of Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 30c2f2c9-ce19-4796-86ea-c2a2cc9622a9 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 28e4e6ba-cd82-4061-bf6e-e8e2816bf183 · inbound
Qwen2-Audio Technical Report Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 13e645a2-f7ef-4f9f-9aaa-37e1ece1f7f5 · inbound
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ef66c2-02cb-486c-9647-a1d08d3bf058 · inbound
Aligning Pre-trained Models for Spoken Language Translation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4887de70-e6f4-4902-baa1-d5453fe20896 · inbound
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429c8830-9c4d-48ff-8009-9d3edca87c9f · inbound
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c420b547-a927-4ac1-bd3b-e5702a8bc775 · inbound
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919c5fa8-8e2c-4e6f-bb69-deb37b4425b7 · inbound
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b249fd0-068e-4b75-bffc-eaebaeb45630 · inbound
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48293247-e7f6-4c2a-bdf6-243e10c9548f · inbound
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16db6505-2dae-4e84-948d-f0a13e6764e3 · inbound
Do Language Models Understand Time? Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e9d748-0816-492a-909a-bb559c1a83cb · inbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e6ff04f-c89d-4cb7-8c50-d03424393eec · inbound
LLaVA-SLT: Visual Language Tuning for Sign Language Translation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 063239ca-0fa4-48cb-92ed-e7111298b7ca · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 286
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046aa89c-25ba-4ade-95a3-63c3f3156b1d · inbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d4f467-c679-43d6-8722-c23eee022818 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 159
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d30c24d8-bc1b-46e8-9270-fae6bbfed8e2 · inbound
On Accelerating Edge AI: Optimizing Resource-Constrained Environments Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ecf23a-b354-4599-b8f9-848ec14743a3 · inbound
Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa34a99f-d3d2-4a41-9b78-3d9ae8641cc2 · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ccf98728-886e-4412-91c8-124834cf70be · inbound
Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b933df-fe53-445c-b865-a4551977dff5 · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f486cbc-df12-485f-8dc2-8c4e544ab8ae · inbound
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30ebd95-91dd-4632-90e8-d9560d30dc93 · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ae0daa-a108-4529-bfcc-022c9b0aa4aa · inbound
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a5d55e2-8499-4997-8b8f-7e832017e2e5 · inbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305d6246-9c3d-4e12-9b2e-2d285b1f1c53 · inbound
Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41bc5a63-d4dc-4a4d-8689-e30e4d35f1c7 · inbound
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f01e68-3340-4ab5-be81-98fc9dd23528 · inbound
Sample-efficient Integration of New Modalities into Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565bc1bb-94cb-4716-b693-4adbfc77deef · inbound
Do Audio-Visual Large Language Models Really See and Hear? Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9ccff55d-a74e-44fa-bf18-f15ea7a9eb21 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 392c5b56-99bb-4e23-8187-94a85e5596c2 · inbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26ee8227-5eea-45f8-9223-3cd80b427a22 · inbound
Probing Cross-modal Information Hubs in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1dffdf43-1dd2-4240-b85f-02f3f24bb3b2 · inbound
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c35675c-944a-4554-a864-3b07115dcaba · inbound
RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation de0d9acc-9854-4560-b490-5baf8474ba19 · inbound
Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.