Pith. sign in

Paper Citation Record · LEDGER

Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.09093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.09093 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:42:55.863310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.858275Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 823baf9b-99e4-45d8-b4a8-ebcf608cccb8 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.232881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:3df05805faa208a04a9207fae89e029ac90d8ca8b5442386cbc00c825913705f

Observation 7ec90a18-15d0-4e14-9d2d-7b78253eb0e9 · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:29:46.301213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:f4b81a4ecedf8742fcc10ebf01e43395f69a093817975a4bdc0d64af483aab80

Observation 0e5f935e-91f7-42dd-90d9-b091d4afbb92 · inbound

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models cites this paper.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.813519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:bdcdd494dbbebcf3454f934b45df13b26aa4a348dbebc746fecbf03732679110

Observation 7b24515a-5acf-4e75-a8b7-b3aa0fb5913f · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:08:01.403062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:bf9c0b6a5fefb7c51b60dfeae6208adf74dc72c9f09af22818d16b96f80fb391

Observation ad33c919-b2e0-490e-b34f-1e843972964e · inbound

A Survey on Knowledge Distillation of Large Language Models cites this paper.

A Survey on Knowledge Distillation of Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 276

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:31:11.548038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-17T23:31:11.213552Z digest=sha256:b5c1ced2e5c8b3de20428d06708024f17376a028e03e004541e452282507eab8

Observation 30c2f2c9-ce19-4796-86ea-c2a2cc9622a9 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.563909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:091544bbbb3f15db1aee6b6b5964656fced9793b6f1680771ac14c41fa89291a

Observation 28e4e6ba-cd82-4061-bf6e-e8e2816bf183 · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:14:45.433296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:f73e125a2b42fbb691b052c9c5b1dbc672e906531487e713086d8ed726c33a84

Observation 13e645a2-f7ef-4f9f-9aaa-37e1ece1f7f5 · inbound

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation cites this paper.

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:55.863310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:55.863310Z digest=sha256:880cd3d53f569aea39c5e3ec9f6821fd34084cbb9fc71095352c49c258a61f39

Observation a8ef66c2-02cb-486c-9647-a1d08d3bf058 · inbound

Aligning Pre-trained Models for Spoken Language Translation cites this paper.

Aligning Pre-trained Models for Spoken Language Translation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:24:15.821992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:24:15.821992Z digest=sha256:b7d0494683833bb1a839c4a910f3fa46e5c7b18f2a3e6d96e40d7c5bbc41db97

Observation 4887de70-e6f4-4902-baa1-d5453fe20896 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.156959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.156959Z digest=sha256:0b29686fa9c3065fb6baf3c461b5cc57156d1081d1b479d6dbda049ebb7ed0bc

Observation 429c8830-9c4d-48ff-8009-9d3edca87c9f · inbound

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos cites this paper.

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T05:53:35.521331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:53:35.521331Z digest=sha256:590d4719366119a0b8e6c6e8d74f0f849170aa473ab2211602b85872a5468114

Observation c420b547-a927-4ac1-bd3b-e5702a8bc775 · inbound

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases cites this paper.

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:37.060329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:37.060329Z digest=sha256:df921fba500a572c20a245050b110139a4042662a90e9ae063f35addb93dbb6a

Observation 919c5fa8-8e2c-4e6f-bb69-deb37b4425b7 · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.909252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.909252Z digest=sha256:73284dee2b0b57ad5197f4124ac5dfb8882a87548df76c4add0ced16f86e8792

Observation 5b249fd0-068e-4b75-bffc-eaebaeb45630 · inbound

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework cites this paper.

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:10:44.701645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:10:44.701645Z digest=sha256:c88435fec300cf3fc868e4de6b772e3eb4adee6d608823bc275845a531af6810

Observation 48293247-e7f6-4c2a-bdf6-243e10c9548f · inbound

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants cites this paper.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.963829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.963829Z digest=sha256:19d6fe9056e46446b3d9b5e1c53e42ab5d626c69e58a4d1ccbb316a5e2e2d1c2

Observation 16db6505-2dae-4e84-948d-f0a13e6764e3 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.600128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.600128Z digest=sha256:56c598cc375ccb55acf6b28bb58c7bbfa9e941025212e3a03b56d09e8956cdb5

Observation 96e9d748-0816-492a-909a-bb559c1a83cb · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.195569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.195569Z digest=sha256:c01cc940a2dba64ac54d23aa7e7b9b8a9d08192577eb1d45aa7e767e07fd67f7

Observation 3e6ff04f-c89d-4cb7-8c50-d03424393eec · inbound

LLaVA-SLT: Visual Language Tuning for Sign Language Translation cites this paper.

LLaVA-SLT: Visual Language Tuning for Sign Language Translation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:19.087751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:33:19.087751Z digest=sha256:92ee285d2ae36edac67dd040c0281eb942f73b3b8b11f10946b365f1a053e3de

Observation 063239ca-0fa4-48cb-92ed-e7111298b7ca · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 286

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.400857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.400857Z digest=sha256:aec53514f66e08fd6b13e7645a52770c1201a3a0a5a32a254985f52020cb2680

Observation 046aa89c-25ba-4ade-95a3-63c3f3156b1d · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.425889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.425889Z digest=sha256:98f26aab8dab0ace08e0706a0c56c4028bc64f6f4d279b0317e955903401d2aa

Observation b1d4f467-c679-43d6-8722-c23eee022818 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.167570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:62b1e716e6717c7b82008a74d4546584c2f76142beda6ae3be61c5715a6c9335

Observation d30c24d8-bc1b-46e8-9270-fae6bbfed8e2 · inbound

On Accelerating Edge AI: Optimizing Resource-Constrained Environments cites this paper.

On Accelerating Edge AI: Optimizing Resource-Constrained Environments Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:38.518573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:46:38.518573Z digest=sha256:faa2781b41fe67bed47c68d045182a4cf842933ce374a703ef1f2edd23afeeff

Observation 73ecf23a-b354-4599-b8f9-848ec14743a3 · inbound

Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure cites this paper.

Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T22:00:02.128524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:00:02.128524Z digest=sha256:72390022e52705badb925ace98bced51e0da2bddf6375080025ecd17dd2066a6

Observation aa34a99f-d3d2-4a41-9b78-3d9ae8641cc2 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.863381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:a9ab535c5e63bc749e4db6fee209a3ecad85c01669f315cddfba0f23207aef88

Observation ccf98728-886e-4412-91c8-124834cf70be · inbound

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling cites this paper.

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:22:51.818382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:22:51.818382Z digest=sha256:93500537740b86c37ca927e9da8e12d36dbde004d1dade47457d1735ec903ab3

Observation 32b933df-fe53-445c-b865-a4551977dff5 · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:13.241354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:13.241354Z digest=sha256:581e99ef426868f2d63c8d946e06952c4d827f6305c2085190c72ca7cdde6396

Observation 5f486cbc-df12-485f-8dc2-8c4e544ab8ae · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.900006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.900006Z digest=sha256:e2624e223b364b12b5bdb600d4f91435afabfd75caa9518e2af5685cefe02178

Observation e30ebd95-91dd-4632-90e8-d9560d30dc93 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.443705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.443705Z digest=sha256:f8071ac6cad835f396009049c5df0389e3ec121653f5049e4b3208da55aed748

Observation 87ae0daa-a108-4529-bfcc-022c9b0aa4aa · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.315938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.315938Z digest=sha256:7a6a46fa1f2c99343d860c7b1eaeda450c1a13442d815aa3ce5d05fcd4d20f8b

Observation 9a5d55e2-8499-4997-8b8f-7e832017e2e5 · inbound

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning cites this paper.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.101479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.101479Z digest=sha256:7e0340127550fa67c0af6444215516289a727f37b1ee3ddecbd990bf8584f166

Observation 305d6246-9c3d-4e12-9b2e-2d285b1f1c53 · inbound

Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild cites this paper.

Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:53:41.142189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:53:41.142189Z digest=sha256:ad90bc0d316b550c41e8699ba77a22f169db5b320450fce7a1ae39d30e23f9b3

Observation 41bc5a63-d4dc-4a4d-8689-e30e4d35f1c7 · inbound

Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition cites this paper.

Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:04.842954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:04.842954Z digest=sha256:fd4ee937b7e3a2d7474afd9719c54eb5ec865ebcd10e1d733d8589a1feb0f507

Observation 99f01e68-3340-4ab5-be81-98fc9dd23528 · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.493319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.493319Z digest=sha256:461b317926acd4cd922d8842029a5106d1a6dd4ec5ea78b8257cfb1ed0fb4f55

Observation 565bc1bb-94cb-4716-b693-4adbfc77deef · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.851160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:6091efc85e1bd0f7779a5102998652f8c73bc0a4201166aa1fdb34eb55645406

Observation 9ccff55d-a74e-44fa-bf18-f15ea7a9eb21 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.005963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:0d93f5d718b1765e83390b263482ba59e07f38423ae85ae47c2ac174e0c512bd

Observation 392c5b56-99bb-4e23-8187-94a85e5596c2 · inbound

Probing Cross-modal Information Hubs in Audio-Visual LLMs cites this paper.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:30.476568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:05:06.936268Z digest=sha256:82dac033674a359e8708bb9f5f7ac91d0115f8afda66bec888cbfd1f77f68518

Observation 26ee8227-5eea-45f8-9223-3cd80b427a22 · inbound

Probing Cross-modal Information Hubs in Audio-Visual LLMs cites this paper.

Probing Cross-modal Information Hubs in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.701574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:c7663a6361d1f3f9cf8af890fcdc21f29a62a7865ea2ffd805686ca01d3b080e

Observation 1dffdf43-1dd2-4240-b85f-02f3f24bb3b2 · inbound

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs cites this paper.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.307495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:143403965f40c866561ef06df39d62b7ada44e7fd7125cda18ac41fb5a7c4433

Observation 5c35675c-944a-4554-a864-3b07115dcaba · inbound

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification cites this paper.

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:37:35.647140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T15:05:31.609683Z digest=sha256:dd06bed6f005456e68b7cc7063fc71234ad8d5cfc8438f84f2603974b05084d0

Observation de0d9acc-9854-4560-b490-5baf8474ba19 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.859688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:a38620e5373c56673bc3fb994260839f304d4241ff936bc5e4535dbc6c7a7677