Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:53.659334Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 7 inbound Pith citation observations for arXiv:2501.02135.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:53.659334Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:38:50.942821Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T10:19:59.955529Z
91 of 91 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6fb6a6f-88c4-470f-8435-e8d282ca3e51 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eef782d7-dfad-4099-991f-2e8ae083af6c · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c782e0-29c5-47ba-9633-2f3bcc71b208 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 696463bf-7541-4d57-a15e-772e0c172831 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Activitynet: A large-scale video bench- mark for human activity understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee34447f-b730-4cc4-8a10-39b953b5709d · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba9c7d69-2a8e-4816-a42d-058aa3640e75 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db535761-dff7-4d9b-a561-24a44171cc38 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33ed591-b22c-4469-9549-ee38a1b3e572 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Vast: A vision-audio-subtitle- text omni-modality foundation model and dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba55ac18-6d33-4f14-9ec9-f4711097467c · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73136e5a-232a-4e01-8692-9f8e92d0acd7 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Gonzalez, Ion Stoica, and Eric P
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a867789-5bfd-4ee7-8af2-1477c6252955 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Meerkat: Audio-visual large language model for grounding in space and time
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b8de55-a451-4c92-8a1f-a99f4a6c7987 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Melfusion: Synthesizing music from image and language cues using diffusion models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe298d5d-8ff4-4b20-82ba-5cd21ae2e6e2 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Scaling instruction- finetuned language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79316ac7-a9e9-4463-99fe-56c004cf567f · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f6076a1-2faf-431a-b201-d432a3924b8c · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95933e1d-a0ed-443f-bc1a-b73fc832abda · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Enhancing Large Vision Language Models with Self-Training on Image Comprehension
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4955511e-2175-49f5-9d7f-1c2c63006fdd · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Learning models with uniform performance via distributionally robust opti- mization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e23ac5b-85f9-4324-8a4d-b005f23fb50d · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Clap learning audio concepts from natural language supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bda884a-bc2d-4d2e-b9a7-62a56a8955a9 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Audio set: An ontology and human-labeled dataset for audio events
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f097dd6-b2f9-4b0d-bcc6-e21e2217f6f1 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Imagebind: One embedding space to bind them all
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b301880-4662-45bf-84c9-aca3195652cb · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs OneLLM: One Framework to Align All Modalities with Language
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bd64c2-7d2f-4489-ae0e-845e50ea02ec · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs ImageBind-LLM: Multi-modality Instruction Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a01f661-e1a5-4c2f-8e79-b5dce552424e · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Measuring Massive Multitask Language Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc4c50c4-fab1-4928-b4bc-83d78ab63276 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Dogs’ responses to visual, auditory, and olfactory cat-related cues
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8c5b6549-bde3-440a-b5c3-d6245b1744fd · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Perceiver: General per- ception with iterative attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8bedab9d-12cb-46de-9443-e3cbc987627a · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Hallucination augmented contrastive learning for multimodal large language model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e5724d0-ce3c-4ccd-bdc1-1b3fafc8f382 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3800269-62b8-4734-bdc6-08e3a4b210ef · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TVQA: Localized, Compositional Video Question Answering
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5e264f-09b4-4cdb-a366-5754356e9394 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fface418-a7d3-416b-baac-bece666f0cb5 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510b7a99-66cb-4573-a058-ef159af51f31 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Learning to answer questions in dynamic audio-visual scenarios
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ff94adcc-faac-4e97-ac44-f7812a0b2232 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce8e1bb-2c42-4e4a-9e91-a9918787bddf · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs VideoChat: Chat-Centric Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118b31ef-01dd-4a1f-88c9-4979b156b259 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab059b0-6c42-4853-bc9d-45afcf1cf908 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Silkie: Preference Distillation for Large Visual Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd3b5391-fe48-4415-b425-0da42fa9e0c2 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Evaluating Object Hallucination in Large Vision-Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9ad60a-928b-4157-9640-b25d85eed6da · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b507ad7-ffb5-453f-a8c3-926016f592bd · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e046ecea-1748-4450-bf00-0df1522eb52c · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Visual instruction tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b80ead4-b1f7-4197-973a-fbee09b55bb6 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Statistical rejection sampling improves preference optimization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 08ff321d-1362-482d-a906-1a98b7441535 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MMBench: Is Your Multi-modal Model an All-around Player?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad0de72-75f8-4fb7-8187-ea511a871abb · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046aa89c-25ba-4ade-95a3-63c3f3156b1d · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc49c4d3-e56b-49f4-bfa0-ab679a64c709 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e065ec-0773-430e-9aaf-436c260e973c · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Audio-visual generalised zero-shot learning with cross-modal attention and language
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bbd32194-8fee-41c4-9768-d7b7d9e0d75e · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0125587d-7cd0-4ab3-b5fd-be82f8e2f3b6 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Hello gpt-4, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2d194b7-840d-40df-92f7-e2d327673f9d · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9ee2955f-56a1-4c70-a346-44c9a35a680e · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18544f24-9ed8-4965-aaed-95a425b655b2 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bfbd38bb-3863-48ba-88df-7add34b3c36e · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aac108e-20a1-4aef-8de7-f8fb82f6f160 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94391aa7-8cd6-4e1a-8420-2a87ba6f0dab · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a466abc-018e-4afb-92bc-f2177dfeed2f · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Learning transferable visual models from natural language supervi- sion
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f733651b-8885-443c-9ce9-b5f0755eb623 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Direct prefer- ence optimization: Your language model is secretly a reward model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 886b8bbe-96b6-47d5-8e41-7983243337a9 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980131d7-c4f1-4d64-a340-6fb74a1ff9e0 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d30c0df-0bb7-4e40-a331-a2302321580b · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b4cf94-e9a4-4f57-b2e6-246bf5423a20 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs PandaGPT: One Model To Instruction-Follow Them All
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d47bb43c-1a55-4c6e-918b-a3548116c8cc · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea8ed703-a1c9-4427-afc9-4c8dc61e29d4 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a424532-6c0f-49ce-8a8a-ce6a0c963f48 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Hashimoto
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 98a6c33f-dff1-4f17-8e65-698d94867b89 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Movieqa: Understanding stories in movies through question-answering
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ecb5f4e-ccb8-4540-a077-9e10b6718e3d · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Winoground: Probing vision and language models for visio- linguistic compositionality
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1f991b8a-3430-4636-be81-fb8e8441bbc6 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40bcfecb-2afb-4dc9-885a-45e2c6ee7e15 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377aaa41-1de1-4f2e-afde-2aef9fdc42b4 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs What Makes for Good Visual Tokenizers for Large Language Models?
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1124e824-98a0-4755-a420-9ab255b17931 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Vision transformer with deformable attention
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b02fd68d-d5a0-4853-885b-4294c3ef7a87 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32dc8df7-cfb2-4296-8d38-38d6549b07f6 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs FunQA: Towards Surprising Video Comprehension
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2070a7eb-04fa-4225-bd49-48c08d1078ec · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Msr-vtt: A large video description dataset for bridging video and language
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec4a0dd-a0e5-4d8b-8be0-95596efb32a7 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d33116-08fb-4d23-8649-ec5344939994 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Avqa: A dataset for audio- visual question answering on videos
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 81f42723-2cf8-4465-b564-255861154816 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f77772-6f7b-420a-a4a7-41991b877ee4 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263a5b8a-9492-4f50-bff1-b94acfbd0a5a · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs A Survey on Multimodal Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 421ebb5c-5ec3-47c5-a0c2-9a8af063be79 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 362cf51a-de64-4401-8e22-33bfc4190d96 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4959690d-1eda-41f4-863a-da5912badaf3 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53739abf-38b7-413f-a609-1a0e732d4b1a · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7c68c3-0f6b-4426-ab90-01903c07f457 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Florence: A New Foundation Model for Computer Vision
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01968fdd-300a-4632-b943-76b243146746 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895dd79a-d9a5-414b-be38-e227518eecfa · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b2b570-59f7-4f18-8830-ad356bb4dfd7 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd54acdd-b01e-4dec-8c59-1306d2deb408 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f7b866-3fc6-473f-a20c-92f3f3de9b62 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6eb1c1-318e-4319-8b65-225047ef36a2 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec54a6a8-e6f2-4f95-a7ee-61ad06ba1e09 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715d3507-7b1e-47e3-abae-e479a023bd11 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc4d8bb-8da8-471c-a2f4-3a13f8abeddb · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs If step 1 fails, we provide GPT-4 with the question, choices, and model prediction
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f83bd79-81e6-4e83-b9f6-6ca70bccb3c4 · outbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs None of the above
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6c576a3d-7bb2-4124-9f5b-6226267d6620 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ae3dda9-073d-4be0-b4ec-65738bd5558c · inbound
Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710a6990-5b09-451f-beab-49c6ecc256c6 · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d4c15d4-49e6-4aea-9529-69d0ef60c2b1 · inbound
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94943c4-2e0a-4714-975b-c1967b14ffec · inbound
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e822f70-1aa7-4fdd-9ad2-ececa9b5e6df · inbound
Do Audio-Visual Large Language Models Really See and Hear? AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 57e8c60e-a833-4a50-80ed-d173abf87baa · inbound
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.