Pith. sign in

Paper Citation Record · LEDGER

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

As of 19 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 7 inbound Pith citation observations for arXiv:2501.02135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02135 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:53.659334Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:38:50.942821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T10:19:59.955529Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved73
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6fb6a6f-88c4-470f-8435-e8d282ca3e51 · outbound

This paper cites GPT-4 Technical Report.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.210029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.210029Z digest=sha256:0106651b0713b4b1eb239696ed55e9c42d938788dfc399c430b575ece5eac192

Observation eef782d7-dfad-4099-991f-2e8ae083af6c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.215905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.215905Z digest=sha256:53af0d15ee8591e3d6c0def87f52b659fb363ae4db51be770289aa083c3d4d24

Observation 28c782e0-29c5-47ba-9633-2f3bcc71b208 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.220744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.220744Z digest=sha256:df10e4e39f20392ca3c9f70bb1fbfbedc68e0beae30360e3cec3c5470642b721

Observation 696463bf-7541-4d57-a15e-772e0c172831 · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Activitynet: A large-scale video bench- mark for human activity understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.226274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.226274Z digest=sha256:7ba2fcec868dc3dbfe83a90432d05e762665550176d36b44baf0d96a47b54ce7

Observation ee34447f-b730-4cc4-8a10-39b953b5709d · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.231406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.231406Z digest=sha256:c8b1e9bef91844a31a46857f2cbf7d7a33d7d4897128211c0d70d7cfd1d873ca

Observation ba9c7d69-2a8e-4816-a42d-058aa3640e75 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.235925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.235925Z digest=sha256:0105d5772051fc0f00ebce6bf4c7a3345c22c02d743297d5e0f2b5b231fa5159

Observation db535761-dff7-4d9b-a561-24a44171cc38 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.240950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.240950Z digest=sha256:d661fb9ca1029a5b59745f4dfee73519e53c2ddccfa61498b8a877816205976a

Observation f33ed591-b22c-4469-9549-ee38a1b3e572 · outbound

This paper cites Vast: A vision-audio-subtitle- text omni-modality foundation model and dataset.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Vast: A vision-audio-subtitle- text omni-modality foundation model and dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.246330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.246330Z digest=sha256:6ac09c4caeec2170971ea84a49af2404a4c54f027139ad507e260a91a0f19536

Observation ba55ac18-6d33-4f14-9ec9-f4711097467c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.251727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.251727Z digest=sha256:6c5a10839dfdb525d6da1c7fc525ba297fd07eee94246045d3072eeb7f1d6242

Observation 73136e5a-232a-4e01-8692-9f8e92d0acd7 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Gonzalez, Ion Stoica, and Eric P

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.257113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.257113Z digest=sha256:dbc0628edd5a2d970dccc20e3df64b01b8d402ae486d75486fc3ba8ff8fb74e1

Observation 9a867789-5bfd-4ee7-8af2-1477c6252955 · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Meerkat: Audio-visual large language model for grounding in space and time

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.262794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.262794Z digest=sha256:167351db5895c0bd9dab51cb8406b4fa44e5c78d5619e9c68085d5753cbcd44b

Observation 30b8de55-a451-4c92-8a1f-a99f4a6c7987 · outbound

This paper cites Melfusion: Synthesizing music from image and language cues using diffusion models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Melfusion: Synthesizing music from image and language cues using diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.267956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.267956Z digest=sha256:cf147ecefba6e7326eaa82e570b8229156e859cdeb7fd4473bcaf1d937412556

Observation fe298d5d-8ff4-4b20-82ba-5cd21ae2e6e2 · outbound

This paper cites Scaling instruction- finetuned language models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Scaling instruction- finetuned language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.273029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.273029Z digest=sha256:62f4c91ccc3c285b0b5b2af9bee8e2dd305a88c185f1f9cab0061bfb71515ad3

Observation 79316ac7-a9e9-4463-99fe-56c004cf567f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.278038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.278038Z digest=sha256:c35399ecb2caf7fd547847d2155d762b7bcfee316b96e9efa6af73183f8c76bd

Observation 4f6076a1-2faf-431a-b201-d432a3924b8c · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.282732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.282732Z digest=sha256:0114d8e62368ab4971e98dfaf62a7a040efee8f909e54f0594181a1ac01c41de

Observation 95933e1d-a0ed-443f-bc1a-b73fc832abda · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.287574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.287574Z digest=sha256:71ac4c2f34bf788358b43225eec1cf993c123a84b0e5e6bcfcbbd72679f99576

Observation 4955511e-2175-49f5-9d7f-1c2c63006fdd · outbound

This paper cites Learning models with uniform performance via distributionally robust opti- mization.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Learning models with uniform performance via distributionally robust opti- mization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.293340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.293340Z digest=sha256:88ca74aa9fbb67bc9eb1ca6a30086a22a7dcbb6e0cabf14eb72888c57f4f80cc

Observation 2e23ac5b-85f9-4324-8a4d-b005f23fb50d · outbound

This paper cites Clap learning audio concepts from natural language supervision.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Clap learning audio concepts from natural language supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.298094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.298094Z digest=sha256:0ee8288e2442ab2724460a46f617d10873fd69a06e9cab035f3656632550dc5f

Observation 6bda884a-bc2d-4d2e-b9a7-62a56a8955a9 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Audio set: An ontology and human-labeled dataset for audio events

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.303507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.303507Z digest=sha256:117bb480aafc9e3758d34d349ab5564af907ede976d27ed4096442f1ea2e0dff

Observation 2f097dd6-b2f9-4b0d-bcc6-e21e2217f6f1 · outbound

This paper cites Imagebind: One embedding space to bind them all.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Imagebind: One embedding space to bind them all

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.308688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.308688Z digest=sha256:155f166ceb2abe644972cb621a78462126f3ed940db6ed6018d047077c616bd2

Observation 4b301880-4662-45bf-84c9-aca3195652cb · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs OneLLM: One Framework to Align All Modalities with Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.313391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.313391Z digest=sha256:530e0744a2261483fbc5c0a21d3d87b3068a9d6025cd480b2b087c897ec9e055

Observation 81bd64c2-7d2f-4489-ae0e-845e50ea02ec · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs ImageBind-LLM: Multi-modality Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.318349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.318349Z digest=sha256:1a9e02b934a99f8799a850c27c6af6a0cb1c7967e8f4283963926a5258a07ca4

Observation 1a01f661-e1a5-4c2f-8e79-b5dce552424e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Measuring Massive Multitask Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.322799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.322799Z digest=sha256:bafe2daec1ac9a53049fc7df5690d5f784c1003e41bb689cde10d917c0244374

Observation cc4c50c4-fab1-4928-b4bc-83d78ab63276 · outbound

This paper cites Dogs’ responses to visual, auditory, and olfactory cat-related cues.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Dogs’ responses to visual, auditory, and olfactory cat-related cues

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.988767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.327027Z digest=sha256:5ffb73c10b6d7c1864bddd1269bc42b8a63e04c8f61429dffd7690695c878ee4

Observation 8c5b6549-bde3-440a-b5c3-d6245b1744fd · outbound

This paper cites Perceiver: General per- ception with iterative attention.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Perceiver: General per- ception with iterative attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.973010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.331590Z digest=sha256:3b6f1b621b4c73fdb7583d001987285feb501749df33103359cd3d33316245d7

Observation 8bedab9d-12cb-46de-9443-e3cbc987627a · outbound

This paper cites Hallucination augmented contrastive learning for multimodal large language model.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Hallucination augmented contrastive learning for multimodal large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.956478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.336326Z digest=sha256:cafd90904c8925548fae6043ddee6d9b3532ea2048d2e9e1052d60be9edc0d02

Observation 0e5724d0-ce3c-4ccd-bdc1-1b3fafc8f382 · outbound

This paper cites FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.341461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.341461Z digest=sha256:58bf4b5643ce28f5e349b08d8282d44b0c57b52dca0a6194451e4a5c0f7dcc82

Observation d3800269-62b8-4734-bdc6-08e3a4b210ef · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TVQA: Localized, Compositional Video Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.346403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.346403Z digest=sha256:b47fb0de723ac530896942d2c7deeca53fe64f084c8a494f551354185a9f496f

Observation 2a5e264f-09b4-4cdb-a366-5754356e9394 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.351626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.351626Z digest=sha256:e72ca4317b3ce1115c3223d064be003ecd4bf9fb7d0870ce65563c1bf4e40a63

Observation fface418-a7d3-416b-baac-bece666f0cb5 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.357487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.357487Z digest=sha256:96f81988abd5d2b80a73a5ea2a1c5313f40e6ce96ae97ab551cdb812ce8b63d4

Observation 510b7a99-66cb-4573-a058-ef159af51f31 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Learning to answer questions in dynamic audio-visual scenarios

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.938787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.362674Z digest=sha256:7054f9f5afd75fde219c0c8dc8dfc4cfd892c39ac6b0df075e518404d8c42fff

Observation ff94adcc-faac-4e97-ac44-f7812a0b2232 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.367395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.367395Z digest=sha256:eea84a8e500ff4e4ff15e6ec0532e8be6b22298e93f55d3ac4432474c167d88f

Observation 5ce8e1bb-2c42-4e4a-9e91-a9918787bddf · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs VideoChat: Chat-Centric Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.372394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.372394Z digest=sha256:94d41680426a28187be4d5c9b93758764332c3d2fb2168512e0a5f1b3eceb2aa

Observation 118b31ef-01dd-4a1f-88c9-4979b156b259 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.377380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.377380Z digest=sha256:5f0a605e945850c9db6451a645a7db6120c26ef36b458cbac7ce04c360303c7f

Observation eab059b0-6c42-4853-bc9d-45afcf1cf908 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.387458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.387458Z digest=sha256:4731feda275d5c71276e758a3680a4c9bc50e7c028cad38c6b9cef15f9fa9f83

Observation cd3b5391-fe48-4415-b425-0da42fa9e0c2 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.392404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.392404Z digest=sha256:8d40105f295c2132ca2a69609c9366c4c22015635649729e5aa1e9004d65cc22

Observation ae9ad60a-928b-4157-9640-b25d85eed6da · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.397488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.397488Z digest=sha256:e4d57dd71af1a44b63c9877a0475c99cf1b6c794147233e6a833d4870c5f1e69

Observation 8b507ad7-ffb5-453f-a8c3-926016f592bd · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.402079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.402079Z digest=sha256:1f549afa58bbcb77c5198808f72fb294679692a9f01f130793e7037ebd38db11

Observation e046ecea-1748-4450-bf00-0df1522eb52c · outbound

This paper cites Visual instruction tuning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Visual instruction tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.406709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.406709Z digest=sha256:3f3fe5118c06de752b28013df9a7a64f4976acd4b640909cd8cd76edf84590bb

Observation 7b80ead4-b1f7-4197-973a-fbee09b55bb6 · outbound

This paper cites Statistical rejection sampling improves preference optimization.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Statistical rejection sampling improves preference optimization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.911036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.410926Z digest=sha256:c5ecea7161d9d20925e3a61b695109c53f3d5a7abc9c2388733285a84531a512

Observation 08ff321d-1362-482d-a906-1a98b7441535 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.415185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.415185Z digest=sha256:4466fb049ed221fadbf563117527a91d0dad8058e2b6b99a78418591de72b215

Observation fad0de72-75f8-4fb7-8187-ea511a871abb · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.420232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.420232Z digest=sha256:87e8b90222a58f8bba3306c13b86379afc0603f5ce27b2be3b43f13c8c4b3d5d

Observation 046aa89c-25ba-4ade-95a3-63c3f3156b1d · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.425889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.425889Z digest=sha256:b1e9edff89cc145b5f7fc3dced6abb237f468657e5877180c42265cde21d25ff

Observation fc49c4d3-e56b-49f4-bfa0-ab679a64c709 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.430790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.430790Z digest=sha256:da1cea6f9cb19dc944c29ea4f1487d6241b8c701742ed865f989cf3e5376cada

Observation 37e065ec-0773-430e-9aaf-436c260e973c · outbound

This paper cites Audio-visual generalised zero-shot learning with cross-modal attention and language.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Audio-visual generalised zero-shot learning with cross-modal attention and language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.895205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.436166Z digest=sha256:6de64cdfb930f81481df33892016b4091a9b78d5e7bfcdb74404a3fd0d9dc852

Observation bbd32194-8fee-41c4-9768-d7b7d9e0d75e · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.440780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.440780Z digest=sha256:0c9e315a274db6372e1c318fbfc78457d16ddb274a149735bfc8ebab9b12dc50

Observation 0125587d-7cd0-4ab3-b5fd-be82f8e2f3b6 · outbound

This paper cites Hello gpt-4, 2024.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Hello gpt-4, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.878752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.446170Z digest=sha256:ea8692b3b9dfee384a29000833d3267ef72cb74e0ccf698f54340804a63b795c

Observation e2d194b7-840d-40df-92f7-e2d327673f9d · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.862430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.451392Z digest=sha256:b86290602faa4381e96cd174d88867bce70593f0b0ee2ec4c4a68dc622718c2f

Observation 9ee2955f-56a1-4c70-a346-44c9a35a680e · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.456596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.456596Z digest=sha256:5825cd6f500bb85ee427c1fa99e18c94df1a977dbad38c443d0c278312b9d6e6

Observation 18544f24-9ed8-4965-aaed-95a425b655b2 · outbound

This paper cites Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Coordinated joint multimodal embeddings for gen- eralized audio-visual zero-shot classification and retrieval of videos

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.846151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.462277Z digest=sha256:8f63579a1496d143c83ca4a741a78fad0c3dccd6aacf6f7e3f5ac976f13d8d06

Observation bfbd38bb-3863-48ba-88df-7add34b3c36e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.467822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.467822Z digest=sha256:4ea1ca07ae7bc7bd46a4034cdf304ae86744ad4249630d7314f0b628d8ebeb01

Observation 9aac108e-20a1-4aef-8de7-f8fb82f6f160 · outbound

This paper cites Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.473214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.473214Z digest=sha256:cee3eb15c2481a6dae4e95718e4d231b4de9084ec225b954e94efe095d6d6481

Observation 94391aa7-8cd6-4e1a-8420-2a87ba6f0dab · outbound

This paper cites Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.478325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.478325Z digest=sha256:b7fe7c4cdb9828214dc01a827822e7e252c0bd3922471846d972987f6d515128

Observation 9a466abc-018e-4afb-92bc-f2177dfeed2f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Learning transferable visual models from natural language supervi- sion

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.483262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.483262Z digest=sha256:f7f417da1045c89472ebd720c37b58b2069fc987468cb149482ab534b4f57acb

Observation f733651b-8885-443c-9ce9-b5f0755eb623 · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Direct prefer- ence optimization: Your language model is secretly a reward model

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.817681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.488113Z digest=sha256:cfa41250f6d64dabb3c0d28b2af7c9956661b3458a7f3f65bb33cce7867f0028

Observation 886b8bbe-96b6-47d5-8e41-7983243337a9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.492799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.492799Z digest=sha256:0cd2b0c104c6f5352d6f2f79002dfeb0856579cf4fb9f9f1e20385b390e3560a

Observation 980131d7-c4f1-4d64-a340-6fb74a1ff9e0 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.497615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.497615Z digest=sha256:0f1782d1c2b3f55e6d51736e8786256f653df146c7610ea7a388bb7ce80c6f18

Observation 9d30c0df-0bb7-4e40-a331-a2302321580b · outbound

This paper cites Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.502799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.502799Z digest=sha256:03283751f1725c461b61c6b797a1f061a206059e8490aec7724f01c34a877c78

Observation 38b4cf94-e9a4-4f57-b2e6-246bf5423a20 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs PandaGPT: One Model To Instruction-Follow Them All

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.507100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.507100Z digest=sha256:26676c3bbfafd8c06145f88fd01d191b671475cfbf043a5983cc86aaa2b9d42a

Observation d47bb43c-1a55-4c6e-918b-a3548116c8cc · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.511386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.511386Z digest=sha256:524730af42d18497d1b91d6cfb051c67c149cd1d5301e4240577b69c9da187a8

Observation ea8ed703-a1c9-4427-afc9-4c8dc61e29d4 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.516206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.516206Z digest=sha256:68e6d8cf25ff2fe3f083dca0e67cf183ecb99aa63c14964999257a1175f8565d

Observation 4a424532-6c0f-49ce-8a8a-ce6a0c963f48 · outbound

This paper cites Hashimoto.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Hashimoto

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.801000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.520651Z digest=sha256:2868e0bd6a2ab1f35f350d7bc148d0a369910e533f5d56d920194cb7877ed125

Observation 98a6c33f-dff1-4f17-8e65-698d94867b89 · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Movieqa: Understanding stories in movies through question-answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.785106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.525220Z digest=sha256:5797cf08841389da48d6ed0c19114900845a33cfedbc909bd90a1e3c2c35003f

Observation 5ecb5f4e-ccb8-4540-a077-9e10b6718e3d · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.768549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.530038Z digest=sha256:a6ce397f362051d64373993015d84b5bb25514e5b583cf53138de9732bac16d4

Observation 1f991b8a-3430-4636-be81-fb8e8441bbc6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.534781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.534781Z digest=sha256:1e10cb6063730726e157756ab15ceed6193ba0ad21a2994aa8a4cc2fd59d51df

Observation 40bcfecb-2afb-4dc9-885a-45e2c6ee7e15 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.539685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.539685Z digest=sha256:222dff5275d1932e0c4de00ba3fbc43804705ac67489c082804282c93a6d9b04

Observation 377aaa41-1de1-4f2e-afde-2aef9fdc42b4 · outbound

This paper cites What Makes for Good Visual Tokenizers for Large Language Models?.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs What Makes for Good Visual Tokenizers for Large Language Models?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.544321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.544321Z digest=sha256:7193b96310afee2c7962ee176aff789377fe2b3a0c2cd508fb3fb90953bcaf75

Observation 1124e824-98a0-4755-a420-9ab255b17931 · outbound

This paper cites Vision transformer with deformable attention.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Vision transformer with deformable attention

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.752232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.549483Z digest=sha256:1f0cae6bf198bbcc173b32b9cb76654171323ac05a71984ad9464e85e1c4b84d

Observation b02fd68d-d5a0-4853-885b-4294c3ef7a87 · outbound

This paper cites Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.554316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.554316Z digest=sha256:7df4ca4b21ae34c7934018fd826c3fe756fe8623680e6ff4cfa92919d792dcd5

Observation 32dc8df7-cfb2-4296-8d38-38d6549b07f6 · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs FunQA: Towards Surprising Video Comprehension

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.559199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.559199Z digest=sha256:fcccb5d3170726b994a1b1345fb993f83db2b37d57783d5c58c922ecf16ffb1a

Observation 2070a7eb-04fa-4225-bd49-48c08d1078ec · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Msr-vtt: A large video description dataset for bridging video and language

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.563821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.563821Z digest=sha256:ce20d3c3499391f076cdca93a5ece206a8381d84adf5441b4a39d4d6760f8a93

Observation bec4a0dd-a0e5-4d8b-8be0-95596efb32a7 · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.568209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.568209Z digest=sha256:13148e06801d0c1ef0022379e26c386aa7a699cab621131640e27fabd387bf08

Observation 84d33116-08fb-4d23-8649-ec5344939994 · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Avqa: A dataset for audio- visual question answering on videos

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.726278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.572969Z digest=sha256:bf45878bf9ed4553c2eaadf82a9726d6927152f2764e5ca5e529dd544e534a86

Observation 81f42723-2cf8-4465-b564-255861154816 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.577696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.577696Z digest=sha256:ad5582419fe19184e3e31e1b6fa86f93874beb5d47bfb46c1d090dd7f0b74627

Observation 85f77772-6f7b-420a-a4a7-41991b877ee4 · outbound

This paper cites CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.582625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.582625Z digest=sha256:6aee0567457bb7a6b0c33403dad78d00b8412edcb2d92817ca70db5142a6c84d

Observation 263a5b8a-9492-4f50-bff1-b94acfbd0a5a · outbound

This paper cites A Survey on Multimodal Large Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs A Survey on Multimodal Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.587409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.587409Z digest=sha256:53c7070aa738d0109331ddee616539ef6d6bb39fdf9f8622e7c9a634d25fd62f

Observation 421ebb5c-5ec3-47c5-a0c2-9a8af063be79 · outbound

This paper cites Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.711306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.592472Z digest=sha256:e4db4795657d85db38071876fa5d9e383bbf29b5643bdeb9343a0716fb9cb66f

Observation 362cf51a-de64-4401-8e22-33bfc4190d96 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.597338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.597338Z digest=sha256:4e22c17b9176cc461c95475d779c4b26469d759239025d9b4c3f694747346563

Observation 4959690d-1eda-41f4-863a-da5912badaf3 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.602129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.602129Z digest=sha256:f07a4c419b46ab9c2cb20cbf1c628dda1a963fbe7961830f5204887c62df8dde

Observation 53739abf-38b7-413f-a609-1a0e732d4b1a · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.606767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.606767Z digest=sha256:d616114fd96dad50322982e496eb8abfce00398a5c1242aa9c9a1747b726f80d

Observation 6b7c68c3-0f6b-4426-ab90-01903c07f457 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Florence: A New Foundation Model for Computer Vision

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.611479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.611479Z digest=sha256:5432f471af3bf75312a3df751a2e01fdf578fbf01e235f557e68efae67ad2683

Observation 01968fdd-300a-4632-b943-76b243146746 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.616669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.616669Z digest=sha256:caac83232297de8801006a1bc9be6efac9d2f7968d5a45c0ff0f33c35eab600f

Observation 895dd79a-d9a5-414b-be38-e227518eecfa · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.621353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.621353Z digest=sha256:836f17189802b7028570127b3761258cac015cd20ea567aebdaa4f6833244810

Observation 26b2b570-59f7-4f18-8830-ad356bb4dfd7 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.625806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.625806Z digest=sha256:976f7fbb894eae8b9cf8f7d3bc2cb8384e20bb31f8c55363d64a55b779b18114

Observation bd54acdd-b01e-4dec-8c59-1306d2deb408 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.630296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.630296Z digest=sha256:4fd1669f3a80bdb7387baaba1b1ab5847432364fdf4d5e9ecf09f426a5f73bc0

Observation 95f7b866-3fc6-473f-a20c-92f3f3de9b62 · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.634569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.634569Z digest=sha256:4878422dd4dc8c387784651ca5b9bf038faac50fb496edd87612df35ea02830d

Observation 8d6eb1c1-318e-4319-8b65-225047ef36a2 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.638650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.638650Z digest=sha256:b70ea9dd1882ade892c043898eb518997ba52caacd6ddee07b05ac8c6b9024a4

Observation ec54a6a8-e6f2-4f95-a7ee-61ad06ba1e09 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.643392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.643392Z digest=sha256:b8d7c1c9ec9d259da5bf79dd9366dc7bad36092030305d1542f7a4febe710ddb

Observation 715d3507-7b1e-47e3-abae-e479a023bd11 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.648345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.648345Z digest=sha256:41c0b349af9f6e5bbd27a92f523fa5064b3f5fea1cf229ab54330f5af843a77a

Observation 1cc4d8bb-8da8-471c-a2f4-3a13f8abeddb · outbound

This paper cites If step 1 fails, we provide GPT-4 with the question, choices, and model prediction.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs If step 1 fails, we provide GPT-4 with the question, choices, and model prediction

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:54.686405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.654185Z digest=sha256:07eb729ee163c2832dfe091f37e2bf823cbb2cce6edbe4048624ca7e56cd4eed

Observation 5f83bd79-81e6-4e83-b9f6-6ca70bccb3c4 · outbound

This paper cites None of the above.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs None of the above

Reference 92

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:54.669572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:19:53.659334Z digest=sha256:2061aa9a7583425fb30171689a4c803d0e7ca4ccaef09fa06f74681a7d321fa2

Pith citing papers

Observation 6c576a3d-7bb2-4124-9f5b-6226267d6620 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.285323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:5fb61073c29628c1677a5840649985ac6301b4caf88ea607b5c29d6c08598d7b

Observation 5ae3dda9-073d-4be0-b4ec-65738bd5558c · inbound

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation cites this paper.

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:38:50.942821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:38:50.942821Z digest=sha256:0aa5db398e54fc021cdc5c72b6f1d52960b4d70320d87c5d42c1a879a12d0cec

Observation 710a6990-5b09-451f-beab-49c6ecc256c6 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.557961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.557961Z digest=sha256:315ffd79dcb609c3de8d684d969512fec72db3583da5e12ec6bf641c7a0ce744

Observation 3d4c15d4-49e6-4aea-9529-69d0ef60c2b1 · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:05.113373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:05.113373Z digest=sha256:032f9d1c818e06659b59254b1f610cd807a88cdec1330ce2bb3935b35f788bdd

Observation f94943c4-2e0a-4714-975b-c1967b14ffec · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.979284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.979284Z digest=sha256:9f1c359dc52c2738300aa98ce3b2a8b3dddda6d62b65c95715a282c4d05ee48f

Observation 7e822f70-1aa7-4fdd-9ad2-ececa9b5e6df · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.817416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:4ef009b747e0fac28cbcc2862cfa5cfc4bda923d97b7f480090226bb9e0c368d

Observation 57e8c60e-a833-4a50-80ed-d173abf87baa · inbound

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments cites this paper.

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:19:59.958876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T10:16:42.095330Z digest=sha256:50ba82d3fd277d9bbc5ed77092a663f9abc543b4752c8ce0d3cf2ac556ae60a4