Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects

T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This survey argues that trust in vision-language models is a relational property built through collaboration, and it offers a taxonomy to map the field.

desk verdict A useful taxonomy for a thin literature, but the central trust claim is stipulated rather than evidenced; worth reviewing with revisions. read the letter →

arxiv 2505.05318 v1 pith:6Y2H4RDA submitted 2025-05-08 cs.CV cs.AIcs.CYcs.HCcs.RO

classification cs.CVcs.AIcs.CYcs.HCcs.RO
keywords visionlanguagemodelsusertrusttrustworthyAIhuman-AIcollaborationsituatedcognitiontheoryofmindmultimodalreasoningstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that user trust in vision-language models (VLMs) is not a static score but a relational property that develops through repeated interaction between a user and an agent perceived as visually intelligent and cooperative. To organize the field, it extends the classic ability-benevolence-integrity model of trust with cognitive-science notions of visual intelligence, collaborative modes of human-AI thought, and agent-behaviour types. A review of 43 papers finds that current research concentrates on integrity violations, such as hallucinations, adversarial attacks, and bias, while studies of perceived ability and direct user collaboration are sparse. Results from a pilot workshop with eight prospective users yield preliminary requirements for future trust studies, including user agency, multi-turn interaction, contextualized trust metrics, and graph-based feedback.

What carries the argument

The central object is the taxonomy shown in Figure 1, an extension of the organizational-trust ABI model into VLM-specific categories. It does the argument's main work by giving each of the three trust factors a concrete VLM interpretation: cognitive-science capabilities for Ability, collaborative thought modes for Benevolence, and agent behaviours plus fairness for Integrity. The taxonomy is used to classify 43 papers and to expose gaps, and it frames the workshop's design: users delegated tasks, compared a text-only LLM with a VLM, and rated mock-up features for a trust-evaluation app.

What would settle it

Conduct a controlled study with a diverse participant pool comparing a text-only LLM, a visually grounded VLM, and a fluent-but-hallucinating VLM on the same video tasks, tracking trust scores and delegation choices over repeated rounds; if the fluent-but-hallucinating VLM earns as much trust as the grounded one, the claim that human-like visual intelligence drives VLM trust fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that trust in a VLM is built through collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles. The authors ground this claim in a taxonomy whose Ability branch decomposes visual intelligence into model building, intuitive physics, intuitive psychology, causality, compositionality, and meta-learning; whose Benevolence branch covers collaborative planning, learning and sensemaking, deliberation, and creation; and whose Integrity branch covers explicability, predictability, legibility with security trade-offs, and fairness. The coverage study shows the field's centre of gravity lies in integrity research, with comparatively little work on the cognitive abilities and collaboration modes that the taxonomy says build trust. The workshop complements the review with user-derived requirements rather than with a test of the taxonomy.

Load-bearing premise

The load-bearing premise is that the eight participants recruited at the authors' own institution, all with design or software backgrounds and little VLM experience, are representative enough of prospective users to ground preliminary requirements for large-scale VLM trust studies; the paper itself acknowledges this scale limits generalizability.

Editorial extensions

If this is right

  • Trust in VLMs should be measured as an evolving, interaction-dependent variable, not as a one-shot performance score.
  • Benchmarks for VLM trust should extend beyond hallucinations and attacks to cover intuitive physics, intuitive psychology, causality, compositionality, and meta-learning.
  • Researchers should run genuine user studies rather than relying mainly on model-in-the-loop preference alignment, because collaboration modes are currently the least-covered branch of the taxonomy.
  • Scene graphs are a promising hybrid modality for making VLM outputs legible and for collecting fine-grained user feedback on individual relational components.
  • Study designs should include ice-breaking rounds and continuous trust tracking, since trust can drop sharply after early failures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: If trust is genuinely relational, single-session benchmark comparisons will understate real-world trust; longitudinal interaction studies are a natural next step the paper points toward but does not run.
  • Extension: The taxonomy could be operationalized as a coding scheme for user-study protocols, with the testable prediction that users' delegation decisions cluster along the ability, benevolence, and integrity dimensions across applications.
  • Extension: The workshop's observation that a text-only LLM sounded more believable than a VLM, even when both erred, suggests language fluency may currently dominate users' trust judgments; if so, visual competence must be made legible to users before the visual-intelligence branch of the taxonomy can matter.
  • Extension: Graph-based annotations could be turned into a trust-measurement instrument, allowing researchers to see which specific relational claims users accept or reject rather than only whether they accept a whole answer.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper surveys research on user trust in Vision Language Models (VLMs) and proposes a multi-disciplinary taxonomy that extends the ABI (Ability, Benevolence, Integrity) framework with cognitive-science capabilities, collaboration modes, and agent behaviors. The authors systematically screened the literature, retaining 43 papers that are mapped onto the taxonomy in Table 1. They also report a pilot workshop with 8 experts in design and development, in which participants compared a text-only LLM (ChatGPT-4o) against a small open-source VLM (Video-LLaMa 7B) on video-understanding tasks, and evaluated mock-ups of a web app for trust measurement. The findings are used to propose preliminary requirements for future user studies on VLM trust.

Significance. If the taxonomy and coverage analysis are taken as a mapping of the field, the paper makes a useful contribution: it concretely quantifies the scarcity of user-centered VLM trust research (only two papers with direct user interaction are identified) and provides a structured vocabulary for describing cognitive abilities, collaboration modes, and integrity factors. The workshop, despite its small scale, offers a transferable study design and generates design requirements that can inform larger investigations. The paper is transparent about its search protocol and limitations, and the coverage table is a valuable resource for researchers entering the area.

major comments (2)
  1. [Section 4.3 and Section 5] The workshop's central comparison is confounded: ChatGPT-4o is a frontier text-only LLM, while Video-LLaMa is a 7B-parameter, open-source VLM. Differences in accuracy and user-perceived trust between the two systems may therefore be driven by model scale, training data, or capability, rather than by the presence or absence of visual input. The interpretation in Section 5, however, generalizes the observation that 'trust can drop sharply after initial failures' into a requirement for VLM trust studies. As reported, this drop was observed only for Video-LLaMa, so it may be an artifact of the specific weak model rather than a property of VLM trust dynamics broadly. Please acknowledge this confound explicitly and either temper the generalizability of the workshop-derived requirements or analyze the data in a way that separates capability effects from modality effects.
  2. [Section 3.2] The sentence 'we argue that trust in VLMs is built through the collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles' is presented as a substantive claim, but it is not an empirical conclusion of the survey. The coverage analysis in Table 1 shows that most reviewed works evaluate model performance or integrity violations, not user-perceived trust, and the paper itself notes that Benevolence coverage is minimal. To avoid overclaiming, please frame this statement as a proposed hypothesis or an organizing lens for the taxonomy, and clarify that the survey maps the existing literature onto this lens rather than providing evidence for it.
minor comments (4)
  1. [Section 3.3] The sentence 'Fewer than 28% of the retrieved papers on trust and TAI keywords focus on Computer Vision and Vision Language Models' is ambiguous: the denominator is unclear (of the 157 candidates or of the 43 retained papers?). Please specify the source of this statistic.
  2. [Table 1 and Section 3.2] The taxonomy heading 'Legibility↔Security' uses a bidirectional arrow that is not explained in the text; the paper discusses legibility and security as counterparts, but the notation should be defined or replaced with a clearer label such as 'Legibility/Security'.
  3. [References] References [36] and [37] appear to refer to the same work (Liu et al., 'Safety of Multimodal Large Language Models on Images and Text'), with one listing the IJCAI publication and the other an arXiv preprint. Please merge or clarify whether these are distinct papers.
  4. [Section 4.1] The description of the preparatory seminar states that participants were informed about 'the diverse stakeholder groups to consider in future research,' but the paper does not report whether this framing influenced the workshop results; a brief note on any observed impact would strengthen the method.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the taxonomy is proposed and then applied, and the pilot findings are explicitly preliminary; no claim reduces to its own input by construction.

full rationale

This is a survey paper, not a derivation with equations or fitted parameters, and no step in its argument reduces to its own input by construction. The proposed taxonomy is introduced in Section 3.2 as an extension of Mayer's ABI framework, and Table 1 then classifies external papers and benchmarks under that taxonomy; while the categories are author-chosen, the resulting coverage claims (e.g., scarce Benevolence coverage) are based on counts of independently published works, not on the taxonomy alone. The central statement that trust in VLMs is built through collaboration with human-like visual intelligence is presented as an argument ('we argue that...'), not as a derived prediction, so it cannot be circular in the technical sense. The workshop in Section 4 is described by the authors as exploratory and limited in generalizability, and its outputs are labeled 'preliminary requirements' rather than validated conclusions. The skepticism about the small, institution-recruited sample and the confounded ChatGPT-4o versus Video-LLaMa comparison is a validity and generalizability concern about the evidence, not a circularity defect, because the paper does not claim those observations follow from the taxonomy. No self-citation chain, imported uniqueness theorem, or fitted-input-called-prediction pattern is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is a qualitative survey, so there are no fitted numbers or invented physical entities. The taxonomy imports several theoretical constructs from prior literature, and the validity of the coverage analysis depends on those imports being appropriate for VLM trust.

assumptions (4)
  • domain assumption The ABI framework (Ability, Benevolence, Integrity) is an appropriate model for trust in human-AI and human-VLM interaction.
    Invoked in Section 3.2 as the foundation for the taxonomy; originally designed for interpersonal trust in organizations (Mayer et al. [40]).
  • domain assumption The cognitive science capabilities from Lake et al. and Collins et al. (model building, intuitive physics, intuitive psychology, causality, compositionality, meta-learning) are the appropriate dimensions for assessing a VLM's perceived ability.
    Used in Section 3.2 to structure the Ability dimension; these categories come from human cognition research and may not map cleanly onto machine perception.
  • domain assumption The one-year publication window and selected keywords capture the relevant literature on user trust in VLMs.
    Section 3.1 restricts the search to works from the last year; the authors note this may exclude earlier studies, which could affect the coverage conclusion that user studies are scarce.
  • domain assumption Trust decisions can be operationalized as a user's choice to delegate a task to an AI model.
    Section 4.1 models trust after Mehrotra et al. [41] as delegation choices; this operationalization may not capture other trust dynamics such as reliance calibration over time.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects." pith.science (2026). https://pith.science/paper/6Y2H4RDA

@misc{pith2026250505318,
  author       = {Pith},
  title        = {Pith review of: Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6Y2H4RDA}},
  note         = {Machine review of arXiv:2505.05318}
}
read the original abstract

The rapid adoption of Vision Language Models (VLMs), pre-trained on large image-text and video-text datasets, calls for protecting and informing users about when to trust these systems. This survey reviews studies on trust dynamics in user-VLM interactions, through a multi-disciplinary taxonomy encompassing different cognitive science capabilities, collaboration modes, and agent behaviours. Literature insights and findings from a workshop with prospective VLM users inform preliminary requirements for future VLM trust studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Can Robots Teach Us About Trust and Reliance? An interdisciplinary dialogue between Social Sciences and Social Robotics

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A sociological reading of human-robot trust concludes that robots are relied on rather than trusted, and outlines a research agenda treating trust as practical engagement.

Reference graph

Works this paper leans on

72 extracted references · 51 canonical work pages · cited by 1 Pith paper

  1. [1]

    K. G. Barman et al. Beyond transparency and explainability: on the need for adequate and contextualized user guidelines for LLM use. Ethics and Information Technology, 26(3), 2024

  2. [2]

    Calegari, G

    R. Calegari, G. G. Castañé, et al. Assessing and enforcing fairness in the AI lifecycle. In IJCAI, 2023. doi: 10.24963/ijcai.2023/735. URL https://doi.org/10.24963/ijcai.2023/735

  3. [3]

    Chander, C

    B. Chander, C. John, et al. Toward Trustworthy Artificial Intelligence (TAI) in the context of explainability and robustness. ACM Computing Surveys, 2024. doi: 10.1145/3675392. URL https://doi.org/10.1145/ 3675392

  4. [4]

    B. Chen, Z. Xu, et al. SpatialVLM: Endowing vision-language models with spatial reasoning capabilities. In CVPR, 2024

  5. [5]

    D. Chen, R. Chen, et al. MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark. In ICML, 2024

  6. [6]

    Chen, D.-Z

    F.-L. Chen, D.-Z. Zhang, et al. VLP: A survey on Vision-Language Pre-training. Machine Intelligence Research, 20(1), 2023

  7. [7]

    Chen, Y .-C

    J.-J. Chen, Y .-C. Liao, et al. ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos. arXiv:2406.19392, 2024

  8. [8]

    Y . Chen, K. Sikka, et al. DRESS: Instructing large vision-language models to align and interact with humans via natural language feedback. In CVPR, 2024

Show all 72 references
  1. [9]

    Cheng, H

    A.-C. Cheng, H. Yin, et al. SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models. In NeurIPS, 2024

  2. [10]

    K. M. Collins, I. Sucholutsky, et al. Building machines that learn and think with people. Nature Human Behaviour, 8(10), 2024

  3. [11]

    Commission

    E. Commission. Ethics guidelines for TAI. https://shorturl.at/yDixx, 2019

  4. [12]

    Commission

    E. Commission. Artificial Intelligence Act. https: //artificialintelligenceact.eu/, 2024

  5. [13]

    A. Deng, Z. Chen, and B. Hooi. Seeing is Believing: Mitigating Hallu- cination in Large Vision-Language Models via CLIP-Guided Decoding,

  6. [14]

    M. A. M. Dona, B. Cabrero-Daniel, et al. Evaluating and Enhanc- ing Trustworthiness of LLMs in Perception Tasks. arXiv:2408.01433, 2024

  7. [15]

    X. Fan, Z. Wu, et al. ContextCam: Bridging Context Awareness with Creative Human-AI Image Co-Creation. In CHI, 2024. ISBN 9798400703300. doi: 10.1145/3613904.3642129. URL https://dl.acm. org/doi/10.1145/3613904.3642129

  8. [16]

    Y . Fang, Z. Yang, et al. From Uncertainty to Trust: Enhancing Reli- ability in Vision-Language Models with Uncertainty-Guided Dropout Decoding, 2024. arXiv:2412.06474 [cs]

  9. [17]

    Y . Gou, K. Chen, et al. Eyes closed, safety on: Protecting Multimodal LLMs via image-to-text transformation. In ECCV, 2024

  10. [18]

    Grunde-McLaughlin, R

    M. Grunde-McLaughlin, R. Krishna, and M. Agrawala. AGQA: A benchmark for compositional spatio-temporal reasoning. In CVPR, 2021

  11. [19]

    Y . Gui, Y . Jin, and Z. Ren. Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees, 2024. arXiv:2405.10301 [stat]

  12. [20]

    Howard, A

    P. Howard, A. Madasu, et al. SocialCounterfactuals: Probing and Mit- igating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples. In CVPR, 2024

  13. [21]

    T. Huai, S. Yang, et al. Debiased Visual Question Answering via the perspective of question types. Pattern Recognition Letters, 178, 2024. ISSN 0167-8655. doi: 10.1016/j.patrec.2024.01.009

  14. [22]

    Huang, X

    Q. Huang, X. Dong, et al. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation. In CVPR, 2024. ISBN 9798350353006. doi: 10.1109/CVPR52733.2024.01274. URL https://ieeexplore.ieee. org/document/10655465/

  15. [23]

    Huang, I

    S. Huang, I. Ponomarenko, et al. ManipVQA: Injecting Robotic Af- fordance and Physically Grounded Information into Multi-Modal Large Language Models. In IROS, 2024

  16. [24]

    Huang, L

    Y . Huang, L. Sun, et al. Trustllm: Trustworthiness in large language models. arXiv:2401.05561, 2024

  17. [25]

    C. M. Islam, S. Salman, et al. Malicious Path Manipulations via Ex- ploitation of Representation Vulnerabilities of Vision-Language Navi- gation Systems. In IROS, 2024

  18. [26]

    J. Jang, C. Kong, et al. Unifying vision-language representation space with single-tower transformer. In AAAI, 2023

  19. [27]

    Kambhampati

    S. Kambhampati. Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1), 2024

  20. [28]

    Khan and Y

    Z. Khan and Y . Fu. Consistency and Uncertainty: Identifying Unreli- able Responses From Black-Box Vision-Language Models for Selective Visual Question Answering. In CVPR, 2024

  21. [29]

    B. M. Lake, T. D. Ullman, et al. Building machines that learn and think like people. Behavioral and Brain Sciences, 40, 2017

  22. [30]

    T. Lee, H. Tu, et al. VHELM: A Holistic Evaluation of Vision Language Models. In NeurIPS, 2024

  23. [31]

    S. Leng, H. Zhang, et al. Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding. In CVPR, 2024

  24. [32]

    Li and G

    J. Li and G. Li. The triangular trade-off between robustness, accu- racy and fairness in deep neural networks: A survey. ACM Computing Surveys, 2024. doi: 10.1145/3645088. URL https://doi.org/10.1145/ 3645088

  25. [33]

    L. H. Li, P. Zhang, et al. Grounded language-image pre-training. In CVPR, 2022

  26. [34]

    B. Liu, G. Li, et al. The gap between Trustworthy AI Research and Trustworthy Software Research: A tertiary study.ACM Computing Sur- veys, 57(3), 2024. doi: 10.1145/3694964. URL https://doi.org/10.1145/ 3694964

  27. [35]

    S. Liu, K. Ying, et al. ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision- Language Models. In NeurIPS, 2024

  28. [36]

    X. Liu, Y . Zhu, et al. Safety of Multimodal Large Language Models on images and text. In IJCAI, 2024. doi: 10.24963/ijcai.2024/901. URL https://doi.org/10.24963/ijcai.2024/901

  29. [37]

    X. Liu, Y . Zhu, et al. Safety of Multimodal Large Language Models on Images and Text. arXiv:2402.00357, 2024

  30. [38]

    H. Luo, J. Bao, et al. SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation. In ICML, 2023

  31. [39]

    Y . Luo, M. Shi, et al. FairCLIP: Harnessing Fairness in Vision- Language Learning. In CVPR, 2024

  32. [40]

    R. C. Mayer, J. H. Davis, and F. D. Schoorman. An Integrative Model of Organizational Trust. Academy of Management Review, 1995

  33. [41]

    Mehrotra, C

    S. Mehrotra, C. C. Jorge, et al. Integrity-based explanations for fostering appropriate trust in AI agents. ACM TIIS, 14(1), 2024

  34. [42]

    Methnani, M

    L. Methnani, M. Chiou, et al. Who’s in charge here? a survey on Trust- worthy AI in variable autonomy robotic systems. ACM Computing Sur- veys, 56(7), 2024. doi: 10.1145/3645090. URL https://doi.org/10.1145/ 3645090

  35. [43]

    Nasiriany, F

    S. Nasiriany, F. Xia, , et al. PIVOT: Iterative visual prompting elicits actionable knowledge for VLMs. arXiv:2402.07872, 2024

  36. [44]

    Novelli, M

    C. Novelli, M. Taddeo, et al. Accountability in artificial intelligence: what it is and how it works. AI & Society, 39(4), 2024

  37. [45]

    Perez-Cerrolaza, J

    J. Perez-Cerrolaza, J. Abella, et al. Artificial intelligence for safety- critical systems in industrial and transportation domains: A survey. ACM Computing Surveys, 56(7), 2024

  38. [46]

    Prabhu, S

    V . Prabhu, S. Purushwalkam, et al. Trust but verify: Programmatic vlm evaluation in the wild. arXiv preprint arXiv:2410.13121, 2024

  39. [47]

    Radford, J

    A. Radford, J. W. Kim, et al. Learning transferable visual models from natural language supervision. In ICML, 2021

  40. [48]

    P. Sahu, K. Sikka, and A. Divakaran. Pelican: Correcting Hallucina- tion in Vision-LLMs via Claim Decomposition and Program of Thought Verification. In EMNLP, 2024

  41. [49]

    Sermanet, T

    P. Sermanet, T. Ding, et al. RoboVQA: Multimodal long-horizon rea- soning for robotics. In ICRA, 2024

  42. [50]

    H. Shao, S. Qian, et al. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of- Thought Reasoning. In NeurIPS, 2024. doi: 10.48550/arXiv.2403. 16999. URL http://arxiv.org/abs/2403.16999

  43. [51]

    Shirai, C

    K. Shirai, C. C. Beltran-Hernandez, et al. Vision-language interpreter for robot task planning. In ICRA, 2024

  44. [52]

    G. A. Sigurdsson, O. Russakovsky, and A. Gupta. What actions are needed for understanding human actions in videos? In Proceedings of the IEEE international conference on computer vision , pages 2137– 2146, 2017

  45. [53]

    Singh, R

    A. Singh, R. Hu, et al. Flava: A foundational language and vision align- ment model. In CVPR, 2022

  46. [54]

    Tocchetti, L

    A. Tocchetti, L. Corti, et al. AI robustness: a human-centered perspec- tive on technological challenges and opportunities. ACM Computing Surveys, 2024. doi: 10.1145/3665926. URL https://doi.org/10.1145/ 3665926

  47. [55]

    Vatsa, A

    M. Vatsa, A. Jain, and R. Singh. Adventures of Trustworthy Vision- Language Models: A Survey. In AAAI, 2024

  48. [56]

    Verma, S

    M. Verma, S. Bhambri, and S. Kambhampati. Theory of mind abili- ties of large language models in human-robot interaction: An illusion? In Companion of the HRI Conference , 2024. ISBN 9798400703232. doi: 10.1145/3610978.3640767. URL https://doi.org/10.1145/3610978. 3640767

  49. [57]

    Villa, J

    A. Villa, J. C. L. Alcázar, et al. Behind the magic, MERLIM multi-modal evaluation benchmark for large image-language models. arXiv:2312.02219, 2023

  50. [58]

    H. Wang, S. Tan, and H. Wang. Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models. In ICML, 2024

  51. [59]

    J. Wang, Y . Ming, et al. Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models. InNeurIPS. arXiv, 2024

  52. [60]

    B. Wu, S. Yu, et al. STAR: A Benchmark for Situated Reasoning in Real-World Videos. In NeurIPS, 2021

  53. [61]

    X. Wu, Y . Wang, et al. Evaluating fairness in large Vision- Language Models across diverse demographic attributes and prompts. arXiv:2406.17974, 2024

  54. [62]

    J. Xiao, A. Yao, et al. Can I trust your answer? Visually grounded Video Question Answering. In CVPR, 2024

  55. [63]

    Y . Xu, J. Yao, et al. Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models. In NeurIPS, 2024

  56. [64]

    Ye-Bin, N

    M. Ye-Bin, N. Hyeon-Woo, et al. BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-Language Models. In ECCV, 2025. ISBN 978-3-031-73247-8

  57. [65]

    T. Yu, H. Zhang, et al. RLAIF-V: Aligning mllms through open-source AI feedback for Super GPT-4V trustworthiness. arXiv:2405.17220, 2024

  58. [66]

    W. Yuan, J. Duan, et al. Robopoint: A vision-language model for spatial affordance prediction for robotics. In CoRL, 2024

  59. [67]

    Zanotti, M

    G. Zanotti, M. Petrolo, et al. Keep trusting! a plea for the notion of trustworthy AI. AI & Society, 2023

  60. [68]

    Zhang, J

    J. Zhang, J. Huang, et al. Vision-language models for vision tasks: A survey. IEEE TPAMI, 2024

  61. [69]

    Zhang, L

    Y . Zhang, L. Chen, et al. SPA-VL: A comprehensive safety preference alignment dataset for vision language model. arXiv:2406.12030, 2024

  62. [70]

    H. Zhao, Z. Cai, et al. Mmicl: empowering vision-language model with multi-modal in-context learning (2023). In ICLR, 2024

  63. [71]

    Zheng, J

    C. Zheng, J. Zhang, et al. Iterated Learning Improves Compositionality in Large Vision-Language Models. In CVPR, 2024. Supplementary material In the following sections, we provide additional details on the struc- ture of the workshop conducted with participants from a Design a...

  64. [2024]

    arXiv:2402.15300 [cs]

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.