REVIEW 2 major objections 4 minor 1 cited by
Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This survey argues that trust in vision-language models is a relational property built through collaboration, and it offers a taxonomy to map the field.
desk verdict A useful taxonomy for a thin literature, but the central trust claim is stipulated rather than evidenced; worth reviewing with revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the taxonomy shown in Figure 1, an extension of the organizational-trust ABI model into VLM-specific categories. It does the argument's main work by giving each of the three trust factors a concrete VLM interpretation: cognitive-science capabilities for Ability, collaborative thought modes for Benevolence, and agent behaviours plus fairness for Integrity. The taxonomy is used to classify 43 papers and to expose gaps, and it frames the workshop's design: users delegated tasks, compared a text-only LLM with a VLM, and rated mock-up features for a trust-evaluation app.
What would settle it
Conduct a controlled study with a diverse participant pool comparing a text-only LLM, a visually grounded VLM, and a fluent-but-hallucinating VLM on the same video tasks, tracking trust scores and delegation choices over repeated rounds; if the fluent-but-hallucinating VLM earns as much trust as the grounded one, the claim that human-like visual intelligence drives VLM trust fails.
Extended reading notes
Core claim
The paper's central claim is that trust in a VLM is built through collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles. The authors ground this claim in a taxonomy whose Ability branch decomposes visual intelligence into model building, intuitive physics, intuitive psychology, causality, compositionality, and meta-learning; whose Benevolence branch covers collaborative planning, learning and sensemaking, deliberation, and creation; and whose Integrity branch covers explicability, predictability, legibility with security trade-offs, and fairness. The coverage study shows the field's centre of gravity lies in integrity research, with comparatively little work on the cognitive abilities and collaboration modes that the taxonomy says build trust. The workshop complements the review with user-derived requirements rather than with a test of the taxonomy.
Load-bearing premise
The load-bearing premise is that the eight participants recruited at the authors' own institution, all with design or software backgrounds and little VLM experience, are representative enough of prospective users to ground preliminary requirements for large-scale VLM trust studies; the paper itself acknowledges this scale limits generalizability.
Editorial extensions
If this is right
- Trust in VLMs should be measured as an evolving, interaction-dependent variable, not as a one-shot performance score.
- Benchmarks for VLM trust should extend beyond hallucinations and attacks to cover intuitive physics, intuitive psychology, causality, compositionality, and meta-learning.
- Researchers should run genuine user studies rather than relying mainly on model-in-the-loop preference alignment, because collaboration modes are currently the least-covered branch of the taxonomy.
- Scene graphs are a promising hybrid modality for making VLM outputs legible and for collecting fine-grained user feedback on individual relational components.
- Study designs should include ice-breaking rounds and continuous trust tracking, since trust can drop sharply after early failures.
Reading between the lines
- Extension: If trust is genuinely relational, single-session benchmark comparisons will understate real-world trust; longitudinal interaction studies are a natural next step the paper points toward but does not run.
- Extension: The taxonomy could be operationalized as a coding scheme for user-study protocols, with the testable prediction that users' delegation decisions cluster along the ability, benevolence, and integrity dimensions across applications.
- Extension: The workshop's observation that a text-only LLM sounded more believable than a VLM, even when both erred, suggests language fluency may currently dominate users' trust judgments; if so, visual competence must be made legible to users before the visual-intelligence branch of the taxonomy can matter.
- Extension: Graph-based annotations could be turned into a trust-measurement instrument, allowing researchers to see which specific relational claims users accept or reject rather than only whether they accept a whole answer.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys research on user trust in Vision Language Models (VLMs) and proposes a multi-disciplinary taxonomy that extends the ABI (Ability, Benevolence, Integrity) framework with cognitive-science capabilities, collaboration modes, and agent behaviors. The authors systematically screened the literature, retaining 43 papers that are mapped onto the taxonomy in Table 1. They also report a pilot workshop with 8 experts in design and development, in which participants compared a text-only LLM (ChatGPT-4o) against a small open-source VLM (Video-LLaMa 7B) on video-understanding tasks, and evaluated mock-ups of a web app for trust measurement. The findings are used to propose preliminary requirements for future user studies on VLM trust.
Significance. If the taxonomy and coverage analysis are taken as a mapping of the field, the paper makes a useful contribution: it concretely quantifies the scarcity of user-centered VLM trust research (only two papers with direct user interaction are identified) and provides a structured vocabulary for describing cognitive abilities, collaboration modes, and integrity factors. The workshop, despite its small scale, offers a transferable study design and generates design requirements that can inform larger investigations. The paper is transparent about its search protocol and limitations, and the coverage table is a valuable resource for researchers entering the area.
major comments (2)
- [Section 4.3 and Section 5] The workshop's central comparison is confounded: ChatGPT-4o is a frontier text-only LLM, while Video-LLaMa is a 7B-parameter, open-source VLM. Differences in accuracy and user-perceived trust between the two systems may therefore be driven by model scale, training data, or capability, rather than by the presence or absence of visual input. The interpretation in Section 5, however, generalizes the observation that 'trust can drop sharply after initial failures' into a requirement for VLM trust studies. As reported, this drop was observed only for Video-LLaMa, so it may be an artifact of the specific weak model rather than a property of VLM trust dynamics broadly. Please acknowledge this confound explicitly and either temper the generalizability of the workshop-derived requirements or analyze the data in a way that separates capability effects from modality effects.
- [Section 3.2] The sentence 'we argue that trust in VLMs is built through the collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles' is presented as a substantive claim, but it is not an empirical conclusion of the survey. The coverage analysis in Table 1 shows that most reviewed works evaluate model performance or integrity violations, not user-perceived trust, and the paper itself notes that Benevolence coverage is minimal. To avoid overclaiming, please frame this statement as a proposed hypothesis or an organizing lens for the taxonomy, and clarify that the survey maps the existing literature onto this lens rather than providing evidence for it.
minor comments (4)
- [Section 3.3] The sentence 'Fewer than 28% of the retrieved papers on trust and TAI keywords focus on Computer Vision and Vision Language Models' is ambiguous: the denominator is unclear (of the 157 candidates or of the 43 retained papers?). Please specify the source of this statistic.
- [Table 1 and Section 3.2] The taxonomy heading 'Legibility↔Security' uses a bidirectional arrow that is not explained in the text; the paper discusses legibility and security as counterparts, but the notation should be defined or replaced with a clearer label such as 'Legibility/Security'.
- [References] References [36] and [37] appear to refer to the same work (Liu et al., 'Safety of Multimodal Large Language Models on Images and Text'), with one listing the IJCAI publication and the other an arXiv preprint. Please merge or clarify whether these are distinct papers.
- [Section 4.1] The description of the preparatory seminar states that participants were informed about 'the diverse stakeholder groups to consider in future research,' but the paper does not report whether this framing influenced the workshop results; a brief note on any observed impact would strengthen the method.
Circularity Check
No significant circularity: the taxonomy is proposed and then applied, and the pilot findings are explicitly preliminary; no claim reduces to its own input by construction.
full rationale
This is a survey paper, not a derivation with equations or fitted parameters, and no step in its argument reduces to its own input by construction. The proposed taxonomy is introduced in Section 3.2 as an extension of Mayer's ABI framework, and Table 1 then classifies external papers and benchmarks under that taxonomy; while the categories are author-chosen, the resulting coverage claims (e.g., scarce Benevolence coverage) are based on counts of independently published works, not on the taxonomy alone. The central statement that trust in VLMs is built through collaboration with human-like visual intelligence is presented as an argument ('we argue that...'), not as a derived prediction, so it cannot be circular in the technical sense. The workshop in Section 4 is described by the authors as exploratory and limited in generalizability, and its outputs are labeled 'preliminary requirements' rather than validated conclusions. The skepticism about the small, institution-recruited sample and the confounded ChatGPT-4o versus Video-LLaMa comparison is a validity and generalizability concern about the evidence, not a circularity defect, because the paper does not claim those observations follow from the taxonomy. No self-citation chain, imported uniqueness theorem, or fitted-input-called-prediction pattern is present.
Assumptions & free parameters
assumptions (4)
- domain assumption The ABI framework (Ability, Benevolence, Integrity) is an appropriate model for trust in human-AI and human-VLM interaction.
- domain assumption The cognitive science capabilities from Lake et al. and Collins et al. (model building, intuitive physics, intuitive psychology, causality, compositionality, meta-learning) are the appropriate dimensions for assessing a VLM's perceived ability.
- domain assumption The one-year publication window and selected keywords capture the relevant literature on user trust in VLMs.
- domain assumption Trust decisions can be operationalized as a user's choice to delegate a task to an AI model.
Cite this review
Pith. "Pith review of Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects." pith.science (2026). https://pith.science/paper/6Y2H4RDA
@misc{pith2026250505318,
author = {Pith},
title = {Pith review of: Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects},
year = {2026},
howpublished = {\url{https://pith.science/paper/6Y2H4RDA}},
note = {Machine review of arXiv:2505.05318}
}
read the original abstract
The rapid adoption of Vision Language Models (VLMs), pre-trained on large image-text and video-text datasets, calls for protecting and informing users about when to trust these systems. This survey reviews studies on trust dynamics in user-VLM interactions, through a multi-disciplinary taxonomy encompassing different cognitive science capabilities, collaboration modes, and agent behaviours. Literature insights and findings from a workshop with prospective VLM users inform preliminary requirements for future VLM trust studies.
Forward citations
Cited by 1 Pith paper
-
What Can Robots Teach Us About Trust and Reliance? An interdisciplinary dialogue between Social Sciences and Social Robotics
A sociological reading of human-robot trust concludes that robots are relied on rather than trusted, and outlines a research agenda treating trust as practical engagement.
Reference graph
Works this paper leans on
-
[1]
K. G. Barman et al. Beyond transparency and explainability: on the need for adequate and contextualized user guidelines for LLM use. Ethics and Information Technology, 26(3), 2024
work page 2024
-
[2]
R. Calegari, G. G. Castañé, et al. Assessing and enforcing fairness in the AI lifecycle. In IJCAI, 2023. doi: 10.24963/ijcai.2023/735. URL https://doi.org/10.24963/ijcai.2023/735
-
[3]
B. Chander, C. John, et al. Toward Trustworthy Artificial Intelligence (TAI) in the context of explainability and robustness. ACM Computing Surveys, 2024. doi: 10.1145/3675392. URL https://doi.org/10.1145/ 3675392
doi:10.1145/3675392 2024
-
[4]
B. Chen, Z. Xu, et al. SpatialVLM: Endowing vision-language models with spatial reasoning capabilities. In CVPR, 2024
work page 2024
-
[5]
D. Chen, R. Chen, et al. MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark. In ICML, 2024
work page 2024
-
[6]
F.-L. Chen, D.-Z. Zhang, et al. VLP: A survey on Vision-Language Pre-training. Machine Intelligence Research, 20(1), 2023
work page 2023
-
[7]
J.-J. Chen, Y .-C. Liao, et al. ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos. arXiv:2406.19392, 2024
arXiv 2024
-
[8]
Y . Chen, K. Sikka, et al. DRESS: Instructing large vision-language models to align and interact with humans via natural language feedback. In CVPR, 2024
work page 2024
Show all 72 references
-
[9]
Cheng, H
A.-C. Cheng, H. Yin, et al. SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models. In NeurIPS, 2024
2024
-
[10]
K. M. Collins, I. Sucholutsky, et al. Building machines that learn and think with people. Nature Human Behaviour, 8(10), 2024
2024
-
[11]
Commission
E. Commission. Ethics guidelines for TAI. https://shorturl.at/yDixx, 2019
2019
-
[12]
Commission
E. Commission. Artificial Intelligence Act. https: //artificialintelligenceact.eu/, 2024
2024
-
[13]
A. Deng, Z. Chen, and B. Hooi. Seeing is Believing: Mitigating Hallu- cination in Large Vision-Language Models via CLIP-Guided Decoding,
-
[14]
M. A. M. Dona, B. Cabrero-Daniel, et al. Evaluating and Enhanc- ing Trustworthiness of LLMs in Perception Tasks. arXiv:2408.01433, 2024
2024 arXiv
-
[15]
X. Fan, Z. Wu, et al. ContextCam: Bridging Context Awareness with Creative Human-AI Image Co-Creation. In CHI, 2024. ISBN 9798400703300. doi: 10.1145/3613904.3642129. URL https://dl.acm. org/doi/10.1145/3613904.3642129
2024
-
[16]
Y . Fang, Z. Yang, et al. From Uncertainty to Trust: Enhancing Reli- ability in Vision-Language Models with Uncertainty-Guided Dropout Decoding, 2024. arXiv:2412.06474 [cs]
2024
-
[17]
Y . Gou, K. Chen, et al. Eyes closed, safety on: Protecting Multimodal LLMs via image-to-text transformation. In ECCV, 2024
2024
-
[18]
Grunde-McLaughlin, R
M. Grunde-McLaughlin, R. Krishna, and M. Agrawala. AGQA: A benchmark for compositional spatio-temporal reasoning. In CVPR, 2021
2021
-
[19]
Y . Gui, Y . Jin, and Z. Ren. Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees, 2024. arXiv:2405.10301 [stat]
2024 arXiv
-
[20]
Howard, A
P. Howard, A. Madasu, et al. SocialCounterfactuals: Probing and Mit- igating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples. In CVPR, 2024
2024
-
[21]
T. Huai, S. Yang, et al. Debiased Visual Question Answering via the perspective of question types. Pattern Recognition Letters, 178, 2024. ISSN 0167-8655. doi: 10.1016/j.patrec.2024.01.009
2024 doi
-
[22]
Huang, X
Q. Huang, X. Dong, et al. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation. In CVPR, 2024. ISBN 9798350353006. doi: 10.1109/CVPR52733.2024.01274. URL https://ieeexplore.ieee. org/document/10655465/
2024
-
[23]
Huang, I
S. Huang, I. Ponomarenko, et al. ManipVQA: Injecting Robotic Af- fordance and Physically Grounded Information into Multi-Modal Large Language Models. In IROS, 2024
2024
-
[24]
Huang, L
Y . Huang, L. Sun, et al. Trustllm: Trustworthiness in large language models. arXiv:2401.05561, 2024
2024 arXiv
-
[25]
C. M. Islam, S. Salman, et al. Malicious Path Manipulations via Ex- ploitation of Representation Vulnerabilities of Vision-Language Navi- gation Systems. In IROS, 2024
2024
-
[26]
J. Jang, C. Kong, et al. Unifying vision-language representation space with single-tower transformer. In AAAI, 2023
2023
-
[27]
Kambhampati
S. Kambhampati. Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1), 2024
2024
-
[28]
Khan and Y
Z. Khan and Y . Fu. Consistency and Uncertainty: Identifying Unreli- able Responses From Black-Box Vision-Language Models for Selective Visual Question Answering. In CVPR, 2024
2024
-
[29]
B. M. Lake, T. D. Ullman, et al. Building machines that learn and think like people. Behavioral and Brain Sciences, 40, 2017
2017
-
[30]
T. Lee, H. Tu, et al. VHELM: A Holistic Evaluation of Vision Language Models. In NeurIPS, 2024
2024
-
[31]
S. Leng, H. Zhang, et al. Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding. In CVPR, 2024
2024
-
[32]
Li and G
J. Li and G. Li. The triangular trade-off between robustness, accu- racy and fairness in deep neural networks: A survey. ACM Computing Surveys, 2024. doi: 10.1145/3645088. URL https://doi.org/10.1145/ 3645088
2024 doi
-
[33]
L. H. Li, P. Zhang, et al. Grounded language-image pre-training. In CVPR, 2022
2022
-
[34]
B. Liu, G. Li, et al. The gap between Trustworthy AI Research and Trustworthy Software Research: A tertiary study.ACM Computing Sur- veys, 57(3), 2024. doi: 10.1145/3694964. URL https://doi.org/10.1145/ 3694964
2024 doi
-
[35]
S. Liu, K. Ying, et al. ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision- Language Models. In NeurIPS, 2024
2024
-
[36]
X. Liu, Y . Zhu, et al. Safety of Multimodal Large Language Models on images and text. In IJCAI, 2024. doi: 10.24963/ijcai.2024/901. URL https://doi.org/10.24963/ijcai.2024/901
2024 doi
-
[37]
X. Liu, Y . Zhu, et al. Safety of Multimodal Large Language Models on Images and Text. arXiv:2402.00357, 2024
2024 arXiv
-
[38]
H. Luo, J. Bao, et al. SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation. In ICML, 2023
2023
-
[39]
Y . Luo, M. Shi, et al. FairCLIP: Harnessing Fairness in Vision- Language Learning. In CVPR, 2024
2024
-
[40]
R. C. Mayer, J. H. Davis, and F. D. Schoorman. An Integrative Model of Organizational Trust. Academy of Management Review, 1995
1995
-
[41]
Mehrotra, C
S. Mehrotra, C. C. Jorge, et al. Integrity-based explanations for fostering appropriate trust in AI agents. ACM TIIS, 14(1), 2024
2024
-
[42]
Methnani, M
L. Methnani, M. Chiou, et al. Who’s in charge here? a survey on Trust- worthy AI in variable autonomy robotic systems. ACM Computing Sur- veys, 56(7), 2024. doi: 10.1145/3645090. URL https://doi.org/10.1145/ 3645090
2024 doi
-
[43]
Nasiriany, F
S. Nasiriany, F. Xia, , et al. PIVOT: Iterative visual prompting elicits actionable knowledge for VLMs. arXiv:2402.07872, 2024
2024 arXiv
-
[44]
Novelli, M
C. Novelli, M. Taddeo, et al. Accountability in artificial intelligence: what it is and how it works. AI & Society, 39(4), 2024
2024
-
[45]
Perez-Cerrolaza, J
J. Perez-Cerrolaza, J. Abella, et al. Artificial intelligence for safety- critical systems in industrial and transportation domains: A survey. ACM Computing Surveys, 56(7), 2024
2024
-
[46]
Prabhu, S
V . Prabhu, S. Purushwalkam, et al. Trust but verify: Programmatic vlm evaluation in the wild. arXiv preprint arXiv:2410.13121, 2024
2024 arXiv
-
[47]
Radford, J
A. Radford, J. W. Kim, et al. Learning transferable visual models from natural language supervision. In ICML, 2021
2021
-
[48]
P. Sahu, K. Sikka, and A. Divakaran. Pelican: Correcting Hallucina- tion in Vision-LLMs via Claim Decomposition and Program of Thought Verification. In EMNLP, 2024
2024
-
[49]
Sermanet, T
P. Sermanet, T. Ding, et al. RoboVQA: Multimodal long-horizon rea- soning for robotics. In ICRA, 2024
2024
- [50]
-
[51]
Shirai, C
K. Shirai, C. C. Beltran-Hernandez, et al. Vision-language interpreter for robot task planning. In ICRA, 2024
2024
-
[52]
G. A. Sigurdsson, O. Russakovsky, and A. Gupta. What actions are needed for understanding human actions in videos? In Proceedings of the IEEE international conference on computer vision , pages 2137– 2146, 2017
2017
-
[53]
Singh, R
A. Singh, R. Hu, et al. Flava: A foundational language and vision align- ment model. In CVPR, 2022
2022
-
[54]
Tocchetti, L
A. Tocchetti, L. Corti, et al. AI robustness: a human-centered perspec- tive on technological challenges and opportunities. ACM Computing Surveys, 2024. doi: 10.1145/3665926. URL https://doi.org/10.1145/ 3665926
2024 doi
-
[55]
Vatsa, A
M. Vatsa, A. Jain, and R. Singh. Adventures of Trustworthy Vision- Language Models: A Survey. In AAAI, 2024
2024
-
[56]
Verma, S
M. Verma, S. Bhambri, and S. Kambhampati. Theory of mind abili- ties of large language models in human-robot interaction: An illusion? In Companion of the HRI Conference , 2024. ISBN 9798400703232. doi: 10.1145/3610978.3640767. URL https://doi.org/10.1145/3610978. 3640767
2024
-
[57]
Villa, J
A. Villa, J. C. L. Alcázar, et al. Behind the magic, MERLIM multi-modal evaluation benchmark for large image-language models. arXiv:2312.02219, 2023
2023 arXiv
-
[58]
H. Wang, S. Tan, and H. Wang. Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models. In ICML, 2024
2024
-
[59]
J. Wang, Y . Ming, et al. Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models. InNeurIPS. arXiv, 2024
2024
-
[60]
B. Wu, S. Yu, et al. STAR: A Benchmark for Situated Reasoning in Real-World Videos. In NeurIPS, 2021
2021
-
[61]
X. Wu, Y . Wang, et al. Evaluating fairness in large Vision- Language Models across diverse demographic attributes and prompts. arXiv:2406.17974, 2024
2024
-
[62]
J. Xiao, A. Yao, et al. Can I trust your answer? Visually grounded Video Question Answering. In CVPR, 2024
2024
-
[63]
Y . Xu, J. Yao, et al. Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models. In NeurIPS, 2024
2024
-
[64]
Ye-Bin, N
M. Ye-Bin, N. Hyeon-Woo, et al. BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-Language Models. In ECCV, 2025. ISBN 978-3-031-73247-8
2025
-
[65]
T. Yu, H. Zhang, et al. RLAIF-V: Aligning mllms through open-source AI feedback for Super GPT-4V trustworthiness. arXiv:2405.17220, 2024
2024
-
[66]
W. Yuan, J. Duan, et al. Robopoint: A vision-language model for spatial affordance prediction for robotics. In CoRL, 2024
2024
-
[67]
Zanotti, M
G. Zanotti, M. Petrolo, et al. Keep trusting! a plea for the notion of trustworthy AI. AI & Society, 2023
2023
-
[68]
Zhang, J
J. Zhang, J. Huang, et al. Vision-language models for vision tasks: A survey. IEEE TPAMI, 2024
2024
-
[69]
Zhang, L
Y . Zhang, L. Chen, et al. SPA-VL: A comprehensive safety preference alignment dataset for vision language model. arXiv:2406.12030, 2024
2024 arXiv
-
[70]
H. Zhao, Z. Cai, et al. Mmicl: empowering vision-language model with multi-modal in-context learning (2023). In ICLR, 2024
2023
-
[71]
Zheng, J
C. Zheng, J. Zhang, et al. Iterated Learning Improves Compositionality in Large Vision-Language Models. In CVPR, 2024. Supplementary material In the following sections, we provide additional details on the struc- ture of the workshop conducted with participants from a Design a...
2024
-
[2024]
arXiv:2402.15300 [cs]
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.