{"id":"167c71c1-c51a-41c0-8567-0539767e7982","arxiv_id":"2505.05318","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review and pilot workshop that organizes VLM trust research into a new taxonomy and finds a shortage of direct user studies.","lead":"This survey maps how researchers study user trust in vision-language models and proposes a taxonomy built from cognitive abilities, collaboration modes, and agent behaviors. It combines a review of 43 papers with a pilot workshop that yields preliminary requirements for future user studies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central claim that trust in VLMs is built through human-like visual intelligence and collaboration is a stipulated framework, not a finding: Table 1 mostly codes capability benchmarks, and the only user data come from 8 participants with a confounded model comparison.","rationale":"The reader's verdict correctly praises the paper's transparent method, novel taxonomy, and honest treatment of the pilot workshop. My concern is not about the paper's integrity or its usefulness as a survey; it is about the epistemic status of the central claim. The sentence in Section 3.2 is presented as an assertion about how trust in VLMs is built, but the survey's own coverage data show that almost no reviewed work measures user trust directly, and the only user-grounded evidence is an n=8 pilot with a confounded model comparison. This means the central claim currently functions as a hypothesis or design brief, not as a conclusion supported by the assembled literature. The proposed re-coding of Table 1 would settle whether the issue is merely imprecise wording or a substantive gap between the taxonomy and the evidence. If the re-coding shows the taxonomy is mostly mapping trustworthiness rather than trust, the paper should condition acceptance on reframing the claim and adding an explicit limitation paragraph. I therefore recommend CONDITIONAL rather than UNCHANGED or REJECT: the paper is a valuable contribution, but its strongest claim needs to be aligned with the evidence it actually provides.","tokens_in":15148,"tokens_out":6088,"duration_ms":70260,"concrete_test":"Re-code all 43 papers in Table 1 using a strict 'user trust evidence' criterion: a paper may be assigned to Ability, Benevolence, or Integrity only if it measures or manipulates user-perceived trust (e.g., self-reported trust, delegation choice, trust rating) or explicitly reports a user study; benchmark-only papers are coded as trustworthiness-only. If a large majority of current Table 1 assignments are trustworthiness-only, the central claim should be reframed as a proposed research agenda rather than a supported finding.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim, stated in Section 3.2 ('trust in VLMs is built through the collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles'), is not supported by the evidence the survey itself presents. In Table 1, the Ability column is populated mostly by performance benchmarks and methods (e.g., SpatialVLM, Pelican, STAR, MERLIM, Visual CoT) that measure model accuracy or robustness, not user-perceived trust. Section 3.3 acknowledges that Benevolence coverage is minimal and identifies only two studies with direct user interaction (Fan et al.; Mehrotra et al.). The only user-grounded evidence is the pilot workshop with 8 participants recruited at one author's institution (Section 4.2), a limitation the authors explicitly acknowledge in Section 4.1 ('the scale of this workshop limits the generalizability of our findings'). More troubling than sample size alone is that Part 1 compares ChatGPT-4o (a frontier text-only model) with Video-LLaMa 7B (a much weaker open-source VLM); users' distrust of Video-LLaMa and the 'trust drops sharply after initial failures' requirement may be artifacts of the weak model chosen, not of VLM trust dynamics generally. The paper therefore asserts a causal/normative structure that the literature mapping and the pilot do not test: whether human-like cognitive abilities and collaborative modes are actual antecedents of user trust, or merely categories borrowed from cognitive science, remains unexamined.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys research on user trust in Vision Language Models (VLMs) and proposes a multi-disciplinary taxonomy that extends the ABI (Ability, Benevolence, Integrity) framework with cognitive-science capabilities, collaboration modes, and agent behaviors. The authors systematically screened the literature, retaining 43 papers that are mapped onto the taxonomy in Table 1. They also report a pilot workshop with 8 experts in design and development, in which participants compared a text-only LLM (ChatGPT-4o) against a small open-source VLM (Video-LLaMa 7B) on video-understanding tasks, and evaluated mock-ups of a web app for trust measurement. The findings are used to propose preliminary requirements for future user studies on VLM trust.","tokens_in":15431,"tokens_out":5087,"duration_ms":55635,"significance":"If the taxonomy and coverage analysis are taken as a mapping of the field, the paper makes a useful contribution: it concretely quantifies the scarcity of user-centered VLM trust research (only two papers with direct user interaction are identified) and provides a structured vocabulary for describing cognitive abilities, collaboration modes, and integrity factors. The workshop, despite its small scale, offers a transferable study design and generates design requirements that can inform larger investigations. The paper is transparent about its search protocol and limitations, and the coverage table is a valuable resource for researchers entering the area.","major_comments":[{"comment":"The workshop's central comparison is confounded: ChatGPT-4o is a frontier text-only LLM, while Video-LLaMa is a 7B-parameter, open-source VLM. Differences in accuracy and user-perceived trust between the two systems may therefore be driven by model scale, training data, or capability, rather than by the presence or absence of visual input. The interpretation in Section 5, however, generalizes the observation that 'trust can drop sharply after initial failures' into a requirement for VLM trust studies. As reported, this drop was observed only for Video-LLaMa, so it may be an artifact of the specific weak model rather than a property of VLM trust dynamics broadly. Please acknowledge this confound explicitly and either temper the generalizability of the workshop-derived requirements or analyze the data in a way that separates capability effects from modality effects.","section":"Section 4.3 and Section 5"},{"comment":"The sentence 'we argue that trust in VLMs is built through the collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles' is presented as a substantive claim, but it is not an empirical conclusion of the survey. The coverage analysis in Table 1 shows that most reviewed works evaluate model performance or integrity violations, not user-perceived trust, and the paper itself notes that Benevolence coverage is minimal. To avoid overclaiming, please frame this statement as a proposed hypothesis or an organizing lens for the taxonomy, and clarify that the survey maps the existing literature onto this lens rather than providing evidence for it.","section":"Section 3.2"}],"minor_comments":[{"comment":"The sentence 'Fewer than 28% of the retrieved papers on trust and TAI keywords focus on Computer Vision and Vision Language Models' is ambiguous: the denominator is unclear (of the 157 candidates or of the 43 retained papers?). Please specify the source of this statistic.","section":"Section 3.3"},{"comment":"The taxonomy heading 'Legibility↔Security' uses a bidirectional arrow that is not explained in the text; the paper discusses legibility and security as counterparts, but the notation should be defined or replaced with a clearer label such as 'Legibility/Security'.","section":"Table 1 and Section 3.2"},{"comment":"References [36] and [37] appear to refer to the same work (Liu et al., 'Safety of Multimodal Large Language Models on Images and Text'), with one listing the IJCAI publication and the other an arXiv preprint. Please merge or clarify whether these are distinct papers.","section":"References"},{"comment":"The description of the preparatory seminar states that participants were informed about 'the diverse stakeholder groups to consider in future research,' but the paper does not report whether this framing influenced the workshop results; a brief note on any observed impact would strengthen the method.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable survey with a useful taxonomy and transparent methods, but the workshop's confounded model comparison and the status of the Section 3.2 claim need to be addressed before publication. If the authors can successfully temper the claims and acknowledge the confound, the paper could become acceptable; however, as written, these issues are load-bearing for the workshop-derived contributions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The taxonomy is the real contribution. Combining ABI with situated cognition, collaboration modes, and agent behavior types to organize VLM trust research is new, and the coverage table gives the field a useful map. The authors are transparent about their search protocol and inclusion criteria, and they're honest that the literature is thin on direct user studies. The workshop is clearly labeled preliminary, with informed consent and an information sheet, which is more than many such papers do.\n\nThe soft spot is the sentence in Section 3.2 suggesting trust in VLMs is built through collaboration with agents that exhibit human-like visual intelligence and comply with cooperation/integrity principles. That's a stipulative framing, not a finding the paper's evidence supports. Table 1 mostly populates Ability with capability benchmarks (SpatialVLM, Pelican, STAR, etc.) rather than user-perceived trust. Benevolence has almost no direct user studies, and the only user data come from 8 participants at one institution, comparing ChatGPT-4o against Video-LLaMa 7B. That comparison confounds model class with capability, so claims like 'trust drops sharply after initial failures' may be an artifact of the weak VLM rather than general trust dynamics. The authors explicitly acknowledge the scale limitation, but they don't acknowledge the model-comparison confound.\n\nThat said, the paper's central value is as an organizing framework, not as an empirical test. Read that way, it holds up. The gap analysis — few user studies, model-in-the-loop dominance, scene graphs as a promising hybrid representation — is grounded in the coverage table and is genuinely useful. For someone planning a user-VLM trust study, the taxonomy and the preliminary requirements are a workable starting point.\n\nI'd send this to peer review, not desk reject. It's a serviceable survey with a novel synthesis and an honest limitations section. I'd ask for revision that either softens the Section 3.2 claim to 'we hypothesize' or adds a paragraph that clearly separates the taxonomy's normative rationale from observed evidence. The workshop sections could also note that the ChatGPT-4o vs Video-LLaMa comparison is not a fair test of VLM trust. Minor stuff otherwise: the coverage table could indicate whether each entry is a user study or a benchmark, which would sharpen the 'gap' story.\n\nWho's the audience? Researchers entering the human-VLM trust area, or those designing user studies. The paper won't change practice, but it will save people time.","headline":"A useful taxonomy for a thin literature, but the central trust claim is stipulated rather than evidenced; worth reviewing with revisions.","tokens_in":15968,"tokens_out":2457,"would_cite":true,"duration_ms":26347,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that trust in vision-language models is a relational property built through collaboration, and it offers a taxonomy to map the field.","keywords":["vision language models","user trust","trustworthy AI","human-AI collaboration","situated cognition","theory of mind","multimodal reasoning","user study"],"falsifier":"Conduct a controlled study with a diverse participant pool comparing a text-only LLM, a visually grounded VLM, and a fluent-but-hallucinating VLM on the same video tasks, tracking trust scores and delegation choices over repeated rounds; if the fluent-but-hallucinating VLM earns as much trust as the grounded one, the claim that human-like visual intelligence drives VLM trust fails.","tokens_in":14987,"feed_emoji":"🤖","tokens_out":7033,"duration_ms":67056,"temperature":0.7,"pith_summary":"The paper tries to establish that user trust in vision-language models (VLMs) is not a static score but a relational property that develops through repeated interaction between a user and an agent perceived as visually intelligent and cooperative. To organize the field, it extends the classic ability-benevolence-integrity model of trust with cognitive-science notions of visual intelligence, collaborative modes of human-AI thought, and agent-behaviour types. A review of 43 papers finds that current research concentrates on integrity violations, such as hallucinations, adversarial attacks, and bias, while studies of perceived ability and direct user collaboration are sparse. Results from a pilot workshop with eight prospective users yield preliminary requirements for future trust studies, including user agency, multi-turn interaction, contextualized trust metrics, and graph-based feedback.","feed_headline":"Vision-language trust work is lopsided: many attacks, few users","feed_subtitle":"A taxonomy of ability, benevolence, and integrity maps the field; a pilot workshop supplies user requirements.","key_machinery":"The central object is the taxonomy shown in Figure 1, an extension of the organizational-trust ABI model into VLM-specific categories. It does the argument's main work by giving each of the three trust factors a concrete VLM interpretation: cognitive-science capabilities for Ability, collaborative thought modes for Benevolence, and agent behaviours plus fairness for Integrity. The taxonomy is used to classify 43 papers and to expose gaps, and it frames the workshop's design: users delegated tasks, compared a text-only LLM with a VLM, and rated mock-up features for a trust-evaluation app.","core_discovery":"The paper's central claim is that trust in a VLM is built through collaboration between users and artificial agents that exhibit human-like visual intelligence and comply with cooperation and integrity principles. The authors ground this claim in a taxonomy whose Ability branch decomposes visual intelligence into model building, intuitive physics, intuitive psychology, causality, compositionality, and meta-learning; whose Benevolence branch covers collaborative planning, learning and sensemaking, deliberation, and creation; and whose Integrity branch covers explicability, predictability, legibility with security trade-offs, and fairness. The coverage study shows the field's centre of gravity lies in integrity research, with comparatively little work on the cognitive abilities and collaboration modes that the taxonomy says build trust. The workshop complements the review with user-derived requirements rather than with a test of the taxonomy.","pith_inferences":["Extension: If trust is genuinely relational, single-session benchmark comparisons will understate real-world trust; longitudinal interaction studies are a natural next step the paper points toward but does not run.","Extension: The taxonomy could be operationalized as a coding scheme for user-study protocols, with the testable prediction that users' delegation decisions cluster along the ability, benevolence, and integrity dimensions across applications.","Extension: The workshop's observation that a text-only LLM sounded more believable than a VLM, even when both erred, suggests language fluency may currently dominate users' trust judgments; if so, visual competence must be made legible to users before the visual-intelligence branch of the taxonomy can matter.","Extension: Graph-based annotations could be turned into a trust-measurement instrument, allowing researchers to see which specific relational claims users accept or reject rather than only whether they accept a whole answer."],"forward_implications":["Trust in VLMs should be measured as an evolving, interaction-dependent variable, not as a one-shot performance score.","Benchmarks for VLM trust should extend beyond hallucinations and attacks to cover intuitive physics, intuitive psychology, causality, compositionality, and meta-learning.","Researchers should run genuine user studies rather than relying mainly on model-in-the-loop preference alignment, because collaboration modes are currently the least-covered branch of the taxonomy.","Scene graphs are a promising hybrid modality for making VLM outputs legible and for collecting fine-grained user feedback on individual relational components.","Study designs should include ice-breaking rounds and continuous trust tracking, since trust can drop sharply after early failures."],"supporting_citations":[{"why":"Supplies the ability-benevolence-integrity model of trust that the paper extends to vision-language models.","marker":"[40]"},{"why":"Provides the core ingredients of human-like intelligence used to define the Ability branch of the taxonomy.","marker":"[29]"},{"why":"Provides the collaborative thought modes used to define the Benevolence branch of the taxonomy.","marker":"[10]"},{"why":"Provides agent-behaviour types (explicability, predictability, legibility) used to define the Integrity branch.","marker":"[56]"},{"why":"Serves as the reference user study for modelling trust as task delegation and supplies the TrustMeter feature used in the workshop.","marker":"[41]"},{"why":"Supplies the STAR benchmark with situated video reasoning tasks and scene graph annotations used in the workshop.","marker":"[60]"},{"why":"Supplies the TrustLLM framework that the proposed VLM taxonomy extends from language models to multimodal models.","marker":"[24]"},{"why":"Provides the background definition of VLMs and zero-shot inference that frames the problem scope.","marker":"[68]"}],"fun_headline_variants":["Trust in VLMs: Research lopsided toward integrity, away from users","VLM trust research: Attack-heavy, user-light, needs rebalancing","User trust in VLMs: More attacks studied than abilities or users","VLM trust is imbalanced: integrity heavy, abilities and users light","Survey finds VLM trust work misses user needs; attacks dominate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the eight participants recruited at the authors' own institution, all with design or software backgrounds and little VLM experience, are representative enough of prospective users to ground preliminary requirements for large-scale VLM trust studies; the paper itself acknowledges this scale limits generalizability.","fun_headline_variants_meta":{"raw":{"variants":["Trust in VLMs: Research lopsided toward integrity, away from users","VLM trust research: Attack-heavy, user-light, needs rebalancing","User trust in VLMs: More attacks studied than abilities or users","VLM trust is imbalanced: integrity heavy, abilities and users light","Survey finds VLM trust work misses user needs; attacks dominate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00074,"raw_usage":{"total_tokens":3213,"prompt_tokens":763,"completion_tokens":2450,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":379,"completion_tokens_details":{"reasoning_tokens":2355}},"tokens_in":379,"tokens_out":2450,"duration_ms":16569,"temperature":1.0,"reasoning_tokens":2355,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:06:09.518392+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a controlled study with a diverse participant pool comparing a text-only LLM, a visually grounded VLM, and a fluent-but-hallucinating VLM on the same video tasks, tracking trust scores and delegation choices over repeated rounds; if the fluent-but-hallucinating VLM earns as much trust as the grounded one, the claim that human-like visual intelligence drives VLM trust fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ability-benevolence-integrity model of trust that the paper extends to vision-language models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the core ingredients of human-like intelligence used to define the Ability branch of the taxonomy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the collaborative thought modes used to define the Benevolence branch of the taxonomy."},{"cited_title":"Mehrotra, C","cited_arxiv_id":null,"evidence_quote":"Serves as the reference user study for modelling trust as task delegation and supplies the TrustMeter feature used in the workshop."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the STAR benchmark with situated video reasoning tasks and scene graph annotations used in the workshop."},{"cited_title":"Zhang, J","cited_arxiv_id":null,"evidence_quote":"Provides the background definition of VLMs and zero-shot inference that frames the problem scope."}],"review_version":1}