Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Vision Language Models as Values Detectors

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that current vision-language models are not aligned with human annotators in detecting the element of relevance in home images, with the best model, LLaVA 34B, reaching only 0.42 alignment, and argues that value-aware…

desk verdict A small, honest exploratory study whose quantitative claims are fragile due to noisy ground truth, but whose qualitative observations about VLM sensitivity to values are worth attention. read the letter →

arxiv 2501.03957 v1 pith:FCUW5LZ4 submitted 2025-01-07 cs.HC cs.CV

classification cs.HCcs.CV
keywords vision-languagemodelsvaluedetectionhuman-AIalignmentimagerelevanceLLaVAGPT-4ohomeenvironmentscenariosannotatoragreement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether vision-language models and human viewers agree on what deserves attention in an everyday scene, and it answers with a measured no. Using twelve generated home images and fourteen annotators, the authors define for each image the element of relevance that most annotators named, then score five models (GPT-4o and four LLaVA variants) against that same reference. The best model, LLaVA 34B, matches the human majority only 42% of the time, and the spread across models is not statistically significant. The authors argue that the failure is partly a descriptive bias in the models and that value-aware fine-tuning and more explicit prompts could turn VLMs into useful detectors of value-laden situations for social robotics and assistive technology.

What carries the argument

The evaluation rests on a small bespoke dataset: twelve diffusion-generated images of domestic situations, ten designed to carry value-relevant content and two controls. Fourteen annotators each named the element needing attention, and the most frequent answer (with random tie-breaking) became the reference element of relevance. The models were given the same request in the same prompt, their first sentence was kept, and each answer received a binary score against the reference. A qualitative response-type taxonomy — descriptive, value-aware, "none", or comment — is what carries the interpretation of the scores, distinguishing mere misalignment from promising value detection.

What would settle it

Compute inter-annotator agreement on the 168 responses, for example with Fleiss' kappa. If agreement is low across the twelve images, the single "element of relevance" per image is not a stable ground truth, and the observed scores from 0.14 to 0.42 would not establish human-model misalignment; conversely, high agreement would support the paper's interpretation of the scores.

Watch

Extended reading notes

Core claim

The central claim is that current vision-language models are not aligned with human annotators in identifying the element of relevance in images: the highest alignment score is 0.42, below the midpoint of the scale. The same data show a potential the paper treats as real: in 16% of disagreements the models gave a more value-aware answer than the human majority, noticing health risks or distress that annotators did not state, and in 29% of disagreements they said "none" where humans found something relevant. The paper's conclusion is that the bottleneck is training and prompting rather than the architecture: with value-laden fine-tuning data and prompts that ask the model to consider values, vision LLMs could become detectors of value-laden scenarios.

Load-bearing premise

The reference answer for each image is whatever the largest group of fourteen annotators happened to say, with ties broken randomly and no check on whether the annotators themselves agreed; if that reference is arbitrary, the alignment scores and model ranking inherit that arbitrariness.

Editorial extensions

If this is right

  • At current capability levels, off-the-shelf VLMs should not be trusted to focus attention on the element a human would consider relevant in a home environment.
  • LLaVA 34B's 0.42 score is the ceiling in this test, so any deployment relying on VLM attention in assistive or robotic systems needs a human-in-the-loop or a more targeted model.
  • Models answer descriptively when a value-aware answer is expected in about a third of disagreements, pointing to training-data bias toward description rather than interpretation.
  • Value-aware prompts and fine-tuning on value-laden images are concrete, untested-in-this-paper paths the authors predict would raise alignment.
  • The two control images produced a desirable outcome in 14% of misaligned answers, where models answered "none" while annotators described irrelevant objects.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the reference element is a single majority answer with random tie-breaking, the alignment scores should be read relative to annotator agreement; a model that agreed with a randomly chosen annotator would also score well below 1.
  • A natural next experiment is to score models against the full distribution of annotator answers rather than the majority, treating any annotator-named element as acceptable; this would separate genuine value detection from disagreement about which object is central.
  • The value-detection hypothesis is testable directly: ask models to label the value at stake (health, safety, family bonds) in each image and compare those labels to annotators' stated concerns, bypassing the ambiguity of a single "element".
  • The same protocol could be run with real photographs and a larger, culturally balanced annotator pool to see whether the 0.42 ceiling is an artifact of generated images or of the prompt.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a small empirical study of whether current vision-language models (VLMs) align with human annotators in identifying the 'element of relevance' in twelve AI-generated home-scene images. Fourteen annotators each named the relevant element for every image, and the most common response was taken as the reference. Five models (GPT-4o and four LLaVA variants) were then prompted with the same task, each producing three answers per image, and each answer was scored as correct/incorrect against the single majority reference. The paper reports alignment scores of 0.14 to 0.42, with LLaVA-34B highest, and a Cochran's Q test giving p = 0.077. The discussion analyzes the response types and argues that the models show potential for detecting value-laden elements despite the low alignment.

Significance. The question asked is timely: VLMs are increasingly deployed in social robotics and assistive technology, where knowing what to attend to matters as much as generating fluent text. The paper contributes a transparently described small benchmark, explicit model versions, and a qualitative error analysis showing that models sometimes produce 'more aware' answers than the majority of human annotators. These qualitative observations are genuinely interesting and could motivate follow-up work. The main quantitative claim, however, is not yet supported: the reference labels are unstable, the binary scoring is too coarse, and the model ranking is statistically indistinguishable. If the measurement issues are fixed, the paper would be a useful preliminary study; in its current form, the headline conclusion outruns the evidence.

major comments (3)
  1. [Section 2; Section 3.2] The single-reference evaluation does not measure model-human alignment reliably. The 'perceived element of relevance' in Table 1 is a majority label from 14 annotators, but no inter-annotator agreement is reported, and Section 2 states that ties are broken randomly. For the dog image, only 5/14 annotators chose the dog's health; the other nine answers were split among four other categories, so the reference is a plurality, not a consensus. Section 3.2 then awards a binary score to every model answer against this one label, which conflates model-human disagreement with label noise. The problem is visible in the controls: a model saying 'none' for the woman-eating dinner image is scored wrong even though the image was designed to contain no element of relevance, and Section 4.2 itself reports that 16% of misaligned answers were 'more aware' than the majority annotation. Please report annotator agreement (e.g., Fleiss' kappa) and evaluate with multiple accepted labels or a soft scoring rule.
  2. [Section 3.3; Abstract] The model ranking in Table 2 and the abstract's phrase 'LLaVA 34B showing the highest performance' are not statistically supported. The paper's own Cochran's Q test returns p = 0.077, so the observed differences among the five models are within sampling noise. The caveat in Section 3.3 is appropriate, but the abstract and conclusion do not carry it. The paper should either remove the best-model claim or support it with pairwise comparisons, effect sizes, and confidence intervals. In addition, the Cochran's Q implementation is underspecified: with three completions per image per model, the independence assumptions of the test are not obvious and should be described.
  3. [Section 4.3; Title] The title and the Section 4.3 discussion frame the contribution as 'VLMs as values detectors,' but the experiment does not operationalize 'value' or 'value-laden.' The prompt asks only for the element that needs attention, the annotation procedure does not ask about values, and the response-type taxonomy in Section 4.1 was constructed after seeing the annotator data. The value-laden interpretation rests on selected examples (smoking near a child, medicines, slipping) and on a general citation about values and perception [8], rather than on a measured value-detection construct. This section should be reframed as an exploratory hypothesis, or a separate validation step for value detection should be added.
minor comments (5)
  1. [Abstract; Section 4.2; Conclusion] The model is called 'LLaVA 34B' in the abstract and Table 2 but 'LLaVA 36B' in Section 4.2 and the Conclusion; please standardize the name.
  2. [Section 3.1] The prompt was chosen after initial testing on one image, which risks prompt overfitting; this should be listed as a limitation, especially because the paper later suggests that better prompts would improve alignment.
  3. [Section 3.2] The paper keeps only the first sentence of each model completion; if any answer was a list, a conditional statement, or a qualification, this truncation could remove the element actually identified, and the paper does not report how often truncation occurred.
  4. [Figure 2] The caption does not state how the 95% confidence intervals were computed (e.g., binomial intervals over the 36 binary responses per model), and the fact that the intervals overlap is itself further evidence that the model ordering is not reliable.
  5. [References] Reference [3] is attributed to 'Daniel, K.'; the correct author is Kahneman, D., and the title is Thinking, Fast and Slow.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the study is an empirical benchmark whose claims rest on independent measurement, not on inputs defined as outputs.

full rationale

This paper is an empirical benchmark, not a derivation. It creates twelve images, collects fourteen annotators' responses, prompts five VLMs with the same question, and reports binary alignment scores. There is no equation-level chain in which a prediction is defined as its input or a fitted parameter is renamed as a result. The only self-citations are [1] in the introduction, citing the authors' own prior work on vision-enabled dialogue, and [2] as the source of scenario events ('The events are taken from the results of previous research [2]'); neither is used to establish the alignment scores, so the self-citations are not load-bearing. The prompt was tuned on one image before evaluation, but this is a design choice and does not force the reported scores. The reliance on a majority-vote 'element of relevance' per image without reporting inter-annotator agreement is a validity threat to the ground truth, not a circularity: the models were not fitted to those labels, and the scoring rule is not equivalent to the conclusion. The qualitative 'more satisfying' comparisons in Section 4.2 are post-hoc interpretation, not a prediction derived from the data. Accordingly, no circular step is identified.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper rests on design assumptions rather than mathematical axioms; the key burden is that the majority-vote ground truth and binary scoring are accepted as meaningful despite evidence that human answers are diverse and model outputs sometimes appear more value-aware.

assumptions (4)
  • domain assumption The most frequent annotator response per image is a valid reference for 'the element that needs attention'.
    Section 2 defines the perceived element as the most common human answer, with random tie-breaking; the entire evaluation compares models to this reference.
  • domain assumption A single image can contain a single well-defined element of relevance.
    The prompt and scoring force a single answer per image; annotators' varied responses (e.g., dog vs. stain vs. none) show this is not necessarily true.
  • domain assumption The binary match between a model's first sentence and the majority human response measures alignment.
    Section 3.2 scores each answer as 0 or 1; the discussion acknowledges that a descriptive 'dog' might share intent with a value-aware 'dog is sick', so phrasing and understanding are conflated.
  • domain assumption The 12 diffusion-generated images are adequate proxies for real home scenarios.
    Section 2 states the images are generated ad-hoc with diffusion models, with field validation deferred to future work.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vision Language Models as Values Detectors." pith.science (2026). https://pith.science/paper/FCUW5LZ4

@misc{pith2026250103957,
  author       = {Pith},
  title        = {Pith review of: Vision Language Models as Values Detectors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FCUW5LZ4}},
  note         = {Machine review of arXiv:2501.03957}
}
read the original abstract

Large Language Models integrating textual and visual inputs have introduced new possibilities for interpreting complex data. Despite their remarkable ability to generate coherent and contextually relevant text based on visual stimuli, the alignment of these models with human perception in identifying relevant elements in images requires further exploration. This paper investigates the alignment between state-of-the-art LLMs and human annotators in detecting elements of relevance within home environment scenarios. We created a set of twelve images depicting various domestic scenarios and enlisted fourteen annotators to identify the key element in each image. We then compared these human responses with outputs from five different LLMs, including GPT-4o and four LLaVA variants. Our findings reveal a varied degree of alignment, with LLaVA 34B showing the highest performance but still scoring low. However, an analysis of the results highlights the models' potential to detect value-laden elements in images, suggesting that with improved training and refined prompts, LLMs could enhance applications in social robotics, assistive technologies, and human-computer interaction by providing deeper insights and more contextually relevant responses.

Figures

Figures reproduced from arXiv: 2501.03957 by the authors.

Figure 1
Figure 1. Four of the twelve images used in the evaluation. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Comparison of the element of focus alignments, with the 95% confi [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 1
Figure 1. The second kind is the most interesting; in this case, the participants display awareness of the situation and the values involved, according to their sensibility. This is the case of the responses referring to the dog’s health. The third category is the none answer, indicating that in the image nothing is particularly relevant; while comments on the image quality and framing form the fourth group of answers. We not… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "Can you be my mum?": Manipulating Social Robots in the Large Language Models Era

    cs.HC 2025-01 conditional novelty 4.0 of 10

    Users trying to make a social robot violate ethical principles use five strategies: Reason, Bargain, Emotion, Gaslight, and Roleplay.

Reference graph

Works this paper leans on

11 extracted references · 11 canonical work pages · cited by 1 Pith paper

  1. [8]

    S., and McGinnies, E

    Postman, L., Bruner, J. S., and McGinnies, E. Personal values as selective factors in perception. The Journal of Abnormal and Social Psychology 43, 2 (1948), 142–154

  2. [1]

    A., and Belpaeme, T

    Abbo, G. A., and Belpaeme, T. I Was Blind but Now I See: Imple- menting Vision-Enabled Dialogue in Social Robots, Nov. 2023

  3. [2]

    A., Marchesi, S., Wykowska, A., and Belpaeme, T

    Abbo, G. A., Marchesi, S., Wykowska, A., and Belpaeme, T. So- cial Value Alignment in Large Language Models. InValue Engineering in Artificial Intelligence(Cham, 2024), N.OsmanandL.Steels, Eds., Springer Nature Switzerland, pp. 83–97

  4. [3]

    Thinking, fast and slow

    Daniel, K. Thinking, fast and slow. 2017

  5. [4]

    Vision-language models as success detectors, 2023

    Du, Y., Konyushkov a, K., Denil, M., Raju, A., Landon, J., Hill, F., de Freitas, N., and Cabi, S. Vision-language models as success detectors, 2023

  6. [5]

    J., and Zou, J

    Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T. J., and Zou, J. A visual–language foundation model for pathology image analysis using medical twitter.Nature medicine 29, 9 (2023), 2307–2316

  7. [6]

    Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved Baselines with Visual Instruction Tuning. InNeurIPS 2023 Workshop on Instruction Tuning and Instruction Following (Nov. 2023). 12

  8. [7]

    Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual Instruction Tuning. In Thirty-Seventh Conference on Neural Information Processing Systems (Nov. 2023)

Show all 11 references
  1. [9]

    W., Hallacy, C., Ramesh, A., Goh, G., Agar w al, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I

    Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agar w al, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning Transferable Visual Models From Natural Language Supervision. InProceedings of the 38th Interna- tional Conference ...

  2. [10]

    H., Cieciuch, J., Vecchione, M., Da vidov, E., Fis- cher, R., Beierlein, C., Ramos, A., Verkasalo, M., Lönnqvist, J.-E., Demirutku, K., Dirilen-Gumus, O., and Konty, M

    Schw artz, S. H., Cieciuch, J., Vecchione, M., Da vidov, E., Fis- cher, R., Beierlein, C., Ramos, A., Verkasalo, M., Lönnqvist, J.-E., Demirutku, K., Dirilen-Gumus, O., and Konty, M. Refining the theory of basic individual values. Journal of Personality and Social Psychology 1...

  3. [11]

    Llama: Open and efficient foundation language models, 2023

    Touvron, H., La vril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Gra ve, E., and Lample, G. Llama: Open and efficient foundation language models, 2023. 13

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.