Pith. sign in

REVIEW 6 cited by

Probing Image-Language Transformers for Verb Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09141 v1 pith:Q4SSTJJK submitted 2021-06-16 cs.CL cs.CV

classification cs.CLcs.CV
keywords datasetimage-languagetransformersverbspretrainedrelytheyunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning (e.g., visual question answering and image retrieval). We are interested in shedding light on the quality of their pretrained representations -- in particular, if these models can distinguish different types of verbs or if they rely solely on nouns in a given sentence. To do so, we collect a dataset of image-sentence pairs (in English) consisting of 421 verbs that are either visual or commonly found in the pretraining data (i.e., the Conceptual Captions dataset). We use this dataset to evaluate pretrained image-language transformers and find that they fail more in situations that require verb understanding compared to other parts of speech. We also investigate what category of verbs are particularly challenging.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens

    cs.CV 2025-09 conditional novelty 7.0 of 10

    Feed-forward neurons in LVLMs encode whether a text token is visually grounded, and a detector built on these neurons can reduce hallucination by overriding or replacing ungrounded tokens.

  2. A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Blind text-only likelihood models match or exceed CLIP on many compositionality benchmarks because positives and negatives differ systematically in length, plausibility, or image style.

  3. ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ACE fine-tunes video-language models with stochastically sampled action synonyms and shadow negatives, improving zero-shot classification of unseen procedural actions by up to 16 percent harmonic mean on ATA, IKEA, and GTEA.

  4. Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CLAVER adds a cross-frame temporal attention mask (Kronecker mask) and LLM-generated interpretive action prompts to CLIP, improving video action recognition.

  5. A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A historical review that organizes multimodal explainability methods into four chronological eras and three explainability types, extending coverage to generative LLMs.

  6. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

Pith tools