Pith. sign in

REVIEW 10 cited by

Vision-Language Models for Vision Tasks: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00685 v2 pith:6TUBGWPC submitted 2023-04-03 cs.CV

classification cs.CV
keywords visualrecognitionmethodstasksmodelspre-trainingsurveyvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most visual recognition studies rely heavily on crowd-labelled data in deep neural networks (DNNs) training, and they usually train a DNN for each single visual recognition task, leading to a laborious and time-consuming visual recognition paradigm. To address the two challenges, Vision-Language Models (VLMs) have been intensively investigated recently, which learns rich vision-language correlation from web-scale image-text pairs that are almost infinitely available on the Internet and enables zero-shot predictions on various visual recognition tasks with a single VLM. This paper provides a systematic review of visual language models for various visual recognition tasks, including: (1) the background that introduces the development of visual recognition paradigms; (2) the foundations of VLM that summarize the widely-adopted network architectures, pre-training objectives, and downstream tasks; (3) the widely-adopted datasets in VLM pre-training and evaluations; (4) the review and categorization of existing VLM pre-training methods, VLM transfer learning methods, and VLM knowledge distillation methods; (5) the benchmarking, analysis and discussion of the reviewed methods; (6) several research challenges and potential research directions that could be pursued in the future VLM studies for visual recognition. A project associated with this survey has been created at https://github.com/jingyi0000/VLM_survey.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    OMG-VLM is a single VLM-based model that handles text-, image-, and multi-attributed graphs through structure-aware adapters, reporting gains on several node/link prediction benchmarks.

  2. Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A robot using GPT-4o labeled hazards in a simulated disaster room, and VR users preferred and rated these annotations highly, though the study lacks a controlled baseline comparison.

  3. When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective

    cs.SE 2025-09 conditional novelty 6.0 of 10

    The first empirical taxonomy of LLM tasks in UAVs, with an academia-industry comparison and survey, shows LLMs are used mainly for planning and interaction, not direct control.

  4. Re:Verse -- Can Your VLM Read a Manga?

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Current VLMs excel at individual manga panel interpretation but systematically fail at temporal causality and cross-panel cohesion in long-form narratives.

  5. Mobile GUI Agents under Real-world Threats: Are We There Yet?

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Introduces an app-content instrumentation framework and benchmark showing that examined GUI agents suffer 42.0% and 36.1% average misleading rates from third-party content in dynamic and static tests respectively.

  6. Robust Onion: Peeling Open Vocab Object Detectors Under Noise

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Empirical study finds OV-OD robustness driven by vision backbone and image domain via layer-wise feature collapse analysis, validated with a low-parameter robustness improvement on real data.

  7. Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    PACTS jointly model action trajectories and predicate belief trajectories in a single generative policy, enabling zero-shot skill composition via symbolic planning without retraining.

  8. RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

    cs.AI 2025-10 conditional novelty 5.0 of 10

    A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.

  9. Robust Onion: Peeling Open Vocab Object Detectors Under Noise

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    Empirical analysis shows open vocabulary object detector robustness is driven mainly by vision backbone and image domain via similar feature collapse patterns, with a lightweight NN & TK0 method improving real-world p...

  10. From Image Captioning to Visual Storytelling

    cs.CL 2025-07 unverdicted novelty 4.0 of 10

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

Pith tools