Pith. sign in

REVIEW 12 cited by

Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19716 v2 pith:O7QSKCDI submitted 2024-05-30 cs.CV cs.CL

classification cs.CVcs.CL
keywords imageself-trainingdatamodelcapabilitycomprehensionimageslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understand image inputs for different queries and conduct subsequent reasoning. Improving this capability requires high-quality vision-language data, which is costly and labor-intensive to acquire. Self-training approaches have been effective in single-modal settings to alleviate the need for labeled data by leveraging model's own generation. However, effective self-training remains a challenge regarding the unique visual perception and reasoning capability of LVLMs. To address this, we introduce Self-Training on Image Comprehension (STIC), which emphasizes a self-training approach specifically for image comprehension. First, the model self-constructs a preference dataset for image descriptions using unlabeled images. Preferred responses are generated through a step-by-step prompt, while dis-preferred responses are generated from either corrupted images or misleading prompts. To further self-improve reasoning on the extracted visual information, we let the model reuse a small portion of existing instruction-tuning data and append its self-generated image descriptions to the prompts. We validate the effectiveness of STIC across seven different benchmarks, demonstrating substantial performance gains of 4.0% on average while using 70% less supervised fine-tuning data than the current method. Further studies investigate various components of STIC and highlight its potential to leverage vast quantities of unlabeled images for self-training. Code and data are made publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LPOI: Listwise Preference Optimization for Vision Language Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    LPOI reduces VLM hallucination by training the model to prefer the original image over progressively masked versions of the same image, using a listwise ranking loss built from pairwise preference data.

  2. Probing Visual Language Priors in VLMs

    cs.CV 2024-12 conditional novelty 7.0 of 10

    ViLP shows that vision-language models often answer from text priors instead of image content, and an image-corruption DPO method partially fixes this.

  3. Improving Large Vision and Language Models by Learning from a Panel of Peers

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A panel of LVLMs that generate, evaluate, and learn from each other's outputs improves average benchmark scores by 9 points across 15 tasks.

  4. Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Mol-LLM combines SELFIES text, 2D molecular graphs, and preference optimization to build a generalist chemistry LLM that outperforms prior generalist models on most property, reaction, and translation benchmarks.

  5. CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs

    cs.CL 2025-01 conditional novelty 6.0 of 10

    CHiP adds image-level and phrase/token-level preference signals to DPO for multimodal LLMs, and the authors report large hallucination-rate reductions on Object HalBench, AMBER, MMHal, and HallusionBench.

  6. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

    cs.CV 2025-01 conditional novelty 6.0 of 10

    IXC-2.5-Reward is an open-source multimodal reward model that achieves 70.0% macro accuracy on VL-RewardBench and improves LVLM chat via PPO.

  7. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.

  8. Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A multimodal LLM can improve itself using only unlabeled images by self-generating questions, self-enhancing answers, and adding a description-alignment loss to DPO.

  9. From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Answer-oriented chain-of-thought prompts that generate both positive and negative reasoning data, combined with iterative DPO, improve multimodal LLM reasoning on several benchmarks.

  10. Diving into Self-Evolving Training for Multimodal Reasoning

    cs.CL 2024-12 conditional novelty 5.0 of 10

    M-STAR, a self-evolving training recipe combining continuous updates, a process-reward-model reranker, and adaptive sampling temperature, improves multimodal reasoning on several benchmarks across three vision-languag...

  11. LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    LeanPO improves Video-LLM alignment by using a reference-free average-likelihood reward, self-generated winning/losing pairs, and dynamic label smoothing, yielding gains on six video benchmarks.

  12. Optimizing Vision-Language Interactions Through Decoder-Only Models

    cs.CV 2024-12 reject novelty 3.0 of 10

    An encoder-free vision-language model that claims state-of-the-art scores; the paper lacks any reproducible evidence.

Pith tools