Pith. sign in

REVIEW 2 cited by

Visual Perturbation-aware Collaborative Learning for Overcoming the Language Prior Problem

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.11850 v1 pith:VHSKLCPN submitted 2022-07-24 cs.CV

classification cs.CV
keywords visuallanguagelearninganswercollaborativepriorproblemrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Several studies have recently pointed that existing Visual Question Answering (VQA) models heavily suffer from the language prior problem, which refers to capturing superficial statistical correlations between the question type and the answer whereas ignoring the image contents. Numerous efforts have been dedicated to strengthen the image dependency by creating the delicate models or introducing the extra visual annotations. However, these methods cannot sufficiently explore how the visual cues explicitly affect the learned answer representation, which is vital for language reliance alleviation. Moreover, they generally emphasize the class-level discrimination of the learned answer representation, which overlooks the more fine-grained instance-level patterns and demands further optimization. In this paper, we propose a novel collaborative learning scheme from the viewpoint of visual perturbation calibration, which can better investigate the fine-grained visual effects and mitigate the language prior problem by learning the instance-level characteristics. Specifically, we devise a visual controller to construct two sorts of curated images with different perturbation extents, based on which the collaborative learning of intra-instance invariance and inter-instance discrimination is implemented by two well-designed discriminators. Besides, we implement the information bottleneck modulator on latent space for further bias alleviation and representation calibration. We impose our visual perturbation-aware framework to three orthodox baselines and the experimental results on two diagnostic VQA-CP benchmark datasets evidently demonstrate its effectiveness. In addition, we also justify its robustness on the balanced VQA benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ICT steers attention-head activations toward visual information using blur-derived trusted and untrusted pairs, improving object hallucination benchmarks in LLaVA-v1.5 and Qwen-VL without slowing generation.

  2. Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DCLA reduces hallucinations by correcting each transformer layer's hidden state toward an exponentially weighted average of earlier layers, gated by a cosine-similarity threshold.

Pith tools