Pith. sign in

REVIEW 2 cited by

SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01754 v3 pith:76KZSPXY submitted 2025-03-03 cs.CV

classification cs.CV
keywords reasoningdiversemodelmodelsself-distillationcapabilitiesframeworkin-context
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reasoning is increasingly crucial for various tasks. While chain-of-thought prompting enables large language models to leverage reasoning effectively, harnessing the reasoning capabilities of Vision-Language Models (VLMs) remains challenging. To solve this problem, we propose a novel self-distillation framework that enhances the reasoning capabilities of the model. The proposed framework introduces several key innovations. We start by employing a prompt library tailored to visual reasoning tasks to generate diverse in-context questions and utilize a two-step reasoning procedure to derive reasoning-guided responses. These responses are then used for self-distillation, enabling the model to internalize the reasoning process. Additionally, we improve the model architecture with several innovative components, including an intervention adapter for efficient parameter updates, a cross-modal skip connection to facilitate information exchange between modalities, and an ensemble learning algorithm to integrate diverse reasoning from multiple in-context questions. Extensive experiments show that our method significantly improves the baseline performance across five VQA datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On-Policy Self-Distillation without Any Supervision

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A language model improves its own math reasoning by distilling its majority-vote consensus into the prefixes of its own disagreeing answers, with no external labels.

  2. Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A self-critique, revision, and verification loop makes small vision-language models produce more detailed and more executable robot plans, beating their own baselines and, on the paper's judge-based evaluation, plans ...

Pith tools