Pith. sign in

REVIEW 2 cited by

X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.18799 v2 pith:5LCYUPST submitted 2023-11-30 cs.CV cs.CL

classification cs.CVcs.CL
keywords reasoningcross-modalmodalitiesdataacrossframeworkprojectionadaptability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent research has achieved significant advancements in visual reasoning tasks through learning image-to-language projections and leveraging the impressive reasoning abilities of Large Language Models (LLMs). This paper introduces an efficient and effective framework that integrates multiple modalities (images, 3D, audio and video) to a frozen LLM and demonstrates an emergent ability for cross-modal reasoning (2+ modality inputs). Our approach explores two distinct projection mechanisms: Q-Formers and Linear Projections (LPs). Through extensive experimentation across all four modalities on 16 benchmarks, we explore both methods and assess their adaptability in integrated and separate cross-modal reasoning. The Q-Former projection demonstrates superior performance in single modality scenarios and adaptability in joint versus discriminative reasoning involving two or more modalities. However, it exhibits lower generalization capabilities than linear projection in contexts where task-modality data are limited. To enable this framework, we devise a scalable pipeline that automatically generates high-quality, instruction-tuning datasets from readily available captioning data across different modalities, and contribute 24K QA data for audio and 250K QA data for 3D. To facilitate further research in cross-modal reasoning, we introduce the DisCRn (Discriminative Cross-modal Reasoning) benchmark comprising 9K audio-video QA samples and 28K image-3D QA samples that require the model to reason discriminatively across disparate input modalities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CHAIR-DPO labels DPO preference pairs with CHAIR hallucinations computed from detector outputs, reducing object hallucinations in LLaVA models without proprietary judges.

  2. SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.

Pith tools