Pith. sign in

REVIEW 24 cited by

Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09093 v1 pith:QSP6PRUV submitted 2023-06-15 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords multi-modalmoduledatallmsmacaw-llmalignmentaudiocapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although instruction-tuned large language models (LLMs) have exhibited remarkable capabilities across various NLP tasks, their effectiveness on other data modalities beyond text has not been fully studied. In this work, we propose Macaw-LLM, a novel multi-modal LLM that seamlessly integrates visual, audio, and textual information. Macaw-LLM consists of three main components: a modality module for encoding multi-modal data, a cognitive module for harnessing pretrained LLMs, and an alignment module for harmonizing diverse representations. Our novel alignment module seamlessly bridges multi-modal features to textual features, simplifying the adaptation process from the modality modules to the cognitive module. In addition, we construct a large-scale multi-modal instruction dataset in terms of multi-turn dialogue, including 69K image instances and 50K video instances. We have made our data, code and model publicly available, which we hope can pave the way for future research in multi-modal LLMs and expand the capabilities of LLMs to handle diverse data modalities and address complex real-world scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sample-efficient Integration of New Modalities into Large Language Models

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A hypernetwork trained on image, audio, and video adapts a shared projector to new, low-resource modalities from as few as 32 examples.

  2. Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

    cs.CV 2025-05 conditional novelty 7.0 of 10

    PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.

  3. Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    CAV-SAM reformulates reference segmentation as pseudo-video object segmentation using diffusion-based semantic transitions and test-time geometric alignment, claiming over 5% improvement over state-of-the-art.

  4. RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.

  5. Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling

    cs.CL 2025-05 conditional novelty 6.0 of 10

    MMER merges multiple multimodal LLMs by averaging task vectors and applying per-modality binary masks, retaining about 99% of original task performance without additional training.

  6. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.

  7. MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new 1.47M-instance biomedical instruction-tuning dataset improves a mixed-modal 7B model's medical VQA accuracy by 18 to 26 percentage points over GPT-4o and Chameleon.

  8. Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases

    cs.CV 2024-12 conditional novelty 6.0 of 10

    By fine-tuning LLaVA on a knowledge-infused agricultural dataset, Agri-LLaVA improves agricultural conversation and VQA over general LMMs, with gains of about 5 points over LLaVA on the new benchmark.

  9. LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LongVALE is a new benchmark of 8,411 long videos with 105,730 omni-modal events, each annotated with temporal boundaries and captions that relate visual, audio, and speech, and it shows that a video LLM trained on thi...

  10. Aligning Pre-trained Models for Spoken Language Translation

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Frozen speech recognition and machine translation models can be aligned by a small connector network to perform end-to-end speech translation, and the connector also serves as a domain adapter.

  11. VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automated arena benchmark that simulates real users asking open-ended video questions, uses GPT-4o as judge, and ranks 11 large multimodal models via ELO ratings.

  12. IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    IntentVCNet uses per-frame object coordinates, red-box visual prompts, and a lightweight box adapter to make video captioning focus on a user-selected object, reporting 225.19 CIDEr on the IntentVC public test set.

  13. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A new benchmark, VF-Eval, measures how well multimodal LLMs check, detect, and reason about errors in AI-generated videos, and shows frontier models remain far below human performance.

  14. Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Prot2Chat fuses protein sequence, structure, and question text in a text-aware adapter before a LoRA-tuned LLM generates answers, reporting large gains on Mol-Instructions but mixed zero-shot results on UniProtQA.

  15. LLaVA-SLT: Visual Language Tuning for Sign Language Translation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    LLaVA-SLT, a three-stage large multimodal model with a hierarchical visual encoder and lightweight MLP connector, achieves state-of-the-art gloss-free sign language translation on CSL-Daily and Phoenix-2014T, approach...

  16. COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A cascaded system that routes only high-entropy videos to a multimodal LLM achieves near-full-accuracy content moderation at a fraction of the GPU cost, and reduced inappropriate video views by 9.9% in an online A/B test.

  17. Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

    cs.CR 2024-11 conditional novelty 5.0 of 10

    An inference-time alignment method using a safety reward model and controlled decoding that reduces jailbreak success rates in multimodal LLMs while preserving utility.

  18. Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

    cs.SD 2025-09 reject novelty 4.0 of 10

    An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.

  19. Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.

  20. Movie2Story: A framework for understanding videos and telling stories in the form of novel text

    cs.CV 2024-12 reject novelty 4.0 of 10

    MSBench evaluates video-plus-audio to novel-style story generation; the M2S pipeline combines existing video, speech, emotion, and speaker tools with an LLM and reportedly beats video-only baselines.

  21. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  22. Do Language Models Understand Time?

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A survey arguing that video-LLMs rely on pretrained encoders and short-biased datasets, leaving them weak at long-term temporal reasoning such as causality and event progression.

  23. MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

    cs.SD 2024-12 reject novelty 3.0 of 10

    MuMu-LLaMA uses LLaMA with pretrained encoders and music decoders to understand and generate music from text, images, and videos, trained on a 167.69 hour machine-annotated dataset.

  24. On Accelerating Edge AI: Optimizing Resource-Constrained Environments

    cs.LG 2025-01 conditional novelty 2.0 of 10

    The paper argues that model compression, neural architecture search, and compiler optimizations work together to accelerate edge AI, but it provides no new experimental evidence.

Pith tools