REVIEW 24 cited by
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Although instruction-tuned large language models (LLMs) have exhibited remarkable capabilities across various NLP tasks, their effectiveness on other data modalities beyond text has not been fully studied. In this work, we propose Macaw-LLM, a novel multi-modal LLM that seamlessly integrates visual, audio, and textual information. Macaw-LLM consists of three main components: a modality module for encoding multi-modal data, a cognitive module for harnessing pretrained LLMs, and an alignment module for harmonizing diverse representations. Our novel alignment module seamlessly bridges multi-modal features to textual features, simplifying the adaptation process from the modality modules to the cognitive module. In addition, we construct a large-scale multi-modal instruction dataset in terms of multi-turn dialogue, including 69K image instances and 50K video instances. We have made our data, code and model publicly available, which we hope can pave the way for future research in multi-modal LLMs and expand the capabilities of LLMs to handle diverse data modalities and address complex real-world scenarios.
Forward citations
Cited by 24 Pith papers
-
Sample-efficient Integration of New Modalities into Large Language Models
A hypernetwork trained on image, audio, and video adapts a shared projector to new, low-resource modalities from as few as 32 examples.
-
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.
-
Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild
CAV-SAM reformulates reference segmentation as pseudo-video object segmentation using diffusion-based semantic transitions and test-time geometric alignment, claiming over 5% improvement over state-of-the-art.
-
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.
-
Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
MMER merges multiple multimodal LLMs by averaging task vectors and applying per-modality binary masks, retaining about 99% of original task performance without additional training.
-
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.
-
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
A new 1.47M-instance biomedical instruction-tuning dataset improves a mixed-modal 7B model's medical VQA accuracy by 18 to 26 percentage points over GPT-4o and Chameleon.
-
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
By fine-tuning LLaVA on a knowledge-infused agricultural dataset, Agri-LLaVA improves agricultural conversation and VQA over general LMMs, with gains of about 5 points over LLaVA on the new benchmark.
-
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
LongVALE is a new benchmark of 8,411 long videos with 105,730 omni-modal events, each annotated with temporal boundaries and captions that relate visual, audio, and speech, and it shows that a video LLM trained on thi...
-
Aligning Pre-trained Models for Spoken Language Translation
Frozen speech recognition and machine translation models can be aligned by a small connector network to perform end-to-end speech translation, and the connector also serves as a domain adapter.
-
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
An automated arena benchmark that simulates real users asking open-ended video questions, uses GPT-4o as judge, and ranks 11 large multimodal models via ELO ratings.
-
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
IntentVCNet uses per-frame object coordinates, red-box visual prompts, and a lightweight box adapter to make video captioning focus on a user-selected object, reporting 225.19 CIDEr on the IntentVC public test set.
-
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
A new benchmark, VF-Eval, measures how well multimodal LLMs check, detect, and reason about errors in AI-generated videos, and shows frontier models remain far below human performance.
-
Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure
Prot2Chat fuses protein sequence, structure, and question text in a text-aware adapter before a LoRA-tuned LLM generates answers, reporting large gains on Mol-Instructions but mixed zero-shot results on UniProtQA.
-
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
LLaVA-SLT, a three-stage large multimodal model with a hierarchical visual encoder and lightweight MLP connector, achieves state-of-the-art gloss-free sign language translation on CSL-Daily and Phoenix-2014T, approach...
-
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework
A cascaded system that routes only high-entropy videos to a multimodal LLM achieves near-full-accuracy content moderation at a fraction of the GPU cost, and reduced inappropriate video views by 9.9% in an online A/B test.
-
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
An inference-time alignment method using a safety reward model and controlled decoding that reduces jailbreak success rates in multimodal LLMs while preserving utility.
-
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.
-
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.
-
Movie2Story: A framework for understanding videos and telling stories in the form of novel text
MSBench evaluates video-plus-audio to novel-style story generation; the M2S pipeline combines existing video, speech, emotion, and speaker tools with an LLM and reportedly beats video-only baselines.
-
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.
-
Do Language Models Understand Time?
A survey arguing that video-LLMs rely on pretrained encoders and short-biased datasets, leaving them weak at long-term temporal reasoning such as causality and event progression.
-
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
MuMu-LLaMA uses LLaMA with pretrained encoders and music decoders to understand and generate music from text, images, and videos, trained on a 167.69 hour machine-annotated dataset.
-
On Accelerating Edge AI: Optimizing Resource-Constrained Environments
The paper argues that model compression, neural architecture search, and compiler optimizations work together to accelerate edge AI, but it provides no new experimental evidence.
Discussion (0). Continue with ORCID to comment.