Pith. sign in

REVIEW 17 cited by

VITA: Towards Open-Source Interactive Omni Multimodal LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05211 v3 pith:JQIYV4L2 submitted 2024-08-09 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords multimodalvitacapabilitiesopen-sourceaudioexperienceinteractivelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research. Project Page: https://vita-home.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free two-stage token compression framework preserves 92.9% of Qwen2.5-Omni-7B's audio-visual understanding accuracy while using 6.8% of the original multimodal-token FLOPs.

  2. Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Lychee-FD resolves modality interference in full-duplex spoken language models by separating acoustic and semantic parameters in deep layers and adding a dense semantic alignment channel, achieving state-of-the-art pe...

  3. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A unified mask-based discrete diffusion model jointly models text, speech, and image tokens and matches or exceeds several any-to-any and specialist multimodal baselines.

  4. EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    The paper introduces an event-level hierarchical control benchmark and a training-free agentic pipeline for video-to-audio generation, claiming 40.7% better controllability and 12.5% better perceptual quality than pri...

  5. EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.

  6. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  7. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  8. FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    FinMME is a new 11,099-sample financial chart benchmark where top AI models average around 50% and FinScore adds penalties for guessing.

  9. Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new multi-image benchmark and graph-based scoring method show that MLLMs produce weak, poorly-grounded chains of thought on visual physics tasks, and that RL post-training can degrade spatial reasoning.

  10. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A human-annotated multimodal preference dataset plus critique-based reward modeling and reward-margin-weighted DPO improves MLLM performance across many benchmarks.

  11. AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

    cs.MM 2025-01 conditional novelty 6.0 of 10

    AGAV-Rater, an LMM fine-tuned in two stages, achieves state-of-the-art quality scores for AI-generated audio-visual content, text-to-audio, and text-to-music.

  12. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...

  13. FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A full-duplex voice system with streaming personalized VAD and semantic end-of-turn detection reports fewer false barge-ins and latencies near commercial benchmarks.

  14. SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities

    cs.CL 2025-06 conditional novelty 5.0 of 10

    SMAR, a KL-based regularizer on per-modality expert routing, retains 86.6% of a Mixtral 8x7B's language score during visual instruction tuning with only 2.5% pure-text data.

  15. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Staged ASR-to-text-response-to-TTS training makes open end-to-end spoken dialogue systems trainable on 300 hours of public human-human data and more coherent than one-step speech-to-speech models.

  16. Ola: Pushing the Frontiers of Omni-Modal Language Model

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Ola, a 7B omni-modal language model, achieves competitive image, video, and audio understanding with progressive modality alignment, though it does not beat all specialized models on every benchmark.

  17. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...

Pith tools