Pith. sign in

REVIEW 27 cited by

Aria: An Open Multimodal Native Mixture-of-Experts Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05993 v4 pith:ZBHLKIDJ submitted 2024-10-08 cs.CV

classification cs.CV
keywords multimodalariamodelnativemodelsunderstandingadaptationsadoptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Information comes in diverse modalities. Multimodal native AI models are essential to integrate real-world information and deliver comprehensive understanding. While proprietary multimodal native models exist, their lack of openness imposes obstacles for adoptions, let alone adaptations. To fill this gap, we introduce Aria, an open multimodal native model with best-in-class performance across a wide range of multimodal, language, and coding tasks. Aria is a mixture-of-expert model with 3.9B and 3.5B activated parameters per visual token and text token, respectively. It outperforms Pixtral-12B and Llama3.2-11B, and is competitive against the best proprietary models on various multimodal tasks. We pre-train Aria from scratch following a 4-stage pipeline, which progressively equips the model with strong capabilities in language understanding, multimodal understanding, long context window, and instruction following. We open-source the model weights along with a codebase that facilitates easy adoptions and adaptations of Aria in real-world applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    Phase-aware expert merging based on routing role profiles preserves more MoE-VLM accuracy than global routing aggregation at matched compression ratios.

  2. Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

    cs.CV 2025-06 conditional novelty 7.0 of 10

    MF2 evaluates long-movie understanding by asking models to classify fact/fib claim pairs; the best model trails humans by 23.5 points in pairwise accuracy.

  3. VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A new dataset of 420 math video-question pairs with step-by-step reasoning annotations shows that current multimodal AI models, including the best proprietary system, answer fewer than half of the multi-binary questio...

  4. Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Multimodal pretraining transfers asymmetrically: language boosts vision, understanding boosts generation, generation is mostly neutral, and early unified training prevents vision laziness.

  5. InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Large multimodal models do far worse when collision videos violate familiar physics, and their small gains come from text exemplars, not the videos.

  6. MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    MESH, a three-layer video hallucination benchmark, shows LVMs ace basic objects and coarse traits but slip badly on fine character details and multi-subject actions in longer clips.

  7. HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical benchmark for multimodal models on human-centric visual understanding finds frontier models average under 60% and miss question-uncued visual evidence, with test-time scaling helping only marginally.

  8. VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    VLM4D benchmarks spatiotemporal reasoning in VLMs and finds large gaps versus humans, with proposed methods showing partial improvement.

  9. "Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    The authors construct the first Chinese youth-toxicity dataset, show that youth and adult perceptions of toxic language diverge, and report that adding contextual meta information improves detection accuracy.

  10. VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A perturbation computed from early attention and value matrices can make LLaVA, Instruct-BLIP, and BLIP2-T5 fail to detect objects inside a specified image region while keeping the rest of the image usable.

  11. AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.

  12. AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Using preference pairs synthesized from the model's own prompt-varied outputs, DPO fine-tuning improves Qwen2.5-VL-7B's video captioning on the VDC benchmark from 43.9 to 51.1 average VDCSCORE.

  13. CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    CAVALRY-V trains a two-stage generator with a semantic-visual loss to produce transferable adversarial video perturbations that reduce video and image MLLM benchmark scores.

  14. IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.

  15. VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A vision-expert-filtered, chain-of-thought-guided, iteratively fine-tuned reward model boosts a compact 7B model's ability to judge vision-language responses, especially detecting hallucinations.

  16. SeqPE: Transformer with Sequential Position Encoding

    cs.LG 2025-06 reject novelty 6.0 of 10

    SeqPE encodes each position as a symbolic digit sequence through a small Transformer, and with contrastive plus distillation losses it reports improved extrapolation in language, QA, and image classification.

  17. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VRBench is a benchmark of 960 long narrative videos with 8,243 human-written multi-step questions, plus a two-level evaluation of answer accuracy and reasoning quality for 31 large models.

  18. Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Training a multimodal LLM on GPT-4o-generated chain-of-thought referring traces, then optimizing with GRPO, improves referring accuracy and abstention on HumanRef.

  19. VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VCapsBench is a video caption quality benchmark with 109,796 QA pairs across 21 fine-grained dimensions on 5,677 videos, evaluating caption accuracy, inconsistency, and coverage.

  20. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.5 of 10

    A monolithic multimodal LLM that cuts pre-training data by 58% and first-token latency by up to 69% while matching or beating its predecessor on 15 benchmarks.

  21. TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    TimeExpert applies dynamic mixture-of-experts routing to video temporal grounding, reporting small state-of-the-art gains over TRACE on dense video captioning, moment retrieval, and highlight detection.

  22. Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    VideoLLaMA2's tennis sequence edit score jumps from 39.7 to 76.0 when text coordinates from detection models are included in the prompt, and a separately fine-tuned CLIP encoder raises single-event accuracy from 0.41 to 0.56.

  23. CyberV: Cybernetics for Test-time Scaling in Video Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A training-free test-time feedback loop, using attention drift to select key frames, improves video MLLM accuracy, with the largest gains on knowledge-heavy VideoMMMU.

  24. SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities

    cs.CL 2025-06 conditional novelty 5.0 of 10

    SMAR, a KL-based regularizer on per-modality expert routing, retains 86.6% of a Mixtral 8x7B's language score during visual instruction tuning with only 2.5% pure-text data.

  25. FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FlexSelect selects a small fraction of query-relevant visual tokens using attention from an intermediate layer, improving long-video accuracy and inference speed across multiple VideoLLMs.

  26. Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

    cs.CV 2025-09 reject novelty 4.0 of 10

    Using SigLIP retrieval to feed a Qwen2-VL or InternVL2 model with similar and dissimilar coordinates yields reported street-level accuracies of 23.2%, 17.1%, and 24.3% on IM2GPS, IM2GPS3k, and YFCC4k.

  27. DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

    cs.CL 2025-06 conditional novelty 4.0 of 10

    DynTok dynamically merges similar adjacent visual tokens into groups, reducing video token counts to 44.4% with comparable or better video understanding accuracy.

Pith tools