Pith. sign in

REVIEW 12 cited by

OneLLM: One Framework to Align All Modalities with Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03700 v2 pith:LZAGGJOD submitted 2023-12-06 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords multimodalonellmmodalitiesimagelanguageprojectionalignencoder
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SensorQA: A Question Answering Benchmark for Daily-Life Monitoring

    cs.CL 2025-01 conditional novelty 7.0 of 10

    SensorQA is a new crowdsourced benchmark showing that current AI models answer only about 28% of daily-life sensor-data questions correctly.

  2. Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    OVODA combines a 3DETR-style detector with a frozen foundation model to detect novel objects and attributes in 3D scenes, and the OVAD dataset adds spatial and motion attribute labels to nuScenes.

  3. Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new benchmark, M3STR, renders knowledge-graph subgraphs as images and shows current MLLMs score near random on anomaly detection and poorly on entity counting.

  4. An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Llambda adapts LVLMs to low-resolution human-behavior videos by generating pseudo-label-guided captions from unlabeled data and fine-tuning with LoRA, reporting Bert-Score gains over LVLM baselines.

  5. SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A three-stage pipeline with LLM decomposition, pretrained embedding retrieval, and LLM assembly outperforms prior sensor QA systems on long-duration, high-frequency data, with caveats on evaluation leakage.

  6. 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    3UR-LLM encodes 3D point clouds with a frozen detector, compresses the features into 32 tokens, and beats 3D-LLM by 7.1 CIDEr on ScanQA with less training time.

  7. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.

  8. Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A 22-dataset benchmark compares existing multimodal AutoML tricks, and an automatic ensemble of those tricks achieves the most robust performance.

  9. Towards Universal Modal Tracking with Online Dense Temporal Token Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A video-level tracker with propagated temporal tokens and gated cross-modal fusion reports state-of-the-art results on RGB, RGB-T, RGB-D, and RGB-E benchmarks after one joint training run.

  10. Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    DCLIP fine-tunes a CLIP student's image encoder to match a YOLO-region, bidirectional cross-attention teacher, improving retrieval while retaining most zero-shot accuracy.

  11. Deploying Foundation Model Powered Agent Services: A Survey

    cs.DC 2024-12 accept novelty 4.0 of 10

    This survey proposes a layered framework (execution, resource, model, agent, application) for deploying foundation-model-powered agent services across edge-cloud environments, and reviews optimization techniques at ea...

  12. PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment

    cs.CV 2024-11 conditional novelty 4.0 of 10

    PSA-VLM adds a safety classifier and prompt-rewriting gate to LLaVA-style vision-language models, reporting state-of-the-art average safety scores on RTVLM and large gains on politics, porn, and cyberbullying datasets.

Pith tools