REVIEW 12 cited by
OneLLM: One Framework to Align All Modalities with Language
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM
Forward citations
Cited by 12 Pith papers
-
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
SensorQA is a new crowdsourced benchmark showing that current AI models answer only about 28% of daily-life sensor-data questions correctly.
-
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
OVODA combines a 3DETR-style detector with a frozen foundation model to detect novel objects and attributes in 3D scenes, and the OVAD dataset adds spatial and motion attribute labels to nuScenes.
-
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
A new benchmark, M3STR, renders knowledge-graph subgraphs as images and shows current MLLMs score near random on anomaly detection and poorly on entity counting.
-
An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
Llambda adapts LVLMs to low-resolution human-behavior videos by generating pseudo-label-guided captions from unlabeled data and fine-tuning with LoRA, reporting Bert-Score gains over LVLM baselines.
-
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
A three-stage pipeline with LLM decomposition, pretrained embedding retrieval, and LLM assembly outperforms prior sensor QA systems on long-duration, high-frequency data, with caveats on evaluation leakage.
-
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
3UR-LLM encodes 3D point clouds with a frozen detector, compresses the features into 32 tokens, and beats 3D-LLM by 7.1 CIDEr on ScanQA with less training time.
-
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.
-
Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data
A 22-dataset benchmark compares existing multimodal AutoML tricks, and an automatic ensemble of those tricks achieves the most robust performance.
-
Towards Universal Modal Tracking with Online Dense Temporal Token Learning
A video-level tracker with propagated temporal tokens and gated cross-modal fusion reports state-of-the-art results on RGB, RGB-T, RGB-D, and RGB-E benchmarks after one joint training run.
-
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
DCLIP fine-tunes a CLIP student's image encoder to match a YOLO-region, bidirectional cross-attention teacher, improving retrieval while retaining most zero-shot accuracy.
-
Deploying Foundation Model Powered Agent Services: A Survey
This survey proposes a layered framework (execution, resource, model, agent, application) for deploying foundation-model-powered agent services across edge-cloud environments, and reviews optimization techniques at ea...
-
PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment
PSA-VLM adds a safety classifier and prompt-rewriting gate to LLaVA-style vision-language models, reporting state-of-the-art average safety scores on RTVLM and large gains on politics, porn, and cyberbullying datasets.
Discussion (0). Continue with ORCID to comment.