REVIEW 66 cited by
NExT-GPT: Any-to-Any Multimodal LLM
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MM-LLMs capable of accepting and delivering content in any modality becomes essential to human-level AI. To fill the gap, we present an end-to-end general-purpose any-to-any MM-LLM system, NExT-GPT. We connect an LLM with multimodal adaptors and different diffusion decoders, enabling NExT-GPT to perceive inputs and generate outputs in arbitrary combinations of text, images, videos, and audio. By leveraging the existing well-trained highly-performing encoders and decoders, NExT-GPT is tuned with only a small amount of parameter (1%) of certain projection layers, which not only benefits low-cost training and also facilitates convenient expansion to more potential modalities. Moreover, we introduce a modality-switching instruction tuning (MosIT) and manually curate a high-quality dataset for MosIT, based on which NExT-GPT is empowered with complex cross-modal semantic understanding and content generation. Overall, our research showcases the promising possibility of building an AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community. Project page: https://next-gpt.github.io/
Forward citations
Showing 60 of 66 Pith papers that cite this
-
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.
-
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
A theoretical analysis claims attention networks fail to learn residual features when time series steps have opposite signs, giving a possible explanation for the known advantage of linear residual models.
-
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
A new audio-visual benchmark shows current multimodal LLMs perform barely above random guessing, with audio perception errors as the dominant failure mode.
-
EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions
EmoteGPT regresses FLAME 3DMM expression parameters from explicit or implicit text using an MLLM with a dedicated <Expr> token, trained on the new Txt2Emote dataset plus image data, outperforming prior text-to-3D face...
-
GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation
GARDRec improves LLM-based next-item ranking by grounding decisions in knowledge-graph embeddings, personalized graph contexts, and late-stage scoring rather than prompt text.
-
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...
-
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
A single 32B multimodal model with task-specific decoders handles roughly 200 scientific tasks across molecules, materials, proteins, spectra, and images, and outperforms general LLMs on most of 66 evaluated tasks.
-
Laguerre Geometry for Interpreting Large Language Models
LLM concepts are Laguerre–Voronoi cells; Geometric Lens reads the exact cell of any hidden vector by isolating residual piecewise-linear flow from cross-token attention transport.
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
-
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...
-
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
Multi-TW is the first Traditional Chinese benchmark to evaluate multimodal models on both image-text and audio-text questions while also measuring inference latency.
-
Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs
A multimodal LLM trained with deliberate, format-constrained reasoning but evaluated with free-form reasoning outperforms models that keep the constraints at test time.
-
NeoBabel: A Multilingual Open Tower for Visual Generation
A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.
-
Let Your Video Listen to Your Music!
MVAA aligns a video's motion peaks to music beats via keyframe re-timing and diffusion-based inpainting, aiming to preserve the original content while improving rhythmic synchronization.
-
DanceChat: Large Language Model-Guided Music-to-Dance Generation
An LLM-generated text choreography, fused with music and beat features, guides a diffusion model to produce more diverse and physically plausible dance motion, with a multi-modal alignment loss intended to bridge musi...
-
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
A training recipe that initializes different depth segments of one transformer from pretrained ViT, LLM, and DiT models, then jointly tunes them to do multimodal understanding and generation.
-
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AJailBench is an open benchmark showing that large audio-language models can be jailbroken through TTS-converted text attacks and through subtle acoustic perturbations that preserve speech semantics.
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.
-
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
TokLIP semanticizes VQ image tokens with a causal CLIP-style encoder, improving multimodal comprehension while preserving autoregressive image generation.
-
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
A prompt-based pipeline with GPT-4o detects hateful memes at state-of-the-art zero-shot accuracy and mitigates them by replacing hateful text or images, with 88% of 631 human-rated mitigated memes judged non-hateful.
-
Computational Reasoning of Large Language Models
TMBench measures LLM computational reasoning by having models simulate m-tag systems step by step, and its pass rates correlate with AIME2024, MATH500, GPQA Diamond, and MMLU Pro scores across 12 leading models.
-
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding
A single medical vision-language model unifying 2D, 3D, and video inputs with rotary-position encoding and token pruning reportedly outperforms task-specific baselines on seven medical VQA benchmarks.
-
UniCoRN: Unified Commented Retrieval Network with LMMs
A frozen multimodal LLM is extended with a retrieval adapter and an entity adapter to retrieve a relevant image and generate a supportive textual comment.
-
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
UniCMs applies consistency distillation to a unified multimodal transformer, treating image mask-diffusion steps and text parallel-decoding steps as one shared denoising trajectory, enabling 2 to 8 step generation and...
-
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
NAVER is a neuro-symbolic pipeline that uses a finite-state automaton with self-correction and probabilistic logic (ProbLog/Scallop) to ground referring expressions, achieving state-of-the-art accuracy on RefCOCO, Ref...
-
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
PhysBench shows that 75 vision-language models, including GPT-4o, score only around 25-50% on physical world understanding, and a tool-augmented PhysAgent framework raises GPT-4o from 49.49% to 58.6%.
-
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
A two-stage VLM fine-tuning approach, coordinate alignment plus chain-of-thought grounding, improves closed-loop navigation and manipulation success rates over prior point-based spatial reasoning methods.
-
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
GVMGen generates background music from video using spatial and temporal cross-attention to condition a MusicGen decoder, reporting state-of-the-art correspondence and diversity.
-
A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following
A multimodal language model, InstructCell, follows natural-language commands to annotate cell types, predict drug sensitivity, and generate realistic single-cell expression profiles.
-
Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation
K-RagRec improves LLM-based recommendation by retrieving and encoding knowledge graph subgraphs as soft prompts, outperforming existing retrieval-augmented LLM recommenders on three datasets.
-
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
HLV-1K is a new benchmark of over one thousand long videos with time-specific question-answer pairs, used to measure how well AI models understand hour-scale video.
-
Olympus: A Universal Task Router for Computer Vision Tasks
Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.
-
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.
-
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
A dual-codebook VQ tokenizer with shared indices reports GenEval 0.55 autoregressive generation, 0.63 reconstruction FID, and a 7.2% average understanding gain over LLaVA-1.5, though the gain is partly inherited from ...
-
EventGPT: Event Stream Understanding with Multimodal Large Language Models
EventGPT adapts a LLaVA-style MLLM to event camera streams via three-stage training (image-language, event-language, instruction tuning) and outperforms RGB-based MLLMs on its own benchmark.
-
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
LongVALE is a new benchmark of 8,411 long videos with 105,730 omni-modal events, each annotated with temporal boundaries and captions that relate visual, audio, and speech, and it shows that a video LLM trained on thi...
-
DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users
DuetML adds multimodal LLM agents, one reactive and one proactive, to an interactive machine learning interface, and a small user study found outside evaluators rated its users' category definitions as more aligned wi...
-
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
A unified multimodal motion LLM that handles nine generation and comprehension tasks across text, audio, and single/multi-agent motion, backed by a new dataset and tokenizer.
-
Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
A modified LLM that predicts the next pixel value in language space achieves state-of-the-art lossless image compression rates on several benchmarks.
-
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
AutoDesign recursively improves a design harness, and the resulting DesignHarness raises PosterBench scores by 5.0 to 19.6 points across seven model configurations.
-
DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
DatasetAgent is an LLM-powered multi-agent pipeline that automatically constructs image classification, detection, and segmentation datasets from web images, with modest downstream gains shown but weak experimental controls.
-
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
A self-improving post-training method that uses a model's own generated images as training data, with SFT and GRPO, improves generation and understanding and reduces task imbalance.
-
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.
-
Large Models in Dialogue for Active Perception and Anomaly Detection
An LLM and a VQA model converse to steer a simulated drone through a scene, improving descriptions and hazard detection over a static baseline.
-
Exploring GPT's Ability as a Judge in Music Understanding
GPT-3.5 detects synthetic annotation errors in beat, chord, and key tasks above random chance, with accuracy partly influenced by musical concepts in the prompt.
-
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
VARGPT combines LLaVA-style next-token visual understanding with VAR-style next-scale visual generation in one autoregressive multimodal model trained in three stages.
-
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.
-
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
CoF improves multimodal LLM benchmark scores by having the model locate an answer region, then reweighting attention toward that region during inference.
-
DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis
DLF improves multimodal sentiment analysis by disentangling shared and specific features and steering cross-modal attention toward the dominant language modality.
-
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
SynerGen-VL introduces token folding and vision expert FFNs to train a 2.4B encoder-free MLLM that matches larger unified models like Emu3 on multiple image understanding and generation benchmarks.
-
VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
An adaptation of MusicGen that adds semantic conditioning from CLIP global features and rhythmic conditioning from CLIP local inter-frame similarity, trained in two stages, generates video-aligned background music.
-
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
X-Prompt compresses in-context image examples into a few learned tokens and adds text-description tasks, enabling a Chameleon-style autoregressive model to handle multiple image generation tasks in one framework.
-
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
Introduces a five-level agriculture benchmark and a 1,784-image multimodal dataset built from EU land-survey photos, with only qualitative model comparisons so far.
-
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
The authors create a 1,000-query multi-modal RAG benchmark with GPT-4o-based metrics, show multi-stage generation beats single-stage, and report fine-tuned 7B-8B models beating GPT-4o only in the single-stage comparison.
-
Effectively obtaining acoustic, visual and textual data from videos
A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.
-
Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge
An emotion-guided LLM prompting framework with bidirectional interaction improves hyperbole and metaphor detection, but the headline gains are measured against a weak BERT baseline.
-
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.
-
Multimodal Representation Alignment for Cross-modal Information Retrieval
Across CLIP, BLIP, Meta-Transformer, and three combined unimodal models on IMDB, Flickr30K, and MS-COCO, cosine similarity gives the best cross-modal retrieval for contrastively trained models, while learned MLP align...
-
A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
A survey that organizes foundation-model recommender systems into feature-based, generative, and agentic paradigms and reviews tasks, empirical results, and open challenges.
-
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
A roadmap paper that organizes LLM code generation into a six-layer architecture and a four-phase human-in-the-loop workflow, and lists open challenges and recommendations.
Discussion (0). Continue with ORCID to comment.