Pith. sign in

REVIEW 66 cited by

NExT-GPT: Any-to-Any Multimodal LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05519 v3 pith:YKO2HDWV submitted 2023-09-11 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords next-gptmodalitiesmultimodalany-to-anycontentonlycapabledecoders
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MM-LLMs capable of accepting and delivering content in any modality becomes essential to human-level AI. To fill the gap, we present an end-to-end general-purpose any-to-any MM-LLM system, NExT-GPT. We connect an LLM with multimodal adaptors and different diffusion decoders, enabling NExT-GPT to perceive inputs and generate outputs in arbitrary combinations of text, images, videos, and audio. By leveraging the existing well-trained highly-performing encoders and decoders, NExT-GPT is tuned with only a small amount of parameter (1%) of certain projection layers, which not only benefits low-cost training and also facilitates convenient expansion to more potential modalities. Moreover, we introduce a modality-switching instruction tuning (MosIT) and manually curate a high-quality dataset for MosIT, based on which NExT-GPT is empowered with complex cross-modal semantic understanding and content generation. Overall, our research showcases the promising possibility of building an AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community. Project page: https://next-gpt.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 66 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 66 Pith citations

  1. Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

    cs.CV 2025-05 conditional novelty 7.0 of 10

    PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.

  2. Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond

    cs.LG 2024-12 reject novelty 7.0 of 10

    A theoretical analysis claims attention networks fail to learn residual features when time series steps have opposite signs, giving a possible explanation for the known advantage of linear residual models.

  3. AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A new audio-visual benchmark shows current multimodal LLMs perform barely above random guessing, with audio perception errors as the dominant failure mode.

  4. EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions

    cs.CV 2026-07 conditional novelty 6.5 of 10

    EmoteGPT regresses FLAME 3DMM expression parameters from explicit or implicit text using an MLLM with a dedicated <Expr> token, trained on the new Txt2Emote dataset plus image data, outperforming prior text-to-3D face...

  5. GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation

    cs.IR 2026-08 conditional novelty 6.0 of 10

    GARDRec improves LLM-based next-item ranking by grounding decisions in knowledge-graph embeddings, personalized graph contexts, and late-stage scoring rather than prompt text.

  6. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  7. S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A single 32B multimodal model with task-specific decoders handles roughly 200 scientific tasks across molecules, materials, proteins, spectra, and images, and outperforms general LLMs on most of 66 evaluated tasks.

  8. Laguerre Geometry for Interpreting Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    LLM concepts are Laguerre–Voronoi cells; Geometric Lens reads the exact cell of any hidden vector by isolating residual piecewise-linear flow from cross-token attention transport.

  9. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  10. Audio-Guided Visual Editing with Complex Multi-Modal Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...

  11. Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Multi-TW is the first Traditional Chinese benchmark to evaluate multimodal models on both image-text and audio-text questions while also measuring inference latency.

  12. Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A multimodal LLM trained with deliberate, format-constrained reasoning but evaluated with free-form reasoning outperforms models that keep the constraints at test time.

  13. NeoBabel: A Multilingual Open Tower for Visual Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.

  14. Let Your Video Listen to Your Music!

    cs.CV 2025-06 reject novelty 6.0 of 10

    MVAA aligns a video's motion peaks to music beats via keyframe re-timing and diffusion-based inpainting, aiming to preserve the original content while improving rhythmic synchronization.

  15. DanceChat: Large Language Model-Guided Music-to-Dance Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An LLM-generated text choreography, fused with music and beat features, guides a diffusion model to produce more diverse and physically plausible dance motion, with a multi-modal alignment loss intended to bridge musi...

  16. HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training recipe that initializes different depth segments of one transformer from pretrained ViT, LLM, and DiT models, then jointly tunes them to do multimodal understanding and generation.

  17. Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

    cs.SD 2025-05 conditional novelty 6.0 of 10

    AJailBench is an open benchmark showing that large audio-language models can be jailbroken through TTS-converted text attacks and through subtle acoustic perturbations that preserve speech semantics.

  18. Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.

  19. TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TokLIP semanticizes VQ image tokens with a causal CLIP-style encoder, improving multimodal comprehension while preserving autoregressive image generation.

  20. Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A prompt-based pipeline with GPT-4o detects hateful memes at state-of-the-art zero-shot accuracy and mitigates them by replacing hateful text or images, with 88% of 631 human-rated mitigated memes judged non-hateful.

  21. Computational Reasoning of Large Language Models

    cs.CL 2025-04 conditional novelty 6.0 of 10

    TMBench measures LLM computational reasoning by having models simulate m-tag systems step by step, and its pass rates correlate with AIME2024, MATH500, GPQA Diamond, and MMLU Pro scores across 12 leading models.

  22. OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding

    cs.CL 2025-04 conditional novelty 6.0 of 10

    A single medical vision-language model unifying 2D, 3D, and video inputs with rotary-position encoding and token pruning reportedly outperforms task-specific baselines on seven medical VQA benchmarks.

  23. UniCoRN: Unified Commented Retrieval Network with LMMs

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A frozen multimodal LLM is extended with a retrieval adapter and an entity adapter to retrieve a relevant image and generate a supportive textual comment.

  24. UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

    cs.CV 2025-02 conditional novelty 6.0 of 10

    UniCMs applies consistency distillation to a unified multimodal transformer, treating image mask-diffusion steps and text parallel-decoding steps as one shared denoising trajectory, enabling 2 to 8 step generation and...

  25. NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NAVER is a neuro-symbolic pipeline that uses a finite-state automaton with self-correction and probabilistic logic (ProbLog/Scallop) to ground referring expressions, achieving state-of-the-art accuracy on RefCOCO, Ref...

  26. PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PhysBench shows that 75 vision-language models, including GPT-4o, score only around 25-50% on physical world understanding, and a tool-augmented PhysAgent framework raises GPT-4o from 49.49% to 58.6%.

  27. SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

    cs.RO 2025-01 conditional novelty 6.0 of 10

    A two-stage VLM fine-tuning approach, coordinate alignment plus chain-of-thought grounding, improves closed-loop navigation and manipulation success rates over prior point-based spatial reasoning methods.

  28. GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

    cs.SD 2025-01 conditional novelty 6.0 of 10

    GVMGen generates background music from video using spatial and temporal cross-attention to condition a MusicGen decoder, reporting state-of-the-art correspondence and diversity.

  29. A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A multimodal language model, InstructCell, follows natural-language commands to annotate cell types, predict drug sensitivity, and generate realistic single-cell expression profiles.

  30. Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation

    cs.IR 2025-01 conditional novelty 6.0 of 10

    K-RagRec improves LLM-based recommendation by retrieving and encoding knowledge graph subgraphs as soft prompts, outperforming existing retrieval-augmented LLM recommenders on three datasets.

  31. HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    HLV-1K is a new benchmark of over one thousand long videos with time-specific question-answer pairs, used to measure how well AI models understand hour-scale video.

  32. Olympus: A Universal Task Router for Computer Vision Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.

  33. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.

  34. TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A dual-codebook VQ tokenizer with shared indices reports GenEval 0.55 autoregressive generation, 0.63 reconstruction FID, and a 7.2% average understanding gain over LLaVA-1.5, though the gain is partly inherited from ...

  35. EventGPT: Event Stream Understanding with Multimodal Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EventGPT adapts a LLaVA-style MLLM to event camera streams via three-stage training (image-language, event-language, instruction tuning) and outperforms RGB-based MLLMs on its own benchmark.

  36. LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LongVALE is a new benchmark of 8,411 long videos with 105,730 omni-modal events, each annotated with temporal boundaries and captions that relate visual, audio, and speech, and it shows that a video LLM trained on thi...

  37. DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users

    cs.HC 2024-11 conditional novelty 6.0 of 10

    DuetML adds multimodal LLM agents, one reactive and one proactive, to an interactive machine learning interface, and a small user study found outside evaluators rated its users' category definitions as more aligned wi...

  38. VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A unified multimodal motion LLM that handles nine generation and comprehension tasks across text, audio, and single/multi-agent motion, backed by a new dataset and tokenizer.

  39. Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A modified LLM that predicts the next pixel value in language space achieves state-of-the-art lossless image compression rates on several benchmarks.

  40. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

    cs.CV 2026-08 conditional novelty 5.0 of 10

    AutoDesign recursively improves a design harness, and the resulting DesignHarness raises PosterBench scores by 5.0 to 19.6 points across seven model configurations.

  41. DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images

    cs.CV 2025-07 reject novelty 5.0 of 10

    DatasetAgent is an LLM-powered multi-agent pipeline that automatically constructs image classification, detection, and segmentation datasets from web images, with modest downstream gains shown but weak experimental controls.

  42. UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A self-improving post-training method that uses a model's own generated images as training data, with SFT and GRPO, improves generation and understanding and reduces task imbalance.

  43. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

    cs.CV 2025-02 reject novelty 5.0 of 10

    HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.

  44. Large Models in Dialogue for Active Perception and Anomaly Detection

    cs.CV 2025-01 conditional novelty 5.0 of 10

    An LLM and a VQA model converse to steer a simulated drone through a scene, improving descriptions and hazard detection over a static baseline.

  45. Exploring GPT's Ability as a Judge in Music Understanding

    cs.IR 2025-01 conditional novelty 5.0 of 10

    GPT-3.5 detects synthetic annotation errors in beat, chord, and key tasks above random chance, with accuracy partly influenced by musical concepts in the prompt.

  46. VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

    cs.CV 2025-01 conditional novelty 5.0 of 10

    VARGPT combines LLaVA-style next-token visual understanding with VAR-style next-scale visual generation in one autoregressive multimodal model trained in three stages.

  47. AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.

  48. CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    CoF improves multimodal LLM benchmark scores by having the model locate an answer region, then reweighting attention toward that region during inference.

  49. DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis

    cs.LG 2024-12 conditional novelty 5.0 of 10

    DLF improves multimodal sentiment analysis by disentangling shared and specific features and steering cross-modal attention toward the dominant language modality.

  50. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

    cs.CV 2024-12 conditional novelty 5.0 of 10

    SynerGen-VL introduces token folding and vision expert FFNs to train a 2.4B encoder-free MLLM that matches larger unified models like Emu3 on multiple image understanding and generation benchmarks.

  51. VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

    cs.SD 2024-12 conditional novelty 5.0 of 10

    An adaptation of MusicGen that adds semantic conditioning from CLIP global features and rhythmic conditioning from CLIP local inter-frame similarity, trained in two stages, generates video-aligned background music.

  52. X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    X-Prompt compresses in-context image examples into a few learned tokens and adds text-description tasks, enabling a Chameleon-style autoregressive model to handle multiple image generation tasks in one framework.

  53. AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Introduces a five-level agriculture benchmark and a 1,784-image multimodal dataset built from EU land-survey photos, with only qualitative model comparisons so far.

  54. Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines

    cs.CL 2024-11 conditional novelty 5.0 of 10

    The authors create a 1,000-query multi-modal RAG benchmark with GPT-4o-based metrics, show multi-stage generation beats single-stage, and report fine-tuned 7B-8B models beating GPT-4o only in the single-stage comparison.

  55. Effectively obtaining acoustic, visual and textual data from videos

    cs.MM 2025-09 conditional novelty 4.0 of 10

    A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.

  56. Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge

    cs.CL 2025-06 conditional novelty 4.0 of 10

    An emotion-guided LLM prompting framework with bidirectional interaction improves hyperbole and metaphor detection, but the headline gains are measured against a weak BERT baseline.

  57. Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.

  58. Multimodal Representation Alignment for Cross-modal Information Retrieval

    cs.IR 2025-06 conditional novelty 4.0 of 10

    Across CLIP, BLIP, Meta-Transformer, and three combined unimodal models on IMDB, Flickr30K, and MS-COCO, cosine similarity gives the best cross-modal retrieval for contrastively trained models, while learned MLP align...

  59. A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms

    cs.IR 2025-04 conditional novelty 4.0 of 10

    A survey that organizes foundation-model recommender systems into feature-based, generative, and agentic paradigms and reviews tasks, empirical results, and open challenges.

  60. Towards Advancing Code Generation with Large Language Models: A Research Roadmap

    cs.SE 2025-01 conditional novelty 4.0 of 10

    A roadmap paper that organizes LLM code generation into a six-layer architecture and a four-phase human-in-the-loop workflow, and lists open challenges and recommendations.

See all 66 Pith citations

Pith tools