REVIEW 24 cited by
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce AnyGPT, an any-to-any multimodal language model that utilizes discrete representations for the unified processing of various modalities, including speech, text, images, and music. AnyGPT can be trained stably without any alterations to the current large language model (LLM) architecture or training paradigms. Instead, it relies exclusively on data-level preprocessing, facilitating the seamless integration of new modalities into LLMs, akin to the incorporation of new languages. We build a multimodal text-centric dataset for multimodal alignment pre-training. Utilizing generative models, we synthesize the first large-scale any-to-any multimodal instruction dataset. It consists of 108k samples of multi-turn conversations that intricately interweave various modalities, thus equipping the model to handle arbitrary combinations of multimodal inputs and outputs. Experimental results demonstrate that AnyGPT is capable of facilitating any-to-any multimodal conversation while achieving performance comparable to specialized models across all modalities, proving that discrete representations can effectively and conveniently unify multiple modalities within a language model. Demos are shown in https://junzhan2000.github.io/AnyGPT.github.io/
Forward citations
Cited by 24 Pith papers
-
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
A new audio-visual benchmark shows current multimodal LLMs perform barely above random guessing, with audio perception errors as the dominant failure mode.
-
POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking
Prompt-optimized suffixes plus synthetic fine-tuning recover ~82% of knowledge that multimodal unlearning methods claim to erase from MLLMs.
-
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
Multi-TW is the first Traditional Chinese benchmark to evaluate multimodal models on both image-text and audio-text questions while also measuring inference latency.
-
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.
-
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
A clue-grounded audio-visual counting benchmark over 497 long videos and an RL-trained counting model, whose headline result is undermined by training on the DVD-Counting evaluation benchmark.
-
Probing the Robustness Properties of Neural Speech Codecs
DAC is the most noise-robust neural codec at high bitrates, but at 3 kbps EnCodec wins, and measured non-linearity correlates with robustness.
-
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.
-
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
Introduces a 700-pair benchmark for temporal causal reasoning in VLMs, revealing large open-source vs. closed-source gaps and strong position bias.
-
Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
A single sequence-based model, pretrained across molecules, proteins, materials, nucleotides and text, outperforms specialist models on several generation tasks and enables cross-domain design.
-
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
Audio LLMs fine-tuned with token-level distillation against an LLM teacher can predict speech quality scores and generate natural-language descriptions, including A/B comparisons.
-
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
CoMT is the first benchmark to ask LVLMs to produce interleaved image and text rationales, and current models perform near random on it.
-
Olympus: A Universal Task Router for Computer Vision Tasks
Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.
-
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
An autoregressive model with group self-attention that separates learning from applying achieves state-of-the-art few-shot image manipulation on unseen instructions.
-
EventGPT: Event Stream Understanding with Multimodal Large Language Models
EventGPT adapts a LLaVA-style MLLM to event camera streams via three-stage training (image-language, event-language, instruction tuning) and outperforms RGB-based MLLMs on its own benchmark.
-
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
A parameter-efficient adapter bridging Whisper and TinyLlama reports relative improvements in speech recognition, named entity recognition, and sentiment analysis on low-resource benchmarks.
-
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Stream-Omni uses CTC-based layer-dimension mapping to align speech with text, achieving vision, speech, and text interaction in one 8B model trained on 23,000 hours of speech.
-
Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
Uncertainty-o estimates uncertainty in large multimodal models by perturbing prompts and computing entropy over semantically clustered answers, improving hallucination detection across five modalities.
-
ALAS: An Automatic Latent Alignment Score for Audio Language Models
ALAS is a reference-based score for audio-text alignment in speech LLMs, computed from frozen hidden states and a Whisper-derived alignment path, with no training or fitted classifier.
-
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.
-
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
An encoder-free vision-language model using separate attention, normalization, and feed-forward weights for image versus text tokens outperforms earlier encoder-free models and narrows the gap to encoder-based VLMs wi...
-
Causal Diffusion Transformers for Generative Modeling
A decoder-only transformer that factors generation over both token order and noise level, coupling autoregressive and diffusion training, achieves competitive ImageNet generation and in-context editing.
-
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
ILLUME unifies visual understanding and generation in one LLM with a semantic vision tokenizer and a self-enhancing alignment scheme, reaching competitive benchmarks with only 15M pretraining pairs.
-
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
Panther improves multimodal LLMs by converting the user's text question into visual prompts that steer a frozen image encoder toward instruction-relevant regions, gaining about 2 to 3 points on several VQA benchmarks ...
-
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.
Discussion (0). Continue with ORCID to comment.