REVIEW 10 cited by
What Makes for Good Visual Tokenizers for Large Language Models?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We empirically investigate proper pre-training methods to build good visual tokenizers, making Large Language Models (LLMs) powerful Multimodal Large Language Models (MLLMs). In our benchmark, which is curated to evaluate MLLMs visual semantic understanding and fine-grained perception capabilities, we discussed different visual tokenizers pre-trained with dominant methods (i.e., DeiT, CLIP, MAE, DINO), and observe that: i) Fully/weakly supervised models capture more semantics than self-supervised models, but the gap is narrowed by scaling up the pre-training dataset. ii) Self-supervised models are better at fine-grained perception, where patch-level supervision is particularly effective. iii) Tuning the visual tokenizer leads to the loss of semantics obtained from large-scale pretraining, which is unfavorable with relatively small-scale instruction-tuning dataset. Given the findings, we reviewed methods that attempted to unify semantics and fine-grained visual understanding, e.g., patch-level feature distillation with semantically-rich targets. We obtain an intriguing insight mask-based strategies that were once all the rage may not be applicable for obtaining good visual tokenizers. Based on this critical observation, we obtain a new MLLM equipped with a tailored Good Visual Tokenizer (GVT), which exhibits strong visual comprehension capability at multiple scales. In particular, without introducing extra parameters and task-specific fine-tuning, GVT achieves superior performance on visual question answering, image captioning, and other fine-grained visual understanding tasks such as object counting and multi-class identification.
Forward citations
Cited by 10 Pith papers
-
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
VTBench evaluates visual tokenizers in isolation across reconstruction, detail, and text tasks, and finds discrete tokenizers lag continuous VAEs.
-
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.
-
Apollo: An Exploration of Video Understanding in Large Multimodal Models
Design decisions for video-LMMs can be made on 2-4B models and datasets and transfer to larger models, yielding efficient Apollo models, though some SOTA claims are contradicted by the paper's own table.
-
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
A new GPT-4-generated dataset of distractors and corrective feedback for visual commonsense reasoning, plus a compact LMM (PEIFG) that produces explainable corrections and beats existing baselines in automatic and hum...
-
Qwen-Audio-VAE Technical Report
A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.
-
ACTLLM: Action Consistency Tuned Large Language Model
ACTLLM trains an LLM to jointly produce structured scene descriptions and actions, and reports improved compositional and zero-shot generalization on CLIPORT and VIMA.
-
AGI Is Coming... Right After AI Learns to Play Wordle
OpenAI's Computer-User Agent solves Wordle only 5.36% of the time and its color perception degrades sharply as the game progresses, showing brittleness in a simple multimodal task.
-
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
Introduces a five-level agriculture benchmark and a 1,784-image multimodal dataset built from EU land-survey photos, with only qualitative model comparisons so far.
-
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
Panther improves multimodal LLMs by converting the user's text question into visual prompts that steer a frozen image encoder toward instruction-relevant regions, gaining about 2 to 3 points on several VQA benchmarks ...
-
MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models
MMRL++ inserts shared, learnable representation tokens into the upper layers of CLIP's image and text encoders and uses low-rank shared aligners, achieving state-of-the-art base-to-novel harmonic mean accuracy on 11 d...
Discussion (0). Continue with ORCID to comment.