Pith. sign in

REVIEW 10 cited by

What Makes for Good Visual Tokenizers for Large Language Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.12223 v2 pith:YY7DSZM4 submitted 2023-05-20 cs.CV

classification cs.CV
keywords visualmodelsfine-grainedgoodtokenizerslanguagelargemethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We empirically investigate proper pre-training methods to build good visual tokenizers, making Large Language Models (LLMs) powerful Multimodal Large Language Models (MLLMs). In our benchmark, which is curated to evaluate MLLMs visual semantic understanding and fine-grained perception capabilities, we discussed different visual tokenizers pre-trained with dominant methods (i.e., DeiT, CLIP, MAE, DINO), and observe that: i) Fully/weakly supervised models capture more semantics than self-supervised models, but the gap is narrowed by scaling up the pre-training dataset. ii) Self-supervised models are better at fine-grained perception, where patch-level supervision is particularly effective. iii) Tuning the visual tokenizer leads to the loss of semantics obtained from large-scale pretraining, which is unfavorable with relatively small-scale instruction-tuning dataset. Given the findings, we reviewed methods that attempted to unify semantics and fine-grained visual understanding, e.g., patch-level feature distillation with semantically-rich targets. We obtain an intriguing insight mask-based strategies that were once all the rage may not be applicable for obtaining good visual tokenizers. Based on this critical observation, we obtain a new MLLM equipped with a tailored Good Visual Tokenizer (GVT), which exhibits strong visual comprehension capability at multiple scales. In particular, without introducing extra parameters and task-specific fine-tuning, GVT achieves superior performance on visual question answering, image captioning, and other fine-grained visual understanding tasks such as object counting and multi-class identification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VTBench evaluates visual tokenizers in isolation across reconstruction, detail, and text tasks, and finds discrete tokenizers lag continuous VAEs.

  2. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.

  3. Apollo: An Exploration of Video Understanding in Large Multimodal Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Design decisions for video-LMMs can be made on 2-4B models and datasets and transfer to larger models, yielding efficient Apollo models, though some SOTA claims are contradicted by the paper's own table.

  4. Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new GPT-4-generated dataset of distractors and corrective feedback for visual commonsense reasoning, plus a compact LMM (PEIFG) that produces explainable corrections and beats existing baselines in automatic and hum...

  5. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  6. ACTLLM: Action Consistency Tuned Large Language Model

    cs.RO 2025-06 conditional novelty 5.0 of 10

    ACTLLM trains an LLM to jointly produce structured scene descriptions and actions, and reports improved compositional and zero-shot generalization on CLIPORT and VIMA.

  7. AGI Is Coming... Right After AI Learns to Play Wordle

    cs.AI 2025-04 conditional novelty 5.0 of 10

    OpenAI's Computer-User Agent solves Wordle only 5.36% of the time and its color perception degrades sharply as the game progresses, showing brittleness in a simple multimodal task.

  8. AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Introduces a five-level agriculture benchmark and a 1,784-image multimodal dataset built from EU land-survey photos, with only qualitative model comparisons so far.

  9. Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Panther improves multimodal LLMs by converting the user's text question into visual prompts that steer a frozen image encoder toward instruction-relevant regions, gaining about 2 to 3 points on several VQA benchmarks ...

  10. MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models

    cs.CV 2025-05 conditional novelty 4.0 of 10

    MMRL++ inserts shared, learnable representation tokens into the upper layers of CLIP's image and text encoders and uses low-rank shared aligners, achieving state-of-the-art base-to-novel harmonic mean accuracy on 11 d...

Pith tools