REVIEW 19 cited by
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fueled a wealth of new models such as LLaVa, InstructBLIP, and PaLI-3. Despite the volume of new releases, key design decisions around image preprocessing, architecture, and optimization are under-explored, making it challenging to understand what factors account for model performance $-$ a challenge further complicated by the lack of objective, consistent evaluations. To address these gaps, we first compile a suite of standardized evaluations spanning visual question answering, object localization, and challenge sets that probe properties such as hallucination; evaluations that provide fine-grained insight VLM capabilities. Second, we rigorously investigate VLMs along key design axes, including pretrained visual representations and training from base vs. instruct-tuned language models, amongst others. We couple our analysis with three resource contributions: (1) a unified framework for evaluating VLMs, (2) optimized, flexible training code, and (3) checkpoints for all models, including a family of VLMs at the 7-13B scale that strictly outperform InstructBLIP and LLaVa v1.5, the state-of-the-art in open VLMs.
Forward citations
Cited by 19 Pith papers
-
An Exam for Active Observers
On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.
-
RoboVista: Evaluating Vision Language Models for Diverse Robot Applications
Expert-curated modular Robot-VQA benchmark of 474 questions across 39 robot tasks shows SOTA VLMs have large gaps that correlate with physical robot execution.
-
Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
A memory-guided LLM planner composes a frozen VLA as a contact-rich primitive with fixed analytic controllers, lifting perturbed manipulation success without VLA finetuning.
-
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
A training-free framework jointly quantizes a VLA model to 4 bits and prunes visual tokens, recovering or exceeding full-precision success rates at 1.93x speedup.
-
Hidden in plain sight: VLMs overlook their visual representations
VLMs perform far worse than their own visual encoders on vision-centric tasks because the language model fails to use accessible visual information and instead follows its language priors.
-
Diffusion Instruction Tuning
Lavender fine-tunes vision-language models by aligning their attention maps with Stable Diffusion's attention targets, improving accuracy on 20 benchmarks with as few as 0.13 million training examples.
-
Mordal: Automated Pretrained Model Selection for Vision Language Models
Mordal automates the selection of pretrained vision encoder and LLM pairs for VLMs, using representation clustering and partial-training scaling predictions to cut search cost by about an order of magnitude.
-
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Eagle2-9B matches or outperforms much larger vision-language models on many benchmarks through a carefully constructed post-training data strategy.
-
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
NoisyEQA benchmarks embodied QA agents against four noise types and claims a self-correcting prompt (NACoT) markedly improves noisy-question accuracy as scored by GPT-4.
-
How to Merge Your Multimodal Models Over Time?
A systematic study of temporal model merging shows that initialization and deployment choices matter far more than the merging technique, with EMA-style weight interpolation as the best practice.
-
Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
Adversarial patches optimized against OpenVLA raise manipulation failure rates from about 23% to 100% in LIBERO simulation and disrupt a physical robot arm in 43% of trials.
-
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
Structural pruning with finetuning plus hidden-state distillation recovers most performance in multimodal LLMs, with 5% of training data sufficient at moderate compression levels.
-
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
EfficientVLA combines LLM layer pruning, task-aware visual token selection, and diffusion-head feature caching to cut CogACT's inference cost to 28.9% of baseline FLOPs with a 0.6% SIMPLER success drop.
-
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
FiS-VLA embeds a diffusion-based action module into the final transformer blocks of a vision-language model, achieving 69% mean success on RLBench and a claimed 117.7 Hz control frequency.
-
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
A new patch-aligned pretraining loss improves fine-grained vision-language alignment and grounding in multimodal LLMs.
-
An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation
A data-driven framework decomposes manipulation tasks into reusable atomic skills, fine-tunes a VLA model per skill, and reports reduced data needs with comparable or better real-robot success rates.
-
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
A frozen-LLM method that steers visual token representations in low-rank subspaces achieves benchmark scores close to LoRA with about 500x fewer trainable parameters.
-
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
StemVLA supervises a GPT-2-based VLA with predicted future 3D-geometry features (VGGT) and temporally aggregated history, reporting 86.0% on LIBERO-Long - but its CALVIN results and equations are placeholders.
-
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
The authors expand the VSR benchmark with 50 question templates and diffusion-augmented images, fuse four vision encoders, and report a VSR-specialist VLLM (VSRE) that substantially outperforms LLaVA1.5 on spatial rea...
Discussion (0). Continue with ORCID to comment.