Pith. sign in

REVIEW 4 cited by

DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08303 v2 pith:TVEIOR7X submitted 2024-07-11 cs.CV cs.AI

classification cs.CVcs.AI
keywords perceptionmllmsvisualcomprehensivedatasetdescriptionsdiverseelements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing Multimodal Large Language Models (MLLMs) increasingly emphasize complex understanding of various visual elements, including multiple objects, text information, and spatial relations. Their development for comprehensive visual perception hinges on the availability of high-quality image-text datasets that offer diverse visual elements and throughout image descriptions. However, the scarcity of such hyper-detailed datasets currently hinders progress within the MLLM community. The bottleneck stems from the limited perceptual capabilities of current caption engines, which fall short in providing complete and accurate annotations. To facilitate the cutting-edge research of MLLMs on comprehensive vision perception, we thereby propose Perceptual Fusion, using a low-budget but highly effective caption engine for complete and accurate image descriptions. Specifically, Perceptual Fusion integrates diverse perception experts as image priors to provide explicit information on visual elements and adopts an efficient MLLM as a centric pivot to mimic advanced MLLMs' perception abilities. We carefully select 1M highly representative images from uncurated LAION dataset and generate dense descriptions using our engine, dubbed DenseFusion-1M. Extensive experiments validate that our engine outperforms its counterparts, where the resulting dataset significantly improves the perception and cognition abilities of existing MLLMs across diverse vision-language benchmarks, especially with high-resolution images as inputs. The dataset and code are publicly available at https://github.com/baaivision/DenseFusion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenRecal: Generation after Recalibration from Large to Small Vision-Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A learnable Recalibrator bridges different tokenizers so that small VLMs can distill knowledge from any large VLM, improving their benchmark scores.

  2. COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A large, human-edited, mask-grounded caption dataset for COCO that improves fine-tuned vision-language and text-to-image models, and defines a panoptic grounded captioning benchmark.

  3. Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A side-by-side mobile benchmark shows VLM runtimes on a OnePlus 13R leave accelerators idle, push CPUs to thermal limits, and achieve order-of-magnitude power savings only when the GPU handles image and language kernels.

  4. EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    An encoder-free vision-language model using separate attention, normalization, and feed-forward weights for image versus text tokens outperforms earlier encoder-free models and narrows the gap to encoder-based VLMs wi...

Pith tools