HyperCap is the first large-scale hyperspectral captioning dataset built from four benchmark HSI datasets using hybrid automated-manual annotations, with evaluations showing classification gains for vision-language models.
mplug: Effective and effi- cient vision-language learning by cross-modal skip- connections
8 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 8representative citing papers
CompART adds a composition loss on decomposed captions to regularize attention sums and improves multi-object grounding plus VQA across four VLM types and six benchmarks.
ELLA introduces a timestep-aware semantic connector to link LLMs with diffusion models for improved dense prompt following, validated on a new 1K-prompt benchmark.
StableSketcher improves text-to-sketch generation by fine-tuning a diffusion VAE and adding a VQA-based RL reward, while releasing the SketchDUO dataset of sketches with captions and QA pairs.
VisChronos is a framework that combines LLMs and dense captioning models to produce event-centric captions from single images, accompanied by the EventCap dataset.
Scene-graph-guided multimodal RAG improves aerial VQA by feeding query-relevant structured visual knowledge to a text LLM instead of dense visual tokens.
GIT achieves new state-of-the-art results on 12 vision-language benchmarks, including surpassing human performance on TextCaps, via a simplified single-encoder single-decoder transformer scaled on large pre-training data.
OpenFlamingo provides open-source autoregressive vision-language models that achieve 80-89% of Flamingo performance on seven vision-language datasets.
citing papers explorer
-
HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models
HyperCap is the first large-scale hyperspectral captioning dataset built from four benchmark HSI datasets using hybrid automated-manual annotations, with evaluations showing classification gains for vision-language models.
-
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
CompART adds a composition loss on decomposed captions to regularize attention sums and improves multi-object grounding plus VQA across four VLM types and six benchmarks.
-
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
ELLA introduces a timestep-aware semantic connector to link LLMs with diffusion models for improved dense prompt following, validated on a new 1K-prompt benchmark.
-
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
StableSketcher improves text-to-sketch generation by fine-tuning a diffusion VAE and adding a VQA-based RL reward, while releasing the SketchDUO dataset of sketches with captions and QA pairs.
-
VisChronos: Revolutionizing Image Captioning Through Real-Life Events
VisChronos is a framework that combines LLMs and dense captioning models to produce event-centric captions from single images, accompanied by the EventCap dataset.
-
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
Scene-graph-guided multimodal RAG improves aerial VQA by feeding query-relevant structured visual knowledge to a text LLM instead of dense visual tokens.
-
GIT: A Generative Image-to-text Transformer for Vision and Language
GIT achieves new state-of-the-art results on 12 vision-language benchmarks, including surpassing human performance on TextCaps, via a simplified single-encoder single-decoder transformer scaled on large pre-training data.
-
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
OpenFlamingo provides open-source autoregressive vision-language models that achieve 80-89% of Flamingo performance on seven vision-language datasets.