Current MLLMs show weak performance on small object understanding tasks, but fine-tuning with the new SOU-Train dataset measurably improves their capabilities.
Title resolution pending
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
dataset 2polarities
use dataset 2representative citing papers
A V-sink/L-sink taxonomy plus a frozen-backbone, NTP-trained layer-wise key gate (LSG) improves LLaVA-1.5-7B by up to +1.55pp on MMStar and +3.08pp on CVBench.
StochasT uses stochastic clustering of language tasks into varying turn depths for the same image to improve LVLMs on both single-turn and multi-turn scenarios without discarding data.
CoLT replaces text chain-of-thought in multimodal LLMs with three supervised latent vectors, improving average accuracy from 75.7 to 79.1 across eight benchmarks while cutting text-decoding time 22.6x.
HyLaR interleaves discrete text generation with continuous visual latent representations and optimizes them via a decoupled RL algorithm using vMF distributions, improving fine-grained visual reasoning.
OpenSpatial supplies a principled open-source data engine and 3-million-sample dataset that raises spatial-reasoning model performance by an average of 19 percent on benchmarks.
citing papers explorer
-
Can Multimodal Large Language Models Truly Understand Small Objects?
Current MLLMs show weak performance on small object understanding tasks, but fine-tuning with the new SOU-Train dataset measurably improves their capabilities.
-
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
A V-sink/L-sink taxonomy plus a frozen-backbone, NTP-trained layer-wise key gate (LSG) improves LLaVA-1.5-7B by up to +1.55pp on MMStar and +3.08pp on CVBench.
-
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
StochasT uses stochastic clustering of language tasks into varying turn depths for the same image to improve LVLMs on both single-turn and multi-turn scenarios without discarding data.
-
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
CoLT replaces text chain-of-thought in multimodal LLMs with three supervised latent vectors, improving average accuracy from 75.7 to 79.1 across eight benchmarks while cutting text-decoding time 22.6x.
-
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization
HyLaR interleaves discrete text generation with continuous visual latent representations and optimizes them via a decoupled RL algorithm using vMF distributions, improving fine-grained visual reasoning.
-
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
OpenSpatial supplies a principled open-source data engine and 3-million-sample dataset that raises spatial-reasoning model performance by an average of 19 percent on benchmarks.