A contextual multimodal document retrieval benchmark (CMDR-Bench) and embedding model (CMDR-Embed) that jointly encodes multiple document pages and splits them into page-level representations, trained with a context-aware contrastive objective, outperforming non-contextual baselines by 13–16 nDCG@5.
Title resolution pending
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 1polarities
background 1representative citing papers
HEE is a training-free, model-agnostic method for high-resolution visual perception in MLLMs using hierarchical entity exploration with dual scoring, detection, clustering, and backtracking.
SOCO is a new benchmark for semantic object correspondence that provides taxonomy, annotations, and language labels to evaluate part-level understanding in vision and multimodal foundation models.
MSCoT uses multi-scale hierarchical token prediction, multi-scale guidance, and a token refiner to deliver SOTA text-to-motion control with 48% FID gain, 61% lower error, and 10x faster inference on HumanML3D.
Ilov3Splat learns view-consistent CLIP and instance feature fields on 3D Gaussians to support open-vocabulary object selection and segmentation without category labels.
Visual attention in MLLMs freezes early during decoding; IVE breaks that inertia by exciting emergent visual tokens and penalizing persistent ones, reducing cognitive relation hallucinations without training.
citing papers explorer
-
CMDR: Contextual Multimodal Document Retrieval
A contextual multimodal document retrieval benchmark (CMDR-Bench) and embedding model (CMDR-Embed) that jointly encodes multiple document pages and splits them into page-level representations, trained with a context-aware contrastive objective, outperforming non-contextual baselines by 13–16 nDCG@5.
-
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
HEE is a training-free, model-agnostic method for high-resolution visual perception in MLLMs using hierarchical entity exploration with dual scoring, detection, clustering, and backtracking.
-
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
SOCO is a new benchmark for semantic object correspondence that provides taxonomy, annotations, and language labels to evaluate part-level understanding in vision and multimodal foundation models.
-
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
MSCoT uses multi-scale hierarchical token prediction, multi-scale guidance, and a token refiner to deliver SOTA text-to-motion control with 48% FID gain, 61% lower error, and 10x faster inference on HumanML3D.
-
Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting
Ilov3Splat learns view-consistent CLIP and instance feature fields on 3D Gaussians to support open-vocabulary object selection and segmentation without category labels.
-
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
Visual attention in MLLMs freezes early during decoding; IVE breaks that inertia by exciting emergent visual tokens and penalizing persistent ones, reducing cognitive relation hallucinations without training.