HIVE enables iterative latent reasoning in multimodal models by injecting hierarchical visual cues from global to fine-grained details directly into latent representations.
URL http://proceedings
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
T3S is a new semantic similarity score for processed images that decomposes semantics into foreground entities, background entities, and relations, outperforming fidelity metrics on COCO and SPA-Data.
citing papers explorer
-
Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
HIVE enables iterative latent reasoning in multimodal models by injecting hierarchical visual cues from global to fine-grained details directly into latent representations.
-
Beyond Fidelity: Semantic Similarity Assessment in Low-Level Image Processing
T3S is a new semantic similarity score for processed images that decomposes semantics into foreground entities, background entities, and relations, outperforming fidelity metrics on COCO and SPA-Data.