Decoupled pose-text fusion on MM-Conv reaches 31.9% top-1 accuracy and shows a learned gate changing policy based on category access in text, serving as a diagnostic against category-representation artifacts.
Bottom up top down detection transformers for language grounding in images and point clouds,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
Decoupled pose-text fusion on MM-Conv reaches 31.9% top-1 accuracy and shows a learned gate changing policy based on category access in text, serving as a diagnostic against category-representation artifacts.