Decoupled pose-text fusion on MM-Conv reaches 31.9% top-1 accuracy and shows a learned gate changing policy based on category access in text, serving as a diagnostic against category-representation artifacts.
MM-Conv: A multimodal dataset and bench- mark for context-aware grounding in 3D dialogue,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
Decoupled pose-text fusion on MM-Conv reaches 31.9% top-1 accuracy and shows a learned gate changing policy based on category access in text, serving as a diagnostic against category-representation artifacts.