A method that compresses a sequence of textual scene graphs into a small set of latent tokens, enabling an LLM to answer situated questions about dynamic scenes with state-of-the-art accuracy on STAR and AGQA2.0.
Search3d: Hierarchical open-vocabulary 3d segmenta- tion,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes
A method that compresses a sequence of textual scene graphs into a small set of latent tokens, enabling an LLM to answer situated questions about dynamic scenes with state-of-the-art accuracy on STAR and AGQA2.0.