A training-free pipeline that summarizes egocentric video clips into a few kilobytes of text per minute and answers multiple-choice episodic memory questions with an LLM reasoner reaches 56.0% accuracy on QAEgo4D-Closed, matching a trained state-of-the-art system.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?
A training-free pipeline that summarizes egocentric video clips into a few kilobytes of text per minute and answers multiple-choice episodic memory questions with an LLM reasoner reaches 56.0% accuracy on QAEgo4D-Closed, matching a trained state-of-the-art system.