REVIEW 6 cited by
LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval. LifelongMemory generates concise video activity descriptions of the camera wearer and leverages the zero-shot capabilities of pretrained large language models to perform reasoning over long-form video context. Furthermore, LifelongMemory uses a confidence and explanation module to produce confident, high-quality, and interpretable answers. Our approach achieves state-of-the-art performance on the EgoSchema benchmark for question answering and is highly competitive on the natural language query (NLQ) challenge of Ego4D. Code is available at https://github.com/agentic-learning-ai-lab/lifelong-memory.
Forward citations
Cited by 6 Pith papers
-
Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks
Pro²Assist uses multimodal egocentric perception from AR glasses to track fine-grained progress in long-horizon procedural tasks and deliver timely proactive assistance, outperforming baselines by over 21% in action u...
-
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
ViSAGE builds entity-centered, self-correcting memories for long-form video understanding and reports state-of-the-art accuracy on M3-Bench and Video-MME-long.
-
Native Active Perception as Reasoning for Omni-Modal Understanding
OmniAgent turns long-video understanding into a query-driven observe-think-act loop with a persistent text memory, outperforming larger passive models on LVBench.
-
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
A 3B agent trained with supervised tool-use traces and reinforcement learning answers week-long egocentric video questions by dynamically selecting hierarchical retrieval, video-LLM, and VLM tools.
-
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
An ensemble of LLMs with confidence filtering and low-confidence re-reasoning reaches 77% accuracy on the EgoSchema benchmark, up from 75% for the prior HCQA system.
-
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles
A training-free ensemble of commercial VLMs with prompt and chain-of-thought engineering reaches 79% on EgoSchema, ranking 2nd in the CVPR 2025 challenge, but the ensemble weights are fit to the test labels.
Discussion (0). Continue with ORCID to comment.